Sound image control apparatus, sound image control method, and program
The sound image control device and method interactively control sound image positions and spatial spread using motion capture and speaker arrays, addressing limitations in existing technologies to enhance interactivity and immersion in virtual reality.
Patent Information
- Application Number
- JP2024137044
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Existing sound image control technologies lack the ability to interactively control the position and spatial spread of sound images in response to user movements, limiting interactivity in virtual reality applications.
A sound image control device and method that uses a speaker array with a virtual sound image generation filter and motion capture to control the spatial spread of sound images based on user movements, employing FIR and IIR filters, delay devices, and amplifiers to generate interactive three-dimensional sound experiences.
Enables real-time interactive control of sound image positions and spatial spread, enhancing user engagement and providing immersive audio experiences in virtual reality environments.
Smart Images

Figure 2026033940000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a sound image control device, a sound image control method, and a program, and in particular to a sound image control device, a sound image control method, and a program that are capable of controlling the generation of a sound image so as to further improve interactivity. [Background technology]
[0002] Conventionally, in linear speaker arrays where multiple speaker units are arranged in a line, wave field synthesis (WFS) or the spectral division method is used to generate sound images in space that include a sense of depth, such as behind or in front of the speakers, while in spherical speaker arrays where multiple speaker units are arranged in a spherical shape, near-field compensated higher order ambisonics (NFC-HOA) is used to generate sound images in space that include a sense of depth, such as behind or in front of the speakers.
[0003] For example, Patent Document 1 discloses a sound field reproduction device that uses a speaker array consisting of a plurality of speaker units to reproduce a sound field formed by a moving sound source. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-122414 Summary of the Invention [Problem to be solved by the invention]
[0005] Incidentally, by combining the wave field synthesis technology, spectrum division method, and focused sound source methods such as NFC-HOA mentioned above with motion capture used in the field of virtual reality, it is thought that it will be possible to interactively control the position of the sound image in space in accordance with the user's body movements.For example, it is thought that further improvements in interactivity can be achieved by interactively controlling the spatial spread of the sound image as well.
[0006] The present disclosure has been made in view of such circumstances, and makes it possible to control the generation of a sound image so as to further improve interactivity. [Means for solving the problem]
[0007] A sound image control device according to one aspect of the present disclosure includes a virtual sound image generation filter that generates a virtual sound image in space using a speaker array consisting of multiple speaker elements, and a control unit that controls the spatial spread of the sound image in accordance with the movements of one or more body parts obtained by motion capture.
[0008] A sound image control method or program according to one aspect of the present disclosure includes generating a virtual sound image in space using a speaker array consisting of a plurality of speaker elements, and controlling the spatial spread of the sound image in accordance with the movements of one or more body parts obtained by motion capture.
[0009] In one aspect of the present disclosure, a virtual sound image is generated in space using a speaker array consisting of multiple speaker elements, and the spatial spread of the sound image is controlled in accordance with the movement of one or more body parts obtained by motion capture. [Effects of the Invention]
[0010] According to one aspect of the present disclosure, it is possible to control the generation of a sound image so as to further improve interactivity.
[0011] The effects described here are not necessarily limited to those described herein, and may be any of the effects described in this disclosure. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a real-time virtual sound image generation system using a linear speaker array. [Figure 2] FIG. 1 is a block diagram showing an example of the configuration of a real-time virtual sound image generation system using a spherical speaker array. [Figure 3] 1A and 1B are diagrams illustrating a primary sound field, which is sound pressure from a point sound source, and a secondary sound field, which is sound pressure reproduced inside a sphere array. [Figure 4] FIG. 1 is a diagram showing an example of an image of a spherical shell distribution f(θ,φ) and a circular ring distribution f(θ,φ). [Figure 5] FIG. 10 is a diagram illustrating rotation in the coefficient domain. [Figure 6] 4 is a flowchart illustrating a sound image control process according to a first sound image control method. [Figure 7] FIG. 10 is a diagram showing an example of two uncorrelated signals generated from one original signal using a Hilbert transformer. [Figure 8] FIG. 10 is a diagram showing an example of two uncorrelated signals generated from one original signal using an uncorrelated FIR filter. [Figure 9] FIG. 10 is a diagram showing an example of two uncorrelated signals generated from one original signal using a comb filter. [Figure 10] 1 is a block diagram illustrating an example of the configuration of an embodiment of a computer to which the present technology is applied. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, specific embodiments to which the present technology is applied will be described in detail with reference to the drawings.
[0014] <Configuration example of a real-time virtual sound image generation system using a linear speaker array> With reference to FIG. 1, a real-time virtual sound image generation system that controls in real time a virtual sound image (focused sound source) generated in space using a linear speaker array will be described.
[0015] The real-time virtual sound image generation system 11 shown in FIG. 1 comprises a sound source 21, a virtual sound image generation filter 22, a linear speaker array 23, and a motion capture device 24.
[0016] The real-time virtual sound image generation system 11 can generate a virtual sound image in space by outputting a sound time signal U(t) output from a sound source 21 through a virtual sound image generation filter 22 from a linear speaker array 23 configured by arranging a plurality of speaker elements 34 (s speaker elements 34-1 to 34-s in the illustrated example) in a line. Here, t in the sound time signal U(t) is a discrete time. The real-time virtual sound image generation system 11 then calculates the position of the virtual sound image generated using the linear speaker array 23 (hereinafter referred to as the virtual sound image position r PS The planar movement of the robot (referred to as "the robot") can be interactively controlled in accordance with the position of the user's hand, which is acquired in real time using the motion capture device 24.
[0017] For example, if the normal derivative of the sound on the linear boundary line along the linear speaker array 23 and the transfer function are known, the subsequent sound field p(r) can be controlled based on the Rayleigh integral of the first kind, as shown in the following equation (1).
[0018]
number
[0019] In this formula (1), r S is the position of each of the speaker elements 34-1 to 34-s, and r PS is the position of the virtual sound image generated by the real-time virtual sound image generation system 11.
[0020] Then, assuming an absorption-type focused sound source in which sound is absorbed into the virtual sound image (which is not actually possible) and sound is expelled from the virtual sound image, the filter W(r S ,r PS , ω) can be derived as shown in the following equation (2).
[0021]
number
[0022] In this equation (2), ω is the frequency of the sound output from the sound source 21, and H1 (1) is the Hankel function of the first kind.
[0023] For example, as shown in equation (2), the filter W(r S ,r PS ,ω) is the frequency ω and the virtual sound image position r PS Since it depends on the time domain filter, a numerical inverse Fourier transform is required, making it somewhat difficult to compute and convolve the time domain filter in real time.
[0024] Therefore, in the real-time virtual sound image generation system 11, the virtual sound image position r is set in accordance with the position of the user's hand. PS In order to control in real time, it is necessary to convert the equation into the time domain. Therefore, by approximating the Hankel function as an exponential function, the above equation (2) can be expressed as the following equation (3).
[0025]
number
[0026] Furthermore, by performing an inverse Fourier transform on the entire equation (3), the filter W(r S ,r PS,ω) is the filter q(t) (=IFFT(√(ω / c / 2πj))), the scalar gain g(r S ,r PS ), fractional delay D(r S ,r PS ,Z) can be expressed as a scalar gain g(r S ,r PS ) and fractional delay D(r S ,r PS , Z) is the speaker position r S and the virtual sound image position r PS Depends on.
[0027]
number
[0028] 1, the virtual sound image generating filter 22 can be configured by an FIR filter 31, delay devices 32-1 to 32-s, and amplifiers 33-1 to 33-s. The FIR filter 31 is a system-dependent fixed filter. The delay amount D of the delay devices 32-1 to 32-s and the gain g of the amplifiers 33-1 to 33-s change every time in accordance with the position of the user's hand acquired by the motion capture device 24.
[0029] The real-time virtual sound image generation system 11 configured as described above interactively generates a virtual sound image position r in accordance with the position of the user's hand acquired using the motion capture device 24. PS It is possible to control the planar movement of the robot in real time.
[0030] <Configuration example of a real-time virtual sound image generation system using a spherical speaker array> With reference to FIG. 2, a real-time virtual sound image generation system that controls in real time a virtual sound image (focused sound source) generated in space using a spherical speaker array will be described.
[0031] The real-time virtual sound image generation system 41 shown in FIG. 2 comprises a sound source 51, a virtual sound image generation filter 52, a spherical speaker array 53, and a motion capture device .
[0032] The real-time virtual sound image generation system 41 can generate a virtual sound image in the air (reproduce stereophonic sound) by outputting a sound time signal U(t) output from a sound source 51 through a virtual sound image generation filter 52 from a spherical speaker array 53 configured by arranging a plurality of speaker elements (in FIG. 2, each black circle represents a speaker element) in a spherical shape. The real-time virtual sound image generation system 41 can generate a virtual sound image in the air (reproduce stereophonic sound) by outputting a virtual sound image position r PS The three-dimensional movement of the virtual sound image can be interactively controlled in accordance with the position of the user's hand, which is acquired in real time using the motion capture device 54. In addition, the real-time virtual sound image generation system 41 uses a spherical coordinate system (r, Ω) = (r, θ, φ), and the virtual sound image position r PS =(r PS ,θ PS ,φ PS ) is the virtual sound image distance r PS and the virtual sound image angle Ω PS =(θ PS ,φ PS )
[0033] For example, when the real-time virtual sound image generation system 41 does not control the sense of distance, it can use a focused sound source method such as VBAP (Vector Based Amplitude Panning) or HOA (Higher Order Ambisonics), and when it controls the sense of distance, it can use a focused sound source method such as NFC-HOA + Realtime processing.
[0034] Here, HOA is a method of decomposing, analyzing, and synthesizing sound using a spherical or circular speaker array and a spatial Fourier transform of the angular directivity pattern. When the sense of distance is not controlled, sound p(r,θ,φ) arriving from outside is expressed by the following equation (5).
[0035]
number
[0036] In this equation (5), j n (·) is the spherical Bessel function, Y n m (·) is a spherical harmonic function, and P n m (·) is the associated Legendre function.
[0037] The spherical harmonic expansion is a spatial Fourier transform of angles, and is expressed as the spherical harmonic function Y of angles θ and φ. n m (θ,φ) is the basis for any property f(θ,φ) on the sphere, and the expansion coefficients f nm The inverse transformation and transformation are shown in the following equation (6).
[0038]
number
[0039] In this equation (6), n is the order, N is the maximum order determined by the number of speaker elements, and S 2 is the unit sphere.
[0040] For example, when reproducing a sound coming from a certain direction using an HOA, if one sound source is placed in the direction of angles θ and φ, the characteristic f(θ,φ) is expressed by the following equation (7), and the expansion coefficient f nm is expressed by the following equation (8).
[0041]
number
number
[0042] In this equation (7), the number δ(·) is the Dirac delta function, and in equation (8), (θ t ,φ t ) is the direction from which the sound is coming.
[0043] Then, in the HOA, the gain G of the l-th speaker driving signal that combines these l is expressed by the following equation (9).
[0044]
number
[0045] In this equation (9), l is the speaker index.
[0046] Next, we derive an NFC-HOA filter to introduce distance control.
[0047] For example, the primary sound field p(α,Ω) is the sound pressure from a point sound source as shown in Figure 3A. α ) is expressed by the following equation (10), and is the secondary sound field p^(α,Ω) which is the sound pressure reproduced inside the sphere array as shown in Figure 3B. α ) is expressed by the following equation (11).
[0048]
number
number
[0049] Then, a filter W that matches equations (10) and (11) is s (ω) can be obtained by expanding equations (10) and (11) into spherical harmonics and matching coefficients of the same order. For example, the NFC-HOA filter W s (ω,r PS ) is expressed by a radial filter and an angular gain as shown in the following equation (12):
[0050]
number
[0051] In this equation (12), r ps is the virtual sound source distance (radius), r is the speaker distance (radius), N is the expansion order, L is the number of speakers, and h n (2) (kr) is the spherical Hankel function, Y n m (Ωs) is a spherical harmonic function, and k=ω / c is the wave number. The radial filter and angular gain depend on the position of the speaker and the virtual sound image. The radial filter is also frequency dependent, so a Fourier transform is required to convert it from a frequency domain filter to an FIR filter.
[0052] Here, the z-domain (time-domain) expression of the radial filter is expressed by converting the frequency-domain filter into a z-domain filter via Laplace transform. For example, the Laplace-domain expression of a third-order radial filter is expressed as shown in the following equation (13).
[0053]
number
[0054] Then, by performing a matched z-transform on equation (13), it can be expressed by one gain, one fractional delay, and an n=3-th order IIR filter, as in the following equation (14).
[0055]
number
[0056] As shown in this equation (14), one gain corresponds to the magnitude relative to the distance to the virtual sound image, one fractional delay corresponds to the arrival delay relative to the distance to the virtual sound image, and the n=3 order IIR filter corresponds to the ratio at which the speaker drive gains of each order are synthesized for each frequency. Also, the n=3 order IIR filter has a fixed denominator, and α l is calculated in advance, and r ps can only be changed.
[0057] Therefore, as shown in FIG. 2, the virtual sound image generating filter 52 can be configured by an amplifier 61, a fractional delay unit 62, and IIR filters 63-0 to 63-N, and the IIR filters 63-0 to 63-N are connected to the speaker elements that make up the spherical speaker array 53 via an (N+1)×L gain matrix 64.
[0058] The real-time virtual sound image generation system 41 configured as described above interactively generates a virtual sound image position r in accordance with the position of the user's hand acquired using the motion capture device 54. PS The three-dimensional movement of the object can be controlled in real time.
[0059] Furthermore, the virtual sound image generating filter 52 is provided with a control unit 65 for interactively controlling the spatial extent of the virtual sound image (e.g., the sense of expanse of the sound source, the apparent width of the sound source, etc.) in accordance with the movements of one or more body parts acquired by the motion capture device 54.
[0060] <Interactive control of the spatial spread of virtual sound images> 4 to 9, interactive control of the spatial spread of a virtual sound image in the real-time virtual sound image generation system 41 will be described.
[0061] For example, the real-time virtual sound image generation system 41 can interactively control the sense of spaciousness of the sound source and the apparent width of the sound source according to the size of the arms of a user spread out inside the spherical speaker array 53. This allows the real-time virtual sound image generation system 41 to spread the sound source to an area surrounding the user inside the spherical speaker array 53, for example.
[0062] Here, the sense of spaciousness of a sound source is defined as the apparent source width (ASW) and the feeling of being surrounded by sound (LEV: Listener Envelopment). For example, the apparent source width is the size of the perceived sound image, and is influenced by factors such as the sound pressure level and the magnitude of low-frequency components. The feeling of being surrounded by sound is the feeling that the surroundings are filled with things other than the apparent sound source, and is influenced by factors such as the front-to-rear energy ratio and the direction of arrival of reflected sound.
[0063] The control unit 65 can then interactively control the spatial spread of the virtual sound image using the first to third sound image control methods described below.
[0064] <First sound image control method> For example, the control unit 65 can interactively control the spatial spread of the virtual sound image using a first sound image control method in which the spherical harmonic coefficients are calculated in advance as the spatial distribution of a single sound source, such as the sense of spread of the sound source or the apparent width of the sound source, and the position of the sound source is also changed in the spherical harmonic coefficient domain.
[0065] With the first sound image control method, the control unit 65 can, for example, control the sound image to expand spatially from moment to moment in response to an action of opening a certain part (for example, a user opening their hands), and can control the sound image to narrow spatially from moment to moment in response to an action of closing a certain part (for example, a user closing their hands). Furthermore, the control unit 65 can control the sound image to expand spatially from moment to moment in response to an action of increasing the distance between two or more parts (for example, a user spreading their hands), and can control the sound image to narrow spatially from moment to moment in response to an action of decreasing the distance between two or more parts (for example, a user bringing their hands closer). At this time, the control unit 65 can change the sense of spaciousness of the sound source, the apparent width of the sound source, and the like, according to the distance between those two or more parts.
[0066] First, the control unit 65 sets a distribution representing the spread of the interactively controlled sound, such as a spherical shell distribution in which sound is distributed in a spherical shell shape, or a circular distribution in which sound is distributed in a circular shape (distribution with a hollow center).
[0067] For example, the spherical shell distribution f(θ, φ) is set as shown in the following equation (15) with the distribution boundary angle being α.
[0068]
number
[0069] Then, the control unit 65 performs a spatial Fourier transform on the equation (15) to obtain the expansion coefficient f nm Calculate.
[0070]
number
[0071] In this equation (16), δ ij is the Kronecker delta, and P n (·) is the Legendre polynomial.
[0072] Figure 4A shows an image of the spherical shell distribution f(θ,φ) in the synthesis result for N=5.
[0073] The circular distribution f(θ, φ) is set as shown in the following equation (17) with the distribution boundary angle being α.
[0074]
number
[0075] Then, the control unit 65 performs a spatial Fourier transform on the equation (17) to obtain the expansion coefficient f nm Calculate.
[0076]
number
[0077] In this equation (18), e jmφ The integral of becomes 0 when m≠0 due to periodicity.
[0078] Figure 4B shows an image of the circular distribution f(θ,φ) in the synthesis result for N=5.
[0079] Here, the expansion coefficient f nm is the zenith center (0,0) as shown in the upper part of Figure 5, and is an arbitrary angle (θ t ,φ t ) to perform a rotation in the coefficient domain using the Wigner-D rotation matrix.
[0080] The Wigner-D function is an element of a matrix that performs rotation using Euler angles (α, β, γ), and is expressed as shown in the following equation (19). For example, when the rotation angle is an arbitrary angle (θ t ,φ t ), the rotation to the Euler angle is (0,2π-θ t ,φ t )
[0081]
number
[0082] The Wigner-D rotation matrix is a submatrix D that arranges the Wigner-D functions shown in equation (19). n The matrix D is a diagonal matrix of wig and the coefficient vector g after rotation N is expressed by the following equation (20).
[0083]
number
[0084] In this equation (20), f N is a column vector in which the coefficients are arranged in ascending order.
[0085] For example, the submatrix D consisting of Wigner-D functions 1 is expressed as the following equation (21).
[0086]
number
[0087] Then, the control unit 65 calculates the coefficient vector g after rotation. N The gain is calculated using, for example, the following equation (22).
[0088]
number
[0089] In this equation (22), g N is a column vector in which the coefficients are arranged in ascending order, and α N is the N-th order Max-rE coefficient.
[0090] As described above, the control unit 65 sets the spherical shell distribution or the circular distribution and calculates the expansion coefficient f nm Calculate the angle (θ t ,φ t ) after rotation in the coefficient domain to the coefficient vector g N By calculating the gain using this, the spatial spread of the virtual sound image can be interactively controlled.
[0091] The sound image control process according to the first sound image control method will be described with reference to the flowchart shown in FIG.
[0092] In step S11, the control unit 65 sets the spherical distribution f(θ, φ) as shown in the above formula (15) or the circular distribution f(θ, φ) as shown in the above formula (17).
[0093] In step S12, when the spherical shell distribution f(θ, φ) is set, the control unit 65 calculates the expansion coefficient f as shown in the above equation (16). nm When the circular distribution f(θ,φ) is set, the expansion coefficient f nm Calculate.
[0094] In step S13, the control unit 65 uses the above-mentioned Wigner-D rotation matrix to rotate the zenith center (0,0) at an arbitrary angle (θ t ,φ t ) in the coefficient domain.
[0095] In step S14, the control unit 65 calculates the coefficient vector g N The gain is calculated using the above equation (22).
[0096] Each step in the sound image control process described above is executed for a predetermined time interval, which causes the sound image distribution to expand and narrow spatially from moment to moment, allowing the control unit 65 to control the apparent width of the sound source and the temporal progression of the sense of spaciousness.
[0097] <Second sound image control method> For example, the control unit 65 can interactively control the spatial spread of the virtual sound image using a second sound image control method in which two or more uncorrelated signals are generated from one sound source signal and presented at spatially separated positions. In this case, the positions at which the two or more uncorrelated signals are presented are the positions of the user's body acquired by the motion capture device 54, and sounds are presented using virtual sound sources created by the focused sound source method.
[0098] For example, in the real-time virtual sound image generation system 41, the degree of uncorrelation can be changed according to the distance between two points acquired by the motion capture device 54 inside the spherical speaker array 53, that is, according to the distance between two positions where two uncorrelated signals are presented. Here, the degree of uncorrelation can be changed by changing the mixing ratio of correlated signals and uncorrelated signals. When the distance between the two points exceeds a certain level, the uncorrelation coefficient is increased so that the respective sounds can be heard separately.
[0099] Furthermore, the control unit 65 can use an all-pass filter, a convolution of uncorrelated FIR (Finite Impulse Response) filter coefficients, a Hilbert transform, or a notch filter, or a combination of these, as a function for generating two or more uncorrelated signals from one sound source signal.
[0100] FIG. 7 shows an example of two uncorrelated signals generated from one original signal using a Hilbert transformer.
[0101] As shown in Figure 7, let x(t) be the original signal, and y(t) be the signal whose phase is shifted by 90 degrees from the original signal generated using a Hilbert transformer. If x(t) and y(t) are plotted on a plane, the result will be as shown on the right side of Figure 7.
[0102] FIG. 8 shows an example of two uncorrelated signals generated from one original signal using a uncorrelated FIR filter.
[0103] Figure 9 shows an example of two uncorrelated signals generated from a single source signal using a comb filter. A comb filter is a filter that prevents the peaks and valleys of its frequency characteristics from overlapping, making the frequencies independent and therefore uncorrelated.
[0104] As described above, the control unit 65 can interactively control the spatial extent of the virtual sound image using the second sound image control method in which the control unit 65 generates two or more uncorrelated signals from one sound source signal and presents them at spatially separated positions.
[0105] For example, the positions of two or more body parts (for example, both hands) are acquired by the motion capture device 54, and the control unit 65 forms two virtual sound images according to the two pieces of position information. Then, the control unit 65 uses, as the original signal output from each virtual sound image, uncorrelated signals created from one sound source signal, not limited to those obtained by the above-mentioned Hilbert transformer, uncorrelated FIR filter, comb filter, etc.
[0106] Furthermore, the control unit 65 may change the degree of correlation depending on the distance between two points acquired by the motion capture device 54. For example, when the distance between two points is short, the control unit 65 adds a high proportion of the original sound to both points, as shown in the following equation (23), in order to strengthen the correlation.
[0107]
number
[0108] For example, in this equation (23), there is no correlation when α = 1, and there is high correlation when α = 0. This allows the control unit 65 to control the virtual sound image so that, for example, when the correlation is high, the sound is heard as a single sound source coming from the center even if the left and right hands are far apart.
[0109] <Third sound image control method> For example, the control unit 65 can interactively control the spatial spread of the virtual sound image using a third sound image control method that spreads uncorrelated signals on a spherical surface.
[0110] For example, the first sound image control method described above is a method of spreading one sound source signal in space, while the second sound image control method described above is a method of arranging two or more uncorrelated signals, one by one, although the number is not limited to two.
[0111] Here, in order to spatially expand the virtual sound image, it is conceivable to create a large number of uncorrelated signals and place them at a large number of sound source positions, but this would result in an enormous amount of calculation. Therefore, in the third sound image control method, the sound output from each speaker element is limited to the positions where the speaker elements of the speaker array are placed, and the volume at which these sounds are output is made into an uncorrelated signal, and the gain (the above-mentioned equation (22)) calculated in the above-mentioned first sound image control method is used. This makes it possible to reduce the amount of calculation.
[0112] On the other hand, the first sound image control method described above does not have distance control, so there is no sense of distance even for the center position of the virtual sound image. Therefore, in the third sound image control method, we propose a method in which a virtual sound source is generated using NFC-HOA for the center position of the virtual sound image, and the sense of spaciousness is generated using HOA.
[0113] That is, in the third sound image control method, when distance control is not performed, the control unit 65 prepares uncorrelated signals for each speaker element constituting the spherical speaker array 53, multiplies the signals by an HOA coefficient, and outputs the results. Also, the control unit 65 performs distance control using NFC-HOA at the center position of the virtual sound image, thereby allowing the sense of spaciousness to be expressed by HOA without distance control.
[0114] By using the first to third sound image control methods described above, the real-time virtual sound image generation system 41 can interactively control the sense of sound source spread and the apparent width of the sound source, and can provide the listener with, for example, a spatial spread of sound. This allows the real-time virtual sound image generation system 41 to provide a sense of envelopment and new experiences such as expanding or separating sounds.
[0115] Furthermore, for example, in the entertainment field, the real-time virtual sound image generation system 41 can present the sensation of spreading or tearing a single sound source. Also, in the field of highly realistic audio reproduction, the real-time virtual sound image generation system 41 can provide the sensation of spreading a sound source as a time transition. Specifically, it can reproduce the sound of fireworks falling or the sound of rain hitting an umbrella in accordance with the movement of the umbrella.
[0116] <Computer configuration example> Next, the above-described series of processes (sound image control method) can be performed by hardware or software. When the series of processes is performed by software, a program constituting the software is installed in a general-purpose computer or the like.
[0117] FIG. 10 is a block diagram showing an example of the configuration of an embodiment of a computer in which a program for executing the above-described series of processes is installed.
[0118] The program can be recorded in advance on the hard disk 105 or ROM 103 as a recording medium built into the computer.
[0119] Alternatively, the program can be stored (recorded) on a removable recording medium 111 driven by the drive 109. Such a removable recording medium 111 can be provided as a so-called package software. Here, examples of the removable recording medium 111 include a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disk, a DVD (Digital Versatile Disc), a magnetic disk, and a semiconductor memory.
[0120] The program can be installed into the computer from the removable recording medium 111 as described above, or can be downloaded to the computer via a communication network or a broadcasting network and installed on the built-in hard disk 105. That is, the program can be transferred to the computer wirelessly from a download site via an artificial satellite for digital satellite broadcasting, or transferred to the computer by wire via a network such as a LAN (Local Area Network) or the Internet.
[0121] The computer includes a CPU (Central Processing Unit) 102 , to which an input / output interface 110 is connected via a bus 101 .
[0122] When a user inputs a command by operating input unit 107 via input / output interface 110, CPU 102 executes a program stored in ROM (Read Only Memory) 103 in accordance with the command. Alternatively, CPU 102 loads a program stored on hard disk 105 into RAM (Random Access Memory) 104 and executes it.
[0123] As a result, CPU 102 performs processing according to the flowchart described above or processing performed by the configuration of the block diagram described above. CPU 102 then outputs the processing results from output unit 106 via input / output interface 110, transmits them from communication unit 108, or records them on hard disk 105, as necessary.
[0124] The input unit 107 is made up of a keyboard, a mouse, a microphone, etc. The output unit 106 is made up of an LCD (Liquid Crystal Display), a speaker, etc.
[0125] In this specification, the processing performed by a computer according to a program does not necessarily have to be performed in chronological order according to the order described in the flowchart. In other words, the processing performed by a computer according to a program also includes processing that is executed in parallel or individually (for example, parallel processing or processing by objects).
[0126] The program may be processed by a single computer (processor), or may be distributed among multiple computers. Furthermore, the program may be transferred to and executed on a remote computer.
[0127] Furthermore, in this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0128] Also, for example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).
[0129] Furthermore, for example, this technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0130] Furthermore, for example, the above-described program can be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and can obtain the necessary information.
[0131] Also, for example, each step described in the above flowchart can be executed by one device or can be shared and executed by multiple devices. Furthermore, if one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices. In other words, multiple processes included in one step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as one step.
[0132] In addition, the processing of the steps of a program executed by a computer may be executed in chronological order according to the order described in this specification, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the processing of each step may be executed in an order different from the order described above. Furthermore, the processing of the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0133] It should be noted that the present technologies described in this specification can be implemented independently and singly, unless a contradiction arises. Of course, any two or more of the present technologies can also be implemented in combination. For example, part or all of the present technologies described in any embodiment can be implemented in combination with part or all of the present technologies described in other embodiments. Furthermore, part or all of any of the present technologies described above can also be implemented in combination with other technologies not described above.
[0134] It should be noted that the present embodiment is not limited to the above-described embodiment, and various modifications are possible within the scope of the gist of the present disclosure. Furthermore, the effects described in this specification are merely examples and are not intended to be limiting, and other effects may also be obtained. [Explanation of symbols]
[0135] 11 Real-time virtual sound image generation system, 21 Sound source, 22 Virtual sound image generation filter, 23 Linear speaker array, 24 Motion capture device, 31 FIR filter, 32-1 to 32-s Delay, 33-1 to 33-s Amplifier, 34-1 to 34-s Speaker element, 41 Real-time virtual sound image generation system, 51 Sound source, 52 Virtual sound image generation filter, 53 Spherical speaker array, 54 Motion capture device, 61 Amplifier, 62 Fractional delay, 63-0 to 63-N IIR filter, 64 Gain matrix, 65 Control unit
Claims
1. a virtual sound image generating filter that generates a virtual sound image in space using a speaker array configured with a plurality of speaker elements; a control unit that controls the spatial spread of the sound image in accordance with the movement of one or more body parts acquired by motion capture; A sound image control device comprising:
2. The control unit determines in advance the spatial distribution of one sound source as a spherical harmonic coefficient, and changes the position of the sound source in the spherical harmonic coefficient domain, thereby controlling the spatial spread of the sound image. The sound image control device according to claim 1 .
3. The control unit controls the sound image so as to expand spatially from moment to moment in response to an opening action of one of the parts, and controls the sound image so as to narrow spatially from moment to moment in response to a closing action of one of the parts. The sound image control device according to claim 2 .
4. The control unit controls the sound image to expand spatially from moment to moment in response to a movement that increases the distance between two or more of the body parts, and controls the sound image to narrow spatially from moment to moment in response to a movement that decreases the distance between two or more of the body parts. The sound image control device according to claim 2 .
5. The control unit generates two or more uncorrelated signals from one sound source signal and presents them at spatially separated positions at the body position acquired by motion capture, thereby controlling the spatial spread of the sound image. The sound image control device according to claim 1 .
6. The control unit changes the degree of uncorrelation of the signal generated from the sound source signal according to the distance between two positions where two uncorrelated signals are presented. The sound image control device according to claim 5 .
7. When the distance between the two positions presenting the two uncorrelated signals exceeds a certain value, the control unit increases the uncorrelation coefficient to separate the respective sounds. The sound image control device according to claim 5 .
8. The control unit uses an all-pass filter, a convolution of uncorrelated FIR (Finite Impulse Response) filter coefficients, a Hilbert transform, or a notch filter, or a combination of these, as a function for generating two or more uncorrelated signals from one sound source signal. The sound image control device according to claim 5 .
9. When distance control is not performed, the control unit prepares uncorrelated signals for each speaker element constituting the spherical speaker array, multiplies the signals by HOA (Higher Order Ambisonics) coefficients, and outputs the uncorrelated signals. The sound image control device according to claim 1 .
10. The control unit performs distance control at the center position of the sound image using NFC-HOA (Near-Field Compensated Higher Order Ambisonics). The sound image control device according to claim 9.
11. The sound image control device Generating a virtual sound image in space using a speaker array configured with a plurality of speaker elements; controlling the spatial spread of the sound image in accordance with the motion of one or more body parts acquired by motion capture; A sound image control method comprising:
12. The sound image control device's computer Generating a virtual sound image in space using a speaker array configured with a plurality of speaker elements; controlling the spatial spread of the sound image in accordance with the motion of one or more body parts acquired by motion capture; A program for executing sound image control processing including the above.
Citation Information
Patent Citations
Sound field reproduction device and program
JP2022122414A