Signal processing device and signal processing method
The signal processing device and method dynamically control the composite value of window functions for each speaker, addressing the mismatch in virtual sound source movement perception by adjusting parameters, thus aligning the listener's experience with the operator's intent, improving audio reproduction accuracy.
Patent Information
- Application Number
- PCT/JP2025/020066
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-06-03
- Publication Date
- 2025-12-26
AI Technical Summary
Conventional audio reproduction techniques struggle to accurately match the perceived movement of a virtual sound source with the intended movement due to the assumption of a constant composite value of window functions, leading to discrepancies between the operator's intent and listener's experience.
A signal processing device and method that apply a window function to each speaker, controlling the composite value of these functions to be non-constant, allowing for dynamic and accurate movement of the virtual sound source by adjusting parameters such as exponents in power and trigonometric functions to align with the operator's intent.
This approach enables precise control over the perceived movement of the virtual sound source, ensuring it aligns with the intended movement, enhancing the audio reproduction experience by applying various window functions to express intended movements effectively.
Smart Images

Figure JP2025020066_26122025_PF_FP_ABST
Abstract
Description
Signal processing device and signal processing method
[0001] The present technology relates to a signal processing device and a signal processing method, and more particularly to a signal processing device and a signal processing method that are capable of suitably expressing the movement of a virtual sound source.
[0002] 2. Description of the Related Art Conventionally, in audio reproduction, a technique called panning is known in which a virtual sound source is localized at an arbitrary position among a plurality of speakers (see, for example, Patent Document 1).
[0003] International Publication No. 2020 / 153092
[0004] When an operator attempts to change the position of a virtual sound source over time by panning, assuming that the composite value of a window function, which is a function that indicates the relationship between time and gain, is constant, it is difficult to match the movement of the virtual sound source perceived by the listener with the movement of the virtual sound source intended by the operator.
[0005] The present technology has been made in consideration of such circumstances, and makes it possible to suitably express the movement of a virtual sound source.
[0006] A signal processing device according to one aspect of the present technology includes a signal output unit that applies a window function, which is a function indicating the relationship between time and gain, to each speaker of an audio signal indicating the sound of content, and outputs the audio signal after the window function has been applied to the plurality of speakers, and a gain control unit that controls the window function applied to the audio signal by the signal output unit so that a composite value of the window functions for each speaker is not constant.
[0007] A signal processing method according to one aspect of the present technology includes applying a window function, which is a function indicating the relationship between time and gain, to each speaker of an audio signal indicating the sound of content; outputting the audio signal after the window function has been applied to the plurality of speakers; and controlling the window function applied to the audio signal so that a composite value of the window function for each speaker is not constant.
[0008] In one aspect of the present technology, a window function, which is a function indicating the relationship between time and gain, is applied to each speaker to an audio signal indicating the sound of content, the audio signal after application of the window function is output to the multiple speakers, and the window function applied to the audio signal is controlled so that the combined value of the window function for each speaker is not constant.
[0009] 1 is a block diagram illustrating an example of a configuration of an audio reproduction system according to an embodiment of the present technology. FIG. 1 is a first diagram illustrating an example of panning. FIG. 2 is a second diagram illustrating an example of panning. FIG. 3 is a diagram illustrating an example of movement of a virtual sound source. FIG. 4 is a diagram illustrating an example of a window function. FIG. 5 is a diagram illustrating an example of an expression of the movement of the virtual sound source. FIG. 6 is a diagram illustrating an example of an input signal when the virtual sound source is moved in a complex manner. FIG. 7 is a flowchart illustrating processing performed by the audio reproduction system. FIG. 8 is a diagram illustrating examples of the movement of the virtual sound source and function values of window functions corresponding to each time. FIG. 9 is a diagram illustrating the movement of the virtual sound source actually perceived by a listener when a window function such that the linear sum of the window functions between speakers is always constant is applied to an input signal. FIG. 10 is a diagram illustrating the movement of the virtual sound source actually perceived by a listener when a window function such that the sum of squares of the window functions between speakers is always constant is applied to an input signal. FIG. 11 is a diagram illustrating the reason why the movement of the virtual sound source perceived by a listener differs from the movement of the virtual sound source intended by an operator. FIG. 12 is a flowchart illustrating processing performed by the audio reproduction system when a gain control unit adds a correction term to the window function. FIG. 13 is a diagram illustrating the movement of a sound image actually perceived by a listener when a window function whose combined value is not constant is applied to an input signal. 10 is a flowchart illustrating processing performed by an audio reproduction system according to an embodiment of the present technology. FIG. 10 is a first diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 11 is a second diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 12 is a third diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 13 is a fourth diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 14 is a fifth diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 15 is a sixth diagram illustrating an example of a window function in which a composite value of the window functions between speakers is not constant. FIG. 16 is a diagram illustrating an example of switching of window functions according to the position of a listener. FIG. 17 is a flowchart illustrating processing performed by an audio reproduction system in a case where an audio amplifier automatically switches the window function for each scene. FIG. 18 is a flowchart illustrating processing performed by an audio reproduction system in a case where an operator manually switches the window function for each scene.
[0010] Hereinafter, embodiments for carrying out the present technology will be described. The description will be made in the following order: 1. Comparative Example Embodiment 2. Embodiment of the Present Technology
[0011] 1. Comparative Example Embodiment FIG. 1 is a block diagram showing an example configuration of an audio playback system according to an embodiment of the present technology.
[0012] The audio reproduction system of FIG. 1 comprises an audio amplifier 1 and speakers 2L and 2R.
[0013] The audio amplifier 1 is a signal processing device that controls the sound output from the speakers 2L and 2R. An audio signal (e.g., a monaural signal) obtained by an audio player (not shown) reproducing music content or the like is supplied to the audio amplifier 1 as an input signal S. The audio amplifier 1 amplifies the input signal S and drives the speakers 2L and 2R based on the amplified input signal S to output the sound of the content.
[0014] The audio amplifier 1 is composed of amplifier sections 11L and 11R, an input section 12, and a gain control section 13.
[0015] The amplifier 11L multiplies the input signal S by a gain GL for the speaker 2L, amplifies the signal, and supplies the amplified input signal S·GL to the speaker 2L as an output signal S′L. The amplifier 11R multiplies the input signal S by a gain GR for the speaker 2R, and supplies the amplified input signal S·GR to the speaker 2R as an output signal S′R.
[0016] The amplifiers 11L and 11R function as signal output units that output audio signals multiplied by gains for each speaker to a plurality of speakers.
[0017] The input unit 12 is configured as a switch, a button, a fader, a touch panel, a mixer, etc., and accepts operations input by an operator of the audio amplifier 1. The input unit 12 supplies information indicating the content of the operator's operation to the gain control unit 13.
[0018] The gain control unit 13 controls the gains GL and GR by which the input signal S is multiplied in the amplifier units 11L and 11R, for example, in accordance with the operation content input using the input unit 12. By controlling the gains GL and GR, the gain control unit 13 can localize a sound image at any position between the speaker 2L and the speaker 2R. Hereinafter, localizing a sound image at any position between the speaker 2L and the speaker 2R will be referred to as panning.
[0019] The speaker 2L is installed, for example, diagonally forward and to the left as viewed from the listener, and the speaker 2R is installed, for example, diagonally forward and to the right as viewed from the listener. The distance between the speaker 2L and the listener is, for example, equal to the distance between the speaker 2R and the listener. The operator of the audio amplifier 1 and the listener may be the same person or different people.
[0020] 1, an example in which two speakers 2L and 2R are provided around the listener has been described, but three or more speakers may be provided. In the following, when there is no need to particularly distinguish between the speaker 2L and the speaker 2R, they will simply be referred to as speaker 2.
[0021] In describing the signal processing performed by the audio amplifier 1 according to the embodiment of the present technology, conventional signal processing will be described below as a comparative example of the signal processing. Figures 2 and 3 are diagrams illustrating an example of panning.
[0022] As shown on the left side of Fig. 2A, if the gain GL for speaker 2L is set to 1.0 and the gain GR for speaker 2R is set to 0.0, the listener will hear the sound of the content coming from the position of speaker 2L. Also, as shown on the right side of Fig. 2A, if the gain GL for speaker 2L is set to 0.0 and the gain GR for speaker 2R is set to 1.0, the listener will hear the sound of the content coming from the position of speaker 2R.
[0023] 2B, if the gain GL for speaker 2L is set to 1.0 and the gain GR for speaker 2R is set to 1.0, the listener will hear the sound of the content from a virtual sound source F that is located, for example, between speaker 2L and speaker 2R and in front of the listener. In other words, the sound image of the sound of the content is localized at the position of the virtual sound source F.
[0024] Here, the gain of the virtual sound source F is greater than 1.0; in other words, the volume of the virtual sound source F is greater than the volume of the speaker 2 when sound is heard from the position of the speaker 2L or the speaker 2R described with reference to A in Figure 2.
[0025] As shown in Figure 2C, if the gain GL for speaker 2L is set to a value less than 1.0 and the gain GR for speaker 2R is set to the same value, the virtual sound source F will be located, for example, between speaker 2L and speaker 2R, directly in front of the listener.
[0026] Here, the gain of the virtual sound source F is 1.0; in other words, the volume of the virtual sound source F is equal to the volume of the speaker 2 when sound is heard from the position of the speaker 2L or the speaker 2R described with reference to Figure 2A.
[0027] When the gain GL and the gain GR are differentiated so that GR<GL<1.0 (there is a volume difference between the speakers 2L and 2R), the virtual sound source F is localized slightly to the left of the front as seen from the listener, as shown in FIG. 3A.
[0028] If GR=GL<1.0, the virtual sound source F is localized in front of the listener, as shown in FIG. 3B.
[0029] When the gain GL and the gain GR are differentiated so that GL<GR<1.0, the virtual sound source F is localized at a position slightly to the right of the front as seen from the listener, as shown in FIG. 3C.
[0030] In this way, by creating a difference between the gain GL and the gain GR, the gain control unit 13 can localize the sound image of the content sound (virtual sound source F) at any position between the speaker 2L and the speaker 2R.
[0031] Here, let us consider the case where the position of the virtual sound source is changed over time by panning.
[0032] FIG. 4 is a diagram showing an example of the movement of a virtual sound source.
[0033] 4, the position of the virtual sound source F is expressed, for example, by an orientation θ relative to the center of the listener's head. Here, an orientation of 0 degrees indicates the position of the speaker 2L, an orientation of φ degrees indicates the front direction of the listener, and an orientation of 2φ degrees indicates the position of the speaker 2R.
[0034] For example, consider a case where a virtual sound source F is moved from the position of speaker 2L to the position of speaker 2R along an arc connecting speaker 2L and speaker 2R, with the center of the listener's head at the center, between time 0 and time T. In this case, the relationship between time t (0≦t≦T) and direction θ (0≦θ≦2φ) is expressed by the following equation (1).
[0035] θ=(2φ / T)t...(1)
[0036] As shown on the right side of FIG. 4, the relationship between time t and direction θ in equation (1) is linear, and the virtual sound source F moves at a constant speed.
[0037] The gain control unit 13 changes the gains GL and GR over time so that the listener perceives the virtual sound source F as moving on an arc connecting the speakers 2L and 2R. Specifically, the gain control unit 13 controls the amplifiers 11L and 11R to apply a window function for the speaker 2L and a window function for the speaker 2R, which are functions indicating the relationship between time t and gain, to the input signal S.
[0038] When the relationship between time t and azimuth θ is linear, the window function can also be said to be a function that indicates the relationship between azimuth θ (position of virtual sound source F) and gain. In other words, when the relationship between time t and azimuth θ is linear, the shape of the graph of the window function that indicates the relationship between time t and gain and the shape of the graph of the window function that indicates the relationship between azimuth θ (position of virtual sound source F) and gain match.
[0039] FIG. 5 is a diagram illustrating an example of a window function.
[0040] Fig. 5A shows a window function in which the linear sum of the window function GL(θ) for speaker 2L and the window function GR(θ) for speaker 2R is always constant. In Fig. 5A, the horizontal axis represents the azimuth θ, and the vertical axis represents the gain. In the example of Fig. 5, the window function GR(θ) is expressed by the following equation (2).
[0041] GR(θ)=p(t)=θ / 2φ=t / T...(2)
[0042] Figure 5B shows window functions such that the sum of the squares of the window function GL(θ) for speaker 2L and the window function GR(θ) for speaker 2R is always constant. In the upper part of Figure 5B, the horizontal axis represents the azimuth θ, and the vertical axis represents the gain. In the lower part of Figure 5B, the horizontal axis represents the azimuth θ, and the vertical axis represents the sound pressure (the square of the gain). If Θ(t) = (π / 2)p(t), then in the example of Figure 5, the window function GR(θ) is expressed by the following equation (3):
[0043] GR(θ)=sin(Θ(t))=sin((π / 2)p(t))=sin(Kp(t))...(3)
[0044] Hereinafter, the linear sum and square sum of the window function GL(θ) for speaker 2L and the window function GR(θ) for speaker 2R are also referred to as a composite value. By applying window functions GL(θ) and GR(θ) to the input signal S such that the composite value of the window functions between the speakers 2 is always constant (either the linear sum or the square sum is constant), the gain control unit 13 is theoretically capable of moving the virtual sound source F along the arc connecting the speakers 2L and 2R.
[0045] 6, at time T1, the audio amplifier 1 multiplies the input signal S by gains GL(T1) and GR(T1), which are function values obtained by substituting t=T1 into a window function in which the linear sum of window functions between the speakers 2 is a constant. At time T1, the virtual sound source F is localized at a position slightly to the left of the front as viewed from the listener.
[0046] 6, at time T2, the audio amplifier 1 multiplies the input signal S by gains GL(T2) and GR(T2), which are function values obtained by substituting t=T2 into a window function in which the linear sum of window functions between the speakers 2 is constant. At time T2, the virtual sound source F is localized at a position directly in front of the listener.
[0047] 6, at time T3, the audio amplifier 1 multiplies the input signal S by gains GL(T3) and GR(T3), which are function values obtained by substituting t=T3 into a window function in which the linear sum of window functions between the speakers 2 is constant. At time T3, the virtual sound source F is localized at a position slightly to the right of the listener.
[0048] In this way, the audio amplifier 1 can change the time substituted into the window function over time, thereby moving the virtual sound source F. By partially applying the window function, the audio amplifier 1 can start moving the virtual sound source F from a position different from the position of the speaker 2L, or finish moving the virtual sound source F at a position different from the position of the speaker 2R.
[0049] Note that when the virtual sound source F is moved in a complex manner, for example, by moving the virtual sound source F closer to or further away from the listener rather than moving the virtual sound source F along an arc, a processed signal obtained by processing the audio signal is input to the audio amplifier 1 as the input signal S, as shown in Fig. 7. The processed signal is generated by convolving, for example, a BRTF (Binaural Room Transfer Function) into the audio signal.
[0050] In this case as well, in the window function applied to the input signal S, the composite value of the window function GL(θ) for the speaker 2L and the window function GR(θ) for the speaker 2R is always constant.
[0051] The processing performed by the audio playback system will be described with reference to the flowchart of FIG.
[0052] In step S1, the input unit 12 of the audio amplifier 1 accepts the selection of a window function configuration by the operator. Here, the operator selects one of two window function configurations: a window function in which the linear sum of the window functions between the speakers 2 is always constant, and a window function in which the square sum of the window functions between the speakers 2 is always constant.
[0053] In step S2, the gain control unit 13 of the audio amplifier 1 calculates the window functions GL(θ) and GR(θ) in accordance with the window function configuration selected by the operator.
[0054] In step S3, the gain control unit 13 stores the window functions GL(θ) and GR(θ) in a memory.
[0055] In step S4, the gain control unit 13 sets the movement of the virtual sound source F (time-series change in the position of the sound image of the sound of the content). The operator may input a desired movement, and the gain control unit 13 may set the movement of the virtual sound source F in accordance with the operator's input, or the gain control unit 13 may automatically set the movement of the virtual sound source F based on the content.
[0056] In step S5, the amplifiers 11L and 11R multiply the input signal S by gains GL and GR, which are function values obtained by substituting the position of the virtual sound source F (sound image) at each time into the window functions GL(θ) and GR(θ), and supply the output signals S'L and S'L to the speakers 2L and 2R.
[0057] FIG. 9 is a diagram showing an example of the movement of a virtual sound source and the function values of the window function corresponding to each time.
[0058] The gain control unit 13 sets the movement of the virtual sound source by determining where to place the virtual sound source for each time. In the example of Fig. 9, the gain control unit 13 sets the movement so that the virtual sound source F is placed at a position (θ1) slightly to the left of the front as seen from the listener at time T1, and at a position (θ2) slightly to the right of the speaker 2L as seen from the listener at time T2.
[0059] At time T1, the amplifiers 11L and 11R multiply the input signal S by gains GL(θ1) and GR(θ1), which are function values obtained by substituting θ=θ1 into a window function with a constant linear sum, for example. At time T2, the amplifiers 11L and 11R multiply the input signal S by gains GL(θ2) and GR(θ2), which are function values obtained by substituting θ=θ2 into a window function with a constant linear sum of window functions between the speakers 2, for example.
[0060] Returning to FIG. 8, in step S6, the speaker 2L outputs a sound indicated by the output signal S'L, and the speaker 2R outputs a sound indicated by the output signal S'R.
[0061] As described above, in the audio amplifier 1, a window function is applied to the input signal S such that the combined value of the window functions between the speakers 2 is always constant, thereby realizing temporal changes in the position of the virtual sound source F.
[0062] FIG. 10 is a diagram illustrating the movement of a virtual sound source F that is actually perceived by a listener when a window function is applied to an input signal S such that the linear sum of the window functions between the speakers 2 is always constant.
[0063] As shown in the upper left of Fig. 10, when the linear sum of the window functions between the speakers 2 is always constant, the sound pressure (sum of squares) of the virtual sound source F decreases as the virtual sound source F approaches the front as seen from the listener, as shown in the upper right of Fig. 10. In this case, even if the operator intends the virtual sound source F to move on the arc connecting the speakers 2L and 2R, as shown in the speech bubble Ba1 of Fig. 10, in reality, the listener perceives the virtual sound source F as moving away from the listener as it approaches the front, as shown in the speech bubble Ba2.
[0064] FIG. 11 is a diagram illustrating the movement of a virtual sound source F that is actually perceived by a listener when a window function is applied to an input signal S such that the sum of squares of the window function between the speakers 2 is always constant.
[0065] As shown in the upper right of Fig. 11, when the sum of squares of the window functions between the speakers 2 is always constant, the gain (linear sum) of the virtual sound source F increases as the virtual sound source F approaches the front as seen from the listener, as shown in the upper left of Fig. 11. In this case, even if the operator intends the virtual sound source F to move on the arc connecting the speakers 2L and 2R, as shown in the speech bubble Ba11 of Fig. 11, in reality, the listener perceives the virtual sound source F to be approaching the listener as it approaches the front, as shown in the speech bubble Ba12.
[0066] In this way, when a window function in which the composite value of the window functions between the speakers 2 is always constant (the linear sum or the square sum is always constant) is applied to the input signal S, it is difficult to match the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator.
[0067] FIG. 12 is a diagram for explaining why the movement of the virtual sound source F perceived by the listener differs from the movement of the virtual sound source F intended by the operator.
[0068] In the calculation of the window function, it is assumed that the head of the listener is a point, as shown in FIG. 12A, and that the sounds output from the speakers 2L and 2R reach the listener's ears directly.
[0069] In reality, as shown in FIG. 12B, the sounds output from the speakers 2L and 2R are diffracted by the listener's head before reaching the listener's ears, and therefore the sound pressures PL(t) and PR(t) that are calculated to reach the listener's ears differ from the sound pressures that actually reach the listener's ears.
[0070] Therefore, the movement of the virtual sound source F perceived by the listener will differ from the movement of the virtual sound source F intended by the operator.
[0071] The gain control unit 13 adds a correction term to the window function, thereby making it possible to match the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator.
[0072] The process performed by the audio reproduction system when the gain control unit 13 adds a correction term to the window function will be described with reference to the flowchart of FIG.
[0073] The processing from step S21 to step S26 is basically the same as the processing from step S1 to S6 in FIG. 8, and therefore a description thereof will be omitted.
[0074] In step S27, the gain control unit 13 determines whether the movement of the virtual sound source F (sound image) perceived by the listener deviates from the movement of the virtual sound source F intended by the operator. Here, for example, the operator acts as a listener and listens to the sound output from the speaker 2, and operates the input unit 12 to input whether the perceived movement of the virtual sound source F deviates from the intended movement of the virtual sound source F.
[0075] If it is determined in step S27 that the movement of the virtual sound source F perceived by the listener deviates from the movement of the virtual sound source F intended by the operator, the gain control unit 13 corrects the window function. Specifically, the gain control unit 13 adds a correction term that mathematically accounts for sound propagation errors due to diffraction at the listener's head, etc., to the window function, assuming that the combined value is constant. This correction makes the combined value of the window functions between the speakers 2 substantially constant at the position of the listener's ears.
[0076] After the window function is corrected, the process returns to step S23, and the subsequent processes are performed.
[0077] On the other hand, if it is determined in step S27 that the movement of the virtual sound source F perceived by the listener does not deviate from the movement of the virtual sound source F intended by the operator, the process ends.
[0078] As described above, the audio amplifier 1 corrects the window function so as to be optimized for each individual listener, thereby making it possible to match the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator.
[0079] 2. Embodiments of the Present Technology In the embodiment of the comparative example, the window function needs to be optimized for each listener in order to match the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator. In addition, the process of calculating the correction term corresponding to each position of the virtual sound source F imposes a large processing load. Considering these optimizations and processing loads, it is not realistic for the audio amplifier 1 to perform the above-described signal processing.
[0080] Therefore, in an embodiment of the present technology, the audio amplifier 1 applies a window function to the input signal S such that the composite value of the window functions between the speakers 2 does not become constant, thereby matching the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator.
[0081] FIG. 14 is a diagram illustrating the movement of a virtual sound source F that is actually perceived by a listener when a window function is applied to the input signal S such that the combined value of the window functions between the speakers 2 is not constant.
[0082] The audio amplifier 1 applies to the input signal S window functions GL(t), GL(t) such that, for example, as the virtual sound source F approaches the front as seen from the listener, the linear sum of the window functions between the speakers 2 becomes larger, as shown by the solid line graph in the upper left of FIG. 14 , and as shown by the solid line graph in the upper right of FIG. 14 , the sum of squares of the window functions between the speakers 2 becomes smaller, as the virtual sound source F approaches the front as seen from the listener.
[0083] When the window functions GL(t) and GL(t) shown in the upper left of Figure 14 are applied to the input signal S, the listener perceives the virtual sound source F as moving on the arc connecting the loudspeakers 2L and 2R, as shown in speech bubble Ba21.
[0084] The window function GR(θ) in which the composite value of the window functions between the speakers 2 is not constant is expressed by, for example, the following equation (4) or (5).
[0085] GR(θ) = (p(t)) α ...(4) GR(θ)=(sin(Kp(t))) α ...(5)
[0086] In equation (4), the window function GR(θ) is formed by a power function, and in equation (5), the window function GR(θ) is formed by a power of a trigonometric function.
[0087] The gain control unit 13 uses the exponent α of the power function that constitutes the window function and the exponent α of the power of the trigonometric function that constitutes the window function as parameters, and by simply controlling the parameter α, can arbitrarily change the curve of the graph of the combined value of the window functions between the speakers 2. The audio amplifier 1 applies to the input signal S window functions GL(t) and GR(t) that are generated by controlling the parameter α and that do not cause the combined value of the window functions between the speakers 2 to be constant, thereby making it possible to match the movement of the virtual sound source F perceived by the listener with the movement of the virtual sound source F intended by the operator.
[0088] The process performed by the audio reproduction system according to the embodiment of the present technology will be described with reference to the flowchart of FIG.
[0089] In step S41, the gain control unit 13 of the audio amplifier 1 sets the movement (time-series change in the position of the sound image) of the virtual sound source F. The operator may input a desired movement, and the gain control unit 13 may set the movement of the virtual sound source F in accordance with the operator's input, or the gain control unit 13 may automatically set the movement of the virtual sound source F based on the content.
[0090] In step S42, the gain control unit 13 selects a window function configuration and parameters suitable for the movement of the virtual sound source F. For example, the gain control unit 13 selects one of two window function configurations, that is, a window function configured by a power function and a window function configured by a power of a trigonometric function, and sets parameter values suitable for the movement of the virtual sound source F.
[0091] In addition, the operator may input the desired window function configuration and parameters, and the gain control unit 13 may select the window function configuration and parameters in accordance with the operator's input, or the gain control unit 13 may automatically select them based on the movement of the virtual sound source F.
[0092] In step S43, the gain control unit 13 controls the parameter α to calculate the window functions GL(θ) and GR(θ).
[0093] In step S44, the gain control unit 13 stores the window functions GL(θ) and GR(θ) in the memory.
[0094] In step S45, the amplifiers 11L and 11R multiply the input signal S by gains GL and GR, which are function values obtained by substituting the position of the virtual sound source F (sound image) at each time into the window functions GL(θ) and GR(θ), and supply the output signals S'L and S'L to the speakers 2L and 2R.
[0095] In step S46, the speaker 2L outputs a sound indicated by the output signal S'L, and the speaker 2R outputs a sound indicated by the output signal S'R.
[0096] In step S47, the gain control unit 13 determines whether the movement of the virtual sound source F (sound image) perceived by the listener deviates from the movement of the virtual sound source F intended by the operator. Here, for example, the operator acts as a listener and listens to the sound output from the speaker 2, and operates the input unit 12 to input whether the perceived movement of the virtual sound source F deviates from the intended movement of the virtual sound source F.
[0097] If it is determined in step S47 that the movement of the virtual sound source F perceived by the listener deviates from the movement of the virtual sound source F intended by the operator, the process returns to step S42, and the subsequent processes are performed.
[0098] On the other hand, if it is determined in step S47 that the movement of the virtual sound source F perceived by the listener does not deviate from the movement of the virtual sound source F intended by the operator, the process ends.
[0099] As described above, in the audio reproduction system of the present technology, the amplifiers 11L and 11R (signal output units) apply a window function, which is a function indicating the relationship between time and gain, to an input signal S (audio signal) indicating the sound of content for each speaker, and output the input signal S after the application of the window function to multiple speakers. In addition, the gain control unit 13 controls the window function applied to the input signal S so that the combined value of the window functions for each speaker is not constant.
[0100] This allows the audio reproduction system to apply various window functions to the input signal S, thereby making it possible to express various movements of the virtual sound source as intended by the operator.
[0101] Next, with reference to FIGS. 16 to 19, examples of window functions in which the combined value of the window functions between the speakers 2 is not constant will be described.
[0102] The upper part of FIG. 16 shows the window function GR(t)=(sin(Kp(t))) 0.5 16 shows the linear sum and the square sum of the window function GL(t), which is symmetric with respect to the window function GR(t) in time. When the window functions GL(t) and GR(t) shown in the upper part of Fig. 16 are applied to the input signal S, the listener perceives the virtual sound source F to be approaching the listener as it approaches the front. This perception is thought to be due to the fact that the window functions are designed so that the linear sum and the square sum become larger as the virtual sound source F approaches the front as seen from the listener.
[0103] At the bottom of FIG. 16, the window function GR(t)=(p(t)) 2 16 shows the linear sum and the square sum of the window function GL(t), which is symmetric with respect to the window function GR(t) in time. When the window functions GL(t) and GR(t) shown in the lower part of Fig. 16 are applied to the input signal S, the listener perceives that the virtual sound source F moves away from the listener quickly as it approaches the front. This perception is thought to be due to the window functions being designed so that the linear sum and the square sum become sharply smaller as the virtual sound source F approaches the front as seen from the listener.
[0104] The upper part of FIG. 17 shows GR(t)=(sin(Kp(t))) 2 +(sin(Kp(t))) 0.5 17 shows the linear sum and the square sum of the window function GL(t) which is symmetric with respect to the window function GR(t) in time. When the window function shown in the upper part of Fig. 17 is applied to the input signal S, the listener perceives, for example, that the virtual sound source F is moving on an arc connecting the loudspeakers 2L and 2R. This perception is thought to be due to the fact that the window function is designed so that the linear sum becomes larger and the square sum becomes smaller as the virtual sound source F approaches the front as seen from the listener.
[0105] At the bottom of FIG. 17, GR(t)=(p(t)) 0.5 ×(sin(Kp(t))) 2 17 shows the linear sum and the square sum of the window function GL(t), which is symmetrical with respect to the window function GR(t) in time. When the window functions GL(t) and GR(t) shown in the lower part of Fig. 17 are applied to the input signal S, the listener perceives, for example, a virtual sound source F as meandering away from the listener as it approaches the front. This perception is thought to be due to the window functions being designed so that the graph of the linear sum and the square sum is a concave arc (a curve with an inflection point).
[0106] As described with reference to FIG. 17, the window function may be configured by performing arithmetic operations on a plurality of functions whose exponents are mutually independent parameters α.
[0107] In FIG. 18, GR(t)=(1 / 3)·(sin(Kp(t))) 2 +2 / 3・(sin(Kp(t))) 0.5 , and the linear sum and square sum of the window function GL(t) that is symmetric with respect to the window function GR(t) with respect to time are shown. In this way, the window function may be configured by weighted addition (or weighted subtraction) at different ratios of multiple functions whose exponents are mutually independent parameter α.
[0108] (sin(Kp(t))) 2 and (sin(Kp(t))) 0.5 Varying the ratio of adding the window function GR(t) changes the sense of speed and size of the virtual sound source F. For example, the window function GR(t) = (1 / 3) · (sin(Kp(t))) 2 +2 / 3・(sin(Kp(t))) 0.5 is applied to the input signal S, the sound image of the sound heard from the position of the virtual sound source F becomes wide and blurred. Because the sound image is wide, the listener perceives, for example, that the virtual sound source F is moving slowly. Also, for example, when the window function GR(t)=(2 / 3)·(sin(Kp(t))) 2 +1 / 3・(sin(Kp(t))) 0.5is applied to the input signal S, the sound image heard from the position of the virtual sound source F becomes narrower and clearer. Because the sound image is narrower, the listener perceives, for example, that the virtual sound source F is moving faster.
[0109] The upper part of FIG. 19 shows the window function GL(t)=(sin(Kp(t))) 4 and the window function GR(t) = (sin(Kp(t))) 0.5 When this window function is applied to the input signal S, the virtual sound source F is localized in front of the listener at a time before time T / 2. In other words, the listener perceives the sound image as moving quickly on the left side and slowly on the right side.
[0110] At the bottom of FIG. 19, the window function GL(t)=(sin(Kp(t))) 0.5 and the window function GR(t) = (sin(Kp(t))) 2 +(sin(Kp(t))) 0.5 When this window function is applied to the input signal S, the virtual sound source F is localized in front of the listener at a time after time T / 2. In other words, the listener perceives the sound image as moving slowly on the left side and quickly on the right side.
[0111] 16 to 18, window functions GL(t) and GR(t) that are symmetrical with respect to the time at which the window function GL(t) for the speaker 2L and the window function GR(t) for the speaker 2R intersect are applied to the input signal S. As described with reference to Fig. 19, window functions GL(t) and GR(t) that are asymmetric with respect to the time at which the window function GL(t) for the speaker 2L and the window function GR(t) for the speaker 2R intersect may also be applied to the input signal S.
[0112] The configuration of the window function and the parameter values described above are selected based on, for example, the content, which includes the movement (horizontal, depth, and height) of the speaker or vocalist reproduced as the virtual sound source F, the speed of the speaker or vocalist, the tempo of the music, the melody of the music, etc.
[0113] For example, the window function GR(t)=(sin(Kp(t))) 0.5 , and a window function GL(t) that is symmetrical with respect to time to the window function GR(t), is applied to the input signal S, the volume of the virtual sound source F increases at time T / 2. As shown in Fig. 20, the audio amplifier 1 applies this window function to the input signal S so that the volume of the virtual sound source F increases at the timing of the beat of the music, thereby making it possible to change the volume of the virtual sound source F in time with the beat.
[0114] When content is divided into a plurality of scenes, the window function applied to the input signal S may be switched for each scene, as shown in FIG.
[0115] For example, in a scene where vocals move from left to right as seen by the listener, a window function is applied to the input signal S such that the virtual sound source F moves at a constant speed on the arc connecting the loudspeakers 2L and 2R, as shown on the left side of Fig. 21. For example, the window function GR(t) for the loudspeaker 2R is GR(t) = (sin(Kp(t))) 2 +(sin(Kp(t))) 0.5 The window function GL(t) for the speaker 2L is a function symmetric with respect to time to the window function GR(t).
[0116] When the scene changes and the vocals move away from the listener, a window function is applied to the input signal S so that the virtual sound source F moves away from the listener when it approaches the front as seen from the listener, as shown in the center of Fig. 21. For example, the window function GR(t) for speaker 2R is GR(t) = (sin(Kp(t))) 3 The window function GL(t) for the speaker 2L is a function symmetric with respect to time to the window function GR(t).
[0117] When the scene changes again and the speaker moves closer to the listener, a window function is applied to the input signal S such that the virtual sound source F approaches the listener as it approaches the front as seen from the listener, as shown on the right side of Fig. 21. For example, the window function GR(t) for speaker 2R is GR(t) = (sin(Kp(t))) 3 +(sin(Kp(t))) 0.5The window function GL(t) for the speaker 2L is a function symmetric with respect to time to the window function GR(t).
[0118] It should be noted that, under the assumption that the composite value of the window function is not constant, the window function applied to the input signal S is not switched for each scene, but rather, in one scene (first scene) among multiple scenes, a window function in which the composite value of the window function between speakers 2 is always constant may be applied to the input signal S, and in another scene (second scene), a window function in which the composite value of the window function between speakers 2 is not constant may be applied to the input signal S.
[0119] Furthermore, the configuration of the window function and the parameter values may be selected based on the position of the listener (the distance between the listener and the speaker 2). When the listener moves, the window function applied to the input signal S is switched according to the change in the listener's position over time.
[0120] FIG. 22 is a diagram showing an example of switching of window functions depending on the position of the listener.
[0121] 22, when the distance between the listener and the speaker 2 is equal to or greater than a predetermined threshold, the sound output from the speaker 2 is hardly affected by diffraction by the listener's head. Therefore, when the distance between the listener and the speaker 2 is equal to or greater than a predetermined threshold, the audio amplifier 1 applies to the input signal S a window function such that the combined value of the window functions between the speakers 2 is always constant.
[0122] When the listener moves and the distance between the listener and the speaker 2 becomes less than a predetermined threshold, as shown on the right side of Figure 22, the audio amplifier 1 switches the window function applied to the input signal S from a window function in which the composite value of the window functions between the speakers 2 is always constant to a window function in which the composite value is not constant.
[0123] With reference to the flowchart in FIG. 23, a process performed by the audio reproduction system when the audio amplifier 1 automatically switches the window function for each scene will be described.
[0124] In step S61, the gain control unit 13 of the audio amplifier 1 sets the movement of the virtual sound source F (time-series change in the position of the sound image) for each scene.
[0125] In step S62, the gain control unit 13 selects, for each scene, the configuration and parameters of a window function suitable for the movement of the virtual sound source F (sound image).
[0126] In step S63, the gain control unit 13 controls the parameter α to calculate the window functions GL(θ) and GR(θ), and generates a sequence in which the window functions GL(θ) and GR(θ) themselves are switched every time the scene changes.
[0127] In step S64, the gain control unit 13 stores the sequence of the window function in the memory.
[0128] In step S65, the amplifiers 11L and 11R multiply the input signal S by gains GL and GR, which are function values obtained by substituting the position of the virtual sound source F (sound image) at each time into the window functions GL(θ) and GR(θ), and supply the results as output signals S'L and S'L to the speakers 2L and 2R.
[0129] In step S66, the speaker 2L outputs a sound indicated by the output signal S'L, and the speaker 2R outputs a sound indicated by the output signal S'R.
[0130] Next, with reference to the flowchart of FIG. 24, a process performed by the audio reproduction system when the operator manually switches the window function for each scene will be described.
[0131] In step S81, the input unit 12 of the audio amplifier 1 accepts adjustment of the parameter α of the window function by the operator. Here, along with the parameter α, the operator inputs the ratio when performing weighted addition (weighted subtraction) of multiple functions, the initial position of the virtual sound source F, the direction of movement of the virtual sound source F, and the like. The configuration of the window function may be fixed or may be selected by the operator.
[0132] In step S82, the gain control unit 13 controls the parameter α to calculate the window functions GL(t) and GR(t).
[0133] In step S83, the gain control unit 13 stores the window functions GL(t) and GR(t) in the memory.
[0134] In step S84, the amplifier units 11L and 11R of the audio amplifier 1 multiply the input signal S by the gains GL and GR, which are function values obtained by substituting the time into the window functions GL(t) and GR(t), and supply the results as output signals S'L and S'L to the speakers 2L and 2R.
[0135] In step S85, the speaker 2L outputs a sound indicated by the output signal S'L, and the speaker 2R outputs a sound indicated by the output signal S'R.
[0136] In step S86, the gain control unit 13 determines whether or not a scene change occurs in the audio amplifier 1. For example, if the parameter α is changed by the operator, it is determined that a scene change occurs.
[0137] If it is determined in step S86 that the scene has changed, the process returns to step S81, and the subsequent processes are carried out.
[0138] On the other hand, if it is determined in step S86 that the scene has not changed, the output of the sound of the content continues, and when the reproduction of the content has finished, the process ends.
[0139] This technology can be used to provide audio for live music performances and conversations in the Metaverse space, and for providing audio in object simulators. The audio amplifier 1 of this technology can switch window functions simply by controlling the parameter α, making it possible to express various movements of the virtual sound source F while reducing the processing load.
[0140] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0141] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0142] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.
[0143] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0144] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0145] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0146] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0147] (1) A signal processing device comprising: a signal output unit that applies a window function, which is a function indicating the relationship between time and gain, to an audio signal indicating sound of content for each speaker, and outputs the audio signal after the application of the window function to the multiple speakers; and a gain control unit that controls the window function applied to the audio signal by the signal output unit so that a composite value of the window functions for each speaker is not constant. (2) The signal processing device described in (1), in which the gain control unit controls the window function applied to the audio signal by the signal output unit so that a linear sum of the window functions for each speaker is not constant and a sum of squares of the window functions for each speaker is not constant. (3) The signal processing device described in (1) or (2), in which the window function whose composite value is not constant is a function constituted by at least one of a trigonometric function and a power function. (4) The signal processing device described in (3), in which the gain control unit uses an exponent of a power of the trigonometric function and an exponent of the power function as parameters and generates the window function by controlling the parameters. (5) The signal processing device according to (4), wherein the window function whose composite value does not become constant is configured by arithmetic operations on a plurality of functions whose exponents are the parameters that are independent of each other. (6) The signal processing device according to (5), wherein the window function whose composite value does not become constant is configured by weighted addition or weighted subtraction at different ratios on a plurality of functions whose exponents are the parameters that are independent of each other. (7) The signal processing device according to any of (1) to (6), wherein the gain control unit generates the window functions for the speakers so that they are symmetrical with respect to a time when the window functions for the speakers intersect. (8) The signal processing device according to any of (1) to (6), wherein the gain control unit generates the window functions for the speakers so that they are asymmetrical with respect to a time when the window functions for the speakers intersect. (9) The signal processing device according to any of (1) to (8), wherein the gain control unit generates the window functions based on the nature of the content.(10) The signal processing device according to (9), wherein the gain control unit switches the window function applied to the audio signal for each scene into which the content is divided. (11) The signal processing device according to (10), wherein the gain control unit controls the window function applied to the audio signal in a first scene so that the combined value of the window functions for the speakers is always constant, and controls the window function applied to the audio signal in a second scene different from the first scene so that the combined value of the window functions for the speakers is not constant. (12) The signal processing device according to any of (1) to (11), wherein the gain control unit switches the window function applied to the audio signal in accordance with a change in distance between a listener of sound output from the speaker and the speaker. (13) The signal processing device according to (12), wherein the gain control unit controls the window function applied to the audio signal so that the combined value of the window functions for the speakers is always constant when a distance between the listener and the speakers is equal to or greater than a threshold, and controls the window function applied to the audio signal so that the combined value of the window functions for the speakers is not constant when the distance between the listener and the speakers is less than a threshold. (14) A signal processing method comprising: applying a window function, which is a function indicating a relationship between time and gain, to an audio signal indicating sound of content, for each speaker; outputting the audio signal after application of the window function to a plurality of speakers; and controlling the window function applied to the audio signal so that the combined value of the window functions for the speakers is not constant.
[0148] 1 Audio amplifier, 2L, 2R Speaker, 11L, 11R Amplification section, 12 Input section, 13 Gain control section
Claims
1. A signal processing device comprising: a signal output unit that applies a window function, which is a function that indicates the relationship between time and gain, to an audio signal that represents the sound of content for each speaker, and outputs the audio signal after the window function has been applied to the multiple speakers; and a gain control unit that controls the window function applied to the audio signal by the signal output unit so that the combined value of the window function for each speaker is not constant.
2. The signal processing device according to claim 1, wherein the gain control unit controls the window function applied to the audio signal by the signal output unit so that the linear sum of the window function for each speaker is not constant and the square sum of the window function for each speaker is not constant.
3. The signal processing device according to claim 1, wherein the window function whose composite value does not become constant is a function formed of at least one of a trigonometric function and a power function.
4. The signal processing device according to claim 3, wherein the gain control section uses the exponent of the power of the trigonometric function and the exponent of the power function as parameters, and generates the window function by controlling the parameters.
5. The signal processing device according to claim 4, wherein the window function, the composite value of which is not constant, is formed by arithmetic operations on a plurality of functions whose exponents are the parameters that are independent of each other.
6. The signal processing device according to claim 5, wherein the window function, the composite value of which is not constant, is formed by weighted addition or weighted subtraction at different ratios of a plurality of functions whose exponents are the parameters that are independent of each other.
7. The signal processing device according to claim 1, wherein the gain control section generates the window functions for the respective speakers so that the window functions are symmetrical with respect to a time at which the window functions for the respective speakers intersect.
8. The signal processing device according to claim 1, wherein the gain control section generates the window function for each speaker so that the window function is asymmetric with respect to a time at which the window functions for each speaker intersect.
9. The signal processing device according to claim 1, wherein the gain control unit generates the window function based on the content.
10. The signal processing device according to claim 9, wherein the gain control unit switches the window function applied to the audio signal for each scene into which the content is divided.
11. The signal processing device described in claim 10, wherein the gain control unit controls the window function applied to the audio signal in a first scene so that the combined value of the window functions for each speaker is always constant, and controls the window function applied to the audio signal in a second scene different from the first scene so that the combined value of the window functions for each speaker is not constant.
12. The signal processing device according to claim 1, wherein the gain control unit switches the window function applied to the audio signal in accordance with a change in the distance between the speaker and a listener of the sound output from the speaker.
13. The signal processing device described in claim 12, wherein the gain control unit controls the window function applied to the audio signal so that the combined value of the window functions for each speaker is always constant when the distance between the listener and the speaker is equal to or greater than a threshold, and controls the window function applied to the audio signal so that the combined value of the window functions for each speaker is not constant when the distance between the listener and the speaker is less than a threshold.
14. A signal processing method comprising: applying a window function, which is a function indicating the relationship between time and gain, to an audio signal representing the sound of content for each speaker; outputting the audio signal after applying the window function to multiple speakers; and controlling the window function applied to the audio signal so that the combined value of the window function for each speaker is not constant.
Citation Information
Patent Citations
Sound image control device and sound image control method
JP2007336184A
Apparatus and method for outputting audio signal, and display apparatus using the same
US20190166419A1