Sound signal processing apparatus and sound signal processing method
By setting up a removal section, a surround processing section, an amplification section, and a synthesis section, and using the first and second coefficients to control the vocal frequency band and amplification rate, the problems of inappropriate surround effect and unclear sound in the prior art are solved, and appropriate surround effect and vocal clarity are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOCIONEXT INC
- Filing Date
- 2021-07-08
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to simultaneously ensure both clarity and stereo sound when adding surround effects, resulting in inappropriate surround effects or unclear sound.
By setting up a removal section, a surround processing section, an amplification section, a synthesis section, and a setting section, the vocal frequency band and amplification rate are controlled by the first and second coefficients respectively, thereby generating an appropriate surround effect and suppressing unclear sound.
It achieves the effect of appropriately adding surround sound while maintaining vocal clarity, thus enhancing the user experience of surround sound.
Smart Images

Figure CN114093378B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an apparatus for processing sound signals and a method for processing sound signals. Background Technology
[0002] Previously, techniques for adding surround effects to sound signals to create a sense of stereo or depth when reproducing sound signals were known. Furthermore, the sound signal used for surround signal processing to add surround effects must exclude vocal components (speech components) such as dialogue and lyrics. Patent Document 1 discloses a sound signal processing apparatus for performing surround signal processing on a sound signal after removing vocal components using a band-stop filter.
[0003] (Existing technical literature)
[0004] (Patent Documents)
[0005] Patent Document 1: Japanese Patent Application Publication No. 9-84198
[0006] However, according to the technology described in Patent Document 1, there may be situations where the surround effect cannot be properly added. Summary of the Invention
[0007] Therefore, sound signal processing devices that can appropriately add surround effects are provided.
[0008] One aspect of this disclosure relates to a sound signal processing apparatus, comprising: a removal unit that generates a first output signal by removing vocal components based on a first channel sound signal and a second channel sound signal, and a first coefficient indicating a vocal frequency band to be removed; a surround processing unit that adds a surround effect to the first output signal to generate a second output signal; an amplification unit connected to a preamplifier of the removal unit or between the removal unit and the surround processing unit, or configured as part of the removal unit or the surround processing unit, for amplifying an input signal based on an amplification rate of a second coefficient; a first synthesis unit that synthesizes the second output signal with one of the first channel sound signal and the second channel sound signal; a second synthesis unit that synthesizes a signal inverted from the second output signal with the other of the first channel sound signal and the second channel sound signal; and a setting unit that sets the first coefficient and the second coefficient, wherein the setting unit sets the second coefficient in such a way that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is a first frequency band.
[0009] One aspect of this disclosure relates to a sound signal processing method, comprising: a removal step, generating a first output signal by removing vocal components based on a first channel sound signal and a second channel sound signal, and a first coefficient indicating a vocal frequency band to be removed; a surround signal processing step, adding a surround effect to the first output signal to generate a second output signal; an amplification step, performed before the removal step or between the removal step and the surround signal processing step, or performed as part of the removal step or the surround signal processing step, to amplify the input signal based on an amplification rate of a second coefficient; a first synthesis step, synthesizing the second output signal with one of the first channel sound signal and the second channel sound signal; a second synthesis step, synthesizing the inverted signal of the second output signal with the other of the first channel sound signal and the second channel sound signal; and a setting step, setting the first coefficient and the second coefficient, wherein the second coefficient is set such that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is a first frequency band.
[0010] According to one embodiment of the present disclosure, a sound signal processing device, etc., can appropriately add a surround effect. Attached Figure Description
[0011] Figure 1 This is a block diagram illustrating the functional structure of the sound signal processing apparatus according to Embodiment 1.
[0012] Figure 2 This is a diagram illustrating an example of the hardware structure of a computer in which the function of the sound signal processing device according to Embodiment 1 is implemented by software.
[0013] Figure 3 This is a diagram illustrating a first example of the relationship between vocal clarity, cutoff frequency, and gain value in Implementation 1.
[0014] Figure 4 This is a second example of the relationship between vocal clarity, cutoff frequency, and gain value involved in Embodiment 1.
[0015] Figure 5 This is a diagram showing the results of a sensory experiment on surround sensation related to Embodiment 1.
[0016] Figure 6 This is a graph showing the results of a sensory experiment on vocal clarity related to Embodiment 1.
[0017] Figure 7 This is a diagram illustrating a third example of the relationship between vocal clarity, cutoff frequency, and gain value in Implementation 1.
[0018] Figure 8 This is a flowchart illustrating the operation of the sound signal processing apparatus according to Embodiment 1.
[0019] Figure 9 This is a block diagram illustrating the functional structure of the sound signal processing apparatus according to Embodiment 2.
[0020] Figure 10 This is a graph illustrating a first example of the relationship between vocal clarity and surround sound, cutoff frequency, and gain value in Embodiment 2.
[0021] Figure 11 This is a graph illustrating a second example of the relationship between vocal clarity and surround sound, cutoff frequency, and gain value in Embodiment 2.
[0022] Symbol Explanation
[0023] 1,100 Sound Signal Processing Device
[0024] 10. Vocal Removal (Removal Section)
[0025] 11. Difference signal generation unit (first signal generation unit)
[0026] 12 Filter Section
[0027] 20 Surrounding Processing Section
[0028] 21. Surround signal generation unit (second signal generation unit)
[0029] 22 Enlarged section
[0030] 30 User Interface
[0031] 40, 140 coefficient determination part
[0032] 50 Synthetic
[0033] 51 First Synthesis Section
[0034] 52 Second Synthesis Section
[0035] 60 Reversal section
[0036] 1000 computers
[0037] 1001 Input Device
[0038] 1002 Output Device
[0039] 1003 CPU
[0040] 1004 Internal Memory
[0041] 1005 RAM
[0042] 1009 bus Detailed Implementation
[0043] (The process by which this disclosure was made)
[0044] Before describing the implementation of this disclosure, the process of achieving the basis of this disclosure will be explained.
[0045] According to the technology in Patent Document 1, a band-stop filter is used to remove vocal components from an additive signal that combines audio signals from the L channel and the R channel. When the band-stop filter is configured to include a low-pass filter (LPF) and a high-pass filter (HPF), the cutoff frequencies of the LPF and HPF are set to frequencies capable of removing vocal components, thereby removing vocal components from the additive signal. Furthermore, the L channel audio signal is the audio signal input to the L-side speaker, and the R channel audio signal is the audio signal input to the R-side speaker. The L-side speaker and the R-side speaker are speakers located at different positions in the same space; for example, the L-side speaker is located to the left of a reference position, and the R-side speaker is located to the right of a reference position.
[0046] Furthermore, if surround signal processing is performed to add surround effects to an additive signal that includes vocal components, a sense of stereo is also added to the vocal components. Therefore, unclear (e.g., fuzzy) speech output may result in a reduced sense of presence or a feeling of dissonance for the user. Therefore, vocal components are removed before performing surround signal processing, as described above.
[0047] Here, the additive signal obtained by the LPF and HPF is a sound signal that, in addition to the vocal component, has had components other than the vocal component in the same frequency band removed. If the cutoff frequency of the LPF is set lower and the cutoff frequency of the HPF is set higher to more accurately remove the vocal component, the amount of components other than the vocal component removed increases. Therefore, the intensity (absolute value) of the additive signal subjected to surround signal processing becomes very small compared to the additive signal before passing through the LPF and HPF. Even after performing surround signal processing on such an additive signal and combining it with the sound signals of the L channel and R channel, the intensity of the additive signal subjected to surround signal processing is smaller than the sound signals of the L channel and R channel. Therefore, the added surround effect is also small. In other words, according to the technology of Patent Document 1, it is difficult to properly add a surround effect.
[0048] Furthermore, components other than vocal elements include, for example, sound effects, instrumental sounds, and background sounds (so-called BGM (background music), which are sound components that do not include speech.
[0049] Furthermore, if the cutoff frequency of the LPF is set higher and the cutoff frequency of the HPF is set lower in order to suppress the decrease in the intensity of the additive signal, it is difficult to remove the vocal components, and therefore the heard speech is unclear. Thus, according to the technology of Patent Document 1, it is difficult to properly add a surround effect, and it is also difficult to suppress the unclear sound.
[0050] Therefore, the inventors of this application, after focusing on studying a sound signal processing device that can appropriately add surround effects to the sound signals of the L channel and the R channel, and further capable of adding surround effects while suppressing unclear sound, invented the sound signal processing device described below.
[0051] One aspect of this disclosure relates to a sound signal processing apparatus, comprising: a removal unit that generates a first output signal by removing vocal components based on a first channel sound signal and a second channel sound signal, and a first coefficient indicating a vocal frequency band to be removed; a surround processing unit that adds a surround effect to the first output signal to generate a second output signal; an amplification unit connected to a preamplifier of the removal unit or between the removal unit and the surround processing unit, or configured as part of the removal unit or the surround processing unit, for amplifying an input signal based on an amplification rate of a second coefficient; a first synthesis unit that synthesizes the second output signal with one of the first channel sound signal and the second channel sound signal; a second synthesis unit that synthesizes a signal inverted from the second output signal with the other of the first channel sound signal and the second channel sound signal; and a setting unit that sets the first coefficient and the second coefficient, wherein the setting unit sets the second coefficient in such a way that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is a first frequency band.
[0052] Accordingly, when the frequency band to be removed becomes wider and the intensity of the first output signal decreases, the amplification factor of the amplification unit in the audio signal processing device increases, thus suppressing the decrease in the intensity of the second output signal. In other words, the audio signal processing device can suppress the decrease in the intensity of the second output signal relative to both the audio signals of the first and second channels, thereby suppressing the weakening of the surround effect in the synthesized signal. Therefore, compared to the case where the amplification factor of the amplification unit remains unchanged even when the frequency band to be removed becomes wider, the audio signal processing device can appropriately add a surround effect.
[0053] Furthermore, for example, the setting unit may set the first coefficient and the second coefficient according to vocal clarity, which indicates the clarity of speech based on the signal synthesized by the first synthesis unit and the second synthesis unit.
[0054] Accordingly, the sound signal processing device is able to generate a speech signal that can output the desired vocal clarity.
[0055] Furthermore, for example, the removal unit may have a high-pass filter, and the setting unit may set the first coefficient such that the cutoff frequency of the high-pass filter is higher as the clarity increases, and set the second coefficient such that the amplification rate is higher as the clarity increases. Alternatively, the removal unit may have a high-pass filter, and the vocal clarity may be represented by a monotonically increasing graph when the cutoff frequency of the high-pass filter is set on the horizontal axis and the amplification rate of the amplification unit is set on the vertical axis. The setting unit sets the first coefficient and the second coefficient based on the vocal clarity and the monotonically increasing graph.
[0056] Accordingly, the sound signal processing device sets the second coefficient in a way that reduces the change in the surround effect caused by the change in the first coefficient. Therefore, it is able to suppress the change in the surround effect and generate a speech signal that can output the vocal clarity.
[0057] Furthermore, for example, the monotonically increasing graph could also be a logarithmic graph.
[0058] Therefore, it is possible to make the variation in the clarity of the output speech equal to the variation in the clarity of the vocal music.
[0059] Furthermore, for example, the monotonically increasing graph could be a graph of straight lines.
[0060] Accordingly, the audio signal processing apparatus can enhance the surround effect when the cutoff frequency of the filter section (e.g., a filter section including a high-pass filter) is set to the high-frequency domain (e.g., 2000Hz or higher), and the amount of signal components removed in the high-frequency domain is less than the amount of signal components removed in the low-frequency domain. Furthermore, the first and second coefficients can be set with simpler calculations, thus reducing the processing load of the audio signal processing apparatus.
[0061] Furthermore, for example, the sound signal processing device may also include a user interface for receiving the vocal clarity from the user.
[0062] Accordingly, the sound signal processing device is further capable of generating a voice signal that can output a voice with a clarity specified by the user.
[0063] Furthermore, for example, the setting unit may further set the second coefficient according to the surround sensation indicated by the user's additional preference for the surround effect.
[0064] Accordingly, the sound signal processing device changes the amplification rate of the amplification unit according to the surround sound effect, thus further generating a signal that can output sound corresponding to the surround sound effect. In other words, the sound signal processing device can further generate a signal that can output the user's preferences.
[0065] Furthermore, for example, the sound signal processing device may also include a user interface for receiving the vocal clarity and the surround sound from the user.
[0066] Accordingly, the coefficient determination unit can determine the second coefficient using the acoustic clarity and surround sound obtained from the user interface. In other words, the sound signal processing device can obtain the acoustic clarity and surround sound used for determining the second coefficient without communicating with external devices, thus reducing the amount of communication required.
[0067] Furthermore, for example, the removal unit may include: a first signal generation unit that generates a difference signal showing the difference between the sound signal of the first channel and the sound signal of the second channel; and a filter unit that removes frequency components of the vocal band based on the first coefficient from the difference signal to generate the first output signal; the surround processing unit includes: a second signal generation unit that adds the surround effect to the first output signal to generate a surround signal; and an amplification unit that amplifies the surround signal with an amplification rate based on the second coefficient to generate the second output signal.
[0068] Therefore, in a sound signal processing device that includes a first signal generation unit, a filter unit, a second signal generation unit, and an amplification unit, a surround effect can be appropriately added.
[0069] One aspect of this disclosure relates to a sound signal processing method, comprising: a removal step, generating a first output signal by removing vocal components based on a first channel sound signal and a second channel sound signal, and a first coefficient indicating a vocal frequency band to be removed; a surround signal processing step, adding a surround effect to the first output signal to generate a second output signal; an amplification step, performed before the removal step or between the removal step and the surround signal processing step, or performed as part of the removal step or the surround signal processing step, to amplify the input signal based on an amplification rate of a second coefficient; a first synthesis step, synthesizing the second output signal with one of the first channel sound signal and the second channel sound signal; a second synthesis step, synthesizing the inverted signal of the second output signal with the other of the first channel sound signal and the second channel sound signal; and a setting step, setting the first coefficient and the second coefficient, wherein the second coefficient is set such that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is a first frequency band.
[0070] Therefore, the same effect as the sound signal processing device is obtained.
[0071] The following description of the implementation method is based on the accompanying drawings.
[0072] Furthermore, the embodiments described below are all general or specific examples. The numerical values, constituent elements, the arrangement and connection of constituent elements, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the scope of the embodiments. Moreover, constituent elements that are not described in the embodiments showing the highest-level concept in the following embodiments are described as arbitrary constituent elements.
[0073] Furthermore, each figure is a schematic diagram and not necessarily a rigorous representation. In each figure, identical structures are marked with the same symbols, and repetitive explanations are omitted or simplified.
[0074] Furthermore, in this specification, terms, values, and ranges indicating the relationship between elements, such as equal, certain, and identical, do not merely represent a strict meaning, but rather imply substantially equivalent ranges, including, for example, differences of approximately a few percent.
[0075] (Implementation Method 1)
[0076] [1-1. Structure of a sound signal processing device]
[0077] First, regarding the structure of the sound signal processing device involved in this embodiment, refer to... Figure 1 as well as Figure 2 Please provide an explanation. Figure 1 This is a block diagram illustrating the functional structure of the sound signal processing apparatus 1 according to this embodiment. The sound signal processing apparatus 1 is a device that generates and outputs sound with a surround sound effect based on the input signal (sound signal) from the L channel and the input signal (sound signal) from the R channel. Furthermore, the audio device equipped with the sound signal processing apparatus 1 includes, for example, two speakers: an L-side speaker and an R-side speaker. Moreover, a surround sound effect is a sound in which the user (listener) listening to the sound can perceive a sense of stereoscopic depth, expansion, or other similar qualities.
[0078] like Figure 1 As shown, the sound signal processing device 1 includes a vocal removal unit 10, a surround processing unit 20, a user interface 30 (UI), a coefficient determination unit 40, a synthesis unit 50, and a reversal unit 60.
[0079] The vocal removal unit 10 removes vocal components from the input signals of the L channel and the R channel. Specifically, the vocal removal unit 10 generates a vocal removal signal by removing vocal components based on the input signals of the L channel and the R channel, and filter coefficients indicating the vocal frequency bands to be removed. More specifically, the vocal removal unit 10 generates a vocal removal signal by removing vocal components from the difference signal between the input signals of the L channel and the R channel, and filter coefficients indicating the vocal frequency bands to be removed. In other words, the vocal removal unit 10 preprocesses the sound signal processed by the surround signal processing unit 20 to suppress unclear speech output due to the addition of stereo effects to the vocal components.
[0080] The input signal of channel L is an example of the audio signal of the first channel, the input signal of channel R is an example of the audio signal of the second channel, and the vocal removal signal is an example of the first audio signal. Furthermore, vocal removal unit 10 is an example of a removal unit.
[0081] The vocal removal unit 10 includes a difference signal generation unit 11 and a filter unit 12.
[0082] The difference signal generation unit 11 receives the input signals from the L channel and the R channel, and generates a difference signal that is the difference between the two input signals. The difference signal is a signal that represents the difference between the input signals from the L channel and the R channel. The difference signal generation unit 11 is an example of a first signal generation unit.
[0083] Here, the input signals for the L channel and the R channel are audio signals used to output stereo sound. The input signal for the L channel includes the sound (speech and other sounds) output from the L-side speaker, and the input signal for the R channel includes the sound (speech and other sounds) output from the R-side speaker. The vocal components (speech signal components) of the input signals for the L channel and the R channel are substantially the same. Furthermore, the components other than the vocal components of the input signals for the L channel and the R channel are signal components that are different between the L channel and the R channel.
[0084] The difference signal generation unit 11 takes the difference between the input signal of the L channel and the input signal of the R channel, thereby eliminating the vocal component (center component) commonly included in both the L channel and R channel input signals. Therefore, the difference signal generated by the difference signal generation unit 11 contains almost no vocal component; however, depending on the content, there may be residual vocal component in the difference signal. For example, if one of the input signals of the L channel and the R channel is intentionally delayed (effect) to stagger the sound output timing, vocal component may be included in the difference signal.
[0085] The filter unit 12 receives the input difference signal and removes the vocal components included in the difference signal to generate a vocal removal signal. The filter unit 12 also removes the frequency components of the vocal frequency band based on the filter coefficients determined by the coefficient determination unit 40 from the difference signal to generate a vocal removal signal.
[0086] The filter unit 12 is configured, for example, to include an IIR (Infinite Impulse Response) filter, but is not limited thereto. In this embodiment, the filter unit 12 is configured, for example, to include a high-pass filter (HPF), but it may also be configured to include a low-pass filter (LPF), or both an HPF and an LPF. For example, in the case of performing surround signal processing on low-frequency speech, the filter unit 12 may be configured to include a low-pass filter. The filter unit 12 may also be configured to include any filter if it can remove vocal components from the difference signal. An example of the filter unit 12 being configured to include an HPF will be described below.
[0087] The filter unit 12 removes vocal components at a cutoff frequency based on the filter coefficients determined by the coefficient determination unit 40. If the cutoff frequency increases, the bandwidth of the removed vocal components widens. In other words, if the cutoff frequency increases, the intensity of the removed vocal signal decreases. Furthermore, the bandwidth of the vocal components is, for example, primarily around 300Hz to 2000Hz, but is not limited to this. And the filter coefficients are an example of the first coefficients representing the vocal bandwidth to be removed.
[0088] The vocal removal unit 10, consisting of the difference signal generation unit 11 and the filter unit 12, is capable of generating a vocal removal signal that has removed most of the vocal components.
[0089] The surround processing unit 20 performs surround signal processing on the vocal removal signal from the vocal removal unit 10 to add surround effects, thereby generating an adjustment signal. The surround processing unit 20 includes a surround signal generation unit 21 and an amplification unit 22.
[0090] The surround signal generation unit 21 performs surround signal processing on the vocal removal signal to generate a surround signal. Alternatively, the surround signal generation unit 21 adds a surround effect to the vocal removal signal to generate a surround signal. Furthermore, regarding surround signal processing, any known processing can be performed if a surround effect can be added to the vocal removal signal. The surround signal generation unit 21 is an example of a second signal generation unit. And the surround signal is an example of a second output signal.
[0091] The amplification unit 22 amplifies the input signal with a gain value (an example of amplification rate) based on the amplification factor determined by the coefficient determination unit 40. In this embodiment, the amplification unit 22 is connected between the surround signal generation unit 21 and the synthesis unit 50. Therefore, the input surround signal is amplified with a gain value based on the amplification factor to generate an adjustment signal. Alternatively, the amplification unit 22 adjusts the intensity of the surround signal synthesized from the input signals of the L channel and the R channel. The intensity of the surround signal is the absolute quantity (integral value) of the signal with added surround effect. Furthermore, the intensity of the surround signal can also be described as the intensity of the stereoscopic, depth, or expansion of sounds other than speech output from the audio device.
[0092] The amplification unit 22 amplifies the surround signal with an amplification factor determined by the coefficient determination unit 40. The amplification unit 22 adjusts the strength of the surround signal by changing the gain value of the surround signal according to the amplification factor from the coefficient determination unit 40. If the gain value increases, the strength of the surround signal increases.
[0093] Thus, in this embodiment, the surround processing unit 20 adds surround effects to the vocal removal signal and adjusts the intensity of the surround signal.
[0094] User interface 30 receives input from the user regarding sound signal processing. User interface 30, for example, obtains information about the user's preferred sound quality and outputs this information to coefficient determination unit 40. In this embodiment, user interface 30 receives input regarding vocal clarity. Vocal clarity indicates the degree of clarity of speech; in this embodiment, it indicates the clarity of speech in the sound output from the L-side speaker and the R-side speaker. Vocal clarity specifies the degree of sound quality the user prefers in the speech. High vocal clarity means, for example, that the speech is clearly audible, i.e., the speech is clear. Vocal clarity is represented by a value from 0 to 100, but it is not limited to this.
[0095] Moreover, the user interface 30 is not a necessary structure for the sound signal processing device 1.
[0096] The coefficient determination unit 40 determines the filtering coefficient of the filter unit 12 and the amplification coefficient of the amplification unit 22. In this embodiment, the coefficient determination unit 40 obtains the vocal clarity from the user interface 30 and determines the filtering coefficient and amplification coefficient based on the obtained vocal clarity. The coefficient determination unit 40 determines the filtering coefficient by associating it with the amplification coefficient. The coefficient determination unit 40 is an example of a setting unit that sets the filtering coefficient and amplification coefficient.
[0097] The coefficient determination unit 40, for example, if the cutoff frequency (cutoff frequency of the HPF) based on the filter coefficient increases, the absolute amount of the vocal signal removed decreases, resulting in a decrease in the intensity of the surround signal. Therefore, it increases the gain value, thereby amplifying the intensity of the surround signal. The coefficient determination unit 40, for example, if the filter coefficient is determined to be a value where the cutoff frequency increases, it also determines the amplification factor to be a value where the gain value increases. The coefficient determination unit 40, for example, if the vocal frequency band removed according to the filter coefficient is a second frequency band wider than the first frequency band, determines the amplification factor in such a way that the gain value in the second frequency band is greater than the gain value in the first frequency band. The coefficient determination unit 40 determines the second coefficient in such a way that it becomes the amplification rate that eliminates the change in the intensity of the vocal signal removed by the filtering process based on the filter unit 12.
[0098] Furthermore, the coefficient determination unit 40 determines the filter coefficient based on the principle that the higher the clarity of the sound, the higher the cutoff frequency of the HPF, and sets the amplification coefficient based on the principle that the higher the gain value of the amplification unit 22, the higher the amplification coefficient.
[0099] The determination of the filtering coefficient and amplification coefficient by the coefficient determination unit 40 will be explained later. Furthermore, the coefficient determination unit 40, for example, determines a set of filtering coefficients and amplification coefficients for each piece of content. That is, the coefficient determination unit 40 does not change the filtering coefficients and amplification coefficients during content reproduction. Moreover, the content is not particularly limited if it includes audio information used for output sound; it can be speech content or moving image content.
[0100] The synthesis unit 50 processes the adjustment signal output from the surround processing unit 20 and returns it to the input signals of the L channel and R channel. The synthesis unit 50 synthesizes the adjustment signal with the input signals of the L channel and R channel, and outputs the synthesized signal to the L-side speaker and the R-side speaker. The synthesis unit 50 includes a first synthesis unit 51 and a second synthesis unit 52. The first synthesis unit 51 and the second synthesis unit 52 are, for example, adders.
[0101] The first synthesis unit 51 combines the adjustment signal with the input signal of the L channel to generate an L-side synthesized signal. The L-side synthesized signal is, for example, the sum of the input signal of the L channel and the adjustment signal. The first synthesis unit 51 outputs the L-side synthesized signal to the L-side speaker. The L-side synthesized signal is an example of the first synthesized signal.
[0102] The second synthesis unit 52 synthesizes the adjustment signal inverted by the inversion unit 60 with the input signal of the R channel to generate an R-side synthesized signal. The R-side synthesized signal is, for example, the sum of the input signal of the R channel and the inverted adjustment signal. The second synthesis unit 52 outputs the R-side synthesized signal to the R-side speaker. The R-side synthesized signal is an example of the second synthesized signal.
[0103] The inversion unit 60 inverts the input signal and outputs it. In this embodiment, the inversion unit 60 inverts the phase of the adjustment signal output from the surround processing unit 20 and outputs it to the second synthesis unit 52. Alternatively, the inversion unit 60 performs a process that delays the adjustment signal by 1 / 2 cycle.
[0104] Furthermore, the inversion unit 60 can be connected to either the surround processing unit 20 and the first synthesis unit 51, or the surround processing unit 20 and the second synthesis unit 52. The inversion unit 60 can be connected in a manner that inverts the phase of either the adjustment signal input to the L channel or the input signal to the R channel. Alternatively, the inversion unit 60 can, for example, invert the phase of the adjustment signal output from the surround processing unit 20 and output it to the first synthesis unit 51.
[0105] Furthermore, while the amplification unit 22 has been described as a component of the surround processing unit 20, it is not limited thereto. The amplification unit 22 may, for example, be connected between the vocal removal unit 10 and the surround processing unit 20, amplifying the vocal removal signal from the filter unit 12 and outputting it to the surround processing unit 20. Also, the amplification unit 22 may, for example, be connected between the difference signal generation unit 11 and the filter unit 12 (configured as part of the vocal removal unit 10), amplifying the difference signal from the difference signal generation unit 11 and outputting it to the filter unit 12. Furthermore, the amplification unit 22 may, for example, be connected between the difference signal generation unit 11 and the signal lines transmitting the input signals of the L channel and the R channel (connected to the preamplifier of the vocal removal unit 10), amplifying the input signals of the L channel and the R channel and outputting them to the difference signal generation unit 11. Thus, the location of the amplification unit 22 is not particularly limited.
[0106] In this case, the amplification unit 22 amplifies any one of the acoustic signal removal signal, the difference signal, or the input signal from the L channel and the input signal from the R channel. However, by amplifying these signals, the intensity of the surround signal is also amplified. Thus, the amplification unit 22 can also indirectly adjust the intensity of the surround signal.
[0107] The hardware structure constituting the aforementioned sound signal processing device 1 is not particularly limited; however, it can, for example, be constructed using a computer. In an example of such a hardware structure, utilizing... Figure 2 Please provide an explanation. Figure 2 This is a diagram illustrating an example of the hardware structure of a computer 1000 in which the functions of the sound signal processing device 1 according to this embodiment are implemented by software.
[0108] like Figure 2 As shown, computer 1000 is a computer having an input device 1001, an output device 1002, a CPU 1003, an internal memory 1004, a RAM 1005, and a bus 1009. The input device 1001, the output device 1002, the CPU 1003, the internal memory 1004, and the RAM 1005 are connected by the bus 1009.
[0109] Input device 1001 is a device that serves as a user interface, such as an input button, touchpad, or touchscreen display, and accepts user operations. Furthermore, input device 1001 can also be configured to accept user touch operations, as well as voice operations and remote operations such as those using a remote control. Input device 1001, for example, is... Figure 1 The user interface 30 shown corresponds to this. Furthermore, the input device 1001, for example, is associated with the input... Figure 1 The device shown corresponds to the input signal of the L channel and the input signal of the R channel.
[0110] Output device 1002 is a device that outputs signals from computer 1000. Besides signal output terminals, it can also be a user interface device such as a speaker or display. Output device 1002, and output... Figure 1 The apparatus shown corresponds to the L-side synthesized signal and the R-side signal. Furthermore, the output device 1002 may also include, equivalent to... Figure 1 The speakers shown are the L-side speaker and the R-side speaker.
[0111] The internal memory 1004 is, for example, flash memory. Furthermore, the internal memory 1004 may also pre-store at least one of the programs for implementing the functions of the sound signal processing device 1 and the applications utilizing the functional structure of the sound signal processing device 1.
[0112] RAM1005 is Random Access Memory, used for storing data and other data when executing programs or applications.
[0113] CPU1003 is the Central Processing Unit, which copies programs and applications stored in internal memory 1004 to RAM 1005, and reads and executes commands included in the program or application sequentially from RAM 1005.
[0114] The computer 1000 may also perform the same processing on the first audio signal (e.g., the input signal of the L channel) and the second audio signal (e.g., the input signal of the R channel) as the vocal removal unit 10, the surround processing unit 20 and the coefficient determination unit 40 according to this embodiment.
[0115] [1-2. Determination of the coefficients in the coefficient determination section]
[0116] Next, for the determination of each coefficient in the coefficient determination section 40, refer to... Figures 3 to 7 Please provide an explanation. Figure 3 This is a graph illustrating a first example of the relationship between vocal intelligibility, cutoff frequency (Fc), and gain value according to this embodiment. Alternatively, it can be said that... Figure 3 The diagram shows the correspondence between the cutoff frequency (Fc) and the gain value corresponding to the vocal intelligibility value.
[0117] like Figure 3As shown, the cutoff frequency and gain value corresponding to the vocal intelligibility value can also have a linear correlation. In this case, if the cutoff frequency increases, the gain value corresponding to that cutoff frequency also increases proportionally to the cutoff frequency. Furthermore, if vocal intelligibility is obtained, the cutoff frequency and gain value corresponding to that vocal intelligibility can be uniquely determined.
[0118] and, Figure 3 The vocal clarity shown is as follows: High vocal clarity (e.g., close to 100) is indicated by setting the HPF cutoff frequency to a high value, and consequently, the gain value is also set to a high value. Therefore, when the intensity of the surround signal decreases due to the filtering process of the filter unit 12, the amplification unit 22 can increase the intensity of the surround signal. Thus, by determining the filter coefficient that increases vocal clarity, the weakening of the surround effect due to the decrease in surround signal intensity can be suppressed.
[0119] and, Figure 3 The vocal clarity shown is as follows: low vocal clarity (e.g., close to 0) determines a low cutoff frequency for the HPF, and consequently, a low gain value.
[0120] Coefficient determination unit 40, for example, using the shown Figure 3 The equation showing the correlation determines the cutoff frequency and gain value. The coefficient determination unit 40, for example, calculates the cutoff frequency according to the following equation 1, thereby determining the cutoff frequency.
[0121] Fc[Hz] = Vocal clarity × A + B (Equation 1)
[0122] A is the skewness, and B is the slice. The skewness A and slice B are determined appropriately according to the content, etc. However, for example, the skewness A can also be 40, and the slice B can also be 200.
[0123] Furthermore, the coefficient determination unit 40 calculates the gain value, for example, according to the following formula 2, thereby determining the gain value.
[0124] Gain value [dB] = (Fc[Hz]) × C + D Equation (2)
[0125] C is the skewness, and D is the slice. The skewness C and slice D are determined appropriately according to the content, etc. However, for example, the skewness C can also be 1 / 350, and the slice D can also be -10 / 7.
[0126] Moreover, the correlation is not limited to linear relationships. Figure 4 This is a second example of the relationship between vocal clarity, cutoff frequency, and gain value involved in this embodiment.
[0127] like Figure 4As shown, the cutoff frequency and gain value relative to vocal intelligibility can also have a non-linear correlation. This correlation can, for example, be represented by an upwardly convex function. Furthermore, the correlation between the cutoff frequency and vocal intelligibility can, for example, be represented by an exponential function as shown in Equation 3 below. Accordingly, it is possible to make the change in speech intelligibility equal to the change in vocal intelligibility. For example, it is possible to make the change in speech intelligibility in the low-frequency domain that results in a specified change in vocal intelligibility equal to the change in speech intelligibility in the high-frequency domain that results in a specified change in vocal intelligibility.
[0128] Fc[Hz]=EXP(vocal clarity × E)×F Equation (3)
[0129] E is the coefficient used to calculate the exponentiation, and F is the slice. The coefficient E and the slice F are determined appropriately according to the content, etc., but, for example, the coefficient E can also be 0.03, and the slice F can also be 200. Moreover, the base of Equation 3 is, for example, Napier's constant.
[0130] Furthermore, the relationship between the cutoff frequency and the gain value can, for example, be represented by an upwardly convex function. The relationship between the cutoff frequency and the gain value can also be represented by, for example, a logarithmic function as shown in Equation 4 below. Accordingly, while maintaining a constant surround sound effect, vocal clarity can be altered. That is, while maintaining a constant surround sound effect, the cutoff frequency and gain value corresponding to vocal clarity can be determined.
[0131] Gain value [dB] = ln(Fc[Hz]) × G + H Equation (4)
[0132] G is the coefficient used to calculate the argument, and H is the slice. The coefficient G and the slice H are determined appropriately according to the content, etc., but, for example, the coefficient G could also be 3.0686, and the slice H could also be -18.327. Moreover, the base of Equation 4 is, for example, Napier's constant.
[0133] Furthermore, surround sound indicates the user's subjective perception of the surround effect. Strong surround sound indicates that the user strongly feels the surround effect (e.g., strongly feels the stereo effect of the sound), while weak surround sound indicates that the user does not feel the surround effect very much.
[0134] like Figure 3 as well as Figure 4 Alternatively, when the cutoff frequency of the filter section 12 (e.g., a high-pass filter) is set on the horizontal axis and the gain value of the amplification section 22 is set on the vertical axis, the vocal clarity can be represented by a monotonically increasing graph. Furthermore, the monotonically increasing graph can specifically be a logarithmic graph or a linear graph. The coefficient determination section 40 utilizes... Figure 3 or Figure 4 The relationship shown in the monotonically increasing graph allows the amplification factor to be determined in conjunction with the filter coefficients. In other words, the coefficient determination unit 40 can determine the intensity of the surround signal in conjunction with the frequency band of the sound removed from the difference signal. Alternatively, the coefficient determination unit 40 can determine the intensity of the surround signal in conjunction with the amount of signal removed from the difference signal (e.g., the integral value of the removed signal).
[0135] Here, for the functional experiments used to derive Equation 4, refer to Figure 5 as well as Figure 6 Please provide an explanation. Figure 5 This is a diagram showing the results of a sensory experiment on surround sensation related to this embodiment. Figure 6 This is a graph showing the results of a sensory experiment on vocal clarity related to this embodiment.
[0136] In the functional experiment, the cutoff frequencies of the filter section 12 were set to 200Hz, 300Hz, 400Hz, 500Hz, 800Hz, 1000Hz, 1500Hz, 2000Hz, 2500Hz, 3000Hz, and 4000Hz, and the gain value of the amplification section 22 at each cutoff frequency varied in 1dB intervals from -5 to +6dB in mode 132. Figure 5 The results show the subjective evaluation of the sense of surround sound in each mode. Figure 6 The results show the subjective evaluation of vocal clarity in various modes. Furthermore, Latin music was used as the sound source in the experiment.
[0137] exist Figure 5 In the diagram, “×1” indicates a condition where the surround sensation is too strong, “△1” indicates a condition where the surround sensation is strong, “○” indicates a condition where the surround sensation is good, “△2” indicates a condition where the surround sensation is weak, and “×2” indicates a condition where the surround sensation is not felt (too weak).
[0138] like Figure 5 The results show that the surround sound tends to be weak when the gain is low and the cutoff frequency is high, and strong when the gain is high and the cutoff frequency is low.
[0139] exist Figure 6 In this text, “○” indicates a condition where the vocal music is clearly audible (condition for clear speech), “△” indicates a condition where the vocal music is vaguely audible, and “×” indicates a condition where the vocal music is unclear. Furthermore, vaguely audible indicates, for example, that the speech is unclear to the extent that the meaning can be understood, and unclear indicates, for example, that the speech is unclear to the extent that at least part of the meaning cannot be understood.
[0140] like Figure 6The results show that vocal clarity tends to become unclear under conditions of high gain and low cutoff frequency.
[0141] Figure 5 as well as Figure 6 The thick box shown indicates the condition where both surround sound and vocal clarity are "○". The coefficient determination unit 40 determines the filter coefficient and amplification coefficient in a manner that becomes the cutoff frequency and gain value within the thick box, thereby achieving both vocal clarity and surround sound simultaneously.
[0142] and then, Figure 7 The diagram shows a set of cutoff frequencies and gain values that provide the same surround sound even when the cutoff frequency is changed, plotted within the thick box for each cutoff frequency. Figure 7 This is a third example of a graph illustrating the relationship between vocal clarity, cutoff frequency, and gain value involved in this embodiment.
[0143] Figure 7 Yes, drawing will Figure 5 as well as Figure 6 The diagram shows the results of evaluating the surround sound at a cutoff frequency of 400Hz and a gain of 0dB as a reference (hereinafter also referred to as the reference surround sound), and obtaining the gain values that are equal to the surround sound at 400Hz at various frequencies other than 400Hz. For example, at a cutoff frequency of 300Hz, the surround sound felt at a gain of -1dB (within the thick box) is equal to the reference surround sound. Furthermore, at a cutoff frequency of 3000Hz, the surround sound felt at a gain of +6dB (within the thick box) is equal to the reference surround sound. Moreover, the reference surround sound is not limited to the surround sound at 400Hz.
[0144] Here, if we calculate an approximation of the drawn data string, then as follows: Figure 7 As shown, it becomes Equation 5 below.
[0145] Gain value [dB] = 3.0686ln(Fc) - 18.327 Equation (5)
[0146] Equation 5 is a function with coefficient G of 3.0686 and slice H of -18.327 in Equation 4. By using this approximation, vocal clarity can be altered while maintaining a certain surround sound effect.
[0147] Moreover, Equations 1 to 5 above are just one example, and are not limited to this. For example, the approximate formula shown in Equation 5 is just one example, and it will vary depending on the type of sound source, the user's attributes (age, gender, etc.).
[0148] Furthermore, any of the described formulas is provided by the storage unit (e.g., of the sound signal processing device 1) Figure 2The internal memory 1004 shown is pre-stored.
[0149] [1-3. Operation of the sound signal processing device]
[0150] Next, regarding the operation of the aforementioned sound signal processing device 1, refer to... Figure 8 Please provide an explanation. Figure 8 This is a flowchart illustrating the operation of the sound signal processing apparatus 1 according to this embodiment. Furthermore, it is conceived below that the storage unit of the sound signal processing apparatus 1 has pre-storage types 3 and 4.
[0151] like Figure 8 As shown, the user interface 30 obtains the vocal clarity from the user (S101). The user interface 30, for example, obtains a value from 0 to 100 as the vocal clarity. Furthermore, the vocal clarity can be obtained during content reproduction or in advance via a storage unit (e.g., provided by the sound signal processing device 1). Figure 2 The internal memory 1004 shown stores the audio. The user interface 30 outputs the obtained vocal clarity to the coefficient determination unit 40.
[0152] Furthermore, the user interface 30 can also obtain the vocal clarity from the user's choice of "high," "medium," "low," etc., instead of a numerical value.
[0153] Next, the coefficient determination unit 40 determines the filter coefficient and the corresponding amplification coefficient based on the vocal clarity (S102). The coefficient determination unit 40 reads Equation 3 from the storage unit, substitutes the vocal clarity into Equation 3 to calculate the cutoff frequency for achieving vocal clarity, and determines the filter coefficient corresponding to the calculated cutoff frequency. Furthermore, the coefficient determination unit 40 reads Equation 4 from the storage unit, substitutes the cutoff frequency corresponding to the determined filter coefficient into Equation 4 to calculate the gain value for achieving the desired surround sound, and determines the amplification coefficient corresponding to the calculated gain value, i.e., the amplification coefficient corresponding to the filter coefficient. Moreover, the coefficient determination unit 40 outputs the determined filter coefficient to the filter unit 12 and the determined amplification coefficient to the amplification unit 22. Step S102 is an example of a setting step.
[0154] Next, the difference signal generation unit 11 generates the difference signal, i.e., the difference signal, between the input signal of the L channel and the input signal of the R channel (S103). The difference signal generation unit 11 outputs the generated difference signal to the filter unit 12.
[0155] Next, the filter unit 12 generates a vocal removal signal based on the difference signal and the filter coefficients (S104). The filter unit 12 extracts high-frequency components from the difference signal based on the cutoff frequency of the filter coefficients, thereby generating the vocal removal signal. The filter unit 12 outputs the vocal removal signal to the surround signal generation unit 21. Step S104 is an example of the removal step.
[0156] Next, the surround signal generation unit 21 performs surround signal processing (S105) on the vocal removal signal to generate a surround signal. The surround signal generation unit 21 outputs the generated surround signal to the amplification unit 22. Step S105 is an example of the surround signal processing step.
[0157] Next, the amplification unit 22 generates an adjustment signal based on the amplification factor and the surround signal (S106). Since the cutoff frequency is set to a high value by the factor determination unit 40, the strength of the surround signal is small (the absolute amount of the surround signal is small). Therefore, the amplification factor is determined to be high. Accordingly, the amplification unit 22 can increase the strength of the surround signal, which has decreased in strength due to the filtering process of the filter unit 12. Step S106 is an example of the amplification step.
[0158] Thus, the amplification unit 22 adjusts the strength of the signal synthesized from the input signals of the L channel and the R channel. The amplification unit 22 then outputs an adjustment signal to the synthesis unit 50.
[0159] Next, the synthesis unit 50 synthesizes the signal based on the adjustment signal with the input signal of the L channel and the input signal of the R channel (S107). In this embodiment, the first synthesis unit 51, as the signal based on the adjustment signal, synthesizes the adjustment signal itself with the input signal of the L channel to generate an L-side synthesized signal. Furthermore, the second synthesis unit 52, as the signal based on the adjustment signal, synthesizes the adjustment signal (after its phase is inverted by the inversion unit 60) with the input signal of the R channel to generate an R-side synthesized signal. The first synthesis unit 51 outputs the generated L-side synthesized signal to the L-side speaker, and the second synthesis unit 52 outputs the generated R-side synthesized signal to the R-side speaker. Step S107 is an example of the first synthesis step and the second synthesis step.
[0160] Accordingly, the signals output from the sound signal processing device 1 to the L-side speaker and the R-side speaker respectively become signals with the desired surround effect intensity. That is, signals that produce the desired surround sound. Therefore, the audio device is able to perform the desired surround reproduction. The audio device can, for example, output sound with the sound image positioned in an area wider than the arrangement positions of the L-side speaker and the R-side speaker.
[0161] (Implementation Method 2)
[0162] [2-1. Structure of a sound signal processing device]
[0163] First, regarding the structure of the sound signal processing device involved in this embodiment, refer to... Figure 9 Please provide an explanation. Figure 9 This is a block diagram illustrating the functional structure of the audio signal processing apparatus 100 according to this embodiment. The main difference between the audio signal processing apparatus 100 according to this embodiment and the audio signal processing apparatus 1 according to Embodiment 1 is that the coefficient determination unit 140 also determines the filtering coefficient and the amplification coefficient based on the surround sound effect. Hereinafter, the audio signal processing apparatus 100 according to this embodiment will be described focusing on the differences between it and the audio signal processing apparatus 1 according to Embodiment 1.
[0164] From now on, for structures that are the same as or similar to the audio signal processing apparatus 1 according to Embodiment 1, the same reference numerals as those used in the audio signal processing apparatus 1 according to Embodiment 1 will be added, and descriptions will be omitted or simplified. Furthermore, the hardware structure constituting the elements of the audio signal processing apparatus 100 is not particularly limited; however, for example, it may be similar to that used in Embodiment 1. Figure 2 The hardware structure of the computer 1000 described is the same.
[0165] like Figure 9 As shown, the sound signal processing apparatus 100, instead of the coefficient determination unit 40 of the sound signal processing apparatus 1 according to Embodiment 1, includes a coefficient determination unit 140. Furthermore, the user interface 30 receives input from the user regarding surround sound in addition to vocal clarity. Surround sound is an example of user preference, indicating the intensity of the surround effect preferred by the user, for example, represented by a numerical value from 0 to 100. For example, a surround sound value of 100 or close to 100 indicates a strong surround effect (e.g., strong stereoscopic, depth, or expansion of sounds other than speech). Conversely, a surround sound value of 0 or close to 0 indicates a weak surround effect (e.g., weak stereoscopic, depth, or expansion of sounds other than speech). Moreover, surround sound is not limited to being represented by a numerical value.
[0166] The coefficient determination unit 140 determines the filtering coefficient and the amplification coefficient based on vocal clarity and surround sound. For example, the coefficient determination unit 140 obtains vocal clarity and surround sound from the user interface 30, determines the filtering coefficient based on the obtained vocal clarity, and determines the amplification coefficient based on the obtained vocal clarity and surround sound.
[0167] [2-2. Determination of the coefficients in the coefficient determination section]
[0168] Next, for the determination of each coefficient in the coefficient determination unit 140, refer to... Figure 10 as well as Figure 11 Please provide an explanation. Figure 10 This is a diagram illustrating a first example of the relationship between vocal clarity and surround sound, cutoff frequency, and gain value involved in this embodiment. Figure 10 The diagram shows the correspondence between the cutoff frequency (Fc) and the gain value relative to the value of vocal clarity, and the correspondence between the gain value and the value of surround sound.
[0169] like Figure 10 The diagram shows that the cutoff frequency and gain value have a linear correlation with vocal intelligibility and a correlation parallel to the gain value's axis with respect to surround sound. In other words, the cutoff frequency is determined by vocal intelligibility, and the gain value is determined by both vocal intelligibility and surround sound. Conversely, surround sound is not used to determine the cutoff frequency.
[0170] and, Figure 10 The surround effect shown is as follows: Elegant surround effect has a small surround effect (e.g., close to 0), which determines a low gain value. Aggresive surround effect has a large surround effect (e.g., close to 100), which determines a high gain value.
[0171] The coefficient determination unit 140 can also, for example, utilize the shown Figure 10 The equation showing the correlation determines the cutoff frequency and the gain value. The coefficient determination unit 140 may also, for example, calculate the gain value according to the following equation 6, thereby determining the gain value. Furthermore, the equation for calculating the cutoff frequency by the coefficient determination unit 140 is the same as equation 1 in Embodiment 1, and its explanation is omitted.
[0172] Gain value [dB] = (Fc[Hz]) × C + D + surround sound effect × E + F Equation (6)
[0173] E is the tilt relative to the surround effect, and F is the slice relative to the surround effect. The tilt C and E, as well as the slices D and F, are appropriately determined according to the content, etc. However, for example, the tilt C could be 1 / 350, the slice D could be -10 / 7, the tilt E could be 1 / 25, and the slice F could be -2. Furthermore, the slice relative to the gain value can be calculated by adding slices D and F together.
[0174] Moreover, the correlation between the cutoff frequency (Fc) and the gain value relative to the vocal intelligibility value is not limited to linearity. Figure 11 This is a second example of the relationship between vocal clarity and surround sound, cutoff frequency, and gain value involved in this embodiment.
[0175] like Figure 11As shown, the cutoff frequency and gain value can also have a non-linear relationship with vocal intelligibility. The relationship between the cutoff frequency and gain value and vocal intelligibility can, for example, be represented by an upwardly convex function.
[0176] The coefficient determination unit 140 can also, for example, utilize the shown Figure 11 The equation showing the correlation determines the cutoff frequency and the gain value. The coefficient determination unit 140 may also, for example, calculate the gain value according to the following equation 7, thereby determining the gain value. Furthermore, the equation for calculating the cutoff frequency by the coefficient determination unit 140 is the same as equation 3 in Embodiment 1, and its explanation is omitted.
[0177] Gain value [dB] = log(Fc[Hz]) × C + D + surround sound × E + F Equation (7)
[0178] The tilt angles C and E, and the slices D and F, are the same as in Equation 6.
[0179] like Figure 10 as well as Figure 11 Alternatively, when the cutoff frequency of the filter section 12 (high-pass filter) is set as the horizontal axis and the gain value of the amplifier section 22 is set as the vertical axis, the surround sound can be represented by a graph parallel to the gain value axis.
[0180] The coefficient determination unit 140 determines the gain value using the cutoff frequency calculated in Equation 3 and Equation 7, and adjusts the surround sound to the user's preference while maintaining a constant vocal clarity. The amplification factor corresponding to this determined gain value is an example of an amplification factor determined based on vocal clarity and surround sound.
[0181] (Other implementation methods)
[0182] The various embodiments (hereinafter also referred to as embodiments, etc.) have been described above. However, this disclosure is not limited to such embodiments. Various modifications that can be conceived by those skilled in the art to the various embodiments, as well as other forms that combine some of the constituent elements of the various embodiments, are also included within the scope of this disclosure, as long as they do not depart from the spirit of this disclosure.
[0183] For example, the various embodiments described above illustrate an example where the coefficient determination unit determines the filter coefficient and amplification coefficient based on the vocal clarity, or vocal clarity and surround sound, obtained from the user interface. However, the method for determining each coefficient is not limited to this. For example, the storage unit of the sound signal processing device may store a table corresponding to information about the sound source or user identification information and the filter coefficient and amplification coefficient. Based on the currently obtained information about the sound source or user identification information and this table, the filter coefficient and amplification coefficient corresponding to the obtained information are determined. The information about the sound source includes the type of sound source, the purpose of the sound source (for movies, karaoke, etc.), etc., but is not limited to this. The user identification information is information used to identify the user. In this case, the filter coefficient and amplification coefficient are matched in the table in such a way that if the filter coefficient increases, the amplification coefficient also increases.
[0184] Furthermore, Equations 2, 4, and 6 of the aforementioned embodiments are examples of equations showing the relationship between the cutoff frequency and the gain value. However, they are not limited to these examples and may also be equations showing the relationship between vocal clarity and the gain value.
[0185] Furthermore, the coefficient determination unit in the aforementioned embodiment may also determine the filtering coefficients without removing the difference signal component when the input signals of the L channel and the R channel do not contain vocal components. In other words, the coefficient determination unit may also determine the filtering coefficients by allowing the difference signal to pass through as is. Alternatively, the coefficient determination unit may obtain information about the reproduced sound via a user interface or the like, determine whether the reproduced sound contains vocal components based on the obtained information, and perform the filtering coefficient determination process according to the determination result.
[0186] Furthermore, all or specific forms of this disclosure can be implemented by systems, apparatuses, methods, integrated circuits, computer programs, or recording media such as computer-readable CD-ROMs. Moreover, it can also be implemented by any combination of systems, apparatuses, methods, integrated circuits, computer programs, and recording media.
[0187] Furthermore, the processing order described in the flowcharts of the aforementioned embodiments is just one example. The order of multiple processes can be changed, and multiple processes can be executed simultaneously.
[0188] A portion of the components constituting the aforementioned audio signal processing device may also be comprised of a system LSI (Large Scale Integration). A system LSI is a highly multifunctional LSI manufactured by integrating multiple components onto a single chip; specifically, it includes a computer system comprising a microprocessor, ROM, RAM, etc. The RAM stores a computer program. The microprocessor operates according to the computer program, thereby enabling the system LSI to perform its functions.
[0189] A portion of the components constituting the aforementioned audio signal processing device may also consist of an IC card or a module that is removable from each device. The IC card or module is a computer system comprised of a microprocessor, ROM, RAM, etc. The IC card or module may also include the multi-functional LSI. The microprocessor operates according to a computer program, thereby enabling the IC card or module to perform its functions. The IC card or module may also possess tamper-proof features.
[0190] Furthermore, a component of the aforementioned sound signal processing apparatus may also include recording the computer program or the digital signal onto a computationally readable recording medium, such as a floppy disk, hard disk, CD-ROM, MO, DVD, DVD-ROM, DVD-RAM, BD (Blu-ray Disc), semiconductor memory, etc. The digital signal may also be recorded on these recording media.
[0191] Furthermore, a component of the aforementioned sound signal processing device may also transmit the computer program or the digital signal via electrical communication lines, wireless or wired communication lines, networks such as the Internet, or data broadcasting.
[0192] This disclosure can also be the methods described above. Furthermore, these methods can be implemented by a computer program, or they can be digital signals composed of said computer program.
[0193] Furthermore, this disclosure can also be a computer system equipped with a microprocessor and a memory, wherein the memory stores the computer program, and the microprocessor operates according to the computer program.
[0194] Furthermore, the program or the digital signal can be recorded onto the recording medium for transmission, or the program or the digital signal can be transmitted via the network or the like, so that it can be executed by an independent other computer system.
[0195] Furthermore, implementation methods can be combined separately.
[0196] This disclosure is applicable to audio devices that perform surround sound reproduction, etc.
Claims
1. A sound signal processing device, The sound signal processing device comprises: The removal unit generates a first output signal obtained by removing vocal components based on the audio signals of the first channel and the second channel, as well as a first coefficient indicating the vocal frequency band to be removed. The surround processing unit adds a surround effect to the first output signal to generate a second output signal; An amplification unit is connected to the pre-stage of the removal unit or between the removal unit and the surround processing unit, or is configured as part of the removal unit or the surround processing unit, to amplify the input signal based on the amplification rate of the second coefficient; The first synthesis unit synthesizes the second output signal and one of the audio signals from the first channel and the second channel. The second synthesis unit combines the inverted signal of the second output signal with the audio signal from the first channel and the other side of the audio signal from the second channel. as well as The setting unit sets the first coefficient and the second coefficient. The setting unit sets the second coefficient in such a way that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is the first frequency band is greater than that when it is the first frequency band.
2. The sound signal processing device as described in claim 1, The setting unit sets the first coefficient and the second coefficient according to vocal clarity, which indicates the clarity of speech based on the signal synthesized by the first synthesis unit and the second synthesis unit.
3. The sound signal processing apparatus as described in claim 2, The removal section has a high-pass filter. The setting unit sets the first coefficient so that the higher the clarity, the higher the cutoff frequency of the high-pass filter, and sets the second coefficient so that the amplification is higher.
4. The sound signal processing device as described in claim 2, The removal section has a high-pass filter. The vocal clarity, when the cutoff frequency of the high-pass filter is set on the horizontal axis and the amplification rate of the amplification unit is set on the vertical axis, is represented by a graph that increases monotonically. The setting unit sets the first coefficient and the second coefficient based on the vocal clarity and the monotonically increasing graph.
5. The sound signal processing apparatus as described in claim 4, The monotonically increasing chart is a logarithmic chart.
6. The sound signal processing apparatus as described in claim 4, The monotonically increasing chart is a straight-line chart.
7. The sound signal processing apparatus as described in any one of claims 2 to 6, The sound signal processing device also includes a user interface for receiving the vocal clarity from the user.
8. The sound signal processing apparatus as described in any one of claims 2 to 6, The setting unit further sets the second coefficient according to the surround feel shown, which represents an additional user preference for the surround effect.
9. The sound signal processing apparatus as described in claim 8, The sound signal processing device also includes a user interface for receiving the vocal clarity and the surround sound from the user.
10. The sound signal processing apparatus as described in any one of claims 1 to 6, 9, The removal section has: A first signal generation unit generates a difference signal that shows the difference between the audio signal from the first channel and the audio signal from the second channel; and The filter section removes the frequency components of the vocal band based on the first coefficient from the difference signal, thereby generating the first output signal. The surrounding processing unit has: The second signal generation unit adds the surround effect to the first output signal to generate a surround signal; and The amplification section amplifies the surround signal by an amplification rate based on the second coefficient, thereby generating the second output signal.
11. A method for processing sound signals, The sound signal processing method includes: In the removal step, a first output signal is generated by removing the vocal components based on the audio signals from the first channel and the second channel, as well as a first coefficient indicating the vocal frequency band to be removed. The surround signal processing step involves adding a surround effect to the first output signal to generate a second output signal; The amplification step is performed either before the removal step or between the removal step and the surround signal processing step, or as part of the removal step or the surround signal processing step, to amplify the input signal based on the amplification rate of the second coefficient. The first synthesis step involves combining the second output signal with one of the audio signals from the first channel and the second channel. The second synthesis step involves combining the inverted signal of the second output signal with the audio signal from the first channel and the other side of the audio signal from the second channel. as well as The setup steps involve setting the first coefficient and the second coefficient. In the setting step, the second coefficient is set such that the amplification rate when the vocal frequency band removed according to the first coefficient is a second frequency band with a bandwidth greater than that when it is the first frequency band.
Citation Information
Patent Citations
Sound signal processor and surround reproducing method
JP1997084198A
Signal processing device
CN101223820A
Sound source separator device, sound source separator method, and program
CN103098132A