Signal processing apparatus, signal processing methods and programs
By using sound source separation and bandwidth expansion processing in the signal processing device, the problem of inaccurate bandwidth expansion in multi-source signals is solved, and appropriate bandwidth expansion and accurate processing of high-resolution audio signals are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2020-07-22
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to perform appropriate bandwidth expansion processing on multi-source mixed signals, especially in high-resolution audio signals. They cannot effectively distinguish and appropriately expand the bandwidth of different sound sources, resulting in inaccurate bandwidth expansion processing during repeated mastering.
The signal processing device includes a sound source separation section and a frequency band extension section. The sound source separation process separates the mixed signal into individual sound source signals, and sets corresponding frequency band extension processing parameters for each sound source type. Combined with frequency envelope shaping or phase rotation processing, it prevents high-frequency components from being unnaturally emphasized.
It achieves appropriate bandwidth expansion for multi-source signals, ensuring the naturalness and accuracy of bandwidth expansion processing. It is suitable for creating repeating masters of high-resolution audio signals, avoiding unnecessary editing.
Smart Images

Figure CN114467139B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to signal processing apparatus, signal processing methods, and procedures. Background Technology
[0002] Sound source separation techniques are known, in which a signal for the sound of a target sound source is extracted from a mixed sound signal comprising sounds from multiple sound sources (e.g., see Patent Document 1). Additionally, bandwidth spreading (expansion) techniques have been proposed, in which a high-frequency component is generated from a signal having low-frequency components, and the resulting high-frequency component is added to a signal having low-frequency components to generate a signal with a wider bandwidth (e.g., see Patent Document 2).
[0003] [List of Citations]
[0004] [Patent Literature]
[0005] [Patent Document 1]
[0006] PCT Patent Publication No. WO2018 / 047643
[0007] [Patent Document 2]
[0008] PCT Patent Publication No. WO2015 / 079946 Summary of the Invention
[0009] [Technical Issues]
[0010] In this field, it is desirable to perform appropriate bandwidth spreading processing, etc.
[0011] The purpose of this disclosure is to provide signal processing apparatus, signal processing methods, and programs that perform appropriate bandwidth spreading processing, etc.
[0012] [Solution to the problem]
[0013] For example, this disclosure provides a signal processing apparatus, including: a sound source separation section configured to apply sound source separation processing to a mixed sound signal comprising signals of multiple sound sources; and a bandwidth extension section configured to apply bandwidth extension processing to a corresponding sound source separation signal obtained by separation by the sound source separation section.
[0014] For example, this disclosure provides a signal processing method, including: applying sound source separation processing to a mixed sound signal comprising signals from multiple sound sources through a sound source separation section; and applying bandwidth spreading processing to a corresponding sound source separated signal obtained by the sound source separation section through a bandwidth spreading section.
[0015] For example, this disclosure provides a program for a computer to perform a signal processing method, which includes: applying sound source separation processing to a mixed sound signal comprising signals from multiple sound sources through a sound source separation section; and applying bandwidth spreading processing to a corresponding sound source separated signal obtained by separation through the sound source separation section through a bandwidth spreading section. Attached Figure Description
[0016] Figure 1 This is a block diagram depicting an example configuration of a signal processing apparatus according to the first embodiment.
[0017] Figure 2 This is a diagram referenced when describing the operation of the bandwidth extension portion according to the first embodiment.
[0018] Figure 3 This is a diagram referenced when describing a configuration example of the signal processing apparatus according to the second embodiment.
[0019] Figure 4 This is a diagram for reference when describing the processing performed in the signal processing apparatus according to the second embodiment.
[0020] Figure 5 This is a diagram referenced when describing a modified example of the signal processing apparatus according to the second embodiment.
[0021] Figure 6 This is a diagram referenced when describing a configuration example of a signal processing apparatus according to a third embodiment.
[0022] Figure 7 These are illustrations referenced when describing a modified example of the signal processing apparatus according to the third embodiment.
[0023] Figure 8 These are illustrations referenced when describing a modified example of the signal processing apparatus according to the third embodiment. Detailed Implementation
[0024] Embodiments of this disclosure will now be described with reference to the accompanying drawings. Note that the descriptions will proceed in the following order.
[0025] <Issues to be considered in the embodiments>
[0026] <First Embodiment>
[0027] <Second Embodiment>
[0028] <Third Embodiment>
[0029] <Modification Example>
[0030] The embodiments described below are suitable specific examples of this disclosure, and the content of this disclosure is not limited to the embodiments, etc.
[0031] <Issues to be considered in the embodiments>
[0032] First, to facilitate understanding of this disclosure, problems to be considered in the embodiments will be described. As mentioned above, apparatus for performing band-expansion processing (hereinafter referred to as band-expansion processing) is known. When it is necessary to expand a finite frequency band of a sound source, it is difficult to perform band-expansion processing correctly because the frequency envelope (spectral envelope) varies depending on the type of sound source, such as a musical instrument. For example, cymbals and other percussion instruments, as well as shakuhachi, shamisen, and koto, produce sounds with extremely high frequency components, while musical instruments, such as pianos and violins, have the characteristic of attenuation increasing uniformly with frequency. When the sound sources do not overlap with each other in time, the type of the sound source can be estimated at each time point, and the behavior of the band-expansion processing (the content of the processing) can be changed according to the type. However, for music and the like, multiple types of sound sources typically emit sound simultaneously, making it difficult to perform appropriate band-expansion processing according to the type of sound source.
[0033] Furthermore, in recent years, high-resolution audio (hereinafter referred to as high-resolution sound sources) with sampling rates exceeding 48kHz has become widespread. When producing high-resolution sound sources, some sounds (such as human voices) are recorded as high-resolution sound sources, but the sounds of many instruments can be recorded as standard-resolution audio (hereinafter referred to as standard-resolution sound sources) with sampling rates of 48kHz or less. Therefore, in this case, it is required that the sounds of all instruments be made high-resolution during the repeat mastering step (repeat mastering). At this time, preferably, band-expansion processing is applied only to sound sources that were not recorded at high resolution, without editing sound sources that were recorded at high resolution. However, during the mixing step, the sounds of all sound sources are mixed, which brings the problem that it is not possible to select whether to perform band-expansion processing for each sound source during the repeat mastering step. This disclosure was developed in view of these circumstances. This disclosure will now be described in detail.
[0034] <First Embodiment>
[0035] [Signal processing apparatus according to the first embodiment]
[0036] (Configuration Example)
[0037] Figure 1This is a block diagram (signal processing apparatus 1) illustrating an example configuration of a signal processing apparatus according to a first embodiment. For example, signal processing apparatus 1 includes a sound source separation section 11, a bandwidth extension section 12, and an addition section 13. In this embodiment, a mixed sound signal x is input to the sound source separation section 11. The mixed sound signal x includes a mixture of sound (signals) from multiple (e.g., N (N is a natural number)) sound sources. Signal processing apparatus 1 includes N bandwidth extension sections (bandwidth extension section 121, bandwidth extension section 122, ..., bandwidth extension section 12...) corresponding to the number of sound sources. N Note that, where it is not necessary to distinguish between the various frequency band extensions, the frequency band extensions may be collectively referred to as frequency band extension 12.
[0038] The sound source separation section 11 applies sound source separation processing to the mixed sound signal x to generate sound source separation signals s1, s2, ..., s1 corresponding to the type of the respective sound source. N The sound source separation signal s1 is provided to the bandwidth extension section 121. The sound source separation signal s2 is provided to the bandwidth extension section 122. N Provided to the bandwidth extension section 12 N .
[0039] The sound source separation process performed by the sound source separation section 11 is not limited to a specific process. For example, in addition to sound source separation processing based on MWF (Multi-channel Wiener Filter) using DNN (Deep Natural Network), the sound source separation process described in Patent Document 1 listed above can also be applied. The sound source separation process described in Patent Document 1 is generally a process in which different sound source separation schemes (specifically, DNN and LSTM (Long Short-Term Memory)) with outputs having different temporal attributes are used to estimate the amplitude spectrum, and in which the estimation results are cascaded using predetermined cascading parameters to generate a sound source separated signal. Needless to say, the sound source separation section 11 can perform sound source separation processes different from the sound source separation processes described above.
[0040] The band extension section 12 applies band extension processing to each of the sound source separation signals s obtained by the sound source separation section 11. For example, the band extension section 12 uses the sound source separation signal s corresponding to the low-frequency signal component as the input signal, applies band extension processing to the sound source separation signal s, and outputs the resulting signal as an output signal j (output signal j1, output signal j2, ..., and output signal j...) that includes the low-frequency component and also includes a high-frequency component with an extended band. NOutput. The band extension section 12 applies a known band extension process to the source separation signal s, for example, the band extension process described in Patent Document 2 listed above. Note that each band extension section 12 is associated with a corresponding type of source separation signal s to be input to the corresponding band extension section 12.
[0041] Note that in the following text, the extended start band refers to the lowest frequency end of the frequency components to be extended through band extension processing. High frequency components refer to signals with a frequency band higher than the extended start band, while low frequency components refer to signals with a frequency band lower than the extended start band.
[0042] The addition section 13 will output the output signal j from the bandwidth extension section 12 (specifically, output signal j1, output signal j2, ..., and output signal j...). N The signals are added together to generate a synthesized output signal S, and the synthesized output signal S is output. In this embodiment, it is assumed that the bandwidth-extended sound source signal corresponding to the output of the signal processing device 1 is the synthesized output signal S.
[0043] (General operation example)
[0044] Now, an example of the operation performed by the signal processing device 1 will be described. A mixed audio signal x is input to the source separation section 11. The source separation section 11 applies source separation processing to the mixed audio signal x to generate a source-separated signal s, and outputs the source-separated signal s. The bandwidth expansion section 12 applies bandwidth expansion processing to the source-separated signal s to generate an output signal j, and outputs the output signal j. The addition section 13 adds the output signals j together to generate a synthesized output signal S, and outputs the synthesized output signal S.
[0045] (Operational example of the bandwidth extension section)
[0046] Incidentally, the band-expansion processing described in Patent Document 2 listed above is based on mixed sounds and does not consider performing optimal band-expansion processing depending on the properties of the sound source, particularly the type of the sound source. For example, cymbals, as percussion instruments, involve an envelope that extends to high frequencies without attenuation. Therefore, in this embodiment, in order to perform optimal band-expansion processing for each type of sound source, a frequency envelope of the high-frequency component (high-frequency band) to be estimated is set for each type of sound source. Specifically, parameters for band-expansion processing corresponding to the type of sound source are set, and band-expansion processing is performed using these parameters. A device for estimating the high-frequency band can be applied as the band-expansion part, allowing the device to learn only the type of sound source (e.g., cymbal sound) as training data.
[0047] Figure 2 An example of the frequency envelope corresponding to the type of sound source is depicted. Figure 2In the diagram, the horizontal axis represents frequency (Hz), and the vertical axis represents sound pressure level (dB). Additionally, in... Figure 2 In this context, f1 indicates the start of the extended frequency band. Furthermore, in... Figure 2 In the diagram, the frequency envelope FE1 following the extended start frequency band f1 schematically represents the frequency envelope of a sound source, such as a human voice, and the frequency envelope FE2 following the extended start frequency band f1 schematically represents the frequency envelope of a sound source, such as a cymbal. For the frequency band extension section 12 corresponding to the human voice, parameters for generating the frequency envelope FE1 are set. Furthermore, for the frequency band extension section 12 corresponding to the cymbal, parameters for generating the frequency envelope FE2 are set. This allows each frequency band extension section 12 to perform appropriate frequency band extension processing corresponding to the properties of the sound source input to the frequency band extension section 12. Note that the parameters are appropriately set according to the content of the frequency band extension processing.
[0048] <Second Embodiment>
[0049] The second embodiment of this disclosure will now be described. Note that, unless otherwise stated, the matters described in the first embodiment may also be applied to the second embodiment. Furthermore, components that are identical or equivalent to those in the first embodiment are denoted by the same reference numerals, and repeated descriptions are appropriately omitted.
[0050] [Overview of the Second Embodiment]
[0051] When band-spreading processing is performed independently on each source-separated signal, the high-frequency components of the synthesized output signal S may be unnaturally emphasized depending on the band-spreading algorithm. For example, if the algorithm used for band-spreading only estimates the amplitude spectrum or the envelope of the amplitude spectrum and replicates the phase in a specific way (e.g., using the same phase as the low-frequency component (low-frequency band), and if the source-separation algorithm also involves insignificant phase changes for each separated source, the high-frequency signals of the source-separated signals with extended bands will all have similar phases. Therefore, even if the amplitude spectrum or the envelope of the amplitude spectrum of each source-separated signal is correctly estimated, the high-frequency components of the synthesized output signal S may be unnaturally emphasized because all high-frequency signals have similar phases. This embodiment is a signal processing apparatus configured to solve the above-mentioned problems.
[0052] [Signal processing apparatus according to the second embodiment]
[0053] (Configuration Example)
[0054] Figure 3This is a block diagram depicting a configuration example of the signal processing apparatus (signal processing apparatus 2) according to the second embodiment. The signal processing apparatus 2 differs from the signal processing apparatus 1 in that it includes a frequency envelope shaping section 21 following the addition section 13. In this embodiment, it is assumed that the output of the frequency envelope shaping section 21 is a band-extended sound source signal.
[0055] The frequency envelope shaping section 21 shapes the frequency envelope of the synthesized output signal S output from the adder 13. For example, if a predetermined discontinuity is detected between the portion of the frequency envelope before the extended start band (the lower limit of the frequency extended by the band extension process) f1 and the portion of the frequency envelope after the extended start band f1, the frequency envelope of the synthesized output signal S is shaped. In this embodiment, the predetermined discontinuity is detected by the frequency envelope shaping section 21. However, this detection can be performed by another functional block. When the frequency envelope shaping section 21 shapes the frequency envelope, it suppresses the amplitude of the extended high-frequency components, allowing the high-frequency components to be prevented from being unnaturally emphasized.
[0056] (Operation Example)
[0057] In this embodiment, discontinuity is detected when the difference between the signal energy before and after the extended start frequency band f1 is equal to or greater than a predetermined value. (Refer to...) Figure 4 Describe a specific example.
[0058] exist Figure 4 In the diagram, the horizontal axis represents frequency (Hz), and the vertical axis represents sound pressure level (dB). Furthermore, in... Figure 4 In this context, f1 indicates the start of the extended frequency band. Furthermore, in... Figure 4 In the diagram, the frequency envelopes following the extended starting frequency band f1 (frequency envelopes FE3 to FE6) show an example of the frequency envelopes of the high-frequency components of the synthesized output signal S.
[0059] For example, such as Figure 4 The description defines predetermined frequency bands (f1-Δf) and (f1+Δf) for portions of the frequency envelope before and after the extended starting frequency band f1, respectively, and determines the energy e of each frequency band for each frequency envelope. Figure 4 (The shaded area in the image). Under the condition that the following Formula 1 is satisfied, a discontinuity is determined between the portions of the frequency envelope before and after the extended start frequency band f1, where e L e represents energy in the low-frequency band. H Th represents the energy in the high-frequency band, and Th represents the threshold used to detect discontinuities.
[0060] (e H / e L )>Th...(1)
[0061] exist Figure 4 In the example shown, when the high-frequency component of the synthesized output signal S forms a frequency envelope FE3, Equation 1 is satisfied, leading to the detection of discontinuities. The frequency envelope FE3 causes the high-frequency component to be unnaturally emphasized, therefore the frequency envelope shaping section 21 performs processing for shaping the frequency envelope, specifically processing for suppressing the amplitude of the high-frequency component. In the amplitude suppression processing, the amplitude of the high-frequency component can be suppressed uniformly, or amplitudes greater than a predetermined threshold can be specifically suppressed.
[0062] On the other hand, Figure 4 In the example shown, if the high-frequency components of the synthesized output signal S form one of the frequency envelopes FE4 to FE6, Equation 1 is not satisfied, leading to the determination that no discontinuity exists. In this case, the high-frequency components are unlikely to be unnaturally emphasized, so the frequency envelope shaping section 21 does not perform processing, wherein the synthesized output signal S is output from the frequency envelope shaping section 21.
[0063] According to the second embodiment described above, when performing bandwidth extension processing, it is possible to prevent high-frequency components after the start of bandwidth extension from being unnaturally emphasized.
[0064] (Modified Example)
[0065] Now, a modified example of the signal processing apparatus according to the second embodiment will be described. Figure 5 This is a block diagram (signal processing device 2A) depicting a configuration example of a signal processing device based on a modified example.
[0066] The signal processing apparatus 2A does not include a frequency envelope shaping section 21, but includes a phase rotation section 22. The phase rotation section 22 is disposed between the bandwidth extension section 12 and the addition section 13. Specifically, the signal processing apparatus 2A includes phase rotation sections 22 (phase rotation sections 221, 222, ..., 22...). N The number of these corresponds to the number of bandwidth extension sections 12. The output signals from the phase rotation section 22 are added together by the addition section 13.
[0067] Phase rotation section 22 uses the frequency band extended by frequency band extension section 12 to rotate (change) the phase of the high-frequency component of output signal j, such that the high-frequency component of output signal j has a different phase depending on the sound source. For example, each of phase rotation sections 22 includes a filter capable of shifting the phase without affecting the amplitude, specifically, an all-pass filter.
[0068] For example, the phase rotation section 22 randomly rotates the phase, thereby preventing the high-frequency components of the bandwidth-extended sound source signal from being unnaturally emphasized. Furthermore, human hearing is not sensitive to high-frequency phase changes, thus preventing the unnatural emphasis of high-frequency components of the bandwidth-extended sound source signal without causing auditory discomfort to the user.
[0069] <Third Embodiment>
[0070] The third embodiment of this disclosure will now be described. Note that, unless otherwise stated, the matters described in the first and second embodiments may also be applied to the third embodiment. Furthermore, components that are identical or equivalent to those in the first and second embodiments are indicated by the same reference numerals, and repeated descriptions are appropriately omitted.
[0071] [Overview of the Third Embodiment]
[0072] As described above, in a sound source comprising a high-resolution sound source (e.g., a sound source containing high-frequency components following the extended start band f1) and a standard-resolution sound source (e.g., a sound source not containing high-frequency components following the extended start band f1) (hereinafter referred to as a hybrid sound source), it is required that band-spreading processing be applied only to the standard-resolution sound source. This embodiment addresses this requirement. Note that the band of the hybrid sound source includes high frequencies following the extended start band f1.
[0073] [Signal processing apparatus according to the third embodiment]
[0074] (Configuration Example)
[0075] Figure 6 This is a block diagram illustrating a configuration example of a signal processing apparatus (signal processing apparatus 3) according to a third embodiment. Similar to signal processing apparatus 1, signal processing apparatus 3 includes a sound source separation section 11, a bandwidth extension section 12 (e.g., bandwidth extension sections 121 and 122), and an adder section 13. A signal from a mixed sound source (hereinafter referred to as mixed sound source signal x1) is input to the sound source separation section 11. The difference between signal processing apparatus 3 and signal processing apparatus 1 is that signal processing apparatus 3 includes a system for inputting the mixed sound source signal x1 to the adder section 13 and the sound source separation section 11.
[0076] (Operation Example)
[0077] Now, an operational example of the signal processing apparatus 3 will be described. The mixed sound source signal x1 is separated into signals of corresponding sound source types by the sound source separation section 11, thereby generating sound source separation signals s. Of the sound source separation signals s of corresponding sound source types, only the sound source separation signals not recorded at high resolution (sound source separation signals s1 and s2 in this example) are provided to the corresponding band extension sections 121 and 122, respectively. The band extension section 121 performs band extension processing to extend the band of the sound source separation signal s1. Furthermore, the band extension section 122 performs band extension processing to extend the band of the sound source separation signal s2.
[0078] For the output signal obtained through the applied bandwidth extension processing, the bandwidth extension section 121 outputs an extended bandwidth signal p1 to the addition section 13, which includes the high-frequency components after the extended start bandwidth f1 in the output signal. Furthermore, for the output signal obtained through the applied bandwidth extension processing, the bandwidth extension section 122 outputs an extended bandwidth signal p2 to the addition section 13, which includes the high-frequency components after the extended start bandwidth f1 in the output signal. In this respect, the bandwidth extension sections 121 and 122 only output extended bandwidth signals to the addition section 13 because the low-frequency components of the sound source separation signals s1 and s2 are included in the mixed sound source signal x1 input to the addition section 13.
[0079] The addition section 13 adds the extended frequency band signals p1 and p2 and the mixed sound source signal x1 together to generate the extended frequency band sound source signal and outputs the extended frequency band sound source signal.
[0080] According to the third embodiment described above, when the high-frequency components of the sound source signal recorded at high resolution remain unchanged, the bandwidth of the sound source signal not recorded at high resolution can be specifically extended. Note that in the above description, the sound source separation signals s1 and s2 are shown as sound source separation signals not recorded at high resolution, but the mixed sound source signal x1 may include more sound source separation signals not recorded at high resolution.
[0081] (Modified Example 1)
[0082] Figure 7 This is a block diagram illustrating a modified example of the signal processing apparatus according to the third embodiment. The above example assumes that the sound source separation section 11 of the signal processing apparatus 3 has the ability to separate sound sources including high-resolution sound sources. However, it is also assumed that the sound source separation section 11 lacks the ability to separate sound sources including high-resolution sound sources.
[0083] In this case, such as Figure 7As shown, the sound source separation section 11 of the signal processing apparatus (signal processing apparatus 3A) according to this modified example includes a downconverter 11A that applies downsampling processing to the mixed sound source signal x1. Performing downsampling on the downconverter 11A enables the sound source separation section 11 to perform sound source separation on the mixed sound source signal x1. In this configuration, for example, the bandwidth extension section 121 includes an upconverter 12. A1 Furthermore, bandwidth spreading is performed after upsampling. Similarly, the bandwidth spreading section 122 includes an up-converter 12. A2 And after upsampling, bandwidth spreading is performed. Upconverter 12 A1 and 12 A2 The processing can be performed in the corresponding preamplifiers of the band extension sections 121 and 122.
[0084] (Modified Example 2)
[0085] Figure 8 This is a block diagram illustrating another modified example of the signal processing apparatus according to the third embodiment. The sound source separation section 11 of the signal processing apparatus (signal processing apparatus 3B) according to this modified example includes a determining section 11B. Note that this example assumes that the sound source separation section 11 of the signal processing apparatus 3B has the ability to separate sound sources including high-resolution sound sources.
[0086] In the signal processing apparatus 3B, the mixed sound source signal x1 is provided only to the sound source separation section 11, and not to the addition section 13. The sound source separation section 11 performs sound source separation processing on the mixed sound source signal x1 to generate sound source separation signals s1, s2, and hm corresponding to the sound source signals recorded at high resolution. The determination section 11B determines whether to apply band-spreading processing to each sound source separation signal in a subsequent stage. If the sound source separation signal contains high-frequency components, the determination section 11B determines that band-spreading processing is not required for the sound source separation signal and outputs the sound source separation signal to the addition section 13. In this modified example, the determination section 11B determines that band-spreading processing is not required for the sound source separation signal hm, and the sound source separation section 11 provides the sound source separation signal hm to the addition section 13.
[0087] Furthermore, if the sound source separation signal does not contain high-frequency components, the determination section 11B determines that bandwidth spreading processing needs to be applied to the sound source separation signal and outputs the sound source separation signal to the bandwidth spreading section 12. In this modified example, the determination section 11B determines that bandwidth spreading processing needs to be applied to the sound source separation signals s1 and s2, and the sound source separation signals s1 and s2 are provided to the bandwidth spreading sections 121 and 122, respectively.
[0088] The bandwidth extension section 121 applies bandwidth extension processing to the sound source separation signal s1 to generate the output signal j1. In the configuration according to the signal processing apparatus 3B, the mixed sound source signal x1 is not provided to the addition section 13; therefore, the bandwidth extension section 121 outputs the output signal j1 containing low-frequency components to the addition section 13, instead of an extended bandwidth signal. Furthermore, the bandwidth extension section 122 applies bandwidth extension processing to the sound source separation signal s2 to generate the output signal j2. In the configuration according to the signal processing apparatus 3B, the mixed sound source signal x1 is not provided to the addition section 13; therefore, the bandwidth extension section 122 outputs the output signal j2 containing low-frequency components to the addition section 13, instead of an extended bandwidth signal. The addition section 13 adds the sound source separation signal hm, the output signal j1, and the output signal j2 together.
[0089] According to the signal processing apparatus 3B of this modified example, an effect similar to that obtained based on the configuration of the signal processing apparatus 3 described above can be produced. Furthermore, according to the signal processing apparatus 3B of this modified example, it automatically determines whether to apply bandwidth spreading processing; therefore, for example, it eliminates the need for the user to know in advance which of the sound source separation signals to which bandwidth spreading processing will be applied and to select whether to apply bandwidth spreading processing during the master copy creation step.
[0090] <Modification Example>
[0091] Several embodiments of this disclosure have been described. However, this disclosure is not limited to the above embodiments, and various modifications can be made to the embodiments without departing from the scope of this disclosure.
[0092] In the above embodiments, the type of sound source is used as an attribute of the sound source. However, another attribute, such as the signaling characteristics of the sound source, can be used.
[0093] When using DNN or LSTM as the sound source separation component, the network input is typically considered to be the amplitude spectrum of the mixed sound signal, and the training data is considered to be the amplitude spectrum of the target sound source. However, the sound source separation signal obtained through sound source separation can be used as training data in the learning process.
[0094] This disclosure can also employ a cloud computing configuration, in which multiple devices perform the processing of a function in a shared and collaborative manner via a network.
[0095] This disclosure can also be implemented in any form, such as an apparatus, method, program, or system. For example, by providing a downloadable program that performs the functions described in the above embodiments, and downloading and installing the program onto an apparatus that does not have the functions described in the above embodiments, the controls described in the embodiments can be performed in that apparatus. This disclosure can also be implemented by a server that distributes such a program. Furthermore, the embodiments and modifications described in the examples can be suitably combined. Moreover, the effects shown herein are not intended to limit the scope of this disclosure.
[0096] This disclosure can be configured as follows. (1)
[0098] A signal processing apparatus, comprising:
[0099] The sound source separation section is configured to apply sound source separation processing to a mixed sound signal comprising signals from multiple sound sources; and
[0100] The bandwidth extension section is configured to apply bandwidth extension processing to the corresponding sound source separation signals obtained by the sound source separation section. (2)
[0102] According to the signal processing apparatus described in (1), wherein,
[0103] Each of the bandwidth extension components applies bandwidth extension processing corresponding to the properties of the sound source separation signal. (3)
[0105] The signal processing apparatus according to (1) or (2) includes:
[0106] The addition section is configured to sum the corresponding outputs of the bandwidth extension provided for each sound source separation signal; and
[0107] The frequency envelope shaping section is configured to shape the frequency envelope of the synthesized output signal to be output from the addition section. (4)
[0109] According to the signal processing apparatus described in (3), wherein,
[0110] Assuming f1 is the lower limit of the frequency extended by the bandwidth spreading process, the frequency envelope shaping part shapes the frequency envelope of the synthesized output signal when a predetermined discontinuity is detected between the part of the frequency envelope before f1 and the part of the frequency envelope after f1. (5)
[0112] According to the signal processing apparatus described in (4), wherein,
[0113] The existence of discontinuity is detected when the signal energy difference between the frequency envelope portion before f1 and the frequency envelope portion after f1 is equal to or greater than a predetermined value. (6)
[0115] The signal processing apparatus according to (1) or (2) includes:
[0116] The phase rotation section is configured to apply processing to rotate the phase of the output signal from the bandwidth extension section. (7)
[0118] According to the signal processing apparatus described in (6), wherein,
[0119] The phase rotation section includes an all-pass filter. (8)
[0121] According to the signal processing apparatus described in (1), wherein,
[0122] The bandwidth extension section only outputs the extended bandwidth signal, which is a signal with a bandwidth extended through bandwidth extension processing. (9)
[0124] The signal processing apparatus according to (8) includes:
[0125] A downconverter is configured to apply downsampling processing to a mixed audio signal, the mixed audio signal including a signal from a sound source containing high-frequency components above a predetermined frequency; and
[0126] The addition section is configured to add the mixed audio signal and the extended frequency band signal together, wherein,
[0127] The sound source separation section applies sound source separation processing to the signal that has already undergone downsampling processing. (10)
[0129] The signal processing apparatus according to (1) includes:
[0130] The addition part is configured to add the source separation signal that has been band-spread processed and the source separation signal that has not been band-spread processed together. (11)
[0132] The signal processing apparatus according to (10) includes:
[0133] The determination section is configured to determine whether to apply bandwidth spreading processing to the sound source separation signal. (12)
[0135] According to the signal processing apparatus described in (11), wherein,
[0136] The determination section determines not to apply bandwidth spreading processing to the sound source separation signal if the sound source separation signal contains high-frequency components equal to or greater than a predetermined frequency, and determines to apply bandwidth spreading processing to the sound source separation signal if the sound source separation signal does not contain high-frequency components equal to or greater than a predetermined frequency. (13)
[0138] A signal processing method, comprising:
[0139] The sound source separation section applies sound source separation processing to a mixed sound signal comprising signals from multiple sound sources; and
[0140] The bandwidth extension section applies bandwidth extension processing to the corresponding sound source separation signals obtained by the sound source separation section. (14)
[0142] A program that enables a computer to perform signal processing methods includes:
[0143] The sound source separation section applies sound source separation processing to a mixed sound signal comprising signals from multiple sound sources; and
[0144] The bandwidth extension section applies bandwidth extension processing to the corresponding sound source separation signals obtained by the sound source separation section.
[0145] Reference Symbol List
[0146] 1,2,2A,3,3A,3B: Signal processing device
[0147] 11: Sound source separation section
[0148] 11A: Downconverter
[0149] 12: Bandwidth extension section
[0150] 13: Addition Part
[0151] 21: Frequency envelope shaping section
[0152] 22: Phase Rotation Part
Claims
1. A signal processing apparatus, comprising: A sound source separation section is configured to apply sound source separation processing to a mixed audio signal comprising signals from multiple sound sources. The sound source separation section includes a down-converter that applies downsampling processing to the mixed audio signal. The down-converter performs downsampling such that the sound source separation section can perform the sound source separation processing on the mixed audio signal. A band-expansion section is configured to apply band-expansion processing to the corresponding sound source separation signal obtained by the sound source separation section, wherein the band-expansion section includes an up-converter that performs upsampling, the band-expansion section performs the band-expansion processing after the up-converter performs upsampling, the processing of the up-converter is performed in the pre-stage of the band-expansion section, and the band-expansion section outputs only an extended band signal, the extended band signal being a signal having a band expanded by the band-expansion processing.
2. The signal processing apparatus according to claim 1, wherein, The bandwidth extension portion applies bandwidth extension processing corresponding to the properties of the sound source separation signal.
3. The signal processing apparatus according to claim 1, comprising: The addition section is configured to sum the corresponding outputs of the bandwidth extension provided for each sound source separation signal; as well as The frequency envelope shaping section is configured to shape the frequency envelope of the synthesized output signal to be output from the addition section.
4. The signal processing apparatus according to claim 3, wherein, Assuming f1 is the lower limit of the frequency extended by the bandwidth extension process, the frequency envelope shaping section shapes the frequency envelope of the synthesized output signal when a predetermined discontinuity is detected between the portion of the frequency envelope before f1 and the portion of the frequency envelope after f1.
5. The signal processing apparatus according to claim 4, wherein, The existence of discontinuity is detected when the signal energy difference between the frequency envelope portion before f1 and the frequency envelope portion after f1 is equal to or greater than a predetermined value.
6. The signal processing apparatus according to claim 1, comprising: The phase rotation section is configured to apply processing for rotating the phase of the output signal from the bandwidth extension section.
7. The signal processing apparatus according to claim 6, wherein, The phase rotation section includes an all-pass filter.
8. The signal processing apparatus according to claim 1, comprising: The mixed sound signal includes a signal from a sound source containing a high-frequency component above a predetermined frequency; as well as The addition section is configured to add the mixed audio signal and the extended frequency band signal together, wherein, The sound source separation section applies the sound source separation processing to the signal that has already undergone the downsampling processing.
9. The signal processing apparatus according to claim 1, comprising: The addition portion is configured to add together the source separation signal that has been processed with the bandwidth extension and the source separation signal that has not been processed with the bandwidth extension.
10. The signal processing apparatus according to claim 9, comprising: The determining component is configured to determine whether to apply the bandwidth spreading process to the sound source separation signal.
11. The signal processing apparatus according to claim 10, wherein, The determining component determines not to apply the bandwidth spreading process to the sound source separation signal if the sound source separation signal contains a high-frequency component equal to or greater than a predetermined frequency, and determines to apply the bandwidth spreading process to the sound source separation signal if the sound source separation signal does not contain a high-frequency component equal to or greater than a predetermined frequency.
12. A signal processing method, comprising: A source separation section applies source separation processing to a mixed audio signal comprising signals from multiple sound sources. The source separation section includes a down-converter that applies downsampling processing to the mixed audio signal. The down-converter performs downsampling so that the source separation section can perform the source separation processing on the mixed audio signal. The band-expansion section applies band-expansion processing to the corresponding sound source separation signal obtained by the sound source separation section. The band-expansion section includes an up-converter that performs upsampling. The band-expansion section performs the band-expansion processing after the up-converter performs upsampling. The processing of the up-converter is performed in the pre-stage of the band-expansion section, and the band-expansion section outputs only the extended band signal, which is a signal having a band expanded by the band-expansion processing.
13. A computer-readable recording medium storing a program for causing a computer to perform a signal processing method, the signal processing method comprising: A source separation section applies source separation processing to a mixed audio signal comprising signals from multiple sound sources. The source separation section includes a down-converter that applies downsampling processing to the mixed audio signal. The down-converter performs downsampling so that the source separation section can perform the source separation processing on the mixed audio signal. The band-expansion section applies band-expansion processing to the corresponding sound source separation signal obtained by the sound source separation section. The band-expansion section includes an up-converter that performs upsampling. The band-expansion section performs the band-expansion processing after the up-converter performs upsampling. The processing of the up-converter is performed in the pre-stage of the band-expansion section, and the band-expansion section outputs only the extended band signal, which is a signal having a band expanded by the band-expansion processing.