Signal processing device, signal processing method and program

The signal processing device addresses high computational and memory costs in sound source separation by applying downsampling, mask generation, and band extension, ensuring effective separation and reduced noise for high-resolution audio sources.

JP7790351B2Active Publication Date: 2025-12-23SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022560683
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-09
Filing Date
2021-10-07
Publication Date
2025-12-23
Estimated Expiration
2041-10-07

AI Technical Summary

Technical Problem

Existing sound source separation technologies struggle with high computational and memory costs when processing high-frequency components in mixed sound signals, leading to reduced performance and difficulty in training effective models for high-resolution audio sources.

Method used

A signal processing device that applies downsampling, mask generation, sound source separation, and band extension to high-frequency mixed sound signals, using a mask processing unit and band extension unit to maintain high-frequency components and reduce computational costs.

Benefits of technology

The solution enables effective sound source separation for high-resolution audio sources, maintaining high-frequency components and reducing computational and memory costs, while ensuring the sum of separation results matches the original sound, minimizing perceptible noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790351000023
    Figure 0007790351000023
  • Figure 0007790351000024
    Figure 0007790351000024
  • Figure 0007790351000025
    Figure 0007790351000025
Patent Text Reader

Abstract

The present invention provides a signal processing device which carries out an appropriate sound source separation process, for example. This signal processing device has: a down converter which applies a down sampling process to a mixed sound signal in which a sound source signal, that contains a harmonic component higher than a prescribed frequency, has been mixed; a mask generation unit which generates a mask on the basis of the down sampling process results by the down converter; and a mask processing unit which applies the mask generated by the mask generation unit to the mixed sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a signal processing device, a signal processing method, and a program. [Background technology]

[0002] A sound source separation technique is known that extracts a signal of a sound from a target sound source (hereinafter referred to as a sound source signal) from a mixed sound signal containing sounds from multiple sound sources (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2018 / 047643 Summary of the Invention [Problem to be solved by the invention]

[0004] In this field, it is desirable to perform effective sound source separation processing on mixed sound signals that contain high-frequency components higher than a predetermined frequency.

[0005] An object of the present disclosure is to provide a signal processing device, a signal processing method, and a program that perform effective sound source separation processing on a mixed sound signal that includes high-frequency components higher than a predetermined frequency. [Means for solving the problem]

[0006] The present disclosure provides, for example, a down-converter that applies down-sampling processing to a mixed sound signal that is mixed with a sound source signal that includes high-frequency components higher than a predetermined frequency; a mask generation unit that generates a mask based on a downsampling processing result by the downconverter; a mask processing unit that applies the mask generated by the mask generation unit to the mixed sound signal; With death, The mask generator is a sound source separation unit that performs sound source separation processing on the mixed sound signal to which downsampling processing has been applied; a band extension unit that applies frequency band extension processing to each of the sound source signals separated by the sound source separation unit; a mask generation processing unit that generates a mask corresponding to each sound source signal based on a relative ratio of the sound source signal to a sum of the sound source signals, based on at least each sound source signal to which frequency band extension processing has been applied; have It is a signal processing device.

[0007] The present disclosure provides, for example, A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on the downsampling processing result by the downconverter; A mask processing unit applies the generated mask to the mixed sound signal. death, a sound source separation unit included in the mask generation unit performs a sound source separation process on the mixed sound signal to which the downsampling process has been applied; a band extension unit included in the mask generation unit applies frequency band extension processing to each sound source signal separated by the sound source separation processing; A mask generation processing unit included in the mask generation unit generates a mask corresponding to each sound source signal based on the relative ratio of the sound source signal to the sum of the sound source signals, based on at least the individual sound source signals to which the frequency band extension processing has been applied. It is a signal processing method.

[0008] The present disclosure provides, for example, A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on the downsampling processing result by the downconverter; A mask processing unit applies the generated mask to the mixed sound signal. death, a sound source separation unit included in the mask generation unit performs a sound source separation process on the mixed sound signal to which the downsampling process has been applied; a band extension unit included in the mask generation unit applies frequency band extension processing to each sound source signal separated by the sound source separation processing; A mask generation processing unit included in the mask generation unit generates a mask corresponding to each sound source signal based on the relative ratio of the sound source signal to the sum of the sound source signals, based on at least the individual sound source signals to which the frequency band extension processing has been applied. It is a program that causes a computer to execute a signal processing method. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a signal processing device according to the first embodiment. [Figure 2]FIG. 2 is a block diagram showing a detailed configuration example of the mask generation unit according to the first embodiment. [Figure 3] FIG. 3 is a flowchart to be referred to when explaining an example of the operation of the signal processing device according to the first embodiment. [Figure 4] FIG. 4 is a block diagram showing an example of the configuration of a signal processing device according to the second embodiment. [Figure 5] FIG. 5 is a block diagram showing a detailed configuration example of a sound source separation unit according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The description will be made in the following order. <Issues to be considered in the implementation> First Embodiment <Second embodiment> <Modification> The embodiments and the like described below are preferred specific examples of the present disclosure, and the contents of the present disclosure are not limited to these embodiments and the like.

[0011] <Issues to be considered in the implementation> First, to facilitate understanding of the present disclosure, issues to be considered in the embodiments will be described.

[0012] The sampling frequency of band-limited audio signals such as telephone signals is generally around 8 kHz, but signals such as music that require high sound quality use sampling rates of 44.1 kHz or 48 kHz. In recent years, high-resolution audio (hereinafter referred to as "high-resolution audio sources") has become popular in pursuit of even higher sound quality, and sampling frequencies have increased to 88.2 kHz to 192 kHz. In other words, mixed sound signals containing high-frequency components higher than a specified frequency (for example, 48 kHz) have come to be used.

[0013] In addition, a technology called sound source separation, which separates individual sound source signals from a mixed sound signal containing various sound source signals, is used in karaoke, sound source remixing, etc. Generally, the memory and computational costs required for sound source separation are proportional to the square of the sampling frequency. In many situations, including embedded systems and cloud services, there is a strong demand to reduce memory and computational costs, while there is also a demand for high-quality sound source separation, which are conflicting requirements.

[0014] In particular, when separating high-resolution audio sources, there is a problem in that the input data becomes too high-dimensional to train a learning model for performing sound source separation using general hardware. Reducing the model size of the learning model to a level that allows learning using general hardware also reduces the performance of the learning model, resulting in a significant decline in sound source separation performance. This is undesirable because the separation results of high-resolution audio sources, which should have higher sound quality than non-high-resolution audio sources (sound sources in normal bands that do not contain high-frequency components higher than a predetermined frequency, hereinafter referred to as non-high-resolution audio sources), become worse than those of non-high-resolution audio sources. Furthermore, even if training-capable hardware is available, high-resolution stem audio sources (individual audio sources before mixing) required for learning sound source separation are difficult to obtain, making it difficult to train a high-resolution audio source separation model in the first place. Furthermore, even if training is possible, the computational cost is too high, making this undesirable. Taking the above points into consideration, the present disclosure will be described in detail using embodiments.

[0015] First Embodiment [Signal processing device according to the first embodiment] (Configuration example) 1 is a block diagram showing an example of the configuration of a signal processing device (signal processing device 1) according to the first embodiment. The signal processing device 1 includes, for example, a mixed sound signal input unit 11, a downconverter 12, a mask generation unit 13, a mask processing unit 14, and a separated sound source signal output unit 15.

[0016] The mixed sound signal input unit 11 is an interface to which a mixed sound signal obtained by mixing multiple sound source signals is input. The multiple sound source signals are sound source signals that include high-frequency components higher than a predetermined frequency. The predetermined frequency is, for example, 48 kHz, but may be another frequency (e.g., 96 kHz). In this way, the mixed sound signal according to this embodiment is a high-resolution sound source. Examples of the mixed sound signal input unit 11 include a drive device that reads the mixed sound signal from a medium (semiconductor memory, magnetic memory, optical memory, etc.) on which the mixed sound signal is recorded, and a communication unit that acquires the mixed sound signal via a network. The mixed sound signal x(h) input to the mixed sound signal input unit 11 is branched and supplied to the downconverter 12 and the masking processing unit 14, respectively. In the following description, the mixed sound signal x(h) will be described as an example of a sound source signal, being a signal obtained by mixing the sound source signals of vocals, drums, and bass.

[0017] The downconverter 12 applies downsampling processing to the mixed sound signal x(h). The downsampling by the downconverter 12 generates a mixed sound signal x(n), which is a mixed sound signal of a non-Hi-Res sound source. The mixed sound signal x(n) is supplied to the mask generation unit 13.

[0018] The mask generation unit 13 generates masks based on the downsampling processing result by the downconverter 12. For example, a mask is generated corresponding to each sound source signal included in the mixed sound signal x(n) and for separating each sound source signal. In this embodiment, the mask generation unit 13 generates a mask MA1 corresponding to vocals, a mask MA2 corresponding to drums, and a mask MA3 corresponding to bass. The masks generated by the mask generation unit 13 are supplied to the mask processing unit 14. A detailed configuration example of the mask generation unit 13 will be described later.

[0019] The mask processing unit 14 applies the mask generated by the mask generation unit 13 to the mixed sound signal x(h). As a result, each sound source signal is separated from the mixed sound signal x(h). For example, when the mask processing unit 14 applies a mask MA1 to the mixed sound signal x(h), a vocal sound source signal s'1(h) is separated from the mixed sound signal x(h). When the mask processing unit 14 applies a mask MA2 to the mixed sound signal x(h), a drum sound source signal s'2(h) is separated from the mixed sound signal x(h). When the mask processing unit 14 applies a mask MA3 to the mixed sound signal x(h), a bass sound source signal s'3(h) is separated from the mixed sound signal x(h).

[0020] The mask processing unit 14 is configured by a filter in which the input of the mask processing unit 14 (in this embodiment, the mixed sound signal x(h)) and the sum of the output (the sum of s'1(h), s'2(h), and s'3(h)) match. An example of such a filter is a Wiener filter. In this case, the mask may also be referred to as a Wiener filter gain, etc.

[0021] The separated sound source signal output unit 15 is an interface that outputs the sound source signals s'1(h), s'2(h), and s'3(h) separated by the mask processing unit 14. The signals output from the separated sound source signal output unit 15 are used according to the application, for example, as object sound sources for remixing (changing the volume, localization, and tone) or generating multi-channel audio.

[0022] Next, a detailed configuration example of the mask generation unit 13 will be described with reference to Fig. 2. The mask generation unit 13 has, for example, a sound source separation unit 131, a band extension unit 132, and a mask generation processing unit 133. The mixed sound signal x(n), which is a non-Hi-Res sound source output from the downconverter 12, is output to both the sound source separation unit 131 and the mask generation processing unit 133.

[0023] The sound source separation unit 131 performs sound source separation processing on the mixed sound signal x(n) that has been downsampled by the downconverter 12. The sound source separation processing is not limited to any particular processing, but for example, the sound source separation processing described in Patent Document 1 can be applied. When sound source separation is realized using a neural network (NN), the mixed sound signal x and its constituent sound source signals si can be used as training data to train the sound source separator f(). Training can be performed using a stochastic gradient method or the like to minimize the error between the separation result f(x, θ) and the correct sound source signal si. The processing by the sound source separation unit 131 obtains a vocal sound source signal s1(n), which is a non-Hi-Res sound source, a drum sound source signal s2(n), which is a non-Hi-Res sound source, and a bass sound source signal s3(n), which is a non-Hi-Res sound source. These sound source signals are supplied to the band expansion unit 132.

[0024] The band extension unit 132 applies frequency band extension processing to the individual sound source signals separated by the sound source separation unit 131, and adds high-frequency components to each sound source signal. The frequency band extension performed by the band extension unit 132 is not limited to a specific processing, but for example, the processing described in Japanese Patent No. 6425097 proposed by the present applicant can be applied. The frequency band extension processing by the band extension unit 132 obtains a band-extended vocal sound source signal s1(h), a band-extended drum sound source signal s2(h), and a band-extended vocal sound source signal s3(h). The obtained sound source signal s1(h), sound source signal s2(h), and sound source signal s3(h) are supplied to the mask generation processing unit 133.

[0025] The sound source signals s1(h), s2(h), and s3(h) obtained at this stage are separated signals with the same bandwidth as the high-resolution sound source, but because the bandwidth extension process is performed on each sound source signal individually, the sum of the band-extended sound source signals does not match the input mixed sound signal x(h). Furthermore, the high-frequency components contained in the mixed sound signal x(h), which is the input high-resolution sound source, are completely ignored, resulting in a fabricated signal.

[0026] The mask generation processing unit 133 generates a mask corresponding to each sound source signal based on each sound source signal to which at least the frequency band extension processing has been applied. In this embodiment, the mask generation processing unit 133 generates a mask corresponding to each sound source signal based on the sound source signal s1(h), the sound source signal s2(h), and the sound source signal s3(h). For example, the mask generation processing unit 133 generates a mask MA1 based on the relative ratio of the sound source signal s1(h) to the sum of the sound source signals. Masks MA2 and MA3 are also generated in a similar manner. The masks MA1, MA2, and MA3 generated by the mask generation processing unit 133 are used in the mask processing unit 14. As described above, the processing by the mask processing unit 14 separates the sound source signals s'1(h), s'2(h), and s'3(h) from the mixed sound signal x(h).

[0027] [Processing flow] Next, an example of the operation of the signal processing device 1 according to this embodiment will be described with reference to the flowchart of FIG.

[0028] When the process starts, in step ST11, a mixed sound signal that is a high-resolution sound source is input. For example, a mixed sound signal x(h) that is a high-resolution sound source is input to the mixed sound signal input unit 11. Then, the process proceeds to step ST12.

[0029] In step ST12, downsampling processing is performed. Specifically, the downconverter 12 performs downsampling processing on the mixed sound signal x(h) input to the mixed sound signal input unit 11. This processing generates a mixed sound signal x(n) that is a non-hi-resolution sound source. Then, the processing proceeds to step ST13.

[0030] In step ST13, sound source separation processing is performed. Specifically, sound source signals s1(n), s2(n), and s3(n) are obtained by performing sound source separation processing on the mixed sound signal x(n) by the sound source separation unit 131. Then, the processing proceeds to step ST14.

[0031] In step ST14, a band extension process is performed. Specifically, the band extension process is performed by the band extension unit 132 on each sound source signal obtained by the sound source separation process by the sound source separation unit 131. As a result, a sound source signal s1(h), a sound source signal s2(h), and a sound source signal s3(h) are obtained. Then, the process proceeds to step ST15.

[0032] In step ST15, a mask generation process is performed. Specifically, the mask generation processing unit 133 generates a mask MA1, a mask MA2, and a mask MA3 based on the sound source signal s1(h), the sound source signal s2(h), and the sound source signal s3(h). Then, the process proceeds to step ST16.

[0033] In step ST16, a mask application process is performed. Specifically, the mask processing unit 14 applies the mask MA1, the mask MA2, and the mask MA3 to the mixed sound signal x(h), thereby separating the sound source signals s'1(h), s'2(h), and s'3(h) from the mixed sound signal x(h). Then, the process proceeds to step ST17.

[0034] In step ST17, a separated sound source signal output process is performed. Specifically, the sound source signals s'1(h), s'2(h), and s'3(h) separated by the mask processing unit 14 are output from the separated sound source signal output unit 15.

[0035] [Effects Obtained by This Embodiment] According to this embodiment, for example, the following effects can be obtained. This technology can perform appropriate source separation for mixed sound signals that are high-resolution audio sources. For example, high-frequency components are calculated based on the input high-resolution audio source using masking processing, so the source separation results retain the high-frequency components of the original high-resolution audio source. The separated audio source signals provide desirable results for content that emphasizes the creator's intent, such as music. In general, in sound source separation, even if there is an error in each separation result and noise is noticeable when listened to individually, it is known that if the sum of the separation results matches the original sound, the noise is less likely to be perceived in situations where all separated sound sources are made to sound at the same time by changing the spatial arrangement or volume balance, such as in upmixing or remixing. According to this embodiment, it is possible to guarantee that the sum of the sound source separation results by the mask processing unit 14 matches the original sound (mixed sound signal that is a high-resolution sound source). Therefore, even if the sound source separation result contains noise, it is possible to obtain a sound source separation result that makes it possible to make the noise less perceptible by changing the spatial arrangement or volume balance, for example. The band extension process by the band extension unit 132 in the above-described embodiment requires much less processing and memory than sound source separation process, and therefore can significantly reduce these costs compared to performing sound source separation in the band of a high-resolution sound source. Furthermore, it is preferable that the sound source separation results for the normal band and the high-resolution sound source are substantially the same in the normal band. However, according to this embodiment, even if the input is a high-resolution sound source, downsampling processing is performed, and the same sound source separation processing (sound source separation processing for the normal band sound source) is applied to the resulting normal band sound source. This eliminates differences in sound quality and separation accuracy between separated sound sources in the normal band, preventing deterioration in sound quality or separation performance even for high-resolution sound sources. Furthermore, there is no need to store parameters for a separate sound source separation model for high-resolution sound sources, which can suppress increases in the number of memories and memory capacity required.

[0036] <Second embodiment> Next, a second embodiment of the present disclosure will be described. Note that, unless otherwise specified, the matters described in the first embodiment can also be applied to the second embodiment. Generally speaking, in the first embodiment, each process is performed on a time domain signal, but in the second embodiment, some of the processes described in the first embodiment are performed on a signal converted into a frequency domain, which is a difference from the first embodiment.

[0037] (Configuration example) 4 is a block diagram showing an example of the configuration of a signal processing device (signal processing device 2) according to the second embodiment. The signal processing device 2 includes a mixed sound signal input unit 21, a downconverter 22, a Short-Term Fourier Transform (STFT) 23, a sound source separation unit 24, an inverse Short-Term Fourier Transform (iSTFT) 25, a band extension unit 26, an STFT 27, a mask generation unit 28, a Multichannel Wiener Filter (MWF) 29, which is a mask processing unit in this embodiment, an iSTFT 30, and a separated sound source signal output unit 31.

[0038] The mixed sound signal input unit 21 has the same configuration as the mixed sound signal input unit 11. A mixed sound signal that is a high-resolution sound source is input to the mixed sound signal input unit 21.

[0039] The downconverter 22 performs downsampling processing on the mixed sound signal in the same manner as the downconverter 12 .

[0040] The STFT 23 performs a short-time Fourier transform process to convert the output signal of the downconverter 22 from a time domain signal to a frequency domain signal.

[0041] The sound source separation unit 24 performs sound source separation processing on the output signal of the STFT 23. An example of the sound source separation processing performed by the sound source separation unit 24 will be described later.

[0042] The iSTFT 25 performs an inverse short-time Fourier transform to convert the output signal of the sound source separation unit 24 from a frequency domain signal to a time domain signal.

[0043] The band extending unit 26 performs band extending processing on the output signal of the iSTFT 25 in the same manner as the band extending unit 132 .

[0044] The STFT 27 performs a short-time Fourier transform to convert the mixed sound signal input to the mixed sound signal input unit 21 and the output signal of the band expanding unit 26 from a time domain signal into a frequency domain signal.

[0045] The mask generation unit 28 generates a mask using the mixed sound signal converted into a frequency domain signal by the STFT 27, etc.

[0046] The MWF 29 separates the sound source signal contained in the mixed sound signal by applying the mask generated by the mask generating unit 28 to the mixed sound signal.

[0047] The iSTFT 30 converts the separation result of the MWF 29 from a frequency domain signal to a time domain signal by performing an inverse short-time Fourier transform.

[0048] The separated sound source signal output unit 31 outputs each sound source signal converted into a time domain signal by the iSTFT 30 .

[0049] (Example of operation) An example of the operation of the signal processing device 2 will now be described in detail. A mixed sound signal x(h), which is a high-resolution sound source, is input to the mixed sound signal input unit 21. The mixed sound signal x(h) is supplied to the downconverter 22 and the STFT 27. The downconverter 22 performs downsampling processing to convert the mixed sound signal x(h) into a mixed sound signal x(n).

[0050] The mixed sound signal x(n) is converted into a mixed sound signal j(n), which is a frequency domain signal, by short-time Fourier transform processing in the STFT 23. Then, the sound source separation processing in the sound source separation unit 24 separates the vocal sound source signal sj1(n), the drum sound source signal sj2(n), and the bass sound source signal sj3(n) contained in the mixed sound signal j(n).

[0051] The subsequent inverse short-time Fourier transform of iSTFT25 converts the sound source signals sj1(n), sj2(n), and sj3(n) into sound source signals s1(n), s2(n), and s3(n), which are time-domain signals.

[0052] The band extension unit 26 performs band extension processing on each sound source signal converted into a time domain signal, thereby obtaining sound source signals s1(h), s2(h), and s3(h) having bands equivalent to those of the high-resolution sound source.

[0053] The mixed sound signal x(h) is converted into a mixed sound signal j(h), which is a frequency domain signal, by the short-time Fourier transform of the STFT 27. Also, the sound source signals s1(h), s2(h), and s3(h) are converted into sound source signals sj1(h), sj2(h), and sj3(h), which are frequency domain signals, respectively.

[0054] The mask generation unit 28 generates a mask corresponding to each sound source signal using the mixed sound signal j(h), sound source signal sj1(h), sound source signal sj2(h), and sound source signal sj3(h). For example, the mask is generated using the ratio of the power spectrum of each sound source signal to the sum. In this example, the mixed sound signal j(h) is used to generate the mask. This makes it possible to preserve the phase component contained in the original signal. Note that the phase component may be restored or adjusted in subsequent processing without using the mixed sound signal j(h) to generate the mask. The mask generation unit 28 generates mask MA1, mask MA2, and mask MA3. The generated masks are supplied to the MWF 29.

[0055] The MWF 29 separates the vocal sound source signal s'j1(h) from the mixed sound signal j(h), for example, by applying a mask MA1 to the mixed sound signal j(h). The MWF 29 also separates the drum sound source signal s'j2(h) from the mixed sound signal j(h), for example, by applying a mask MA2 to the mixed sound signal j(h). The MWF 29 also separates the bass sound source signal s'j3(h) from the mixed sound signal j(h), for example, by applying a mask MA3 to the mixed sound signal j(h).

[0056] Then, the sound source signal s'j1(h), the sound source signal s'j2(h), and the sound source signal s'j3(h) are converted into sound source signal s'1(h), the sound source signal s'2(h), and the sound source signal s'3(h), which are time domain signals, respectively, by the inverse short-time Fourier transform of the iSTFT30. The converted signals are output from the separated sound source signal output unit 31.

[0057] (Example of operation of the sound source separation unit) 5 is a diagram showing a detailed configuration example of the sound source separation unit 24. The sound source separation unit 24 has a DNN (Deep Neural Network) 241A, a DNN 241B, a DNN 241C, and an MWF 242. The following describes a DNN- and MWF-based sound source separation method performed by the sound source separation unit 24 having such a configuration. In the following description, signals are expressed in the STFT domain.

[0058] The I-channel mixed sound signal at frequency bin k and frame m is TIFF0007790351000001.tif7163jth source signal Assuming TIFF0007790351000002.tif7163, MWF assumes the signal model as shown in the following equation (1). TIFF0007790351000003.tif23163Here, in equation (1), TIFF0007790351000004.tif7163 is the power spectral density, TIFF0007790351000005.tif7163 is the spatial correlation matrix.

[0059] From equation (1), the mixed sound signal is the jth source signal and complex Gaussian noise. TIFF0007790351000006.tif7163. Furthermore, by assuming that each source signal is independent of the others, we can obtain the least mean square error (LMSE) TIFF0007790351000007.tif7163 It can be estimated from TIFF0007790351000008.tif7163. Minimum mean square error estimate TIFF0007790351000009.tif7163 can be calculated as follows: TIFF0007790351000010.tif9163

[0060] To find the source signal from equation (2), TIFF0007790351000011.tif7163 and You need to find TIFF0007790351000012.tif7163.

[0061] In Patent Document 1, the spatial correlation matrix is ​​assumed to be time-invariant (the sound source position does not change) and is calculated using a DNN. If you set it to TIFF0007790351000013.tif7163, TIFF0007790351000014.tif7163 and TIFF0007790351000015.tif7163 can be calculated using the following formulas (3) and (4). TIFF0007790351000016.tif10163TIFF0007790351000017.tif13163

[0062] The above-mentioned equation (2) can be expressed as follows using a mixed sound signal: TIFF0007790351000018.tif7163 in this case TIFF0007790351000019.tif7163 and TIFF0007790351000020.tif7163 can be calculated using the following formulas (5) and (6). TIFF0007790351000021.tif10163TIFF0007790351000022.tif12163

[0063] The present embodiment described above can also provide the same effects as the first embodiment.

[0064] <Modification> Although several embodiments of the present disclosure have been described above, the present disclosure is not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present disclosure.

[0065] In the above-described embodiment, a mask processor other than a Wiener filter may be used. For example, a complex ratio mask described in Donald S. Williamson, et al. "Complex Ratio Masking for Monaural Speech Separation," IEEE Trans. ASLP, 2016, Vol. 24, No. 3 may be used as the mask processor. Furthermore, the mask applied to the Wiener filter may be generated by other known methods.

[0066] In the above-described embodiment, the band extension processing may be performed individually on each sound source signal, or may be performed on a predetermined sound source signal while referring to other sound source signals. In the latter case, it is not necessary to provide a band extension unit for each sound source signal.

[0067] In the second embodiment described above, the configuration related to the STFT 27 and the iSTFT 30 may be omitted. Furthermore, the processing subsequent to the band extender 26 may be performed using a time domain signal. In this way, the configuration of the device may be changed as appropriate without departing from the gist of the present disclosure.

[0068] The present disclosure can also employ a cloud computing configuration in which a single function is shared and processed collaboratively by multiple devices via a network.

[0069] The present disclosure may also be realized in any form, such as a device, a method, a program, or a system. For example, a program that performs the functions described in the above-described embodiments may be made downloadable, and a device that does not have the functions described in the embodiments may download and install the program, thereby enabling the device to perform the control described in the embodiments. The present disclosure may also be realized by a server that distributes such a program. The matters described in each embodiment and modified example may be combined as appropriate. The contents of the present disclosure should not be interpreted as being limited to the effects exemplified in this specification.

[0070] The present disclosure may also have the following configurations. (1) a down-converter that applies down-sampling processing to a mixed sound signal that is mixed with a sound source signal that includes high-frequency components higher than a predetermined frequency; a mask generation unit that generates a mask based on a result of the downsampling process performed by the downconverter; a mask processing unit that applies the mask generated by the mask generation unit to the mixed sound signal; A signal processing device comprising: (2) The mask generation unit a sound source separation unit that performs sound source separation processing on the mixed sound signal to which the downsampling processing has been applied; a band extension unit that applies frequency band extension processing to each of the sound source signals separated by the sound source separation unit; a mask generation processing unit that generates the mask corresponding to each sound source signal based on at least each sound source signal to which the frequency band extension processing has been applied; have The signal processing device according to (1). (3) The mask generation processing unit further generates the mask using the mixed sound signal. A signal processing device according to (2). (4) The mask processing unit is configured by a filter in which the sum of the input and output of the mask processing unit matches. A signal processing device according to any one of (1) to (3). (5) The mask processing unit is configured by a Wiener filter. A signal processing device according to (4). (6) The mask processing unit separates and outputs a sound source signal included in the mixed sound signal. A signal processing device according to any one of (1) to (5). (7) The band extension unit applies the frequency band extension process to each sound source signal. A signal processing device according to (2) or (3). (8) The band extension unit applies the frequency band extension process to a predetermined sound source signal by referring to other sound source signals. A signal processing device according to (2) or (3). (9) A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on a downsampling processing result by the downconverter; A mask processing unit applies the generated mask to the mixed sound signal. Signal processing methods. (10) A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on a downsampling processing result by the downconverter; A mask processing unit applies the generated mask to the mixed sound signal. A program that causes a computer to execute a signal processing method. [Explanation of symbols]

[0071] 1, 2 Signal processing device 12, 22... Down converter 13, 28... Mask generation section 14 Mask processing section 24...Sound source separation section 29···MWF 131...Sound source separation section 26, 132... Bandwidth expansion section 133 Mask generation processing unit

Claims

1. a down-converter that applies down-sampling processing to a mixed sound signal that is mixed with a sound source signal that includes high-frequency components higher than a predetermined frequency; a mask generation unit that generates a mask based on a result of the downsampling process performed by the downconverter; a mask processing unit that applies the mask generated by the mask generation unit to the mixed sound signal; and The mask generation unit a sound source separation unit that performs sound source separation processing on the mixed sound signal to which the downsampling processing has been applied; a band extension unit that applies frequency band extension processing to each of the sound source signals separated by the sound source separation unit; a mask generation processing unit that generates the mask corresponding to each sound source signal based on a relative ratio of the sound source signal to a sum of the sound source signals, based on at least each sound source signal to which the frequency band extension processing has been applied; have Signal processing device.

2. The mask generation processing unit further generates the mask using the mixed sound signal. The signal processing device according to claim 1 .

3. The mask processing unit is configured by a filter in which the sum of the input and output of the mask processing unit matches. The signal processing device according to claim 1 .

4. The mask processing unit is configured by a Wiener filter. The signal processing device according to claim 3 .

5. The mask processing unit separates and outputs a sound source signal included in the mixed sound signal. The signal processing device according to claim 1 .

6. The band extension unit applies the frequency band extension process to each sound source signal. The signal processing device according to claim 1 .

7. The band extension unit applies the frequency band extension process to a predetermined sound source signal by referring to other sound source signals. The signal processing device according to claim 1 .

8. A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on a downsampling processing result by the downconverter; a mask processing unit that applies the generated mask to the mixed sound signal; a sound source separation unit included in the mask generation unit performs a sound source separation process on the mixed sound signal to which the downsampling process has been applied; a band extension unit included in the mask generation unit applies frequency band extension processing to each of the sound source signals separated by the sound source separation processing; A mask generation processing unit included in the mask generation unit generates the mask corresponding to each sound source signal based on a relative ratio of the sound source signal to a sum of the sound source signals, based on at least the individual sound source signals to which the frequency band extension processing has been applied. Signal processing methods.

9. A downconverter applies downsampling processing to the mixed sound signal into which a sound source signal containing high-frequency components higher than a predetermined frequency is mixed, a mask generation unit that generates a mask based on a downsampling processing result by the downconverter; a mask processing unit that applies the generated mask to the mixed sound signal; a sound source separation unit included in the mask generation unit performs a sound source separation process on the mixed sound signal to which the downsampling process has been applied; a band extension unit included in the mask generation unit applies frequency band extension processing to each of the sound source signals separated by the sound source separation processing; A mask generation processing unit included in the mask generation unit generates the mask corresponding to each sound source signal based on a relative ratio of the sound source signal to a sum of the sound source signals, based on at least the individual sound source signals to which the frequency band extension processing has been applied. A program that causes a computer to execute a signal processing method.

Citation Information

Patent Citations

  • Device and method for sound source separation, and program

    WO2018047643A1