Sound quality improvement device, sound quality improvement method and program

The combination of independent component analysis and semi-supervised non-negative matrix factorization enhances sound source separation and quality by using one sound source as a teacher signal to estimate basis and activation matrices, addressing the limitations of existing methods.

JP2025122810APending Publication Date: 2025-08-22HIROSHIMA CITY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024018478
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Existing sound source separation methods like ICA and semi-supervised non-negative matrix factorization face challenges in achieving high-quality separation, particularly when sound sources lack distribution distortion or result in incomplete separation of sounds.

Method used

A sound quality improvement device and method that combines independent component analysis with semi-supervised non-negative matrix factorization, using one sound source as a teacher signal to estimate basis matrices and activation matrices for each sound source, thereby enhancing sound separation and quality.

Benefits of technology

Improves the quality of separated sounds by effectively separating and reconstructing individual sound sources from a mixed audio signal, even when distribution distortion is absent or separation is incomplete.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025122810000001_ABST
    Figure 2025122810000001_ABST
Patent Text Reader

Abstract

To provide a sound quality improvement device, a sound quality improvement method and a program, which can improve sound quality of sound separated and extracted from an acoustic signal in which a plurality of sounds are mixed.SOLUTION: A sound quality improvement device 1 includes a separation unit 10 and an estimation unit 20. The separation unit 10 performs independent component analysis on first acoustic signals SA1(t) and SB1(t), which are obtained by recording composite sound of sounds emitted from a plurality of sound sources (a first sound source O1 and a second sound source O2) at different positions (microphones 2A and 2B) and separates the signals into second acoustic signals SA2(t) and SB2(t) for each sound source. The estimation unit 20 performs non-negative value matrix factorization on the first acoustic signal SB1(t) and the second acoustic signal SA2(t) and estimates a third acoustic signal S3(t) corresponding to the sound source on at least one sound source in the plurality of sound sources (first sound source O1 and second sound source O2).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound quality improvement device, a sound quality improvement method, and a program. [Background technology]

[0002] Attempts have been made to perform sound source separation using ICA (Independent Component Analysis) to separate an audio signal corresponding to one sound source from an audio signal containing sounds emitted from multiple sound sources.Furthermore, attempts have been made to generate audio signals with even higher sound quality by separating an audio signal corresponding to one sound source from an audio signal containing sounds emitted from multiple sound sources using semi-supervised non-negative matrix factorization (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Daichi Kitamura, Nobutaka Ono, Hiroshi Saruwatari, Yu Takahashi, and Kazunobu Kondo, "Effective Basis Learning Method for Improving Source Separation Performance in Semi-Supervised Nonnegative Matrix Factorization," IEICE Technical Report, EA2015-130, vol. 115, no. 521, pp. 355-360, March 2016. Summary of the Invention [Problem to be solved by the invention]

[0004] While ICA can achieve high-quality sound source separation, it requires a certain distribution distortion for each sound source. Conversely, ICA cannot effectively separate acoustic signals from sound sources that do not have a distribution distortion. On the other hand, even with semi-supervised non-negative matrix factorization, some sounds may remain unremoved or may become inseparable.

[0005] The present invention has been made in light of the above-mentioned circumstances, and aims to provide a sound quality improvement device, a sound quality improvement method, and a program that can improve the sound quality of sounds that are separated and extracted from an acoustic signal containing a mixture of multiple sounds. [Means for solving the problem]

[0006] In order to achieve the above object, a sound quality improving device according to a first aspect of the present invention comprises: a separation unit that performs independent component analysis on a first acoustic signal obtained by recording a composite sound of sounds emitted from a plurality of sound sources at different positions, and separates the first acoustic signal into second acoustic signals for each sound source; an estimation unit that performs nonnegative matrix factorization on the first acoustic signal and the second acoustic signal, respectively, to estimate a third acoustic signal corresponding to a sound source for at least one of the plurality of sound sources; Equipped with.

[0007] In this case, the non-negative matrix factorization is a process of decomposing the first acoustic signal or the second acoustic signal into a basis matrix representing a spectral pattern thereof and an activation matrix representing a temporal intensity change of the spectral pattern, The estimation unit performing non-negative matrix factorization using the second acoustic signal corresponding to a first sound source among the plurality of sound sources as a teacher signal to estimate the basis matrix corresponding to the first sound source; performing non-negative matrix factorization on the first acoustic signal, and estimating an activation matrix corresponding to the first sound source and a basis matrix and an activation matrix corresponding to a second sound source among the plurality of sound sources based on the estimated basis matrix; generating the third acoustic signal corresponding to the first sound source based on the basis matrix and activation matrix corresponding to the estimated first sound source, or generating the third acoustic signal corresponding to the second sound source based on the basis matrix and activation matrix corresponding to the estimated second sound source; This may also be the case.

[0008] A sound quality improving method according to a second aspect of the present invention comprises: A sound quality improvement method executed by an information processing device, A composite sound of sounds emitted from each of a plurality of sound sources is recorded at different positions, and a first acoustic signal is obtained by performing independent component analysis on the first acoustic signal to separate it into second acoustic signals for each sound source; A third acoustic signal corresponding to a sound source is estimated for at least one of the plurality of sound sources by performing nonnegative matrix factorization on the first acoustic signal and the second acoustic signal, respectively.

[0009] A program according to a third aspect of the present invention comprises: Computer, a separation unit that performs independent component analysis on a first acoustic signal obtained by recording a composite sound of sounds emitted from a plurality of sound sources at different positions, and separates the first acoustic signal into second acoustic signals for each sound source; an estimation unit that performs nonnegative matrix factorization on the first acoustic signal and the second acoustic signal, respectively, to estimate a third acoustic signal corresponding to a sound source for at least one of the plurality of sound sources; Function as. [Effects of the Invention]

[0010] According to the present invention, it is possible to improve the quality of sound that is separated and extracted from an audio signal in which a plurality of sounds are mixed. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing a functional configuration of a sound quality improving device according to an embodiment of the present invention; [Figure 2] FIG. 2 is a schematic diagram illustrating a configuration of a separation unit. [Figure 3] FIG. 2 is a block diagram showing the configuration of an estimation unit. [Figure 4] FIG. 2 is a block diagram showing the hardware configuration of the sound quality improving device of FIG. 1. [Figure 5] 2 is a flowchart of the processing of the sound quality improving device of FIG. 1. [Figure 6] FIG. 2 is a block diagram showing the operation of the sound quality improving device of FIG. [Figure 7] FIG. 10 is a block diagram showing a modified example of the functional configuration of the sound quality improving device when there are three sound sources. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each drawing, the same or equivalent parts are denoted by the same reference numerals. In the following embodiments, the terms "have," "include," or "contain" also mean "consist of" or "consist of."

[0013] [Overall configuration] The sound quality improving device 1 shown in FIG. 1 separates a composite sound of sounds emitted from a plurality of sound sources, a first sound source O1 and a second sound source O2, into a sound emitted from the first sound source O1 and a sound emitted from the second sound source O2, and reproduces and outputs the sounds with improved sound quality.

[0014] The sound quality improving device 1 includes microphones 2A and 2B and a speaker 3. The microphone 2A receives the synthesized sound and outputs a first acoustic signal S A1 The microphone 2B receives the synthesized sound and outputs a first acoustic signal S B1 Since the microphone 2A and the microphone 2B are installed in different locations, the first acoustic signal S A1 (t) and the first acoustic signal S B1 In the sound signals O1 and O2 (t), the loudness of the sounds emitted from the first sound source O1 and the second sound source O2 differs. The speaker 3 outputs a third acoustic signal S3(t) for each sound source with improved sound quality. Here, at least one of the third acoustic signal S3(t) corresponding to the sound emitted from the first sound source O1, i.e., the first sound source O1, and the third acoustic signal S3(t) corresponding to the sound emitted from the second sound source O2, i.e., the second sound source O2, is output from the speaker 3.

[0015] 1, the sound quality improving device 1 includes a separation unit 10 and an estimation unit 20. The separation unit 10 separates a first acoustic signal S output from a microphone 2A into a first sound signal S A1 (t) and the first acoustic signal S output from the microphone 2B. B1 (t) and the estimation unit 20 outputs a third acoustic signal S3(t) corresponding to the first sound source O1 or the second sound source O2.

[0016] [Separation section] The separation unit 10 separates a first acoustic signal S obtained by recording a composite sound of sounds emitted from a plurality of sound sources, i.e., a first sound source O1 and a second sound source O2, at different positions (microphones 2A and 2B). A1 (t), S B1 (t) is subjected to independent component analysis to obtain a second acoustic signal for each sound source, i.e., a second acoustic signal S corresponding to the first sound source O1. A2 (t) and the second acoustic signal S corresponding to the second sound source O2. B2 (t) and are separated.

[0017] Independent component analysis (ICA) is a method for separating a composite signal into individual signals when the composite signal is observed. For example, consider an unknown signal s generated from n independent signal sources. i (t) (i, n is a natural number greater than or equal to 2; t = 0, 1, 2, ...; in Figure 2, n = 2), and these signals are converted into signals x by a regular matrix called H. i If these signals are expressed as vectors s(t) and x(t), respectively, we get the following equations: x(t)=Hs(t) (1)

[0018] To convert vector x(t) into vector s(t), it is sufficient to find the inverse matrix of H. However, in reality, since the matrix H is unknown, the following equation is set using vector y(t): y(t)=Wx(t) (2) In independent component analysis, each component y of the vector y(t) iFind the matrix W such that (t) is independent. Each component y of the vector y(t) i The independence of (t) is expressed by the generalized Kullback-Leibler divergence shown in the following equation. In the following equation, y(t)=y, y i (t)=y i is shown as:

number

[0019] In this embodiment, since there are two sound sources, n can be set to 2. For example, as shown in FIG. 2, the first acoustic signal S A1 (t) is denoted by x1(t), and the first acoustic signal S sent from the microphone 2B is B1 Let (t) be x2(t). The signals x1(t) and x2(t) are composite signals that are a mixture of the sound signal s1(t) emitted from the first sound source O1 and the sound signal s2(t) emitted from the second sound source O2. In independent component analysis, the matrix that converts the signals s1(t) and s2(t) into the composite signals x1(t) and x2(t) is represented as matrix H, with elements h11, h21, h12, and h22.

[0020] The separation unit 10 multiplies the signals x1(t) and x2(t) by a matrix W having elements w11, w21, w12, and w22 to estimate a signal y1(t) of a sound emitted from the first sound source O1 and a signal y2(t) of a sound emitted from the second sound source O2. The separation unit 10 sets an objective function in the matrix W, and updates the values ​​of the elements w11, w21, w12, and w22 of the matrix W so that the objective function is minimized. This brings the signals y1(t) and y2(t) closer to the signals s1(t) and s2(t). The signals y1(t) and y2(t) are converted into the second acoustic signal S shown in FIG. A2 (t), S B2 Corresponding to (t).

[0021] More specifically, the separator 10, for example, divides the matrix W into a whitening matrix and a unitary rotation matrix, determines the value of each element of the whitening matrix from the correlation matrix of the signals x1(t) and x2(t) by eigenvalue decomposition, and then determines the value of each element of the unitary rotation matrix by repeating Newton's method, Gram-Schmidt orthogonalization, and norm normalization until convergence is achieved. In this way, the matrix W can be estimated. Alternatively, the matrix W may be determined using maximum likelihood estimation or natural gradient. In this embodiment, there are no limitations on how the matrix W is determined.

[0022] The separator 10 separates the first acoustic signal S sent from the microphone 2A. A1 (t)(x1(t)) and the first acoustic signal S sent from microphone 2B. B1 (t)(x2(t)), the second acoustic signal S corresponding to the first sound source O1 is calculated using the above equation (2). A2 (t)(y1(t)), and the second acoustic signal S corresponding to the second sound source O2. B2 (t)(y2(t)) and output it.

[0023] In this embodiment, the generalized Kullback-Leibler divergence is used, but the squared Euclidean distance, the Itakura-Saito divergence, or the Alpha-divergence may also be used.

[0024] [Estimation part] The estimation unit (SS-NMF; Semi-Supervised Nonnegative Matrix Factorization) 20 estimates the first acoustic signal S B1 (t) and the second acoustic signal S A2 (t) is subjected to non-negative matrix factorization to estimate a third acoustic signal S3(t) corresponding to at least one of the multiple sound sources (first sound source O1, second sound source O2).

[0025] Nonnegative matrix factorization is a process that decomposes a matrix of a signal's amplitude spectrum into a basis matrix that represents the signal's spectral pattern and an activation matrix that represents the temporal intensity change of the spectral pattern. The estimation unit 20 generates an amplitude spectrogram matrix by performing frequency transformation on the input signal and estimates the basis matrix and activation matrix by approximating the amplitude spectrogram matrix to the product of the basis matrix and activation matrix. Control conditions are set for all matrices so that their elements are nonnegative. The basis matrix and activation matrix are estimated using a distance function between the amplitude spectrogram matrix obtained from the signal and the product of the basis matrix and activation matrix. Various distance functions can be used, including the generalized Kullback-Leibler divergence, squared Euclidean distance, Itakura-Saito divergence, and alpha-divergence.

[0026] As shown in FIG. 3, the estimation unit 20 includes a frequency conversion unit 21, a learning unit 22, a frequency conversion unit 23, an approximation calculation unit 24, an inverse frequency conversion unit 25, and an output signal selection unit .

[0027] The frequency conversion unit 21 converts the second acoustic signal S output from the separation unit 10 into A2 (t) is subjected to frequency conversion, resulting in a second acoustic signal S A2 The matrix S of the amplitude spectrogram of (t) A2 It should be noted that a power spectrogram may be obtained instead of the amplitude spectrogram.

[0028] 2nd acoustic signal S A2 (t) is an acoustic signal corresponding to the sound emitted from the first sound source O1 in the separation unit 10. A2 (t) into the matrix S of the amplitude spectrogram A2 The model is expressed as follows using the basis matrix F and the activation matrix W: Note that this activation matrix W is different from the matrix W in the separation unit 10. S A2 =FW …(4)

[0029] The learning unit 22 learns the matrix S of the amplitude spectrogram. A2 and the basis matrix F and the activation matrix W are estimated by performing an approximation operation so that a distance function between the product of the basis matrix F and the activation matrix W is minimized. A2 Using (t) as a teacher signal, non-negative matrix factorization is performed to estimate a basis matrix F corresponding to the first sound source O1. The estimated basis matrix F is output to the approximation calculation unit 24.

[0030] The frequency converter 23 converts the first acoustic signal S B1 (t) is subjected to frequency conversion, resulting in the first acoustic signal S B1 The matrix S of the amplitude spectrogram of (t) B1 is obtained.

[0031] 1st acoustic signal S B1 (t) is an acoustic signal in which the sound emitted from the first sound source O1 and the sound emitted from the second sound source O2 are mixed in the separation unit 10. Therefore, this first acoustic signal S B1 (t) into the matrix S of the amplitude spectrogram B1 The model is expressed as follows using the basis matrix F' and activation matrix W corresponding to the sound emitted from the first sound source O1, and the basis matrix T and activation matrix V corresponding to the sound emitted from the second sound source O2. S B1 =F'W+TV …(5) Here, F=F'. That is, the approximation calculation unit 24 sets the basis matrix F estimated by the learning unit 22 as the basis matrix F'.

[0032] The approximation calculation unit 24 calculates the matrix S of the amplitude spectrogram B1and the approximation calculation unit 24 estimates the basis matrix F and the activation matrix W by performing an approximation calculation so that the distance function between the sum of the product of the basis matrix F and the activation matrix W and the product of the basis matrix T and the activation matrix V is minimized. In other words, the approximation calculation unit 24 estimates the basis matrix F and the activation matrix W by approximating the first acoustic signal S B1 (t) is subjected to non-negative matrix factorization, and based on the basis matrix F estimated by the learning unit 22, an activation matrix W corresponding to the first sound source O1, and a basis matrix T and activation matrix V corresponding to the second sound source O2 are estimated.

[0033] The approximation calculation unit 24 generates a product of a basis matrix F corresponding to the estimated first sound source O1 and an activation matrix W, or generates a product of a basis matrix T corresponding to the estimated second sound source O2 and an activation matrix V.

[0034] The inverse frequency transform unit 25 generates a time series signal by inverse frequency transforming at least one of the product of a basis matrix F and an activation matrix W and the product of a basis matrix T and an activation matrix V. That is, the inverse frequency transform unit 25 generates an acoustic signal corresponding to the first sound source O1 based on the basis matrix F and activation matrix W corresponding to the estimated first sound source O1, or generates an acoustic signal corresponding to the second sound source O2 based on the basis matrix T and activation matrix V corresponding to the estimated second sound source O2.

[0035] The output signal selection unit 26 selects at least one of the acoustic signal corresponding to the first sound source O1 and the acoustic signal corresponding to the second sound source O2, and outputs it as a third acoustic signal S3(t). The speaker 3 outputs the selected third acoustic signal S3(t) as sound.

[0036] As described above, in the estimation unit 20, the learning unit 22 A2Using (t) as a teacher signal, the basis matrix F corresponding to the first sound source O1 is learned and obtained, and the basis matrix is ​​used to estimate W+TV, i.e., the activation matrix W of the first sound source O1, the basis matrix T of the second sound source O2, and the activation matrix V. That is, the estimation unit 20 estimates the third acoustic signal S3(t) by semi-supervised non-negative matrix factorization. Since only the basis matrix F of the first sound source O1 is learned and the activation matrix W, the basis matrix T, and the activation matrix V are estimated without learning the basis matrix T of the second sound source O2, this is semi-supervised learning.

[0037] [Hardware configuration of sound quality improvement device 1] The sound quality improving device 1 as an information processing device shown in Fig. 1 is realized, for example, by a computer having the hardware configuration shown in Fig. 4 executing a software program. Specifically, the sound quality improving device 1 includes a CPU (Central Processing Unit) 31 that controls the entire device, a main memory 32 that operates as a work area or the like for the CPU 31, an external memory 33 that stores the operating program or the like for the CPU 31, an operation unit 34, a display 35, a device interface 36, and an internal bus 30 that connects these.

[0038] As will be described later, the CPU 31 executes a program 39 stored in the main memory 32 to realize various functions of the sound quality improving device 1. The CPU 31 has a timer that measures time, and executes processing according to the time measured by the timer.

[0039] The main memory 32 is composed of RAM (Random Access Memory) and the like. A program 39 to be executed by the CPU 31 is loaded into the main memory 32 from the external memory 33. The main memory 32 is also used as a work area (temporary data storage area) for the CPU 31. The functions of the separation unit 10 and the estimation unit 20 shown in FIGS. 1, 2, and 3 are realized by the CPU 31 executing the program.

[0040] The external memory 33 is configured by a non-volatile memory such as a flash memory, a hard disk, etc. The external memory 33 stores in advance a program 39 to be executed by the CPU 31.

[0041] The operation unit 34 is composed of devices such as a keyboard and a mouse, and an interface device that connects these devices to the internal bus 30. The display 35 is a display device such as a CRT (Cathode Ray Tube) or a liquid crystal monitor. Note that a touch panel can be used as the operation unit 34 and the display 35.

[0042] The device interface 36 connects to external hardware devices. The computer is connected to the microphones 2A and 2B and the speaker 3 shown in FIG.

[0043] The functions of the sound quality improving device 1 can be implemented in a computer system consisting of one or more computers including one or more processors and one or more storage devices including non-transitory storage media. The multiple computers realize the functions of the sound quality improving device 1 while communicating with each other via a communication network connected to each other. For example, some of the multiple functions of the sound quality improving device 1 may be implemented in one computer, and other parts may be implemented in other computers.

[0044] [Sound quality improvement processing] Next, a description will be given of the sound quality improvement process executed by the sound quality improvement device 1. As shown in Fig. 5, a separation unit 10 constituting the sound quality improvement device 1 separates a first acoustic signal S obtained by recording a synthesized sound of sounds emitted from a first sound source O1 and a second sound source O2 at different positions. A1 (t), S B1 (t) is subjected to independent component analysis to obtain a second acoustic signal S corresponding to the first sound source O1. A2 (t) and a second acoustic signal S corresponding to a second sound source O2. B2 (t) and are generated (step S1).

[0045] Next, in the estimation unit 20, the frequency conversion unit 21 converts the second acoustic signal S corresponding to the first sound source O1 into A2 (t) is frequency transformed to obtain the amplitude spectrogram matrix S A2 , while the frequency converter 23 generates the first acoustic signal S B1 (t) is frequency transformed to obtain the amplitude spectrogram matrix S B1 is generated (step S2).

[0046] Next, the learning unit 22 learns the second acoustic signal S corresponding to the first sound source O1. A2 By learning using (t) as a teacher signal, the amplitude spectrum matrix S A2 The basis matrix F corresponding to the above is estimated (step S3). The estimated basis matrix F is set as the basis matrix F′ of the approximation calculation unit 24.

[0047] Next, the approximation calculation unit 24 calculates the first acoustic signal S input from the microphone 2B. B1 The matrix S of the amplitude spectrogram of (t) B1 The activation matrix W corresponding to the first sound source O1 and the basis matrix T and activation matrix V corresponding to the second sound source O2 are estimated based on the basis matrix F estimated by the learning unit 22 by performing non-negative matrix factorization on the matrix O (step S4).

[0048] Next, the inverse frequency transform unit 25 performs an inverse frequency transform on the product of the estimated basis matrix F and activation matrix W to obtain a third acoustic signal S3(t) corresponding to the first sound source O1, and also performs an inverse frequency transform on the product of the estimated basis matrix T and activation matrix V to obtain a third acoustic signal S3(t) corresponding to the second sound source O2 (step S5).

[0049] Next, the output signal selection unit 26 selects and outputs either the third acoustic signal S3(t) corresponding to the first sound source O1 or the third acoustic signal S3(t) corresponding to the second sound source O2 (step S6). When both are output, it is preferable to output them at different times.

[0050] Here, a case will be described in which the first sound source O1 is a musical instrument and the second sound source O2 is a person. In this case, as shown in Fig. 6, a synthesized sound of the instrument sound and a voice is input to the separation unit 10. The separation unit 10 performs a process of separating the instrument sound and the voice from the synthesized sound by independent component analysis.

[0051] If this separation is successful (if the instrument sound and the voice are clearly separated), for example, the separated voice is decomposed into basis matrices F and T and activation matrices W and V corresponding to the first sound source O1 and the second sound source O2, respectively, by nonnegative matrix factorization in the estimation unit 20, and is estimated with high accuracy. The instrument sound and voice reconstructed using the basis matrices F and T and the activation matrices W and V are narrowed down to the sounds emitted from the first sound source O1 and the second sound source O2, respectively, thereby achieving high sound quality. The estimation unit 20 learns the basis matrix F corresponding to the instrument sound (first sound source O1), and estimates the voice (second sound source O2) based on the learned basis matrix F. Therefore, this estimation process is equivalent to obtaining high-quality voice by subtracting the instrument sound from the synthesized sound (instrument + voice).

[0052] Furthermore, even if the separation unit 10 fails to separate (if the separated sound contains both instrument sound components and voice components), the estimation unit 20 further promotes separation of the two, thereby obtaining audio with higher sound quality.

[0053] The first sound source O1 may be a person and the second sound source O2 may be a musical instrument. Both may be musical instruments, or both may be people. Alternatively, the sound source may be other than a musical instrument or a person. For example, it is also possible to separate sounds using noise from the outside world as a sound source.

[0054] As described above in detail, the sound quality improving device 1 according to this embodiment generates a first acoustic signal S obtained by recording, at different positions, a synthesized sound of sounds emitted from a first sound source O1 and a second sound source O2 among a plurality of sound sources. A1 (t), S B1 (t) is subjected to independent component analysis to obtain a second acoustic signal S corresponding to the first sound source O1 and the second sound source O2 for each sound source. A2(t), S B2 a separation unit 10 for separating a first acoustic signal S B1 (t) and the second acoustic signal S A2 and an estimation unit 20 that performs nonnegative matrix factorization on each of the first sound source O1 and the second sound source O2 to estimate a third sound signal S3(t) corresponding to the sound source for at least one of the first sound source O1 and the second sound source O2. The sound quality improvement device 1 according to this embodiment further performs nonnegative matrix factorization on the sounds separated by independent component analysis to separate the sounds, thereby improving the sound quality when separating and extracting any sound from an audio signal in which multiple sounds are mixed.

[0055] Furthermore, according to the sound quality improvement device 1 of this embodiment, the non-negative matrix factorization performed by the estimation unit 20 is A2 The estimation unit 20 decomposes the second acoustic signal S (t) corresponding to the first sound source O1 into a basis matrix F that indicates the spectral pattern of the second acoustic signal S (t) and an activation matrix W that indicates the temporal intensity change of the spectral pattern. A2 The estimation unit 20 performs non-negative matrix factorization using the first acoustic signal S (t) as a teacher signal to estimate a basis matrix F corresponding to the first sound source O1. B1 Semi-supervised non-negative matrix factorization is performed on (t) to estimate an activation matrix W corresponding to the first sound source O1 and a basis matrix T and activation matrix V corresponding to the second sound source O2 based on the estimated basis matrix F. In this way, both sounds can be improved to high quality without learning basis matrices for all sound sources.

[0056] In the above embodiment, there are two sound sources. However, there may be three or more sound sources. For example, with reference to FIG. 7, consider a case where a composite sound generated from a first sound source O1, a second sound source O2, and a third sound source O3 is separated into acoustic signals for each sound source. In this case, three microphones, namely, microphones 2A, 2B, and 2C, are required. Microphone 2A receives a first acoustic signal S corresponding to the received composite sound. A1 (t), and the microphone 2B outputs a first acoustic signal S corresponding to the received synthesized sound.B1 (t), and the microphone 2C outputs a first acoustic signal S corresponding to the received synthesized sound. C1 The separation unit 10 outputs the first acoustic signal S A1 (t), first acoustic signal S B1 (t) and the first acoustic signal S C1 (t), the synthesized sound is generated by combining the second acoustic signal S corresponding to the first sound source O1 with the A2 (t) and the second acoustic signal S corresponding to the second sound source O2. B2 (t) and the second acoustic signal S corresponding to the third sound source. C2 (t) and are separated.

[0057] The estimation unit 20 estimates the second acoustic signal S corresponding to the first sound source O1 separated by the separation unit 10. A2 The estimation unit 20 performs non-negative matrix factorization using the training signal t to estimate a basis matrix F corresponding to the first sound source O1. Furthermore, the estimation unit 20 estimates a second acoustic signal S corresponding to the second sound source O2. B2 The estimation unit 20 performs non-negative matrix factorization using the first acoustic signal S (t) as a teacher signal to estimate a basis matrix T corresponding to the second sound source O2. C1 The estimation unit 20 performs non-negative matrix factorization on (t) to estimate an activation matrix W corresponding to the first sound source O1, an activation matrix V corresponding to the second sound source O2, and a basis matrix and activation matrix corresponding to the third sound source O3 based on the estimated basis matrix F corresponding to the first sound source O1 and the basis matrix T corresponding to the second sound source O2. The estimation unit 20 can reconstruct a third acoustic signal S3(t) of the first sound source O1, the second sound source O2, and the third sound source O3 by multiplying the basis matrix and activation matrix corresponding to each sound source.

[0058] The estimation unit 20 estimates a third acoustic signal S3(t) corresponding to the first sound source O1 based on a basis matrix F and an activation matrix W corresponding to the first sound source O1. Furthermore, the estimation unit 20 estimates a third acoustic signal S3(t) corresponding to the second sound source O2 based on a basis matrix T and an activation matrix V corresponding to the second sound source O2. Furthermore, the estimation unit 20 estimates a third acoustic signal S3(t) corresponding to the third sound source O3 based on a basis matrix and an activation matrix corresponding to the third sound source O3. The estimation unit 20 selectively outputs at least one of these third acoustic signals S3(t).

[0059] In this way, recursive processing can separate sounds that are a composite of three or more sounds and improve the quality of the sound.

[0060] As described above, in the above embodiment, the estimation unit 20 estimates the third acoustic signal S3(t) by semi-supervised non-negative matrix factorization. However, the estimation unit 20 may learn the basis matrix T in addition to the basis matrix F, and estimate the third acoustic signal S3(t) for each sound source by fully supervised non-negative matrix factorization.

[0061] Furthermore, the hardware configuration and software configuration of the sound quality improving device 1 are merely examples and can be changed and modified as desired.

[0062] As described above, the core processing portion of the sound quality improving device 1, which is composed of the CPU 31, main memory 32, external memory 33, operation unit 34, display 35, device interface 36, internal bus 30, etc., can be realized using an ordinary computer system rather than a dedicated system. For example, the sound quality improving device 1 that executes the above-described processing may be configured by storing and distributing a computer program for executing the above-described operations on a computer-readable recording medium (such as a flexible disk, CD-ROM, or DVD-ROM), and installing the computer program on a computer. Alternatively, the sound quality improving device 1 may be configured by storing the computer program in a storage device of a server device on a communication network such as the Internet, and downloading the program to an ordinary computer system.

[0063] When the functions of a computer are realized by dividing the functions between an OS (operating system) and an application program, or by the OS and the application program working together, only the application program portion may be stored on a recording medium or storage device.

[0064] It is also possible to superimpose a computer program on a carrier wave and distribute it over a communications network. For example, the computer program may be posted on a bulletin board system (BBS) on the communications network and distributed over the network. The computer program may then be started and executed under the control of an operating system in the same way as any other application program, thereby enabling the above-mentioned processing to be performed.

[0065] This invention allows various embodiments and modifications without departing from the broad spirit and scope of this invention. Furthermore, the above-described embodiments are intended to explain this invention and do not limit the scope of this invention. That is, the scope of this invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and the meaning of the invention equivalent thereto are considered to be within the scope of this invention. [Industrial Applicability]

[0066] The present invention can be applied to separating sounds from an audio signal containing a mixture of multiple sounds and improving the sound quality. [Explanation of symbols]

[0067] 1 sound quality improvement device (information processing device), 2A, 2B, 2C microphones, 3 speaker, 10 separation unit, 20 estimation unit, 21 frequency conversion unit, 22 learning unit, 23 frequency conversion unit, 24 approximation calculation unit, 25 inverse frequency conversion unit, 26 output signal selection unit, 30 internal bus, 31 CPU, 32 main memory, 33 external memory, 34 operation unit, 35 display, 36 device interface, 39 program, O1 first sound source, O2 second sound source, O3 third sound source

Claims

1. a separation unit that performs independent component analysis on a first acoustic signal obtained by recording a composite sound of sounds emitted from a plurality of sound sources at different positions, and separates the composite sound into second acoustic signals for each sound source; an estimation unit that performs nonnegative matrix factorization on the first acoustic signal and the second acoustic signal, respectively, to estimate a third acoustic signal corresponding to a sound source for at least one of the plurality of sound sources; A sound quality improvement device comprising:

2. non-negative matrix factorization is a process of decomposing the first acoustic signal or the second acoustic signal into a basis matrix representing a spectral pattern thereof and an activation matrix representing a temporal intensity change of the spectral pattern; The estimation unit performing non-negative matrix factorization using the second acoustic signal corresponding to a first sound source among the plurality of sound sources as a teacher signal to estimate the basis matrix corresponding to the first sound source; performing non-negative matrix factorization on the first acoustic signal, and estimating an activation matrix corresponding to the first sound source and a basis matrix and an activation matrix corresponding to a second sound source among the plurality of sound sources based on the estimated basis matrix; generating the third acoustic signal corresponding to the first sound source based on the basis matrix and activation matrix corresponding to the estimated first sound source, or generating the third acoustic signal corresponding to the second sound source based on the basis matrix and activation matrix corresponding to the estimated second sound source; The sound quality improving device according to claim 1 .

3. A sound quality improvement method executed by an information processing device, performing independent component analysis on a first acoustic signal obtained by recording a composite sound of sounds emitted from a plurality of sound sources at different positions, and separating the first acoustic signal into second acoustic signals for each sound source; performing non-negative matrix factorization on the first acoustic signal and the second acoustic signal, respectively, to estimate a third acoustic signal corresponding to a sound source for at least one of the plurality of sound sources; How to improve sound quality.

4. Computer, a separation unit that performs independent component analysis on a first acoustic signal obtained by recording a composite sound of sounds emitted from a plurality of sound sources at different positions, and separates the composite sound into second acoustic signals for each sound source; an estimation unit that performs nonnegative matrix factorization on the first acoustic signal and the second acoustic signal, respectively, to estimate a third acoustic signal corresponding to a sound source for at least one of the plurality of sound sources; A program that functions as a