Learning method, sound source separation method, learning device, sound source separation system, and program
By training a model with both acoustic and video signals to incorporate speaker characteristics, the method enhances sound source separation accuracy by reducing distortion and residual interference.
Patent Information
- Application Number
- JP2024104541
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2041-02-08
AI Technical Summary
Conventional sound source separation methods, including permutation invariant training and multimodal speech separation, fail to consider speaker characteristics in the training process, leading to residual speech and speech distortion in the separated signal, thereby degrading separation accuracy.
A model is trained using a mixed acoustic signal and a sound source video signal to estimate a separated signal, incorporating speaker characteristics such as timing and phonetic information by learning the difference between features of the separated signal and teacher sound source video signal.
This approach improves sound source separation accuracy by considering speaker characteristics, reducing distortion and residual interfering sounds in the separated signal.
Smart Images

Figure 0007803372000010 
Figure 0007803372000011 
Figure 0007803372000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to sound source separation technology, and more particularly to multimodal sound source separation. [Background technology]
[0002] In single-channel source separation techniques, which estimate the pre-mixing speech signals of each speaker from a mixed signal of speech from multiple speakers observed with a single microphone, it is common to simultaneously estimate all source signals contained in the mixed signal using a neural network. The estimated source signals are called the separated signals. In this framework, the output order of signals corresponding to each speaker in the separated signal is arbitrary, so subsequent processing such as speaker identification is required to extract the speech of a specific speaker. Furthermore, when training the neural network model parameters, it is necessary to calculate the error between the separated signal and the pre-mixing source signal for each speaker and evaluate the overall error from these. Here too, there is a problem in that the error cannot be determined unless the correspondence between the separated signal and the source signal for each speaker is established. This problem is known as the permutation problem.
[0003] In response to this, permutation invariant training (PIT) has been proposed, which calculates the error between all correspondences between the source signal corresponding to each speaker and the separated signal elements, and optimizes the network model parameters based on this to minimize the overall error (see, for example, Non-Patent Document 1). Multimodal speech separation has also been proposed, in which facial images of each speaker are input simultaneously with a mixed speech signal, and the output order of the signals corresponding to each speaker contained in the separated signal is uniquely determined from each speaker's video (see, for example, Non-Patent Documents 2 and 3). Multimodal speech source separation solves the permutation problem by using each speaker's video, while taking into account the timing and content of speech during separation, and has been confirmed to demonstrate higher performance than speech separation that uses only sound. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] D. Yu, M. Kolbak, Z. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multitalker speech separation,” in Proc. ICASSP, 2017, pp. 241-245. [Non-patent document 2] R. Lu, Z. Duan, and C. Zhang, “Audio-visual deep clustering for speech separation,” IEEE / ACM Trans. ASLP, vol. 27, no. 11, pp. 1697-1712, 2019. [Non-patent document 3] A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, WT Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,” ACM Trans. Graph., vol. 37, no. 4, pp. 112:1-112:11, 2018. Summary of the Invention [Problem to be solved by the invention]
[0005] However, conventional PIT and multimodal source separation train model parameters by only considering the distance between the source signal and the separated signal in the sound domain. This training method cannot directly consider the speaker characteristics contained in the separated signal (e.g., speaker identity and phonetic information). This leads to residual speech from other speakers and speech distortion in the separated signal, degrading separation accuracy.
[0006] Such a problem is not limited to the case of performing sound source separation, but is common to the case of performing sound source separation of any sound.
[0007] The present invention has been made in view of the above points, and has an object to improve the separation accuracy of sound sources. [Means for solving the problem]
[0008] A mixed acoustic signal representing a mixed sound of sounds emitted from multiple sound sources and a sound source video signal representing images of at least some of the multiple sound sources are applied to a model, and a separated signal including a signal representing a target sound emitted from a certain sound source among the multiple sound sources is estimated. This model is obtained by learning based on the difference between the features of the separated signal and the features of the teacher sound source video signal, which are obtained by applying at least a teacher mixed acoustic signal, which is teacher data for the mixed acoustic signal, and a teacher sound source video signal, which is teacher data for the sound source video signal, to the model. [Effects of the Invention]
[0009] This allows the characteristics of the sound source contained in the separation signal, which are expressed in the characteristics of the sound source video signal, to be taken into consideration in the sound source separation, thereby improving the separation accuracy of the sound source separation. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram illustrating a functional configuration of a sound source separation device according to an embodiment. [Figure 2] FIG. 2 is a block diagram illustrating the functional configuration of the learning device according to the embodiment. [Figure 3] FIG. 3 is a block diagram illustrating the hardware configuration of the device. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [First embodiment] This embodiment introduces a function for performing multimodal sound source separation that takes into account the characteristics of the separated signals. This reduces distortion and residual interfering sounds contained in the separated signals, thereby improving the accuracy of sound source separation. A key feature of this embodiment is that a model for estimating the separated signals is trained based on at least the differences between the characteristics of the separated signals and the characteristics of a teacher sound source video signal, which is training data for a sound source video signal representing the video of the sound source. The video of the sound source contains elements that are closely related to the characteristics of the sound emitted from each sound source. For example, if the sound source is a speaker, the video of the speaker (e.g., a video including a video of the speaker's face) contains elements such as the timing of speech, phonetic information estimated from the mouth, and speaker information such as gender and age, which are closely related to the characteristics of the sound source signal. Furthermore, the video of the sound source is not affected by ambient sounds (e.g., noise), and these elements do not deteriorate even in high-noise environments. Therefore, in this embodiment, the features of the separated signals are matched to the features of the training sound source video signal, and a model for estimating the separated signals is trained based on the differences between these features, and sound source separation is performed using this model. In other words, the speaker information obtained from the image signal and the throat movements and mouth used for speaking are used to estimate the likelihood of the signal being emitted and used for sound source separation. In other words, the process of generating a sound signal is obtained from the image signal to perform sound source separation.
[0012] That is, in this embodiment, a mixed audio signal representing a mixed sound of sounds emitted from multiple audio sources and an audio source video signal representing images of at least some of the multiple audio sources are applied to a model to estimate a separated signal including a signal representing a target sound emitted from one of the multiple audio sources. However, the model is obtained by learning based on the difference between the features of the separated signal and the features of the audio source video signal, which are obtained by applying at least the teacher mixed audio signal, which is training data for the mixed audio signal, and the teacher audio source video signal, which is training data for the audio source video signal. In this way, by explicitly incorporating the relationship between the features of the separated signal and the features of the audio source video signal into the model learning for audio source separation, it becomes possible to perform audio source separation taking into account factors such as, for example, phonological information estimated from the timing of speech and the mouth, and speaker information such as gender and age. This makes it possible to consider features that were previously not possible in multimodal audio source separation, thereby reducing, for example, distortion and residual interfering sounds in the separated signal and improving the separation accuracy of audio source separation.
[0013] The present embodiment will be described in detail below. In the present embodiment, the model is a neural network, the sound source is a speaker, and the sound is speech. However, this does not limit the present invention. <Configuration> As illustrated in FIG. 1, the sound source separation device 11 of this embodiment has a storage unit 110, an audio stream processing unit 111, a video stream processing unit 112, a fusion unit 113, a separated signal estimation unit 114, and a control unit 116, and performs each of the processes described below based on the control of the control unit 116. Although a detailed description will be omitted, data obtained in the sound source separation process is stored in the storage unit 110 and is read out and used as needed. As illustrated in FIG. 2, the learning device 12 has a storage unit 120, an audio stream processing unit 121, a video stream processing unit 122, a fusion unit 123, a separated signal estimation unit 124, a separated signal feature estimation unit 125, a control unit 126, and a parameter update unit 127, and performs each of the processes described below based on the control of the control unit 126. Although a detailed description will be omitted, data obtained in the learning process is stored in the storage unit 120 and is read out and used as needed.
[0014] <Sound source separation processing (multimodal sound source separation processing)> The sound source separation process of this embodiment will be described with reference to FIG. Input: Mixed acoustic signal X={x1,...,x T} Sound source video signal V={V1,...,V N} Model parameter θ a ,θ v ,θ f ,θ s Output: Separated signal Y={Y1,...,Y N} The sound source separation device 11 of this embodiment separates a mixed sound signal X={x1,...,x T} and a sound source video signal V={V1,...,V N} and model parameters (sound stream model parameters θ a , model parameters for video streams θ v , fusion model parameters θ f , model parameters θ for estimating separated signals s ) and the model parameters θ a ,θ v ,θ f ,θ s The mixed audio signal X and the sound source video signal V are applied to a neural network (model) determined based on the above, and a separated signal Y={Y1,...,Y N} is estimated and output. The neural network is obtained by learning based on the difference between the feature Y of the separated signal, which is obtained by applying at least the teacher mixed audio signal X', which is the teacher data of the mixed audio signal X, and the teacher audio-source video signal V', which is the teacher data of the audio-source video signal V, to the neural network, and the feature Y of the teacher audio-source video signal, which is the teacher data of the audio signal S. (That is, the neural network is obtained by learning the model parameters θ a , θv , θ f , θ s The details of this learning process will be described later.
[0015] The multiple sound sources include sound sources that present an appearance correlated with the sound they emit. Examples of sound sources that present an appearance correlated with the sound they emit include speakers, animals, plants, natural objects, natural phenomena, machines, etc. As an example, this embodiment illustrates a case where the multiple sound sources include multiple speakers that are different from each other. All of the multiple sound sources may be speakers, or only some of them may be speakers. When a sound source is a speaker, the sound emitted from the speaker is a voice, the mixed audio signal X includes an audio signal representing the voice, and the sound source video signal V represents a video of the speaker.
[0016] The mixed acoustic signal X may be, for example, a time waveform signal (i.e., a time domain signal) obtained by digitally converting an acoustic signal obtained by observing a mixed sound with an acoustic sensor such as a microphone, or may be a time-frequency domain signal obtained by converting the time waveform signal into the frequency domain for each predetermined time interval (for example, an interval determined by a window function multiplied by the time waveform signal). Examples of time-frequency domain signals include an amplitude spectrogram obtained by converting a time waveform using a short-time Fourier transform or a logarithmic Mel filter bank output. Since amplitude spectrograms and logarithmic Mel filter banks are well known, a description thereof will be omitted. In this embodiment, the mixed acoustic signal X is expressed as X={x1,...,x T} where T is a positive integer representing the time frame length, and x t is an element of the mixed acoustic signal X of the t-th frame, where t=1,...,T is a positive integer representing the frame index. In other words, the mixed acoustic signal X is a time-series discrete acoustic signal.
[0017] The sound source video signal V is a video signal obtained by capturing an image of a sound source using a video sensor such as a webcam or a smartphone camera. For example, the sound source video signal V represents a video of each of multiple sound sources. The sound source video signal V may represent all of the images of the multiple sound sources described above, or may represent only the images of some of the sound sources. For example, the sound source video signal V may represent a video of one or more sound sources emitting a target sound among the multiple sound sources described above, or may represent a video of the sound source emitting the target sound and other sound sources, respectively. For example, if the sound source is a speaker, the sound source video signal V may represent a video of one or more speakers emitting the target sound among the multiple speakers, or may represent a video of the speaker emitting the target sound and other speakers, respectively. The sound source video signal V represents a video of the sound source that has an appearance correlated with the sound it emits. For example, if the sound source is a speaker, the sound source video signal V represents a video including a video of the speaker's face. In this embodiment, the sound source video signal V is expressed as V={V1,...,V N} where V n ={v n1 ,...,v nF} represents the video signal of the nth sound source (e.g., a speaker), and v nf represents the video signal of the fth frame of the nth sound source. n=1,...,N is a positive integer representing the index of the sound source, and N is an integer equal to or greater than 1 representing the number of sound sources (for example, sound source separation processing is practical when N is an integer equal to or greater than 2), and f=1,...,F is a positive integer representing the index of the video frame, and F is a positive integer representing the number of video frames. nf The number of channels, number of pixels, and fps are arbitrary. For example, a grayscale image with one channel and a resolution of the entire face of 224 pixels x 224 pixels at 25 fps is nf In this example, grayscale is used to reduce the resources used for calculations, but RGB images are also perfectly acceptable.
[0018] In this embodiment, the acoustic signals representing the sounds emitted from the multiple sound sources before mixing are expressed as S={S1,...,S N} where S n ={s n1,...,s nT} represents the acoustic signal emitted from the nth sound source, and s nt represents the acoustic signal of the t-th frame of the sound emitted from the n-th sound source. The separated signal Y is an estimated signal of the acoustic signal S. In this embodiment, the separated signal Y is expressed as Y={Y1,...,Y N} where Y n ={y n1 ,...,y nT} is the acoustic signal S emitted from the nth sound source n ={s n1 ,...,s nT}, and y nt is the acoustic signal s of the t-th frame of the sound emitted from the n-th sound source. nt is the estimated signal of Y1,...,Y N Any or all of Y1,...,Y correspond to the aforementioned "signal representing the target sound (signal representing the target sound emitted from a certain sound source among multiple sound sources)." N Which of the signals represents the target sound depends on the application of the separated signals. Note that the acoustic signal S and the separated signal Y may be time waveform signals (i.e., time domain signals) or time-frequency domain signals such as amplitude spectrograms or logarithmic Mel filter bank outputs.
[0019] <Overall flow of sound source separation processing> As shown in FIG. 1, the storage unit 110 stores model parameters θ obtained by a learning process described later. a ,θ v ,θ f ,θ s The model parameters for the sound stream θ a is input to the audio stream processing unit 111, and the video stream model parameters θ v is input to the video stream processing unit 112, and the fusion model parameters θ f is input to the fusion unit 113, and the separated signal estimation model parameters θ s is input to the separated signal estimation unit 114. The sound stream processing unit 111 calculates the sound stream model parameters θ aBased on this, the input mixed acoustic signal X is converted into the mixed acoustic signal embedding vector C a The video stream processing unit 112 obtains and outputs the video stream model parameters θ v Based on this, the input audio-video signal V is converted into an embedding vector C v The fusion unit 113 obtains and outputs the fusion model parameters θ f Based on this, the embedding vector C of the input mixed acoustic signal is a and the embedding vector C of the audio and video signals v The separated signal estimation unit 114 obtains and outputs the embedding vector M of the sound source signal from the separated signal estimation model parameters θ s Based on this, a separated signal Y is obtained from the embedding vector M of the input sound source signal and the mixed acoustic signal X, and is output. This will be described in detail below.
[0020] <Processing of the sound stream processing unit 111 (step S111)> Input: Mixed acoustic signal X={x1,...,x T} Model parameters for sound streams θ a Output: Embedding vector C of the mixed acoustic signal a The sound stream processing unit 111 receives the mixed sound signal X and sound stream model parameters θ a The model parameters for the sound stream θ a may be input and set in advance to the sound stream processing unit 111, or may be input each time the mixed sound signal X is input. The sound stream processing unit 111 receives the mixed sound signal X and the sound stream model parameters θ a From the embedded vector C of the mixed acoustic signal a is estimated and output. This embedding vector C a represents the features of the mixed acoustic signal X, and for example, k is a manually determined number of dimensions greater than or equal to one. a It is expressed as a sequence of vectors with continuous or discrete values. a For example, the embedding vector C aThe sequence length of the embedded vector C is the same as that of the mixed acoustic signal X. a For example, T×k a or k a The audio stream processing unit 111 calculates the embedding vector C according to, for example, the following equation (1): a Estimate.
number
[0021] <Processing of Video Stream Processing Unit 112 (Step S112)> Input: Audio source video signal V = {V1,...,V N} Model parameter θ for video stream v Output: Embedding vector C of audio source and video signal V ={C V 1,...,C V N} The video stream unit 112 receives the sound source video signal V and the video stream model parameters θ v The model parameters for the video stream θ v may be input and set in advance to the video stream processing unit 112, or may be input each time the sound source video signal V is input. The video stream processing unit 112 receives the sound source video signal V and the video stream model parameters θ v From the embedding vector C of the audio source and video signal V={C V 1,...,C V N} is estimated and output. This embedding vector C V represents the characteristics of the audio-video signal V, and C V n (where n=1,...,N) represents the video characteristics of the nth sound source (the speaker in this embodiment). For example, C V n is a manually determined arbitrary number of dimensions k V n k V 1,...,k V N may be the same as each other, or at least some may be different from each other. V n For example, 1792. Also, C V n The sequence length of C is the same as that of the mixed acoustic signal, T. V n For example, T×k V n or k V n ×T matrix. β γ In this case, the subscript "γ" should be placed directly below the superscript "β", but in this specification, due to limitations on notation, the subscript "γ" may be written diagonally below and to the right of the superscript "β". The video stream unit 112 may calculate this embedding vector C V Estimate.
number
[0022] <<Processing of Fusion Unit 113 (Step S113)>> Input: Embedding vector C of mixed acoustic signals a Embedding vector C of audio and video signals V Fusion model parameter θ f Output: Embedding vector of the source signal M={M1,...,M N} The fusion unit 113 receives the embedded vector C of the mixed acoustic signal. a , the embedding vector C of the audio source and video signal V , and the fusion model parameters θ f The fusion model parameters θ f may be input and set in advance to the fusion unit 113, or may be input and set in advance to the embedding vector C a and the embedding vector C of the audio and video signals V The fusion unit 113 may input the embedding vector C of the mixed acoustic signal. a , the embedding vector C of the audio source and video signal V , and the fusion model parameters θ f From the source signal embedding vector M={M1,...,M N} is estimated and output. This embedding vector M is the embedding vector C of the mixed acoustic signal. a and the embedding vector C of the audio and video signals V Here, M n (where n=1,...,N) represents the element corresponding to the n-th sound source (speaker in this embodiment) of the embedding vector M of the sound source signal. For example, M n is a manually determined arbitrary number of dimensions km n k m 1,...,k m N may be the same as each other, or at least some may be different from each other. m n For example, 1792. Also, M n The sequence length of M is the same as that of the mixed acoustic signal. n For example, T×k m n or k m n ×T matrix. The fusion unit 113 estimates the embedding vector M of this sound source signal, for example, according to the following equation (3).
number
[0023] <<Processing of Separated Signal Estimation Unit 114 (Step S114)>> Input: Source signal embedding vector M={M1,...,M N} Mixed acoustic signal X={x1,...,x T} Model parameters θ for estimating separated signals s Output: Separated signal Y={Y1,...,Y N} The separated signal estimation unit 114 receives the embedding vector M of the sound source signal, the mixed acoustic signal X, and the separated signal estimation model parameters θ s The model parameters for estimating the separated signals θ s may be input and set in advance to the separated signal estimation unit 114, or may be input each time the embedding vector M of the sound source signal and the mixed acoustic signal X are input. The separated signal estimation unit 114 receives the embedding vector M of the sound source signal, the mixed acoustic signal X, and the separated signal estimation model parameters θ s From the separated signal Y={Y1,...,Y N The separated signal estimation unit 114 estimates and outputs the separated signal Y, for example, according to the following equation (4).
number
[0024] <Learning processing (multimodal learning processing)> The learning process of this embodiment will be described with reference to FIG. Input: Supervised mixed acoustic signal X' = {x1',...,x T '} Teacher sound source video signal V'={V1',...,V N '} Teacher sound source signal S={S1,...,S N} Output: Model parameters θ for the sound stream a , model parameters for video streams θ v , fusion model parameters θ f , model parameters θ for estimating separated signals s , and the model parameters θ for estimating the separated signal features avc The learning device 12 of this embodiment at least generates training data for the mixed acoustic signal X, which is a training mixed acoustic signal X′={x1′,...,x T '} and the teacher sound-source video signal V'={V1',...,V N '} to the neural network (model), the separated signal Y'={Y1',...,Y N '} and the teacher audio and video signals V'={V1',...,V N '} and the model parameters θ a ,θ v ,θ f ,θ s ,θ avc For example, the learning device 12 obtains and outputs a teacher sound source video signal S corresponding to at least a first sound source among the plurality of sound sources. n and the separated signal Y corresponding to the second sound source different from the first sound source. n' The degree of similarity between the teacher sound source video signal S corresponding to the first sound source becomes smaller. n and the separated signal Y corresponding to the first sound source. n In this embodiment, learning is performed so that the degree of similarity between the elements representing the features of the separated signal Y' and the features of the teacher sound-source video signal V' is increased. In addition, in addition to the features of the separated signal Y' and the features of the teacher sound-source video signal V', a model parameter θ is obtained by learning based on the difference between the separated signal Y' obtained by applying the teacher mixed sound signal X' and the teacher sound-source video signal V' to a neural network (model) and the teacher sound-source signal S which is teacher data of the separated signal corresponding to the teacher mixed sound signal X' and the teacher sound-source video signal V'. a ,θ v ,θ f ,θ s ,θ avc The following example shows how to obtain and output the above data. However, this does not limit the present invention.
[0025] Teacher mixed acoustic signal X'={x1',...,x T '} is the mixed acoustic signal X={x1,...,x T}, and the training mixed acoustic signal X'={x1',...,x T'} is the data format of the mixed acoustic signal X={x1,...,x T There are multiple teacher mixed acoustic signals X', and the multiple teacher mixed acoustic signals X' may or may not include the mixed acoustic signal X that is the input for the sound source separation processing.
[0026] Teacher sound source video signal V'={V1',...,V N '} is the audio source video signal V={V1,...,V N}, and the teacher audio and video signal V'={V1',...,V N '} is the data format of the audio / video signal V={V1,...,V N There are multiple teacher sound source video signals V', and the multiple teacher sound source video signals V' may or may not include the sound source video signal V that is the input for the sound source separation processing.
[0027] Teacher sound source signal S={S1,...,S N} is the training mixed acoustic signal X'={x1',...,x T '} and teacher audio and video signals V'={V1',...,V N The training sound source signal S is an acoustic signal representing the unmixed sounds emitted from multiple sound sources corresponding to the training sound source signal S = {S1,...,S N} is the training mixed acoustic signal X'={x1',...,x T '} and teacher audio and video signals V'={V1',...,V N In the learning process, the following process is performed on each of the corresponding pairs of teacher mixed acoustic signal X', teacher sound source video signal V', and teacher sound source signal S.
[0028] <Overall learning process flow> As shown in FIG. 2, the learning device 12 receives corresponding teacher mixed acoustic signals X′={x1′,...,x T '} and teacher audio and video signals V'={V1',...,V N '} and the teacher sound source signal S={S1,...,S N} are input. The teacher mixed audio signal X' is input to the sound stream processing unit 121 and the separated signal estimation unit 124, the teacher sound source video signal V' is input to the video stream processing unit 122, and the teacher sound source signal S is input to the parameter update unit 127. The sound stream processing unit 121 calculates the sound stream model parameters θ a The provisional model parameters θ a Based on the input mixed acoustic signal X', we obtain the mixed acoustic signal embedding vector C a The video stream processing unit 122 obtains and outputs the video stream model parameters θ v The provisional model parameters θ v Based on the input teacher audio-video signal V', the audio-video signal embedding vector C v The fusion unit 123 obtains and outputs the fusion model parameters θ f The provisional model parameters θ f ', the embedding vector C of the input mixed acoustic signal a ' and the embedding vector C of the audio and video signals v The separated signal estimation unit 124 obtains and outputs an embedding vector M' of the sound source signal from the separated signal estimation model parameters θ s The separated signal feature estimation unit 125 obtains and outputs a separated signal Y' from the embedding vector M' of the input sound source signal and the teacher mixed acoustic signal X' based on the separated signal feature estimation unit 125. avc Based on the input source signal embedding vector M', the separated signal embedding vector C avc The parameter update unit 127 obtains and outputs the error between the teacher sound source signal S and the separated signal Y′, and the embedding vector C of the sound source video signal. v ' and the embedding vector C of the separated signal avc The provisional model parameters θ a ',θ v ',θ f ',θ s ',θ avc By repeating this process, the provisional model parameters θ a ',θ v',θ f ',θ s ',θ avc ' is the model parameter θ a ,θ v ,θ f ,θ s ,θ avc This is explained in detail below.
[0029] <<Initial Setting Process of Parameter Update Unit 127 (Step S1271)>> The parameter update unit 127 updates the model parameters θ a ,θ v ,θ f ,θ s ,θ avc The provisional model parameters θ a ',θ v ',θ f ',θ s ',θ avc The initial value of the provisional model parameter θ′ is stored in the storage unit 120. a ',θ v ',θ f ',θ s ',θ avc The initial value of ' can be anything.
[0030] <Processing of the Sound Stream Processing Unit 121 (Step S121)> Input: Supervised mixed acoustic signal X' = {x1',...,x T '} Provisional model parameters θ for the sound stream a ' Output: Embedding vector C of the mixed acoustic signal a ' The sound stream processing unit 121 receives the input teacher mixed sound signal X′ and the sound stream provisional model parameters θ a The sound stream processing unit 121 receives the teacher mixed sound signal X′ and the sound stream temporary model parameters θ a ', the embedding vector C of the mixed acoustic signal a ' is estimated and output. This estimation process is a ,C a is X',θa ',C a ', this is the same as the process (equation (1)) of the sound stream processing unit 111 (step S111) described above.
[0031] <Processing of Video Stream Processing Unit 122 (Step S122)> Input: Teacher audio and video signal V' = {V1',...,V N '} Provisional model parameters θ for video stream v ' Output: Embedding vector C of audio source and video signal V '={C V 1',...,C V N '} The video stream unit 122 receives the input teacher sound source video signal V′ and the video stream provisional model parameters θ v The video stream unit 122 receives the teacher sound source video signal V′ and the video stream provisional model parameters θ v ', the embedding vector C of the audio source and video signal V '={C V 1',...,C V N '} is estimated and output. This estimation process is v ,C V is V',θ v ',C V ' are replaced with ', the process is the same as the process (Equation (2)) of the video stream processing unit 112 described above (step S112).
[0032] <<Processing of Fusion Unit 123 (Step S123)>> Input: Embedding vector C of mixed acoustic signals a ' Embedding vector C of audio and video signals V ' Provisional model parameters for fusion θ f ' Output: Embedding vector of the source signal M'={M1',...,M N '} The fusion unit 123 includes an embedding vector C of the mixed acoustic signal. a ', the embedding vector C of the audio source and video signal V ', and the temporary fusion model parameters θ f The fusion unit 123 receives the embedded vector C a ', the embedding vector C of the audio source and video signal V ', and the model parameters for fusion θ f ', the embedding vector of the sound source signal M'={M1',...,M N The data format of the embedding vector M' of the sound source signal is the same as the data format of the embedding vector M of the sound source signal. a ,C V ,θ f ,M={M1,...,M N} is C a ',C V ',θ f ',M'={M1',...,M N '}, this is the same as the process (Equation (3)) of the merging unit 113 (step S113) described above.
[0033] <<Processing of Separated Signal Estimation Unit 124 (Step S124)>> Input: Embedding vector of source signal M'={M1',...,M N '} Teacher mixed acoustic signal X'={x1',...,x T '} Temporary model parameters θ for estimating separated signals s ' Output: Separated signal Y'={Y1',...,Y N '} The separated signal estimation unit 124 receives the embedding vector M′ of the sound source signal, the teacher mixed acoustic signal X′, and the provisional model parameters θ for estimating the separated signals read from the storage unit 120. s The separated signal estimation unit 114 receives the embedding vector M′ of the sound source signal, the teacher mixed acoustic signal X′, and the provisional model parameters θ s ', the separated signal Y'={Y1',...,YN '} is estimated and output. Separation signal Y'={Y1',...,Y N '} data format is the above-mentioned separated signal Y={Y1,...,Y N}. This estimation process is performed for M={M1,...,M N},X={x1,...,x T},θ s ,Y={Y1,...,Y N} is M'={M1',...,M N '},X'={x1',...,x T '},θ s ',Y'={Y1',...,Y N '}, this is the same as the process (equation (4)) of the separated signal estimation unit 114 (step S114) described above.
[0034] <Processing of Separated Signal Feature Estimation Unit 125 (Step S125)> Input: Embedding vector of source signal M'={M1',...,M N '} Temporary model parameters θ for estimating separated signal features avc ' Output: Embedding vector C of the separated signals avc '={C avc 1',...,C avc N '} The separated signal feature estimation unit 125 receives the embedding vector M′ of the sound source signal and the provisional model parameters θ for estimating the separated signal features read from the storage unit 120. avc The separated signal feature estimation unit 125 receives the embedding vector M′ of the sound source signal and the provisional model parameters θ avc ', the embedding vector C of the separated signal avc '={C avc 1',...,C avc N '} is estimated and output. Here, the embedding vector C avc ' represents the characteristics of the separated signal Y', and C avc n ' (where n=1,...,N) is the nth separated signal Y n'. For example, C avc n ' is a manually determined arbitrary number of dimensions k avc n k avc 1,...,k avc N may be the same as each other, or at least some may be different from each other. avc n For example, 1792. Also, C avc n The sequence length of C' is the same as that of the mixed acoustic signal, T. avc n ' is, for example, T×k avc n or k avc n ×T matrix. The separated signal feature estimation unit 125 calculates the embedding vector C avc ' is estimated.
number
[0035] <<Processing of the parameter update unit 127 (step S1272)>> The parameter update unit 127 receives the input teacher sound source signal S={S1,...,S N}, the separated signal Y′={Y1′,...,Y N '}, the embedding vector C of the audio source video signal obtained in step S122 V '={C V 1',...,C V N'}, and the embedding vector C of the separated signal obtained in step S125 avc '={C avc 1',...,C avc N The parameter update unit 127 calculates the error (difference) between the teacher sound source signal S and the separated signal Y′, and the embedding vector C V ' and the embedding vector C of the separated signal avc Based on the error (inter-modal error) between ' and ', the provisional model parameters Θ = {θ a ',θ v ',θ f ',θ s ',θ avc Update '}.
number
number
number
number
[0036] <<End Condition Determination Process (Step S126)>> Next, the control unit 126 determines whether a predetermined termination condition is satisfied. There is no limitation on the termination condition, but the termination condition may be, for example, that the number of updates of the provisional model parameter Θ has reached a predetermined number, or that the update amount of the provisional model parameter Θ is within a predetermined range. If it is determined that the termination condition is not satisfied, the learning device 12 generates new teacher mixed acoustic signals X'={x1',...,x T '} and teacher audio and video signals V'={V1',...,V N '} and the teacher sound source signal S={S1,...,S N} is input, and the processes of steps S121, S122, S123, S124, S125, and S1272 are executed again. On the other hand, if it is determined that the termination condition is satisfied, the updated provisional model parameters Θ={θ a ',θ v',θ f ',θ s ',θ avc θ out of '} a ',θ v ',θ f ',θ s ' are the model parameters θ a ,θ v ,θ f ,θ s The output model parameters θ a ,θ v ,θ f ,θ s are stored in the storage unit 110 of the sound source separation device 11 (FIG. 1) and are used in the sound source separation process. a ,θ v ,θ f ,θ s The separated signals Y={Y1,...,Y N} is associated with the video of each sound source (speaker) and reflects the characteristics of the sound source (for example, speech timing, phoneme information, speaker information such as gender and age, etc.). Therefore, in this embodiment, multimodal sound source separation with high separation accuracy can be achieved.
[0037] [Hardware configuration] The sound source separation device 11 and the learning device 12 in each embodiment are devices configured by a general-purpose or dedicated computer having a processor (hardware processor) such as a CPU (central processing unit) and memories such as RAM (random-access memory) and ROM (read-only memory) executing a predetermined program. That is, the sound source separation device 11 and the learning device 12 each have processing circuitry configured to implement each unit possessed by the device. This computer may have one processor and memory, or may have multiple processors and memories. This program may be installed on the computer or may be pre-recorded in a ROM or the like. Furthermore, some or all of the processing units may be configured using electronic circuits that independently realize processing functions, rather than electronic circuits that realize functional configuration by loading a program like a CPU. Furthermore, the electronic circuits constituting one device may include multiple CPUs.
[0038] FIG. 3 is a block diagram illustrating the hardware configuration of the sound source separation device 11 and the learning device 12 in each embodiment. As illustrated in FIG. 3, the sound source separation device 11 and the learning device 12 in this example include a central processing unit (CPU) 10a, an input unit 10b, an output unit 10c, a random access memory (RAM) 10d, a read-only memory (ROM) 10e, an auxiliary storage device 10f, and a bus 10g. The CPU 10a in this example includes a control unit 10aa, a calculation unit 10ab, and a register 10ac, and executes various calculation processes according to various programs loaded into the register 10ac. The input unit 10b is an input terminal to which data is input, a keyboard, a mouse, a touch panel, or the like. The output unit 10c is an output terminal to which data is output, a display, a LAN card controlled by the CPU 10a that has loaded a predetermined program, or the like. The RAM 10d is a static random access memory (SRAM), a dynamic random access memory (DRAM), or the like, and has a program area 10da where a predetermined program is stored and a data area 10db where various data are stored. The auxiliary storage device 10f is a hard disk, a magneto-optical disc (MO), a semiconductor memory, or the like, and has a program area 10fa where a predetermined program is stored and a data area 10fb where various data are stored. The bus 10g connects the CPU 10a, the input unit 10b, the output unit 10c, the RAM 10d, the ROM 10e, and the auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f to the program area 10da of the RAM 10d in accordance with the loaded OS (Operating System) program. Similarly, the CPU 10a writes various data stored in the data area 10fb of the auxiliary storage device 10f to the data area 10db of the RAM 10d. The address on the RAM 10d where this program or data is written is stored in the register 10ac of the CPU 10a.The control unit 10aa of the CPU 10a sequentially reads these addresses stored in the register 10ac, reads programs and data from the areas on the RAM 10d indicated by the read addresses, causes the calculation unit 10ab to sequentially execute the calculations indicated by the programs, and stores the calculation results in the register 10ac. With this configuration, the functional configuration of the sound source separation device 11 and the learning device 12 is realized.
[0039] The above-mentioned program can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include non-transitory recording media. Examples of such recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories.
[0040] This program may be distributed, for example, by selling, transferring, or lending a portable recording medium, such as a DVD or CD-ROM, on which the program is recorded. Furthermore, the program may be distributed by storing the program in a storage device of a server computer and transferring the program from the server computer to other computers via a network. As described above, a computer that executes such a program may, for example, first temporarily store the program recorded on a portable recording medium or transferred from the server computer in its own storage device. Then, when executing a process, the computer reads the program stored in its own storage device and executes processing in accordance with the read program. Alternatively, the program may be executed by a computer that reads the program directly from a portable recording medium and executes processing in accordance with the program. Furthermore, the computer may execute processing in accordance with the received program each time a program is transferred from the server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the processing function simply by issuing an execution instruction and obtaining the results, thereby executing the processing described above through a so-called ASP (Application Service Provider) type service. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).
[0041] In each embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0042] It should be noted that the present invention is not limited to the above-described embodiment. For example, the sound source separation device 11 and the learning device 12 may be configured separately and connected via a network such as the Internet, and the model parameters may be provided from the learning device 12 to the sound source separation device 11 via the network. Alternatively, the model parameters may be provided from the learning device 12 to the sound source separation device 11 via a portable recording medium such as a USB memory without going via a network. Alternatively, the sound source separation device 11 and the learning device 12 may be configured integrally, and the model parameters obtained by the learning device 12 may be provided to the sound source separation device 11.
[0043] In addition, although a neural network is used as the model in this embodiment, this does not limit the present invention, and a probabilistic model such as a hidden Markov model or other models may also be used as the model.
[0044] In addition, in the present embodiment, the sound source is a speaker and the sound is speech, but this does not limit the present invention, and the sound source may include animals other than humans, plants, natural objects, natural phenomena, machines, etc., and the sound may include cries, friction sounds, vibration sounds, rain sounds, thunder sounds, engine sounds, etc.
[0045] Furthermore, the various processes described above may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capacity of the device executing the processes or as necessary. Needless to say, other modifications are possible within the scope of the present invention. [Explanation of symbols]
[0046] 11 Sound source separation device 12 Learning Device 111,121 Sound stream processing unit 112,122 Video Stream Processing Unit 113,123 Fusion part 114,124 Separated signal estimation unit 125 Separated signal feature estimation unit
Claims
1. A learning method using a learning device, A learning method comprising: a learning step of obtaining the model by learning based on the difference between the features of a separated signal including a signal representing a target sound emitted from a sound source among the multiple sound sources, and the features of the teacher sound source video signal, which is obtained by applying at least a teacher mixed sound signal, which is teacher data of a mixed sound signal representing a mixed sound of sounds emitted from multiple sound sources, and a teacher sound source video signal, which is teacher data of a sound source video signal representing an image of at least a part of the multiple sound sources, to a model.
2. The learning method of claim 1, The learning step obtains the model by learning such that at least a degree of similarity between an element representing a feature of the teacher sound source video signal corresponding to a first sound source among the plurality of sound sources and an element representing a feature of the separated signal corresponding to a second sound source different from the first sound source decreases, and a degree of similarity between an element representing a feature of the teacher sound source video signal corresponding to the first sound source and an element representing a feature of the separated signal corresponding to the first sound source increases.
3. The learning method according to claim 1 or 2, the learning step further comprises obtaining the model by learning based on a difference between the separated signal obtained by applying the teacher mixed acoustic signal and the teacher sound-source video signal to the model, and a teacher sound-source signal which is teacher data of the separated signal corresponding to the teacher mixed acoustic signal and the teacher sound-source video signal.
4. A learning method according to any one of claims 1 to 3, The sound source image signal represents an image of each of the plurality of sound sources.
5. 5. A learning method according to claim 1, The method of training, wherein the plurality of sound sources comprises a plurality of different speakers, the mixed acoustic signal comprises a speech signal, and the sound source video signal represents a video of the speakers.
6. The learning method of claim 5, The method of learning, wherein the source video signal represents a video including a facial video of the speaker.
7. 7. A learning method according to claim 1, A learning method, wherein the separated signals include a signal representing a target sound emitted from one of the plurality of sound sources and a signal representing a sound emitted from another sound source.
8. In a learning device, a learning step is performed to obtain the model by learning based on the difference between the features of a separated signal including a signal representing a target sound emitted from a sound source among the multiple sound sources and the features of the teacher sound source video signal, which is obtained by applying at least a teacher mixed sound signal, which is teacher data of a mixed sound signal representing a mixed sound of sounds emitted from multiple sound sources, and a teacher sound source video signal, which is teacher data of a sound source video signal representing an image of at least a part of the multiple sound sources, to a model; an estimation step in which, in the sound source separation device, a third mixed acoustic signal representing a mixed sound of sounds emitted from a plurality of third sound sources and a third sound source video signal representing an image of at least a part of the plurality of third sound sources are applied to the model, and a third separation signal including a signal representing a target sound emitted from a sound source among the plurality of third sound sources is estimated; A sound source separation method having the following.
9. A learning device having a learning unit that obtains the model by learning based on the differences between the features of a separated signal including a signal representing a target sound emitted from a sound source among a plurality of sound sources, and the features of the teacher sound source video signal, which is obtained by applying at least a teacher mixed sound signal that is teacher data for a mixed sound signal representing a mixed sound of sounds emitted from a plurality of sound sources and a teacher sound source video signal that is teacher data for a sound source video signal representing an image of at least a part of the plurality of sound sources to a model.
10. A learning device having a learning unit that obtains the model by learning based on the difference between the features of a separated signal including a signal representing a target sound emitted from a sound source among the multiple sound sources and the features of the teacher sound source video signal, the separated signal being obtained by applying at least a teacher mixed sound signal, which is teacher data of a mixed sound signal representing a mixed sound of sounds emitted from multiple sound sources, and a teacher sound source video signal, which is teacher data of a sound source video signal representing an image of at least a part of the multiple sound sources, to the model; a sound source separation device having an estimation unit that applies a third mixed acoustic signal representing a mixed sound of sounds emitted from a plurality of third sound sources and a third sound source video signal representing an image of at least a part of the plurality of third sound sources to the model, and estimates a third separated signal including a signal representing a target sound emitted from a sound source among the plurality of third sound sources; A sound source separation system having:
11. A program for causing a computer to execute the process of the learning method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image / sound processing system, information processing apparatus, image / sound processing method, and image / sound processing program
JP2015177490A
Speech dialization method and apparatus based on audio visual data
JP2020187346A
Method and system for enhancing a speech signal of a human speaker in a video using visual information
US20190005976A1
Audiovisual source separation and localization using generative adversarial networks
WO2020217165A1