Signal filtering device, signal filtering method and program
The signal filtering device enhances the accuracy of extracting a target audio signal from mixed audio by employing cross-modal learning and sound source separation techniques, effectively isolating the desired signal using concept embedding vectors and mask information generation.
Patent Information
- Application Number
- JP2023570545
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-10-29
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing signal filtering devices struggle to accurately extract a target audio signal from a mixed audio signal that includes signals other than the target audio signal.
A signal filtering device that employs cross-modal representation learning, target speaker extraction methods, and sound source separation techniques to enhance the accuracy of extracting a target audio signal by using concept embedding vectors and mask information generation.
Improves the accuracy of extracting a target audio signal from a mixed audio signal by effectively isolating the desired signal from other audio signals, utilizing techniques such as deep metric learning and audiovisual embedding networks.
Smart Images

Figure 0007761851000008 
Figure 0007761851000009 
Figure 0007761851000010
Abstract
Description
[Technical Field]
[0001] The present invention relates to a signal filtering device, a signal filtering method, and a program. [Background technology]
[0002] When multiple speakers speak, their voices may be mixed. The effect of being able to hear the voice of a selected speaker from a mixed voice is known as the cocktail party effect. Research is being conducted to realize this cocktail party effect using signal filtering devices.
[0003] Hereinafter, the audio signal may be a signal corresponding to a spoken language or a signal corresponding to the sound of a musical instrument or the like (acoustic signal). The signal filtering device extracts or removes a specific part or element of the audio signal input to the signal filtering device by filtering the audio signal. That is, the signal filtering device extracts or removes the audio signal targeted for extraction (hereinafter referred to as the "target audio signal") from the mixed audio signal.
[0004] The signal filtering device disclosed in Non-Patent Document 1 performs filtering based on the physical characteristics of the target speech signal, which are the direction of the sound source, the harmonic structure of the speech frequency components, the statistical independence of the speech signal, and the timbre proximity or consistency of the target speaker. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] K. Zmolikova, M. Delcroix, K. Kinoshita, T. Ochiai, T. Nakatani, L. Burget, and J. Cernocky, "SpeakerBeam: Speaker Aware Neural Network for Target Speaker Extraction in Speech Mixtures", IEEE Journal of Selected Topics in Signal Processing, vol.13, no.4, pp.800-814, 2019. Summary of the Invention [Problem to be solved by the invention]
[0006] However, there is a problem in that it is not possible to improve the accuracy of extracting a target audio signal from an audio signal in which the target audio signal is mixed with an audio signal other than the target audio signal.
[0007] In view of the above circumstances, the present invention aims to provide a signal filtering device, a signal filtering method, and a program that can improve the accuracy of extracting a target audio signal from an audio signal in which the target audio signal is mixed with audio signals other than the target audio signal. [Means for solving the problem]
[0008] One aspect of the present invention is a signal filtering device comprising: a separation unit that separates a predetermined number of candidate signals from a mixed signal as candidates for a target signal; an encoding unit that encodes related information of the target signal into a first feature vector and encodes the predetermined number of candidate signals into the predetermined number of second feature vectors; and a selection unit that derives a similarity between the first feature vector and the second feature vector for each of the candidate signals and selects, from the predetermined number of candidate signals, the candidate signal with the highest similarity as the target signal.
[0009] One aspect of the present invention is a signal filtering method executed by a signal filtering device, the signal filtering method including the steps of separating a predetermined number of candidate signals from a mixed signal as candidates for a target signal, encoding related information of the target signal into a first feature vector and encoding the predetermined number of candidate signals into a predetermined number of second feature vectors, and deriving a similarity between the first feature vector and the second feature vector for each of the candidate signals, and selecting, from the predetermined number of candidate signals, the candidate signal with the highest similarity as the target signal.
[0010] One aspect of the present invention is a program for causing a computer to function as the signal filtering device described above. [Effects of the Invention]
[0011] According to the present invention, it is possible to improve the accuracy of extracting a target audio signal from an audio signal in which the target audio signal is mixed with an audio signal other than the target audio signal. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a signal filtering device in a first embodiment. [Figure 2] 5 is a flowchart showing an example of the operation of the signal filtering device in the first embodiment. [Figure 3] FIG. 10 is a diagram illustrating an example of the configuration of a signal filtering device in a second embodiment. [Figure 4] FIG. 10 is a diagram showing an example of a similarity contour in the second embodiment. [Figure 5] 10 is a flowchart showing an example of the operation of the signal filtering device in the second embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of the configuration of a signal filtering device in a third embodiment. [Figure 7] 10 is a flowchart showing an example of the operation of the signal filtering device in the third embodiment. [Figure 8]4 shows examples of averaged signal-to-distortion ratio scores for a target speech signal in the first and second embodiments. [Figure 9] 10 shows an example of extraction of a target audio signal in the second embodiment. [Figure 10] 10 shows an example of signal-to-distortion ratio scores for each overlap rate in the second and third embodiments. [Figure 11] FIG. 2 is a diagram illustrating an example of a hardware configuration of a signal filtering device in each embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] (overview) Hereinafter, an audio signal in which the target audio signal is mixed with an audio signal other than the target audio signal will be referred to as a "mixed audio signal." Hereinafter, a function for extracting a target audio signal from a mixed audio signal based on a concept specified by a predetermined method will be referred to as a "Concept Beam." The predetermined method is not limited to a specific method, but may be, for example, a method of specification using an audio signal, a still image signal, a video signal (image signal), or a text signal (explanatory text signal). Furthermore, the target audio signal is a specific part or element of the mixed audio signal.
[0014] For example, a mixed speech signal from multiple speakers talking about different topics is input to the signal filtering device, and a signal for specifying a concept to be extracted (hereinafter referred to as a "concept specifying signal") is also input to the signal filtering device.
[0015] The signal filtering device extracts semantic information in the form of a multidimensional vector, i.e., concept information in the form of a multidimensional vector (hereinafter referred to as a "concept embedding vector"), from the concept designation signal. Speech language related to a concept (latent semantic information) designated using this concept designation signal may be included in the mixed speech signal. For example, waveform data (speech language) of the word "bicycle" related to a bicycle image in a frame of a still image as the concept designation signal may be included in the mixed speech signal.
[0016] The signal filtering device extracts a target speech signal by a speaker who is talking about the concept that is the target of extraction from the mixed speech signal. For example, when an image signal of a bicycle is input to the signal filtering device, the signal filtering device extracts a target speech signal by a speaker who is talking about the concept that is the target of extraction, "bicycle," from the mixed speech signal.
[0017] In the first and second embodiments described below, the signal filtering device applies cross-modal representation learning (Reference 1: D. Harwath, A. Recasens, D. Suris, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” International Journal of Computer Vision, 2019.) In this way, the signal filtering device represents a concept specified using a concept specification signal using a concept embedding vector (concept vector).
[0018] In the first and second embodiments described below, the signal filtering device applies a target speaker extraction method (Reference 2: M. Delcroix, K. Zmolikova, T. Ochiai, K. Kinoshita, and T. Nakatani, “Speaker activity driven neural speech extraction,” in Proc. ICASSP, 2021.) to extract a target speech signal from a mixed speech signal based on a concept represented using a concept embedding vector.
[0019] In a third embodiment described below, the signal filtering device applies a sound source separation technique (Reference 3: M. Kolbak, D. Yu, Z.-H. Tan, and J. Jensen, “Multi-talker Speech Separation with Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks,” IEEE / ACM Transactions on Audio, Speech and Language Processing, vol. 25, no. 10, pp. 1901-1913, 2017.) to extract a target speech signal from a mixed speech signal.
[0020] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail with reference to the drawings. In the following, symbols placed above letters in mathematical formulas are written immediately before the letters. For example, the symbol "^" placed above the letter "X" in a mathematical formula is written immediately before the letter "X" as "^X". For example, the symbol "-" placed above the letter "I" in a mathematical formula is written immediately before the letter "I" as "(-)I".
[0021] (First embodiment) FIG. 1 is a diagram illustrating an example of the configuration of a signal filtering device 1a according to the first embodiment. The signal filtering device 1a is a device that extracts a target audio signal from a mixed audio signal. The signal filtering device 1a extracts a target audio signal from a mixed audio signal that includes the target audio signal and audio signals other than the target audio signal by performing a filtering process on the mixed audio signal. In the first embodiment, the signal filtering device 1a uses, as an example, a concept embedding vector (image embedding vector) obtained using an audiovisual (image and audio) embedding network (neural network) as a clue for extracting the target audio signal from the mixed audio signal.
[0022] The signal filtering device 1a includes an acquisition unit 11, an information generation unit 12a, an extraction unit 13, and a mask processing unit 14. The information generation unit 12a includes an encoding unit 121a and a linear transformation unit 122. The extraction unit 13 includes a first extraction layer 131, a combination processing unit 132, and a second extraction layer 133a.
[0023] <Learning stage> The image embedding vector and the audio embedding vector are obtained based on a large amount of pair data of images and audio that describes the content of the images. In the learning stage before the estimation stage, the encoding unit 121a performs deep metric learning so that the image embedding vector and the audio embedding vector are located close to each other in the latent space (audiovisual embedding space).
[0024] The process of extracting a target speech signal from a mixed speech signal according to the speech of a speaker who is talking about the concept to be extracted (the content of a still image as a concept-specifying signal) is formulated as shown in Equation (1).
[0025]
number
[0026] Here, "Y∈C T×F " represents the mixed audio signal (input signal) in the short-time Fourier transform domain. "T" represents the number of frames per time in the mixed audio signal. "F" represents the number of frequency bins in the mixed audio signal. "^X k ∈C T×F ” represents the target speech signal of the kth speaker. “f(·)” is the concept-based target speech signal “^X k is a function that represents the process (ConceptBeam) of extracting " from the mixed audio signal "Y".
[0027] The parameters of the encoding unit 121a and the parameters of the extraction unit 13 may be learned simultaneously, but are learned independently to ensure stability. In order for the information generation unit 12a and the extraction unit 13 to perform deep learning, a set "{Y, X k ,C k} K k=1 " is required, where "X k " represents the reference speech signal associated with the target speech signal of the kth speaker. k " represents a concept-specific signal (e.g., a still image). "K" represents the total number of speakers associated with the mixed speech signal.
[0028] The information generator 12a has an audiovisual embedding network (see, for example, Reference 1). The information generator 12a generates an image feature vector (image feature information) based on a concept designation signal (still image) input to the audiovisual embedding network. The information generator 12a generates a concept embedding vector based on the image feature vector.
[0029] The encoding unit 121a uses an audiovisual embedding network to associate a time interval (segment) of an audio signal in which the name or appearance (concept to be extracted) representing an object in an image frame is described in audio language with the object through unsupervised learning.
[0030] In the first embodiment, the encoding unit 121a encodes the image “C k The globally pooled image feature vector "(-)I" as visual information obtained from an image encoder that encodes "(-)I" is used as the concept embedding vector "e". Here, the linear transformation unit 122 performs a linear transformation on the globally pooled image feature vector "(-)I". The linear transformation unit 122 generates a d'-dimensional vector obtained by the linear transformation as the concept embedding vector "e".
[0031] The information generator 12a may represent concepts that cross both the modalities of image (visual) and audio (auditory) using a concept embedding vector. That is, a cross-modal embedding vector may be used as the concept embedding vector. The cross-modal embedding vector may be, for example, an image and audio embedding vector.
[0032] The acquisition unit 11 acquires a mixed speech signal (input signal). The extraction unit 13 generates mask information based on the mixed speech signal and the concept embedding vector. The mask processing unit 14 extracts a target speech signal "X k The loss function in deep metric learning is the estimated target speech signal "^X k ” and the reference audio signal “X k " is a function that represents the mean square error between
[0033] <Estimation stage> The parameters of the information generator 12a (audiovisual embedding network) trained in the training stage are fixed in the estimation stage, and the parameters of the extractor 13 (each extraction layer) trained in the training stage are fixed in the estimation stage.
[0034] The acquisition unit 11 acquires a mixed audio signal (input signal). The encoding unit 121a acquires a concept-specific signal. The encoding unit 121a generates an image feature vector from the concept-specific signal using an audiovisual embedding network (see, for example, Reference 1).
[0035] The encoding unit 121a may convert information of different modalities (image and audio) into vectors in an embedding space (hereinafter referred to as a "shared embedding space") that can express the features of the different modalities. For example, if the concept designation signal is a still image or a video, the encoding unit 121a encodes the input concept designation signal into an image feature vector. For example, if the concept designation signal is audio, the encoding unit 121a encodes the input concept designation signal into an audio feature vector (audio feature information).
[0036] "I∈R H×W×d " represents the image feature map output from the image encoder of the encoding unit 121a. "A∈R T’×d " represents the audio feature map output from the audio encoder of the encoding unit 121a. When the concept designation signal is a still image, "(-)I" shown in equation (2) represents an image feature vector globally pooled in the spatial direction. When the concept designation signal is a moving image, "(-)I" represents an image feature vector globally pooled in the spatial direction or the temporal direction.
[0037]
number
[0038] Here, "I h,w,: " represents a d-dimensional vector (image feature vector) indicating the coordinates (h, w) in the image feature map. "H" represents the height of the image downsampled by the image encoder. "W" represents the width of the image downsampled by the image encoder. "(-)A" shown in equation (3) represents the audio feature vector globally pooled in the time direction.
[0039]
number
[0040] Here, "A t’,: " represents a d-dimensional vector (audio feature vector) indicating the t'th frame in the audio feature map. "T'" represents the number of time frames of the audio signal downsampled by the audio encoder.
[0041] The extraction unit 13 uses the concept embedding vectors derived based on these feature vectors in the shared embedding space for filtering the mixed audio signal.
[0042] The extraction unit 13 extracts a desired element or region from the mixed speech signal based on a concept embedding vector generated in response to a concept designation signal (target concept designator). The extraction unit 13 has an extraction network (a neural network for extraction). Based on the mixed speech signal "Y" and the concept embedding vector "e" input to the extraction network, the extraction unit 13 extracts a "time-frequency mask" representing the desired element or region using mask information "M k ∈R T×F " is generated.
[0043] The formation of mask information is, for example, k =g(Y, e)" where "g(·)" represents the extraction network. The first extraction layer 131 is the first bidirectional long short-term memory (BLSTM) layer (hidden layer) of the extraction network. The combination processor 132 multiplies the output of the first extraction layer 131 by the concept embedding vector "e" element by element. This results in a multiplication combination (see Reference 2) of the extraction result by the first extraction layer 131 and the concept embedding vector "e". The second extraction layer 133a extracts mask information from the result of the multiplication combination by the combination processor 132.
[0044] The mask processing unit 14 extracts the mask information “M k ” and the mixed audio signal “Y” element by element. k " is estimated.
[0045] Next, an example of the operation of the signal filtering device 1a will be described. 2 is a flowchart showing an example of the operation of the signal filtering device 1a in the first embodiment. The encoding unit 121a encodes the concept-specific signal into an image feature vector (d-dimensional vector) (step S101). The linear transformation unit 122 generates the linear transformation result of the image feature vector as a concept embedding vector (step S102). The extraction unit 13 extracts mask information from the mixed audio signal including the target audio signal based on the concept embedding vector (step S103). The mask processing unit 14 estimates the target audio signal from the mixed audio signal using the mask information (step S104).
[0046] As described above, the information generating unit 12a generates a concept embedding vector (feature information) of a concept designation signal (related information) of a target speech signal (target signal). The extracting unit 13 extracts mask information from a mixed speech signal (mixed signal) including the target speech signal based on the concept embedding vector. The mask processing unit 14 estimates the target speech signal from the mixed speech signal using the mask information.
[0047] Here, the information generating unit 12a encodes the concept designation signal (related information) into a d-dimensional vector (multidimensional vector). The information generating unit 12a generates the linear transformation result of the d-dimensional vector as a concept embedding vector (feature information).
[0048] This makes it possible to improve the accuracy of extracting the target audio signal from an audio signal (mixed audio signal) in which the target audio signal is mixed with audio signals other than the target audio signal.
[0049] (Second embodiment) The second embodiment differs from the first embodiment in that mask information is extracted from a mixed speech signal using concept activity information. The second embodiment will be described focusing on the differences from the first embodiment.
[0050] 3 is a diagram showing an example of the configuration of a signal filtering device 1b in the second embodiment. The signal filtering device 1b is a device that extracts a target audio signal from a mixed audio signal. The signal filtering device 1b extracts the target audio signal from the mixed audio signal by filtering the mixed audio signal including the target audio signal and audio signals other than the target audio signal.
[0051] The signal filtering device 1b includes an acquisition unit 11, an information generation unit 12b, an extraction unit 13, and a mask processing unit 14. The information generation unit 12b includes an encoding unit 121b, a similarity derivation unit 123, an auxiliary unit 124, and a weighted sum unit 125. The extraction unit 13 includes a first extraction layer 131, a combination processing unit 132, and a second extraction layer 133b.
[0052] The information generating unit 12b generates a similarity profile for the concept designation signal. The similarity profile is information that represents an audiovisual correspondence. For example, the similarity profile is information that represents the similarity between an image feature and an audio feature in time series. The similarity profile is expressed as the inner product of an image feature vector "I" and an audio feature vector "A", as shown in equation (4).
[0053]
number
[0054] The information generating unit 12b generates concept activity information based on the similarity contour. Since the concept activity information is generated based on the similarity contour, it represents a time period in which the concept to be extracted appears in the mixed speech signal. For example, the concept activity information is information representing a time period in the mixed speech signal that includes the spoken word "bicycle" uttered regarding the concept "bicycle" to be extracted.
[0055] The information generating unit 12b generates a concept embedding vector based on the concept activity information. The extracting unit 13 extracts mask information from the mixed speech signal based on the concept embedding vector.
[0056] <Learning stage> Instead of using a mixed audio signal for training, for example, oracle concept activity information is used for training. Oracle concept activity information is information obtained as the output of an audiovisual embedding network (see, for example, Reference 1) by inputting a reference audio signal of a target audio signal into the audiovisual embedding network.
[0057] By using the oracle concept activity information (time-series data) to generate a concept embedding vector, it is expected that the extraction unit 13 will be able to accurately extract the characteristics of a specific concept in the target speech signal. In supervised learning, which extracts the target speech signal from a mixed speech signal, a concept embedding vector close to the vector indicating the speaker of the target speech signal is generated.
[0058] <Estimation stage> FIG. 4 is a diagram showing an example of a similarity contour in the second embodiment. Audiovisual correspondences are used to generate concept embedding vectors. The similarity contour "s t’ " is used to identify regions (segments) of the audio signal where words related to the concept represented in the image are spoken (see Reference 1).
[0059] For example, when a speaker is talking about a concept designation signal 100 (a still image), the similarity contour represents the degree of similarity between the content of the concept designation signal 100 and the content of the speaker's voice. The concept designation signal 100 illustrated in FIG. 4 includes, for example, an image of a bicycle. Therefore, the similarity contour for a time interval in which the speaker's voice includes, for example, the word "bicycle" is relatively high compared to the similarity contour for a time interval in which the word "bicycle" is not included.
[0060] In the following, it is assumed that the speech sections of the speakers in the mixed speech signal partially overlap. The information generating unit 12b derives a similarity contour based on the concept designation signal 100 and the mixed speech signal. For example, the encoding unit 121b generates an image feature map of the concept designation signal 100. The encoding unit 121b may generate an image feature vector in the image feature map of the concept designation signal 100. The encoding unit 121b may generate a speech feature vector in the speech feature map of the mixed speech signal.
[0061] The similarity derivation unit 123 derives a similarity contour between an image feature vector in the image feature map and an audio feature vector in the audio feature map as shown in equation (4). In addition, the similarity derivation unit 123 scales the similarity contour into a value that varies between 0 and 1 using a sigmoid function as shown in equation (5).
[0062]
number
[0063] Here, "b" is a predetermined parameter that can be learned. t’ The time series of " is the concept activity information. That is, the similarity contour scaled to a value that varies between 0 and 1 is the concept activity information.
[0064] The auxiliary unit 124 includes an auxiliary network. The auxiliary unit 124 receives the mixed audio signal "y t " from the acquisition unit 11. The weighted sum unit 125 obtains the output "h(y t The weighted sum of the concept activity information and the concept embedding vector is calculated as follows:
[0065]
number
[0066] Here, "h(·)" represents the auxiliary network. To enable the concept embedding vector to be derived from the mixed audio signal, the auxiliary network synchronizes the concept activity information with the mixed audio signal. "y t " represents the t-th frame in the mixed audio signal "Y". The length "T'" of the series of concept activity information "p t’ " and the length "T" of the series of the t-th frame "y t " satisfy the relationship "T' < T". The auxiliary unit 124 linearly interpolates the concept activity information "p t’ ". The weighted sum unit 125 derives the concept activity information "p t " of the series with length "T" based on the linearly interpolated concept activity information. The weighted sum unit 125 is related to the activity-driven extraction network (ADEnet) (see Reference 2). This activity-driven extraction network utilizes information representing the time interval during which the speaker spoke to extract the target audio signal.
[0067] Note that instead of using the time-series data of the concept activity information exemplified in Equation (4), the weighted sum unit 125 may use the similarity profile exemplified in Equation (4) to derive the concept embedding vector exemplified in Equation (6).
[0068] Next, an operation example of the signal filtering device 1b will be described. FIG. 5 is a flowchart showing an operation example of the signal filtering device in the second embodiment. The encoding unit 121b encodes the concept specification signal into an image feature vector (step S201). The encoding unit 121b encodes the mixed audio signal into an audio feature vector (step S202). The similarity derivation unit 123 derives the similarity profile between the image feature vector and the audio feature vector (step S203).
[0069] The auxiliary unit 124 outputs the mixed audio signal to the weighted sum unit 125 (step S204). The weighted sum unit 125 generates a result of the weighted sum of the similarity contour and the mixed audio signal as a concept embedding vector (step S205). The extraction unit 13 extracts mask information from the mixed audio signal including the target audio signal based on the concept embedding vector (step S206). The mask processing unit 14 estimates the target audio signal from the mixed audio signal using the mask information (step S207).
[0070] As described above, the information generating unit 12b generates a concept embedding vector (feature information) of a concept designation signal (related information) of a target speech signal (target signal). The extracting unit 13 extracts mask information from a mixed speech signal (mixed signal) including the target speech signal based on the concept embedding vector. The mask processing unit 14 estimates the target speech signal from the mixed speech signal using the mask information.
[0071] Here, the information generation unit 12b encodes the concept designation signal (related information) into an image feature vector (first multidimensional vector). The information generation unit 12b encodes the mixed audio signal (mixed signal) into an audio feature vector (second multidimensional vector). The information generation unit 12b derives a similarity contour (time-series similarity) between the image feature vector and the audio feature vector. The information generation unit 12b generates a weighted sum of the similarity contour and the mixed audio signal (mixed signal) as a concept embedding vector.
[0072] This makes it possible to improve the accuracy of extracting the target audio signal from an audio signal (mixed audio signal) in which the target audio signal is mixed with audio signals other than the target audio signal.
[0073] (Third embodiment) The third embodiment differs from the first and second embodiments in that the audio signals in the mixed audio signal are separated for each speaker (sound source). The third embodiment will be described focusing on the differences from the first and second embodiments.
[0074] 6 is a diagram showing an example of the configuration of a signal filtering device 1c in the third embodiment. The signal filtering device 1c is a device that extracts a target audio signal from a mixed audio signal. The signal filtering device 1c extracts the target audio signal from the mixed audio signal by filtering the mixed audio signal that includes the target audio signal and audio signals other than the target audio signal.
[0075] The signal filtering device 1c includes a separation unit 15, an encoding unit 121c, and a selection unit 126. The separation unit 15 includes a first extraction layer 131 and a second extraction layer 133c. The encoding unit 121c or the selection unit 126 includes an audiovisual embedding network (see, for example, Reference 1).
[0076] The architecture of the separation network provided in the separation unit 15 is the same as the extraction network provided in the extraction unit 13. When the number of speakers (number of sound sources) of the sound signals in the mixed sound signal is known, the sound signals can be separated for each speaker (sound source). In the third embodiment, the number of sound sources is L. The L sound signals in the mixed sound signal are divided into {(~)X1, ..., (~)X L}
[0077] The second extraction layer 133c (output layer) extracts the speech signal in the mixed speech signal from the speaker's speech signal "(~)X l The second extraction layer 133c separates the speech signals in the mixed speech signal for each speaker using a method such as PIT (Permutation Invariant Training). The second extraction layer 133c outputs the speech signals of each speaker to the encoding unit 121c.
[0078] The encoding unit 121c receives the still image “C k The encoding unit 121c receives the speech signals of each speaker from the second extraction layer 133c. The encoding unit 121c uses an audiovisual embedding network to encode the still image "C k ” image feature vector “(-)I kThe encoding unit 121c derives the speech feature vector "(-)A l " is derived.
[0079] The encoding unit 121c encodes the globally pooled image feature vector “(−)I k ” and the global pooled speech feature vector “(-)A” of each speaker’s speech signal and the speech signal “(~)X l " is output to the selection unit 126.
[0080] The selection unit 126 selects the concept designation signal “C k " based on the global pooled image feature vector "(-)I k ” and the global pooled speech feature vector of each speaker’s speech signal “(-)A l " Similarity with "(-)I k (-)A l The selection unit 126 derives the speech signals of each speaker "(~)X l " is input from the separating unit 15 or the encoding unit 121c. The selecting unit 126 selects the speech signals "(~)X l ” among them, the most similar audio signal “(~)X l ” into the target audio signal “^X k " is selected as shown in equation (7).
[0081]
number
[0082] Next, an example of the operation of the signal filtering device 1c will be described. 7 is a flowchart showing an example of the operation of the signal filtering device in the third embodiment. The separation unit 15 separates L candidates of the target audio signal from the mixed audio signal (step S301). The encoding unit 121c encodes the concept-designating signal into an image feature vector (step S302). The encoding unit 121c encodes the L candidates of the target audio signal into L audio feature vectors (step S303).
[0083] The selection unit 126 derives the similarity (inner product) between the globally pooled image feature vector and the globally pooled audio feature vector for each candidate target audio signal (step S304).The selection unit 126 selects the target audio signal with the highest similarity from among the L candidates for the target audio signal (step S305).
[0084] As described above, the separation unit 15 separates L (predetermined number) candidates (candidate signals) of the target sound signal from the mixed sound signal (mixed signal) as candidates for the target sound signal to be selected. The L candidates of the target sound signal are sound signals associated with L predetermined sound sources (for example, speakers). The separation unit 15 separates the candidates of the target sound signal in the mixed sound signal for each sound source using a method such as PIT.
[0085] The encoding unit 121c encodes a concept designation signal (related information) related to the target audio signal into an image feature vector (first feature vector). The encoding unit 121c encodes L candidates (candidate signals) of the target audio signal into L audio feature vectors (second feature vectors).
[0086] The selection unit 126 derives the similarity between the global-pooled image feature vector and the global-pooled audio feature vector for each candidate target audio signal (candidate signal). The selection unit 126 derives the inner product of the image feature vector and the audio feature vector as the similarity. The selection unit 126 selects the target audio signal (candidate signal) with the highest similarity from among the L candidates for the target audio signal as the final target audio signal (target signal).
[0087] This makes it possible to improve the accuracy of extracting the target audio signal from an audio signal (mixed audio signal) in which the target audio signal is mixed with audio signals other than the target audio signal.
[0088] (Example of effect) An example of evaluation results regarding the performance of the above-described signal filtering device in extracting a target audio signal will be described below.
[0089] The Places spoken caption dataset, which contains images of various scenes and locations and is annotated with audio captions, was used as training data to create a mixed audio signal from two speakers. This audio caption dataset consists of an image dataset and audio captions in English and Japanese. The images in the image dataset are classified into 205 different scene classes. In addition, 97,555 pairs of images and audio captions were extracted from each language dataset. Only the Japanese audio captions are labeled with the gender of the speaker.
[0090] To evaluate the effectiveness of the signal filtering device in both languages, the signal was divided into a training set of 90,000 pairs, a validation set of 4,000 pairs, and an evaluation set of 3,555 pairs for each language. The training set was then used to pre-train an audiovisual embedding network (deep metric learning).
[0091] Image-audio caption pairs belonging to different image classes were selected, and audio captions were mixed at signal-to-noise ratios ranging from 0 to 5 dB to create mixed audio signals from two speakers. As a result, the training set contained 90,000 mixed audio signals, the validation set contained 4,000 mixed audio signals, and the evaluation set contained 3,555 mixed audio signals. The frequency of the audio captions was downsampled to 8 kHz to reduce computational and memory costs.
[0092] A 258-dimensional vector, which combines the real and imaginary parts of the complex spectrum, was used as the input speech feature. This complex spectrum was obtained by short-time Fourier transform with a window length of 32 ms and a window shift length of 8 ms.
[0093] As image preprocessing, the image dimensions were resized so that the minimum image dimension was 256 pixels. The resized images were then center-cropped to 224 × 224. The pixels of the center-cropped images were normalized according to the global pixel mean and variance.
[0094] The audiovisual embedding network used was ResNet-ResDAVEnet (see Reference 1). The image encoder was ResNet 50. When a 224x224x3 image was input, the image encoder output a 7x7x1,024 image feature map. Here, the height (H) and width (W) of the image feature map are both 7.
[0095] The speech encoder is "ResDAVEnet". When a 40-dimensional logarithmic mel filterbank spectrogram is input, the speech encoder outputs a "T' x 1,024" speech feature map. This filterbank spectrogram is calculated from the input speech features. The dimension "d" is 1,024. The time resolution "T'" is finally "T / 16".
[0096] The linear transformation unit 122 illustrated in FIG. 1 has a fully connected layer with 896 units (d'=896). The auxiliary unit 124 (auxiliary network) illustrated in FIG. 3 has two fully connected layers. These two fully connected layers have 200 hidden units, 896 hidden units, and a ReLU (Rectified Linear Unit) activation function. Therefore, the dimension of the concept embedding vector is 896.
[0097] The extraction network of the extraction unit 13 and the separation network of the separation unit 15 each have four bidirectional long short-term memory layers consisting of 896 units. The extraction network of the extraction unit 13 and the separation network of the separation unit 15 each have an 896-unit linear mapping layer after each bidirectional long short-term memory layer. This linear mapping layer connects the forward output of the LSTM (Long Short Term Memory) with the backward output of the LSTM.
[0098] The extraction unit 13 used one fully connected layer and a ReLU activation function to estimate the mask information (time-frequency mask). The combination processor 132 combined the output of the first bidirectional long short-term memory layer in the extraction unit 13 (extraction network) with the concept embedding vector.
[0099] In training the separation network of the separation unit 15, the number of sound sources "L" is 2. The total number of speakers "K" is 2. The initial learning rate is 0.0001. "Adam" is used as the optimization method for training, and gradient clipping is performed. The target speech signals extracted by the signal filtering device were evaluated using the signal-to-distortion ratio (SDR). The SDR represents the performance of extracting each speaker's target speech signal from the mixed speech signal. The SDR scores were averaged across all experimental results.
[0100] 8 shows examples of averaged signal-to-distortion ratio (SDR) scores (dB) for target audio signals in the first and second embodiments. The values in the column of the item "opposite-sex mixed audio" indicate the signal-to-distortion ratio scores for opposite-sex mixed audio signals. The values in the column of the item "same-sex mixed audio" indicate the signal-to-distortion ratio scores for same-sex mixed audio signals. The values in the column of the item "opposite-sex and same-sex mixed audio" indicate the signal-to-distortion ratio scores for opposite-sex and same-sex mixed audio signals.
[0101] The item "Image feature vector" indicates the signal-to-distortion ratio score in the signal filtering device 1a of the first embodiment. The item "Similarity contour" indicates the signal-to-distortion ratio score when the similarity derivation unit 123 of the signal filtering device 1b of the second embodiment outputs the similarity contour to the weighted sum unit 125. The item "Concept activity information" indicates the signal-to-distortion ratio score when the similarity derivation unit 123 of the signal filtering device 1b of the second embodiment outputs the concept activity information to the weighted sum unit 125.
[0102] In order to determine which of the image feature vector, similarity contour, and concept activity information is the best configuration for generating a concept embedding vector, mixed speech from two speakers with no overlapping time intervals was used to evaluate signal filtering device 1a and signal filtering device 1b.
[0103] As a result of each evaluation, the performance of target speech signal extraction was highest when the concept embedding vector generated using concept activity information was used for extraction.In the following, in the target speech signal extraction method, the concept embedding vector is generated using concept activity information.
[0104] FIG. 9 shows an example of extraction of a target audio signal in the second embodiment (extraction method). The concept designation signal 101 is an image of a scene in which a man wearing glasses is playing the guitar inside a bookstore. The concept designation signal 102 is an image of a night scene with blue pillars and a roller coaster. A first speaker (not shown) is speaking about the concept designation signal 101. The first target speech signal is the speech signal of the first speaker. A second speaker (not shown) is speaking about the concept designation signal 102. The second target speech signal is the speech signal of the second speaker.
[0105] The signal filtering device 1b can extract the first and second target speech signals even in a time period in which the speech of the first speaker and the speech of the second speaker overlap in the mixed speech signal. In particular, each time when the value of the "concept activity information" becomes 1 corresponds to a concept (e.g., the spoken language "glasses" and the spoken language "man") associated with a prominent object (e.g., a man wearing glasses) in the image of the concept designation signal 101. The same is true for the image of the concept designation signal 102. Each time when a concept appears in the concept designation signal serves as a clue to derive a concept embedding vector. The concept embedding vector is used to generate mask information for extracting the target speech signal from the mixed speech signal.
[0106] The extraction performance (SDR score) of the first target speech signal is 17.7 dB. The extraction performance (SDR score) of the second target speech signal is 17.0 dB. As can be seen from these results, speech from two speakers can be successfully extracted from mixed speech.
[0107] Figure 10 shows an example of signal-to-distortion ratio scores for each overlap rate in the second embodiment (extraction method) and the third embodiment (separation method). Using mixed speech signals from two speakers, the extraction performance of the signal filtering device 1b (extraction method using concept activity information) is compared with that of the signal filtering device 1c (separation method). The mixed speech signals from two speakers were obtained by mixing Japanese speech captions with five different overlap rates.
[0108] The extraction performance of signal filtering device 1b and the extraction performance of signal filtering device 1c tend to be similar as the overlap rate decreases, whereas the extraction performance of signal filtering device 1b and the extraction performance of signal filtering device 1c decrease as the overlap rate increases.
[0109] The extraction performance of the signal filtering device 1c is 10 dB or more even when the overlap rate is 100%. However, the signal filtering device 1c needs to acquire in advance information indicating the number of speakers (number of sound sources) of the target speech signals contained in the mixed speech signal. It is effective to selectively use the signal filtering device 1b and the signal filtering device 1c depending on whether the number of speakers is known or not and the overlap rate between the target speech signals.
[0110] Next, examples of each possible usage scenario will be described. A first usage scenario is assumed to be a situation in which a presenter is explaining the contents of a poster (the concept targeted for extraction) at a booth in a poster venue at an academic conference or exhibition. Due to unrelated sounds and noise, the voice of the target presenter (target audio signal) is difficult to hear. However, the signal filtering device of each of the above embodiments utilizes the contents of the poster (image) as a concept-designating signal (auxiliary information). The signal filtering device of each of the above embodiments extracts the presenter's voice from audio containing a mixture of various sounds. This makes it possible to make the presenter's voice easier to hear.
[0111] A second usage scenario is assumed to be a situation in which a target video content (a concept to be extracted) is searched for from a large amount of video content, such as television broadcasts and video streaming. The signal filtering device of each of the above embodiments utilizes still images and videos including an image representing the search target concept (the concept to be extracted) as a concept designation signal (auxiliary information). For example, the signal filtering device utilizes still images and videos including an image representing the search target bicycle as the concept designation signal. The signal filtering device extracts a target audio signal explaining the search target concept from a mixed audio signal associated with a large amount of video content. For example, the signal filtering device extracts a target audio signal "bicycle" explaining a bicycle from a mixed audio signal associated with a large amount of video content including a bicycle video. This makes it possible to search for the target video content (e.g., a bicycle video) associated with the extracted mixed audio signal.
[0112] A third usage scenario is assumed in which speech recognition is performed on target speech to add subtitles to instructional content in television broadcasts and video streaming. Instructional content is content that uses still images and videos to explain a concept targeted for extraction, such as a video explaining cooking, a video explaining crafting methods, or a video teaching material. In instructional content, the target speech is often buried in background sounds and noise, making it difficult to perform speech recognition on the target speech. The signal filtering device of each of the above embodiments utilizes still images and videos explaining the concept targeted for explanation as a concept designation signal (auxiliary information). Extracting the speaker's target speech signal improves speech recognition performance.
[0113] A fourth usage scenario is assumed to be a situation in which the signal filtering device is used for music. Hereinafter, a mixed audio signal in which the audio signal to be extracted is mixed with an audio signal other than the audio signal to be extracted is referred to as a "mixed audio signal." For example, an audio signal in which the sounds of multiple types of musical instruments are mixed may be input to the signal filtering device as the mixed audio signal of each of the above embodiments. The signal filtering device utilizes a still image or video containing an image of the target instrument as a concept designation signal (auxiliary information). The audio signal extracted as the sound of the target instrument becomes easier to hear.
[0114] A fifth usage scenario is assumed in which an acoustic signal associated with a concept to be extracted is searched for from a mixed acoustic signal. The mixed acoustic signal is, for example, an acoustic signal recorded by a microphone (e.g., a surveillance microphone) installed outdoors. The mixed acoustic signal includes, for example, environmental sounds such as the sound of cars. Still images and videos associated with the concept to be extracted are used as concept designation signals (auxiliary information).
[0115] In a sixth usage scenario, the concept designation signal (auxiliary information) may be an audio signal instead of an image signal. When the concept designation signal is an audio signal, the signal filtering device may extract, from the mixed audio signal, a target audio signal of a speaker who is talking about content similar to the topic content (concept). When a first speaker who speaks English and a second speaker who speaks Japanese are talking about the same concept (e.g., the content of the same image), the signal filtering device may extract, from the mixed audio signal, one of the English audio signal of the first speaker and the Japanese audio signal of the second speaker by using the language used in the target audio signal as the concept designation signal. The signal filtering device may remove, from the mixed audio signal, one of the English audio signal of the first speaker and the Japanese audio signal of the second speaker by using the language used in the target audio signal or a language not used in the target audio signal as the concept designation signal.
[0116] (Example of hardware configuration) FIG. 11 is a diagram illustrating an example of the hardware configuration of a signal filtering device 1 in each embodiment. The signal filtering device 1 corresponds to a signal filtering device 1a, a signal filtering device 1b, and a signal filtering device 1c. Some or all of the functional units of the signal filtering device 1 are implemented as software by a processor 111, such as a central processing unit (CPU), executing a program stored in a storage device 112 having a non-volatile recording medium (non-transitory recording medium) and a memory 113. The program may be recorded on a computer-readable non-transitory recording medium. Examples of the computer-readable non-transitory recording medium include portable media such as a flexible disk, a magneto-optical disk, a read-only memory (ROM), and a compact disc read-only memory (CD-ROM), and storage devices such as a hard disk built into a computer system. A communication unit 114 executes a predetermined communication process. The communication unit 114 may acquire data and programs.
[0117] Some or all of the functional units of the signal filtering device 1 may be realized using hardware including electronic circuits (electronic circuits or circuitry) using, for example, an LSI (Large Scale Integrated circuit), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0118] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Industrial Applicability]
[0119] The present invention is applicable to systems that filter signals. [Explanation of symbols]
[0120] 1, 1a, 1b, 1c...signal filtering device, 11...acquisition unit, 12a, 12b...information generation unit, 13...extraction unit, 14...mask processing unit, 15...separation unit, 100...concept designation signal, 101...concept designation signal, 102...concept designation signal, 111...processor, 112...storage device, 113...memory, 114...communication unit, 121a, 121b...encoding unit, 123...similarity derivation unit, 124...auxiliary unit, 125...weighted sum unit, 126...selection unit, 131...first extraction layer, 132...combination processing unit, 133a, 133b, 133c...second extraction layer
Claims
1. a separation unit that separates a predetermined number of candidate signals from the mixed signal as candidates for a target signal; an encoding unit that encodes the related information of the target signal into a first feature vector and encodes the predetermined number of candidate signals into the predetermined number of second feature vectors; a selection unit that derives a similarity between the first feature vector and the second feature vector for each of the candidate signals, and selects the candidate signal with the highest similarity as the target signal from among the predetermined number of candidate signals; Equipped with the first feature vector is an image feature vector; the second feature vector is a speech feature vector; Signal filtering device.
2. the selection unit derives an inner product of the first feature vector and the second feature vector as the similarity.
2. The signal filtering device of claim 1.
3. the predetermined number of candidate signals are speech signals associated with the predetermined number of sound sources, the separation unit separates the candidate signals in the mixed signal for each sound source.
3. A signal filtering device according to claim 1 or claim 2.
4. A signal filtering method performed by a signal filtering device, comprising: Separating a predetermined number of candidate signals from the mixed signal as candidates for a target signal; encoding related information of the target signal into a first feature vector and encoding the predetermined number of candidate signals into the predetermined number of second feature vectors; deriving a similarity between the first feature vector and the second feature vector for each of the candidate signals, and selecting the candidate signal with the highest similarity as the target signal from among the predetermined number of candidate signals; Including, the first feature vector is an image feature vector; the second feature vector is a speech feature vector; Signal filtering methods.
5. A program for causing a computer to function as the signal filtering device according to any one of claims 1 to 3.
Citation Information
Patent Citations
Device and method for sound source selection
JP2004287311A
Method for distinguishing one or more components of a signal
JP2018502319A
Signal processing device, signal processing method and program
JP2021152623A
Method, apparatus and computer-readable storage medium for recognizing mixed speech
JP2021516369A
Speech extraction using attention network
US20200335119A1