Sound extraction system and sound extraction method

The sound extraction system accurately isolates user-defined sounds from mixed signals using text-embedded models and time-frequency masks, overcoming the limitations of conventional technologies that rely on predefined event types.

JP7726757B2Active Publication Date: 2025-08-20HITACHI LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021192632
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-08-20
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

Conventional sound extraction technologies fail to accurately extract sounds at a finer granularity than predefined event types, leading to incorrect extraction of sounds like 'bang' when the user wants 'sound of metal being struck', as they classify sounds into predefined categories.

Method used

A sound extraction system that uses a learning subsystem to generate models for extracting specific sounds based on text representing the desired sound range, employing feature extraction, text-embedded extraction, and time-frequency mask generation to accurately isolate the target sound from a mixed signal.

Benefits of technology

Enables high-accuracy extraction of user-defined sounds from mixed signals, even when the desired sound range does not match predefined event types, by utilizing onomatopoeia-based text embeddings and time-frequency masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726757000001
    Figure 0007726757000001
  • Figure 0007726757000002
    Figure 0007726757000002
  • Figure 0007726757000003
    Figure 0007726757000003
Patent Text Reader

Abstract

To provide a sound extraction system and a sound extraction method that can accurately extract signals corresponding to sounds that a user wants to extract from a mixed signal.SOLUTION: A sound extraction system includes a sound extraction device that extracts a signal corresponding to a sound to be extracted from a mixed signal that include a signal corresponding to the sound to be extracted. The sound extraction device is configured to extract the signal corresponding to the sound to be extracted from the mixed signal based on text representing a range of the sound to be extracted and the mixed signal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sound extraction system and a sound extraction method. [Background technology]

[0002] It is important to extract sounds with specific characteristics from a mixture of sounds from multiple sound sources. For example, sound is recorded and abnormalities or signs of abnormalities in equipment or machinery are automatically detected from the sounds (abnormal sounds). However, there are cases where environmental noise is loud, and in such cases the accuracy of abnormal sound detection can be significantly reduced. Therefore, in order to improve the accuracy of abnormal sound detection and to analyze the abnormal sounds themselves, it is necessary to extract (or emphasize) the sounds of the target equipment or machinery from the recorded sounds.

[0003] Additionally, surrounding situations are recognized based on sounds recorded by microphones in surveillance cameras, dashcams, monitoring robots, smart speakers, etc. However, there are cases where the environmental noise is loud, and in such cases the accuracy of situation recognition can be significantly reduced. Therefore, in order to improve the accuracy of situation recognition and analyze the recorded sounds, it is necessary to extract sounds that provide clues for situation recognition from the recorded sounds.

[0004] Known technologies for sound extraction include those described in Patent Document 1 and Non-Patent Document 1. Non-Patent Document 1 states, "The waveforms of mixed environmental sounds are transformed into a spectrogram. By using this input, Mask U-Net consists of sound event detection CNN and segmentation U-Net predicts masks for separating out each class from the input spectrogram. Inverse STFT is applied to reconstruct the time domain signal." The technology described in Non-Patent Document 1 (hereinafter also referred to as "conventional technology") classifies sounds into a finite number of predefined types of events and extracts sounds for each type of event. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-178886 [Non-patent literature]

[0006] [Non-Patent Document 1] Y. Sudo, “Environmental sound segmentation Mask utilizing U-Net,” in IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), 2019. Summary of the Invention [Problem to be solved by the invention]

[0007] However, conventional technologies cannot extract the sound a user wants to extract if the range of sounds the user wants to extract does not match a predefined type of event. For example, conventional technologies cannot extract sounds at a finer granularity than the predefined event types. For example, even if an event type is defined as the sound of metal being struck, sounds such as "bang" and "clang" may be included in the sound of metal being struck. Therefore, for example, if a user wants to extract the sound "bang" as the "sound of metal being struck," the conventional technology may extract the sound "clang" as the sound of metal being struck simply by classifying the event as the "sound of metal being struck."

[0008] As described above, in the conventional technology, it may be impossible to accurately extract the sound that the user wants to extract. The present invention has been made to solve the above problem. That is, one of the objects of the present invention is to provide a sound extraction system and a sound extraction method that can accurately extract a signal corresponding to the sound that the user wants to extract from a mixed signal. [Means for solving the problem]

[0009] In order to solve the above problem, the sound extraction system of the present invention is a sound extraction system that includes a sound extraction device that extracts a signal corresponding to a sound to be extracted from a mixed signal including the signal corresponding to the sound to be extracted, and the sound extraction device is configured to extract the signal corresponding to the sound to be extracted from the mixed signal based on text that represents the range of the sound to be extracted and the mixed signal.

[0010] The sound extraction method of the present invention uses a sound extraction device that extracts a signal corresponding to a sound to be extracted from a mixed signal containing the signal corresponding to the sound to be extracted, and the sound extraction device extracts the signal corresponding to the sound to be extracted from the mixed signal based on text that represents the range of the sound to be extracted and the mixed signal. [Effects of the Invention]

[0011] According to the present invention, a signal corresponding to a sound that a user wants to extract can be extracted from a mixed signal with high accuracy. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram showing an example of the schematic configuration of a sound extraction system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of an information processing device. [Figure 3] FIG. 3 is a block diagram for explaining an example of the configuration of the learning subsystem for each function. [Figure 4] FIG. 4 is a flowchart showing an example of a processing flow of the learning subsystem. [Figure 5] FIG. 5 is a block diagram for explaining an example of the configuration of the sound extraction subsystem for each function. [Figure 6] FIG. 6 is a flowchart showing an example of the processing flow of the sound extraction subsystem. [Figure 7] FIG. 7 shows data showing an example of the extraction results obtained by the sound extraction system. [Figure 8] FIG. 8 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the second embodiment of the present invention. [Figure 9] FIG. 9 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the third embodiment of the present invention. [Figure 10] FIG. 10 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the fourth embodiment of the present invention. [Figure 11] FIG. 11 is a block diagram for explaining an example of the configuration of the learning subsystem for each function. [Figure 12] FIG. 12 is a flowchart showing an example of the processing flow of the learning subsystem. [Figure 13] FIG. 13 is a block diagram for explaining an example of the configuration of the sound extraction subsystem for each function. [Figure 14] FIG. 14 is a flowchart showing an example of the processing flow of the sound extraction subsystem. [Figure 15] FIG. 15 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the fifth embodiment of the present invention. [Figure 16] FIG. 16 is a block diagram for explaining an example of the configuration of the learning subsystem for each function. [Figure 17] FIG. 17 is a flowchart showing an example of the processing flow of the learning subsystem. [Figure 18] FIG. 18 is a block diagram for explaining an example of the configuration of the sound extraction subsystem for each function. [Figure 19] FIG. 19 is a flowchart showing an example of the processing flow of the sound extraction subsystem. [Figure 20] FIG. 20 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the sixth embodiment of the present invention. [Figure 21] FIG. 21 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the seventh embodiment of the present invention. [Figure 22] FIG. 22 is a block diagram showing an example of the schematic configuration of a sound extraction system according to the eighth embodiment of the present invention. [Figure 23] FIG. 23 is a block diagram for explaining an example of the configuration of the learning subsystem for each function. [Figure 24] FIG. 24 is a flowchart showing an example of the processing flow of the learning subsystem. [Figure 25] FIG. 25 is a block diagram for explaining an example of the configuration of the sound extraction subsystem for each function. [Figure 26] FIG. 26 is a flowchart showing an example of the processing flow of the sound extraction subsystem. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, a sound extraction system according to each embodiment of the present invention will be described with reference to the drawings.

[0014] <<First Embodiment>> <Summary of the Invention> FIG. 1 is a block diagram showing an example of the schematic configuration of a sound extraction system 100 according to a first embodiment of the present invention. As shown in FIG. 1, the sound extraction system 100 includes a learning subsystem 110, a feature extraction model database 120, a text-embedded extraction model database 130, a time-frequency mask generation model database 140, a sound extraction subsystem 150, and a training dataset database 160. The learning subsystem 110 may also be referred to as a "learning device." The sound extraction subsystem 150 may also be referred to as a "sound extraction device."

[0015] The sound extraction system 100 first reads from the training dataset database 160 a set of three pairs (sometimes referred to as a "training dataset"): a time waveform of a target signal (a signal corresponding to a sound to be extracted), a time waveform of a mixed signal obtained by mixing the time waveform of the target signal (target signal) with a noise other than the sound to be extracted (a time waveform of a signal corresponding to the noise (signal corresponding to the noise)), and a variable-length onomatopoeia text (onomatopoeia text corresponding to the sound to be extracted), and inputs these to the learning subsystem 110.

[0016] Here, instead of reading out the "time waveform of a mixed signal obtained by mixing the time waveform of the target signal (target signal) and noise other than the sound to be extracted (the time waveform of a signal corresponding to the noise (signal corresponding to the noise))" from the training dataset database 160, it is also possible to read out the "noise other than the sound to be extracted (the time waveform of the signal corresponding to the noise (signal corresponding to the noise))" before mixing and mix it with the time waveform of the target signal (signal corresponding to the sound to be extracted) to generate a "time waveform of a mixed signal obtained by mixing the time waveform of the target signal (target signal) and noise other than the sound to be extracted (the time waveform of the signal corresponding to the noise (signal corresponding to the noise))," thereby generating a set of triplet. Mixing signals after reading them in this way has two advantages. First, by creating a training dataset by mixing signals at a signal-to-noise ratio expected when the sound extraction system 100 is used, it is possible to train a model to enable extraction appropriate for the signal-to-noise ratio for each usage scenario. Second, it is possible to reduce the storage capacity required for the training dataset database 160.

[0017] Hereinafter, the set of triplets generated by reading from the training dataset database 160 and then mixing them will be referred to as the set of triplets or the learning dataset.

[0018] The learning subsystem 110 performs a learning process based on the set of triples, and outputs a feature extraction model, a text-embedded extraction model, and a time-frequency mask generation model, which are stored in the respective databases. That is, the learning subsystem 110 stores the feature extraction model in the feature extraction model database 120, the text-embedded extraction model in the text-embedded extraction model database 130, and the time-frequency mask generation model in the time-frequency mask generation model database 140.

[0019] The sound extraction subsystem 150 reads the feature extraction model, text-embedded extraction model, and time-frequency mask generation model from the databases (feature extraction model database 120, text-embedded extraction model database 130, and time-frequency mask generation model database 140) and performs sound extraction processing based on (using) them. As a result, the sound extraction subsystem 150 extracts the time waveform of an extracted signal from the time waveform of the mixed signal and the variable-length onomatopoeia text. Furthermore, the sound extraction subsystem 150 outputs the time waveform of the extracted signal.

[0020] With this basic configuration, the sound extraction system 100 can accurately extract (extract or emphasize) a signal corresponding to the sound that the user wants to extract from the mixed signal, even if the range of sound that the user wants to extract as a certain type of event cannot be strictly defined in advance.

[0021] It should be noted that the definition of event types used in the above-mentioned Non-Patent Document 1 varies depending on the application site, and it is rare that the range of sounds that the user wants to extract matches the predefined types of events. If the range of sounds that the user wants to extract does not match the predefined types of events, the sound that the user wants to extract cannot be extracted. In contrast, onomatopoeia are relatively versatile and therefore have a high possibility of being usable across application sites.

[0022] Furthermore, the technology described in the aforementioned Patent Document 1 is known as a process for outputting environmental sounds in response to input of onomatopoeia. Patent Document 1 states that "the system includes a voice input unit that inputs a voice signal, a voice recognition unit that performs voice recognition processing on the voice signal input to the voice input unit to generate onomatopoeia, a sound data storage unit that stores environmental sounds and the onomatopoeia corresponding to the environmental sounds, a correspondence storage unit that stores correspondence information that associates a first onomatopoeia, a second onomatopoeia, and the frequency with which the second onomatopoeia is given when the first onomatopoeia is recognized by the voice recognition unit, a conversion unit that uses the correspondence information to convert the first onomatopoeia recognized by the voice recognition unit into a second onomatopoeia corresponding to the first onomatopoeia recognized by the voice recognition unit, and a search and extraction unit that extracts environmental sounds that correspond to the second onomatopoeia converted by the conversion unit from the sound data storage unit, and ranks and presents the extracted multiple environmental sound candidates based on the frequency with which the extracted multiple environmental sound candidates are given."

[0023] However, the technology of Patent Document 1 cannot extract sounds from mixed sounds. The term "extraction" in Patent Document 1 means searching a database to extract environmental sounds that meet certain conditions. Furthermore, the technology of Patent Document 1 only has a mapping from environmental sounds to onomatopoeia, but does not have a mapping from onomatopoeia to environmental sounds, so the only sounds that are output are those that exist in the database. Unless a mixed sound that is exactly the same as a sound that exists in the database is input, it is impossible to extract a sound from a mixed sound. In the extraction of sounds from mixed sounds, which is the subject of the present invention, it is almost impossible to input a mixed sound that is exactly the same as a sound that exists in the database. Therefore, Patent Document 1 cannot extract sounds from mixed sounds.

[0024] <Hardware configuration> The sound extraction system 100 can be configured, for example, by a computer (information processing device). FIG. 2 is a block diagram showing an example of the configuration of the information processing device. As shown in FIG. 2, the information processing device 200 includes a CPU 201, a ROM 202, a RAM 203, a non-volatile storage device (HDD) 204 that can read and write data, a network interface 205, and an input / output interface 206. These are communicably connected to each other via a bus 207. The CPU 201 loads various programs (not shown) stored in the ROM 202 and / or the storage device 204 into the RAM 203 and executes the programs loaded into the RAM 203 to realize various functions. As described above, various programs executed by the CPU 201 are loaded into the RAM 203, and data used when the CPU 201 executes the various programs is temporarily stored in the RAM 203. The ROM 202 and / or the storage device 204 are non-volatile storage media that store various programs. The network interface 205 is an interface for connecting the information processing device 200 to a network. The input / output interface 206 is an interface for connecting to operating devices such as a keyboard and a mouse, audio devices such as a microphone, and display devices such as a display.

[0025] For example, the training dataset database 160 of the sound extraction system 100 is configured as a database stored in a storage device 204 provided in the information processing device 200. The learning subsystem 110 of the sound extraction system 100 is configured as an information processing device 200. The feature extraction model database 120, text-embedded extraction model database 130, and time-frequency mask generation model database 140 of the sound extraction system 100 are configured as databases stored in a storage device 204 provided in the information processing device 200. The sound extraction subsystem 150 of the sound extraction system 100 is configured as an information processing device 200. Note that the information processing device 200 that constitutes one system may be multiple information processing devices or a virtual information processing device built on the cloud.

[0026] <Learning Subsystem> (Functions of the learning subsystem) The configuration of the learning subsystem 110 will be described below, mainly for each function. Fig. 3 is a block diagram for describing an example configuration of the learning subsystem 110 for each function. As shown in Fig. 3, the learning subsystem 110 includes a target signal frame segmentation processing unit 111, a target signal window function multiplication unit 112, a target signal frequency domain signal generation unit 113, a mixed signal frame segmentation processing unit 114, a mixed signal window function multiplication unit 115, a mixed signal frequency domain signal generation unit 116, a feature extraction unit 117, a phoneme conversion unit 118, a text embedding extraction unit 119, a time-frequency mask generation unit 119a, a time-frequency mask multiplication unit 119b, and a learning unit 119c. The target signal frame segmentation processing unit 111, the target signal window function multiplication unit 112, the target signal frequency domain signal generation unit 113, the mixed signal frame segmentation processing unit 114, the mixed signal window function multiplication unit 115, the mixed signal frequency domain signal generation unit 116, the feature extraction unit 117, the phoneme conversion unit 118, the text embedding extraction unit 119, the time-frequency mask generation unit 119a, the time-frequency mask multiplication unit 119b, and the learning unit 119c are configured by various programs (not shown) stored in the ROM 202 and / or the storage device 204 of the information processing device 200.

[0027] The target signal frame division processing unit 111 divides the time waveform D10 of the target signal into frames and outputs a frame-divided signal of the target signal (not shown). The target signal window function multiplication unit 112 performs window function multiplication and converts the frame-divided signal of the target signal into a window function-multiplied signal of the target signal (not shown).

[0028] The target signal frequency domain signal generator 113 performs a short-time Fourier transform to convert the window function-multiplied signal of the target signal into a time-frequency domain representation D11 of the target signal. Note that the target signal frequency domain signal generator 113 can also use a frequency transform method such as a "constant Q transform (CQT)" instead of the short-time Fourier transform.

[0029] The mixed signal frame division processing unit 114 divides the time waveform D20 of the mixed signal into frames, and outputs a frame division signal (not shown) of the mixed signal.

[0030] The mixed signal window function multiplication unit 115 performs window function multiplication to convert the frame-divided mixed signal into a window function-multiplied mixed signal (not shown). The mixed signal frequency domain signal generation unit 116 performs a short-time Fourier transform to convert the window function-multiplied mixed signal into a time-frequency domain representation D21 of the mixed signal. Note that the mixed signal frequency domain signal generation unit can also use a frequency transformation method such as a "constant Q transform (CQT)" instead of the short-time Fourier transform.

[0031] The feature extraction unit 117 converts the time-frequency domain representation D21 of the mixed signal into a sound feature vector D22. In this example, the feature extraction unit 117 uses a feature extraction model that is a neural network with variable weighting coefficient parameters. The feature extraction unit 117 inputs the time-frequency domain representation D21 of the mixed signal to the latest feature extraction model that has been updated immediately before, and outputs the sound feature vector D22. The feature extraction model may be, for example, a neural network in which multiple convolutional layers, activation functions, and pooling layers are stacked, with skip connections sandwiched between them.

[0032] The sound feature vector D22 may be an amplitude spectrogram of the time-frequency domain representation D21. In this case, the feature extraction unit 117 calculates an amplitude spectrogram (vector) of the time-frequency domain representation D21. In this case, the weighting coefficient parameters used in the calculation are invariant. For example, the sound feature vector may be a power spectrogram of the time-frequency domain representation D21. In this case, the feature extraction unit 117 calculates a power spectrogram (vector) of the time-frequency domain representation D21. In this case, the weighting coefficient parameters used in the calculation are invariant. For example, the sound feature vector D22 may be a logarithmic mel power spectrogram of the time-frequency domain representation D21. In this case, the feature extraction unit 117 calculates a power spectrogram of the time-frequency domain representation D21, multiplies the obtained power spectrogram by a mel filter bank to calculate a mel power spectrogram, and outputs the logarithmic mel power spectrogram (vector) by taking a logarithm on the obtained mel power spectrogram. In this case, the weighting coefficient parameters used in the calculation are unchanged. Instead of the Mel filter bank, a filter bank such as a 1 / 3 octave band filter may be used.

[0033] Furthermore, the sound feature vector D22 may be a time series of Mel-Frequency Cepstral Coefficients (MFCCs) instead of a logarithmic Mel-Power Spectrogram. In this case, the feature extraction unit 117 calculates the logarithmic value of the power spectrogram, multiplies it by a filter bank, performs a discrete cosine transform, and outputs a time series (vector) of MFCCs. In this case, the weighting coefficient parameters used in the calculation remain unchanged.

[0034] The sound feature vector D22 may be a time difference of a logarithmic Mel Power spectrogram or a time series of MFCCs, a time series of time derivatives (delta), or a concatenated vector of these. In any of these cases, the weighting coefficient parameters used in the calculation remain unchanged.

[0035] The phoneme conversion unit 118 outputs a variable-length phoneme sequence D31 by phoneme conversion processing from the variable-length onomatopoeia text D30. For example, if the onomatopoeia text D30 is "kankan," " / ka N ka N / " is output as the phoneme sequence D31. If the onomatopoeia text D30 is "katakata don," " / katakatado: N / " is output as the phoneme sequence D31.

[0036] The text embedding extraction unit 119 uses the latest text embedding extraction model to output a text embedding vector D32 (text embedding vector D32) from the phoneme sequence D31. The text embedding vector D32 is a vector with a certain number of dimensions D. First, the text embedding extraction unit 119 assigns a one-hot vector to each phoneme in the input phoneme sequence D31 to create a one-hot vector sequence. The one-hot vector here is a vector in which 1 is assigned only to the dimension corresponding to the type of phoneme to be converted (such as " / a / , / i / , / u / , / e / , / o / , / k / , / s / , / N / ") and 0 is assigned to the other dimensions.

[0037] Next, the one-hot vector sequence is input to a text embedding extraction model, which outputs a text embedding vector D32. The text embedding extraction model may be a well-known Transformer model or a recurrent neural network with layers such as a Long-Short-Term Memory (LSTM), a bidirectional LSTM, a gated recurrent unit (GRU), or a bidirectional GRU.

[0038] The time-frequency mask generation unit 119a generates a time-frequency mask from the sound feature vector D22 and the text-embedded vector D32 using the latest time-frequency mask generation model.

[0039] The time-frequency mask is an estimate of what proportion of the amplitude of the mixed signal is the extracted signal at each time frequency in the time-frequency domain representation. That is, the time-frequency mask takes a value greater than 0 and less than 1 at each time frequency. The closer to 1, the more of the amplitude of the mixed signal is the extracted signal, and the closer to 0, the more of the amplitude of the mixed signal is components other than the extracted signal.

[0040] The time-frequency mask generation model is a neural network that generates a time-frequency mask using a sound feature vector D22 and a text embedding vector D32 as input. For example, it may be a neural network with multiple convolutional layers, activation functions, and pooling layers stacked with skip connections in between. In particular, when the sound feature vector D22 is an amplitude spectrogram or a power spectrogram, the time-frequency mask generation unit 119a may be, for example, a conditional U-Net. That is, the sound feature vector, which is an amplitude spectrogram or a power spectrogram, is considered as an image, and input to a U-Net encoder configured by stacking K convolutional layers to calculate a U-Net feature map. The obtained U-Net feature map and the text embedding vector D32 are input to a U-Net decoder configured by stacking K convolutional layers to output a time-frequency mask, which is an image with the same number of pixels as the sound feature vector D22. Each convolutional layer k=1,...,K of the U-Net encoder outputs a feature map V_k with a different time-frequency resolution corresponding to that layer. The U-Net feature map is the concatenation of the feature maps V_1, V_2, , V_K of all convolutional layers. Each convolutional layer k=1, , K of the U-Net decoder accepts V_K-k+1 and the text embedding vector D32 as inputs.

[0041] Alternatively, only convolutional layer k=1 can receive the feature map V_K and the text embedding vector D32 as input, while other convolutional layers k≠1 can receive only the feature map V_K-k+1 as input without receiving the text embedding vector D32. If each convolutional layer of the U-Net decoder receives the text embedding vector D32 as input, the benefit is that accuracy will be higher if the training dataset is abundant. If only convolutional layer k=1 receives the text embedding vector D32 as input, the benefit is that the number of weight coefficients in the neural network can be kept small.

[0042] The time-frequency mask multiplication unit 119b generates a time-frequency domain representation D41 of the extracted signal by multiplying the time-frequency domain representation D21 of the mixed signal by the time-frequency mask.

[0043] The learning unit 119c learns the parameters of the feature extraction model, the text-embedded extraction model, and the time-frequency mask generation model by minimizing the value of a loss function defined by the distance between the time-frequency domain representation D41 of the extracted signal and the time-frequency domain representation D11 of the target signal.

[0044] The learning unit 119c further calculates a convergence determination function used to determine whether convergence has occurred. For example, the convergence determination function is defined by the magnitude of change in the value of the loss function from the previous iteration (learning). Alternatively, the convergence determination function may be defined by the magnitude of change in the parameters of the feature extraction model from the previous iteration (learning), the magnitude of change in the parameters of the text embedding extraction model from the previous iteration (learning), the magnitude of change in the parameters of the time-frequency mask generation model from the previous iteration (learning), or the product of these. If the change is sufficiently small, it is determined that convergence has occurred. For example, if the convergence determination function is smaller than a predetermined threshold, it is determined that convergence has occurred.

[0045] (Overview of operation) An overview of the operation of the learning subsystem 110 will be described. The learning subsystem 110 reads a set of triplet data from the training dataset database 160: a time waveform D10 of a target signal, a time waveform D20 of a mixed signal obtained by mixing the time waveform of the target signal with a signal corresponding to noise other than the target signal, and a variable-length onomatopoeia text D30.

[0046] A time waveform D10 of the target signal is input to a target signal frame division processing unit 111, a target signal window function multiplication unit 112, and a target signal frequency domain signal generation unit 113 in this order, and converted into a time-frequency domain representation D11 of the target signal.

[0047] The mixed signal time waveform D20 is input to the mixed signal frame division processing unit 114, the mixed signal window function multiplication unit 115, and the mixed signal frequency domain signal generation unit 116 in this order, and converted into a mixed signal time frequency domain representation D21.

[0048] The time-frequency domain representation D21 of the mixed signal is input to the feature extraction unit 117 and converted into a sound feature vector D22.

[0049] The variable-length onomatopoeia text D30 is input to the phoneme conversion unit 118 and converted into a variable-length phoneme string D31. The variable-length phoneme string D31 is input to the text embedding extraction unit 119 and converted into a text embedding vector D32.

[0050] The sound feature vector D22 and the text-embedded vector D32 are input to the time-frequency mask generation unit 119a, which generates a time-frequency mask.

[0051] The time-frequency domain representation D21 of the mixed signal and the time-frequency mask are multiplied in a time-frequency mask multiplier 119b to generate a time-frequency domain representation D41 of the extracted signal. The time-frequency domain representation D41 of the extracted signal is input to a training unit 119c.

[0052] The time-frequency domain representation D11 of the target signal and the time-frequency domain representation D41 of the extracted signal are input to a learning unit 119c. The learning unit 119c learns and updates the parameters of the feature extraction model, the text-embedded extraction model, and the time-frequency mask generation model. The feature extraction model, the text-embedded extraction model, and the time-frequency mask generation model with their updated parameters are stored in a feature extraction model database 120, a text-embedded extraction model database 130, and a time-frequency mask generation model database 140, respectively. For convenience, the feature extraction model, the text-embedded extraction model, and the time-frequency mask generation model with their updated parameters are also referred to as "trained models."

[0053] (Specific operation) The specific operation of the learning subsystem 110 will now be described. FIG. 4 shows an example of a processing flow of the learning subsystem 110. The learning subsystem 110 executes the processing flow of FIG. 4. The learning subsystem 110 reads a set of triplet data, namely, a time waveform D10 of a target signal, a time waveform D20 of a mixed signal, and a variable-length onomatopoeia text D30, from the training dataset database 160, and starts processing from step 400 in FIG. 4 and proceeds to step 401. In step 401, the learning subsystem 110 calculates a variable-length phoneme string D31 by phoneme conversion processing from the variable-length onomatopoeia text D30 using the phoneme conversion unit 118 (converting the onomatopoeia text D30 into the phoneme string D31).

[0054] Thereafter, the learning subsystem 110 proceeds to step 402 and determines whether the learning termination condition is met. The learning termination condition is met when either condition 1 or condition 2 described below is met. Condition 1 is met when a predetermined convergence condition is met (for example, when the convergence determination function is smaller than a predetermined threshold value). Condition 2 is met when the counter C1 is greater than the threshold value ThC (C1>ThC). Note that the learning termination condition may be condition 2 alone.

[0055] If the learning termination condition is not met, the learning subsystem 110 determines "NO" in step 402, executes the processes of steps 403 to 415 described below in order, and then returns to step 402.

[0056] Step 403: The learning subsystem 110 uses the text embedding extraction unit 119 to calculate (extract) the text embedding vector D32 from the phoneme sequence D31 using the latest text embedding extraction model.

[0057] Step 404: The learning subsystem 110 causes the mixed signal frame division processing unit 114 to divide the time waveform D20 of the mixed signal into frames, and calculates (outputs) frame division signals of the mixed signal.

[0058] Step 405: The learning subsystem 110 performs window function multiplication using the mixed signal window function multiplication unit 115 to convert the frame-divided mixed signal into a window function-multiplied mixed signal.

[0059] Step 406: The learning subsystem 110 converts the window function-multiplied signal of the mixture signal into a time-frequency domain representation D21 of the mixture signal by the mixture signal frequency domain signal generator 116.

[0060] Step 407: The learning subsystem 110 calculates the sound feature vector D22 from the time-frequency domain representation D21 of the mixed signal by the feature extraction unit 117. In this example, the learning subsystem 110 calculates the sound feature vector D22 from the time-frequency domain representation D21 of the mixed signal by the feature extraction unit 117 using the latest feature extraction model.

[0061] Step 408: The learning subsystem 110 uses the latest time-frequency mask generation model by the time-frequency mask generation unit 119a to generate a time-frequency mask from the sound feature vector D22 and the text embedding vector D32.

[0062] Step 409: The learning subsystem 110 generates a time-frequency domain representation D41 of the extracted signal by multiplying the time-frequency domain representation D21 of the mixed signal by the time-frequency mask using the time-frequency mask multiplication unit 119b.

[0063] Step 410: The learning subsystem 110 causes the target signal frame division processing unit 111 to divide the time waveform D10 of the target signal into frames, and calculates (outputs) frame-divided signals of the target signal.

[0064] Step 411: The learning subsystem 110 performs window function multiplication using the target signal window function multiplication unit 112 to convert the frame-divided signal of the target signal into a window function-multiplied signal of the target signal.

[0065] Step 412: The learning subsystem 110 performs a short-time Fourier transform using the target signal frequency domain signal generator 113 to convert the window function multiplied signal of the target signal into a time-frequency domain representation D11 of the target signal.

[0066] Step 413: The learning subsystem 110, using the learning unit 119c, learns each parameter of the feature extraction model, the text embedding extraction model, and the time-frequency mask generation model (each parameter of the neural network (NN)) by minimizing the value of a loss function defined by the distance between the time-frequency domain representation D41 of the extracted signal and the time-frequency domain representation D11 of the target signal (i.e., updates each model).

[0067] Step 414: The learning subsystem 110, using the learning unit 119c, calculates a convergence condition indicating whether convergence has occurred. The convergence condition is defined, for example, by the magnitude of change in the loss function from the previous iteration (learning). Alternatively, the convergence condition is defined by the magnitude of change in each parameter of the feature extraction model, text embedding extraction model, and time-frequency mask generation model from the previous iteration (learning). If the change is sufficiently small, it is determined that convergence has occurred (step 402).

[0068] Step 415: The learning subsystem 110 increases the current value of the counter C1 by "1".

[0069] If the learning termination condition is met in step 402, the learning subsystem 110 determines "YES" in step 402 and proceeds to step 416, where it stores the feature extraction model, text-embedded extraction model, and time-frequency mask generation model (each parameter of the neural network (NN)) in each database (the feature extraction model database 120, the text-embedded extraction model database 130, and the time-frequency mask generation model database 140). After that, the learning subsystem 110 proceeds to step 495, where it temporarily ends this processing flow.

[0070] <Sound extraction subsystem> (Sound extraction subsystem function) The configuration of the sound extraction subsystem 150 will be explained mainly for each function. Figure 5 is a block diagram for explaining an example of the configuration of the sound extraction subsystem 150 for each function.

[0071] 5 , the sound extraction subsystem 150 includes a mixed signal frame segmentation processing unit 151, a mixed signal window function multiplication unit 152, a mixed signal frequency domain signal generation unit 153, a feature extraction unit 154, a time-frequency mask generation unit 155, a phoneme conversion unit 156, a text embedding extraction unit 157, a time-frequency mask multiplication unit 158, and a phase restoration unit 159. Note that the mixed signal frame segmentation processing unit 151, the mixed signal window function multiplication unit 152, the mixed signal frequency domain signal generation unit 153, the feature extraction unit 154, the time-frequency mask generation unit 155, the phoneme conversion unit 156, the text embedding extraction unit 157, the time-frequency mask multiplication unit 158, and the phase restoration unit 159 are configured by various programs (not shown) stored in the ROM 202 and / or the storage device 204 of the information processing device 200.

[0072] The mixed signal frame division processing unit 151 divides the time waveform D50 of the mixed signal into frames, and calculates (outputs) frame division signals (not shown) of the mixed signal. The mixed signal window function multiplication unit 152 performs window function multiplication, and converts the frame division signals of the mixed signal into window function-multiplied signals of the mixed signal (not shown).

[0073] The mixed signal frequency domain signal generator 153 performs a short-time Fourier transform to convert the window function-multiplied mixed signal into a time-frequency domain representation D51 of the mixed signal. A frequency transform method such as CQT can be used instead of the short-time Fourier transform, but the same processing as that of the learning subsystem 110 is performed.

[0074] In this example, similar to the learning subsystem 110, the feature extraction unit 154 converts the time-frequency domain representation D51 of the mixed signal into a sound feature vector using the latest feature extraction model (a feature extraction model that is a neural network with variable weighting coefficient parameters). The time-frequency domain representation D51 of the mixed signal is input to the latest feature extraction model that has been updated immediately before in the learning subsystem 110, and the sound feature vector is calculated.

[0075] Note that when a logarithmic mel-power spectrogram, a time series of MFCCs, a delta or delta-delta concatenation thereof, or the like is used in the learning subsystem 110, the feature vector here may be a logarithmic mel-power spectrogram, a time series of MFCCs, a delta or delta-delta concatenation thereof, or the like corresponding to those used in the learning subsystem 110. In this case, the feature extraction unit 154 performs the same processing as the feature extraction unit 117 of the learning subsystem 110.

[0076] The phoneme conversion unit 156 outputs a variable-length phoneme string D61 by phoneme conversion processing from the variable-length onomatopoeia text D60 (converts the onomatopoeia text D60 into the phoneme string D61).

[0077] The text embedding extraction unit 157 uses the latest text embedding extraction model to calculate (extract) a text embedding vector D62 (text embedding vector D62) from the phoneme sequence D61.

[0078] The time-frequency mask generation unit 155 generates a time-frequency mask from the sound feature vector D52 and the text-embedding vector D62 using the latest time-frequency mask generation model.

[0079] The time-frequency mask multiplication unit 158 generates a time-frequency domain representation D71 of the extracted signal by multiplying the time-frequency domain representation D51 of the mixed signal by the time-frequency mask.

[0080] The phase restoration unit 159 generates a time waveform D72 of the extracted signal from the time-frequency domain representation D71 of the extracted signal using a known Griffin-Lim algorithm or the like.

[0081] (Overview of operation) An overview of the operation of the sound extraction subsystem 150 will now be described. As shown in Fig. 5, a time waveform D50 of a mixed signal is input in order to a mixed signal frame division processing unit 151, a mixed signal window function multiplication unit 152, and a mixed signal frequency domain signal generation unit 153, and converted into a time-frequency domain representation D51 of the mixed signal. The time-frequency domain representation D51 of the mixed signal is input to a feature extraction unit 154, and converted into a sound feature vector D52.

[0082] The variable-length onomatopoeia text D60 is input to the phoneme conversion unit 156 and converted into a variable-length phoneme sequence D61. The variable-length phoneme sequence D61 is input to the text embedding extraction unit 157 and converted into a text embedding vector D62. The sound feature vector D52 and the text embedding vector D62 are input to the time-frequency mask generation unit 155, which generates a time-frequency mask.

[0083] The time-frequency domain representation D51 of the mixed signal and the time-frequency mask are multiplied in a time-frequency mask multiplication unit 158 to generate a time-frequency domain representation D71 of the extracted signal. The time-frequency domain representation D71 of the extracted signal is input to a phase restoration unit 159 to generate a time waveform D72 of the extracted signal.

[0084] (Specific operation) The specific operation of the sound extraction subsystem 150 will be described. Fig. 6 shows an example of a processing flow of the sound extraction subsystem 150. The sound extraction subsystem 150 executes the processing flow of Fig. 6. When the time waveform D50 of the mixed signal and the variable-length onomatopoeia text D60 are input, the sound extraction subsystem 150 starts processing from step 600 in Fig. 6 and executes the processing of steps 601 to 609 described below in order, and then proceeds to step 695, where the processing flow is temporarily ended.

[0085] Step 601: The sound extraction subsystem 150 outputs a variable-length phoneme string D61 by phoneme conversion processing from a variable-length onomatopoeia text D60 using the phoneme conversion unit 156 (converts the onomatopoeia text D60 into a phoneme string D61).

[0086] Step 602: The sound extraction subsystem 150 uses the text embedding extraction unit 157 to calculate (extract) the text embedding vector D62 from the phoneme sequence D61 using the latest text embedding extraction model.

[0087] Step 603: The sound extraction subsystem 150 causes the mixed signal frame division processing unit 151 to divide the time waveform D50 of the mixed signal into frames, and calculates (outputs) frame division signals of the mixed signal.

[0088] Step 604: The sound extraction subsystem 150 performs window function multiplication using the mixed signal window function multiplication unit 152 to convert the frame-divided signal of the mixed signal into a window function-multiplied signal of the mixed signal.

[0089] Step 605: The sound extraction subsystem 150 performs a short-time Fourier transform using the mixed signal frequency domain signal generator 153 to convert the window function multiplied signal of the mixed signal into a time-frequency domain representation D51 of the mixed signal.

[0090] Step 606: The sound extraction subsystem 150 calculates a sound feature vector D52 from the time-frequency domain representation D51 of the mixed signal using the feature extraction unit 154. In this example, the sound extraction subsystem 150 calculates the sound feature vector D52 from the time-frequency domain representation D51 of the mixed signal using the latest feature extraction model using the feature extraction unit 154.

[0091] Step 607: The sound extraction subsystem 150 causes the time-frequency mask generation unit 155 to generate a time-frequency mask from the sound feature vector D52 and the text embedding vector D62 using the latest time-frequency mask generation model.

[0092] Step 608: The sound extraction subsystem 150 generates a time-frequency domain representation D71 of the extracted signal by multiplying the time-frequency domain representation D51 of the mixed signal by the time-frequency mask using the time-frequency mask multiplier 158.

[0093] Step 609: The sound extraction subsystem 150 uses the phase restoration unit 159 to generate a time waveform D72 of the extracted signal from the time-frequency domain representation D71 of the extracted signal, using the known Griffin-Lim algorithm or the like.

[0094] <Example> FIG. 7 shows an example of an extraction result by the sound extraction system 100. The top row is a power spectrogram of the mixed signal. The horizontal axis represents time (seconds), and the vertical axis represents frequency (kHz). White indicates time frequencies with high power, and black indicates time frequencies with low power. To make the effect of the embodiment easier to understand, the input mixed signals are all signals that are a mixture of multiple acoustic events of the same type, and since the range of sounds that the user wants to extract cannot be predefined as a certain type of event, this is an example that cannot be extracted using a conventional extraction method based on the type of event (i.e., a conventional method corresponding to the conventional technology).

[0095] The first column represents the task of extracting only the target signal corresponding to the onomatopoeia "pop" ( / poq / ) from a mixed signal containing multiple metallic sounds. The second column represents the task of extracting only the target signal corresponding to the onomatopoeia "piririririn" ( / piriririri N / ) from a mixed signal containing multiple bell sounds. The third column represents the task of extracting only the target signal corresponding to the onomatopoeia "pururururururu" ( / pururururururu / ) from a mixed signal containing multiple telephone ringing sounds. The fourth column represents the task of extracting only the target signal corresponding to the onomatopoeia "tichichichichichi" ( / ti ch i ch i ch i ch i ch i / ) from a mixed signal containing multiple percussion sounds. The fifth column represents the task of extracting only the target signal corresponding to the onomatopoeia "tottututu" ( / toqtututu / ) from a mixed signal containing the sound of rolling dice.

[0096] In each column, the first row ("Mixture sound") shows the input mixed signal, the second row ("Subclass-conditioned method") shows the result of the extraction method based on the type of event, the third row ("Onomatopoeia-conditioned method") shows the result of the extraction method of this embodiment, and the fourth row ("Ground truth") shows the target signal assumed to be correct. Comparing the mixed signal (row 1) and the conventional method (row 2), it can be seen that the extraction result of this embodiment (row 3) is similar to the correct target signal (row 4). This suggests the effectiveness of this embodiment in extracting target sounds even when the user cannot predefine the range of sounds they want to extract as a certain type of event.

[0097] <Effects> As described above, the sound extraction system 100 according to the first embodiment of the present invention can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wishes to extract from a mixed signal. Furthermore, the sound extraction system 100 according to the first embodiment can provide an infinite number of texts and can specify an infinite number of sound ranges. Therefore, even if the range of sounds that the user wishes to extract as a certain type of event cannot be defined in advance, the sound extraction system 100 according to the first embodiment can accurately extract a signal corresponding to the sound that the user wishes to extract from a mixed signal by providing text corresponding to the sound that the user wishes to extract.

[0098] <<Second embodiment>> A sound extraction system 800 according to a second embodiment of the present invention will now be described. FIG. 8 is a block diagram showing a schematic configuration example of the sound extraction system 800 according to the second embodiment of the present invention. As shown in FIG. 8, the sound extraction system 800 differs from the sound extraction system 100 according to the first embodiment only in the following respects. The sound extraction system 800 does not include the learning subsystem 110 of the sound extraction system 100 according to the first embodiment, and uses a feature extraction model database 820, a text-embedded extraction model database 830, and a time-frequency mask generation model database 840, which store feature extraction models, text-embedded extraction models, and time-frequency mask generation models that have been trained in advance based on a correspondence database between general environmental sounds and onomatopoeia. The following description will focus on these differences.

[0099] 8, the sound extraction system 800 includes a sound extraction subsystem 150, a feature extraction model database 820, a text-embedded extraction model database 830, and a time-frequency mask generation model database 840. When the time waveform of a mixed signal and a variable-length onomatopoeia text are input, the sound extraction subsystem 150 outputs the time waveform of the extracted signal using an existing feature extraction model, text-embedded extraction model, and time-frequency mask generation model. Note that the details of this processing are the same as those in the first embodiment except for the use of an existing feature extraction model, text-embedded extraction model, and time-frequency mask generation model, and therefore will not be described again.

[0100] <Effects> As described above, the sound extraction system 800 according to the second embodiment of the present invention, like the first embodiment, can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wants to extract from a mixed signal. Furthermore, the sound extraction system 800 according to the second embodiment can use a feature extraction model, a text-embedded extraction model, and a time-frequency mask generation model that have been trained in advance based on a correspondence database between general environmental sounds and onomatopoeia, and therefore does not require new learning processing by the learning subsystem 110 as in the sound extraction system 100 according to the first embodiment. This has the advantage of not requiring the construction of a new training dataset for each site.

[0101] <<Third Embodiment>> A sound extraction system 900 according to a third embodiment of the present invention will now be described. FIG. 9 is a block diagram showing an example of the schematic configuration of the sound extraction system 900 according to the third embodiment of the present invention. As shown in FIG. 9, in the sound extraction system 900, a learning subsystem 110 uses a feature extraction model database 920, a text-embedded extraction model database 930, and a time-frequency mask generation model database 940, which store existing feature extraction models, text-embedded extraction models, and time-frequency mask generation models that have been trained in advance based on a database of correspondences between general environmental sounds and onomatopoeia. The learning subsystem 110 learns using a learning dataset (training dataset) for each site, thereby optimizing the model to suit the site and improving accuracy. The sound extraction system 900 according to the third embodiment differs from the sound extraction system 100 according to the first embodiment only in the above respects. Therefore, the following description will focus on these differences.

[0102] 9, the sound extraction system 900 has a configuration in which a feature extraction model database 920, a text-embedded extraction model database 930, and a time-frequency mask generation model database 940 are added to the sound extraction system 100 according to the first embodiment. For convenience, the existing models stored in the feature extraction model database 920, the text-embedded extraction model database 930, and the time-frequency mask generation model database 940 are also referred to as "initial feature extraction models, initial text-embedded extraction models, and initial time-frequency mask generation models," and these are also referred to as "initial trained models."

[0103] The learning subsystem 110 learns using a learning dataset (training dataset) for each site, thereby optimizing (updating) the models (existing feature extraction model, text-embedded extraction model, and time-frequency mask generation model) to suit the site, and stores each optimized model in the feature extraction model database 120, the text-embedded extraction model database 130, and the time-frequency mask generation model database 140, respectively.

[0104] When the sound extraction subsystem 150 receives the time waveform of the mixed signal and the variable-length onomatopoeia text, it outputs the time waveform of the extracted signal using models that are optimized from existing feature extraction models, text-embedded extraction models, and time-frequency mask generation models. Note that the details of this process are the same as those in the first embodiment except for the use of optimized existing models for feature extraction models, text-embedded extraction models, and time-frequency mask generation models, and therefore will not be described here.

[0105] <Effects> As described above, the sound extraction system 900 according to the third embodiment of the present invention can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wants to extract from a mixed signal, similar to the first embodiment. Furthermore, the sound extraction system 900 according to the third embodiment has the advantage that it is possible to improve the accuracy of the model according to the site, while using an existing model, so that only a small number of training data sets need to be newly constructed for each site.

[0106] <<Fourth Embodiment>> A sound extraction system 1000 according to a fourth embodiment of the present invention will be described. Fig. 10 is a block diagram showing a schematic configuration example of the sound extraction system 1000 according to the fourth embodiment of the present invention. As shown in Fig. 10, the sound extraction system 1000 differs from the sound extraction system 100 according to the first embodiment only in that it uses explanatory text (for example, "A clatter sound is followed by a boom," "A shocking sound is followed by a clatter sound," etc.) instead of onomatopoeia as text representing the range of sounds. Therefore, the following description will mainly focus on this difference.

[0107] <Learning Subsystem> (Functions of the learning subsystem) Fig. 11 is a block diagram for explaining each function of an example configuration of the learning subsystem 110 of the sound extraction system 1000. As shown in Fig. 11, the learning subsystem 110 includes a target signal frame segmentation processing unit 111, a target signal window function multiplication unit 112, a target signal frequency domain signal generation unit 113, a mixed signal frame segmentation processing unit 114, a mixed signal window function multiplication unit 115, a mixed signal frequency domain signal generation unit 116, a feature extraction unit 117, a text embedding extraction unit 119, a time-frequency mask generation unit 119a, a time-frequency mask multiplication unit 119b, and a learning unit 119c.

[0108] (Overview of operation) An overview of the operation of the learning subsystem 110 will be given below. The learning subsystem 110 reads from the training dataset database 160 a set of triplet data: a time waveform D10 of a target signal (a signal corresponding to a sound to be extracted), a time waveform D20 of a mixed signal obtained by mixing the time waveform of the target signal with a signal corresponding to noise other than the target signal (noise other than the sound to be extracted), and a variable-length explanatory text D1100 (an explanatory text corresponding to a sound to be extracted).

[0109] A time waveform D10 of the target signal is input to a target signal frame division processing unit 111, a target signal window function multiplication unit 112, and a target signal frequency domain signal generation unit 113 in this order, and converted into a time-frequency domain representation D11 of the target signal.

[0110] The mixed signal time waveform D20 is input to the mixed signal frame division processing unit 114, the mixed signal window function multiplication unit 115, and the mixed signal frequency domain signal generation unit 116 in this order, and converted into a mixed signal time frequency domain representation D21.

[0111] The time-frequency domain representation D21 of the mixed signal is input to the feature extraction unit 117 and converted into a sound feature vector D22.

[0112] The variable-length explanatory text D1100 is input to the text embedding extraction unit 119, and converted into an embedding vector D1101 of the explanatory text D1100 (text embedding vector D1101).

[0113] The sound feature vector D22 and the text-embedded vector D1101 are input to the time-frequency mask generation unit 119a, which generates a time-frequency mask.

[0114] The time-frequency domain representation D21 of the mixed signal and the time-frequency mask are multiplied in a time-frequency mask multiplier 119b to generate a time-frequency domain representation D41 of the extracted signal.

[0115] When the time-frequency domain representation D11 of the target signal and the time-frequency domain representation D41 of the extracted signal are input to the learning unit 119c, the parameters of the feature extraction model, the text-embedded extraction model, and the time-frequency mask generation model are learned and updated.

[0116] The feature extraction model, text-embedded extraction model, and time-frequency mask generation model with their parameters updated are stored in the feature extraction model database 120, the text-embedded extraction model database 130, and the time-frequency mask generation model database 140, respectively.

[0117] (Specific operation) The specific operation of the learning subsystem 110 will be described. FIG. 12 shows an example of a processing flow of the learning subsystem 110. The learning subsystem 110 executes the processing flow of FIG. 12. After reading a set of triplet data, namely, a time waveform D10 of a target signal, a time waveform D20 of a mixed signal, and a variable-length explanatory text D1100, from the training dataset database 160, the learning subsystem 110 starts processing at step 1200 in FIG. 12 and proceeds to step 1201, where it determines whether a learning termination condition is met. The learning termination condition is met when either condition 1 or condition 2 described below is met. Condition 1 is met when a predetermined convergence condition is met (for example, when a convergence determination function is smaller than a predetermined threshold value). Condition 2 is met when a counter C1 is greater than a threshold value ThC (C1>ThC). Note that the learning termination condition may be condition 2 alone.

[0118] If the learning termination condition is not met, the learning subsystem 110 determines "NO" in step 1201, executes the processes of steps 1202 to 1214 described below in order, and then returns to step 1201.

[0119] Step 1202: The learning subsystem 110 uses the text embedding extraction unit 119 to calculate (extract) the text embedding vector D1101 from the variable-length explanatory text D1100 using the latest text embedding extraction model.

[0120] Step 1203: The learning subsystem 110 causes the mixed signal frame division processing unit 114 to divide the time waveform of the mixed signal into frames, and calculates (outputs) frame division signals of the mixed signal.

[0121] Step 1204: The learning subsystem 110 performs window function multiplication using the mixed signal window function multiplication unit 115 to convert the frame-divided mixed signal into a window function-multiplied mixed signal.

[0122] Step 1205: The learning subsystem 110 converts the window function-multiplied signal of the mixture signal into a time-frequency domain representation D21 of the mixture signal by the mixture signal frequency domain signal generator 116.

[0123] Step 1206: The learning subsystem 110 calculates the sound feature vector D22 from the time-frequency domain representation D21 of the mixed signal by the feature extraction unit 117. In this example, the learning subsystem 110 calculates the sound feature vector D22 from the time-frequency domain representation D21 of the mixed signal by the feature extraction unit 117 using the latest feature extraction model.

[0124] Step 1207: The learning subsystem 110 uses the latest time-frequency mask generation model by the time-frequency mask generation unit 119a to generate a time-frequency mask from the sound feature vector D22 and the text-embedding vector D1101.

[0125] Step 1208: The learning subsystem 110 generates a time-frequency domain representation D41 of the extracted signal by multiplying the time-frequency domain representation D21 of the mixed signal by the time-frequency mask using the time-frequency mask multiplication unit 119b.

[0126] Step 1209: The learning subsystem 110 causes the target signal frame division processing unit 111 to divide the time waveform D10 of the target signal into frames, and calculates (outputs) frame-divided signals of the target signal.

[0127] Step 1210: The learning subsystem 110 performs window function multiplication using the target signal window function multiplication unit 112 to convert the frame-divided signal of the target signal into a window function-multiplied signal of the target signal.

[0128] Step 1211: The learning subsystem 110 performs a short-time Fourier transform using the target signal frequency domain signal generator 113 to convert the target signal multiplied by the window function into a time-frequency domain representation D11 of the target signal.

[0129] Step 1212: The learning subsystem 110, using the learning unit 119c, learns (updates) each parameter of the feature extraction model, the text embedding extraction model, and the time-frequency mask generation model (each parameter of the neural network (NN)) by minimizing the value of a loss function defined by the distance between the time-frequency domain representation D41 of the extracted signal and the time-frequency domain representation D11 of the target signal.

[0130] Step 1213: The learning subsystem 110 calculates a convergence condition indicating whether convergence has occurred. The convergence condition is defined, for example, by the magnitude of change in the loss function from the previous iteration (learning). Alternatively, the convergence condition is defined by the magnitude of change in each parameter of the feature extraction model, text embedding extraction model, and time-frequency mask generation model from the previous iteration (learning). If the change is sufficiently small, it is determined that convergence has occurred (step 1201).

[0131] Step 1214: The learning subsystem 110 increases the current value of the counter C1 by "1".

[0132] If the learning termination condition is met in step 1201, the learning subsystem 110 determines "YES" in step 1201 and proceeds to step 1215, where it saves the feature extraction model, text embedding extraction model, and time-frequency mask generation model (each parameter of the neural network (NN)) in each database. Thereafter, the learning subsystem 110 proceeds to step 1295, where it temporarily ends this processing flow.

[0133] <Sound extraction subsystem> (Sound extraction subsystem function) The configuration of the sound extraction subsystem 150 will be explained below, mainly for each function. Fig. 13 is a block diagram for explaining an example of the configuration of the sound extraction subsystem 150 for each function.

[0134] As shown in FIG. 13 , the sound extraction subsystem 150 includes a mixed signal frame segmentation processing unit 151, a mixed signal window function multiplication unit 152, a mixed signal frequency domain signal generation unit 153, a feature extraction unit 154, a time-frequency mask generation unit 155, a text embedding extraction unit 157, a time-frequency mask multiplication unit 158, and a phase restoration unit 159.

[0135] (Overview of operation) 13, a time waveform D50 of a mixed signal is input to a mixed signal frame division processing unit 151, a mixed signal window function multiplication unit 152, and a mixed signal frequency domain signal generation unit 153 in that order, and converted into a time-frequency domain representation D51 of the mixed signal. The time-frequency domain representation D51 of the mixed signal is input to a feature extraction unit 154, and converted into a feature vector D52 of the sound.

[0136] The variable-length description text D1300 is input to the text embedding extraction unit 157, which converts it into an embedding vector D1301 (text embedding vector D1301) of the description text D1300. The sound feature vector D52 and the text embedding vector D1301 are input to the time-frequency mask generation unit 155, which generates a time-frequency mask.

[0137] The time-frequency domain representation D51 of the mixed signal and the time-frequency mask are multiplied in a time-frequency mask multiplication unit 158 to generate a time-frequency domain representation D71 of the extracted signal. The time-frequency domain representation D71 of the extracted signal is input to a phase restoration unit 159 to generate a time waveform D72 of the extracted signal.

[0138] (Specific operation) Fig. 14 shows an example of a processing flow of the sound extraction subsystem 150. The sound extraction subsystem 150 executes the processing flow of Fig. 14. When the time waveform D50 of the mixed signal and the variable-length explanatory text D1300 are input, the sound extraction subsystem 150 starts processing from step 1400 in Fig. 14 and sequentially executes the processing of steps 1401 to 1408 described below, and then proceeds to step 1495, where the processing flow is temporarily ended.

[0139] Step 1401: The sound extraction subsystem 150 uses the latest text embedding extraction model in the text embedding extraction unit 157 to calculate (extract) the text embedding vector D1301 from the variable-length explanatory text D1300.

[0140] Step 1402: The sound extraction subsystem 150 causes the mixed signal frame division processing unit 151 to divide the time waveform of the mixed signal into frames, and calculates (outputs) frame-divided signals of the mixed signal.

[0141] Step 1403: The sound extraction subsystem 150 performs window function multiplication using the mixed signal window function multiplication unit 152 to convert the frame-divided mixed signal into a window function-multiplied mixed signal.

[0142] Step 1404: The sound extraction subsystem 150 performs a short-time Fourier transform using the mixed signal frequency domain signal generator 153 to convert the window function multiplied signal of the mixed signal into a time-frequency domain representation D51 of the mixed signal.

[0143] Step 1405: The sound extraction subsystem 150 calculates a sound feature vector D52 from the time-frequency domain representation D51 of the mixed signal using the feature extraction unit 154. In this example, the sound extraction subsystem 150 calculates the sound feature vector D52 from the time-frequency domain representation D51 of the mixed signal using the latest feature extraction model using the feature extraction unit 154.

[0144] Step 1406: The sound extraction subsystem 150 causes the time-frequency mask generation unit 155 to generate a time-frequency mask from the sound feature vector D52 and the text-embedding vector D1301 using the latest time-frequency mask generation model.

[0145] Step 1407: The sound extraction subsystem 150 generates a time-frequency domain representation D71 of the extracted signal by multiplying the time-frequency domain representation D51 of the mixed signal by the time-frequency mask using the time-frequency mask multiplier 158.

[0146] Step 1408: The sound extraction subsystem 150 generates a time waveform D72 of the extracted signal from the time-frequency domain representation D71 of the extracted signal using, for example, the well-known Griffin-Lim algorithm.

[0147] <Effects> As described above, the sound extraction system 1000 according to the fourth embodiment of the present invention can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wishes to extract from a mixed signal. With such a basic configuration, the sound extraction system 1000 according to the fourth embodiment can extract sound even when the range of sounds that the user wishes to extract as a certain type of event cannot be defined in advance. The explanatory text is relatively general-purpose and can be used across a variety of application sites.

[0148] <<Fifth Embodiment>> A sound extraction system 1500 according to a fifth embodiment of the present invention will be described. FIG. 15 is a block diagram showing an example of the schematic configuration of the sound extraction system 1500 according to the fifth embodiment of the present invention. As shown in FIG. 15, the sound extraction system 1500 includes a learning subsystem 110, a signal extraction model database 1510, a text-embedded extraction model database 130, a sound extraction subsystem 150, and a training dataset database 160. Differences from the first embodiment shown in FIG. 1 will be described. The learning subsystem 110 executes a learning process, outputs a signal extraction model and a text-embedded extraction model, and stores them in the respective databases. That is, the learning subsystem 110 stores the signal extraction model in the signal extraction model database 1510, and stores the text-embedded extraction model in the text-embedded extraction model database 130.

[0149] The sound extraction subsystem 150 reads the signal extraction model and the text-embedded extraction model from the databases (the signal extraction model database 1510 and the text-embedded extraction model database 130) and performs sound extraction processing based on (using) them. As a result, the sound extraction subsystem 150 extracts the time waveform of the extracted signal from the time waveform of the mixed signal and the variable-length onomatopoeia text. Furthermore, the sound extraction subsystem 150 outputs the time waveform of the extracted signal.

[0150] <Learning Subsystem> (Functions of the learning subsystem) The configuration of the learning subsystem 110 will be described below, mainly for each function. Fig. 16 is a block diagram for explaining an example configuration of the learning subsystem 110 for each function. As shown in Fig. 16, the learning subsystem 110 includes a phoneme conversion unit 118, a text-embedded extraction unit 119, a signal extraction unit 1600, and a learning unit 119c. Note that the phoneme conversion unit 118, the text-embedded extraction unit 119, the signal extraction unit 1600, and the learning unit 119c are configured by various programs (not shown) stored in the ROM 202 and / or the storage device 204 of the information processing device 200.

[0151] The signal extraction unit 1600 generates a time waveform D1600 of an extracted signal from the time waveform D20 of the mixed signal and the text-embedding vector D32 using the latest signal extraction model.

[0152] The signal extraction model is a neural network that receives the time waveform D20 of the mixed signal and the text embedding vector D32 as input and outputs the time waveform D1600 of the extracted signal. The signal extraction model may be, for example, a neural network consisting of only a fully connected layer, or a neural network in which multiple convolutional layers, activation functions, and pooling layers are stacked, with a self-attention layer or skip connections sandwiched between them. When using a time-frequency mask as in the first embodiment, a time-frequency representation is required. However, the time-frequency representation is not necessarily suitable in terms of extraction accuracy. In contrast, the signal extraction model here directly inputs the time waveform into the neural network, which has the advantage that a representation with high extraction accuracy can be obtained if the training dataset is sufficiently large.

[0153] Furthermore, neural networks that are composed only of fully connected layers have the advantage of high extraction accuracy when the training dataset is large, while neural networks that have multiple convolutional layers, activation functions, and pooling layers stacked together, with self-attention layers and skip connections sandwiched between them, have the advantage of high extraction accuracy even when the training dataset is small.

[0154] The signal extraction model may be a model such as the well-known Conv-TasNet, which includes an encoder that inputs the time waveform D20 of the mixed signal and outputs a feature vector time series; a time feature mask generation neural network that inputs the feature vector time series and the embedding vector D32 and calculates a two-dimensional mask (time feature mask) of the time axis and the feature axis; a multiplication mechanism that multiplies the feature vector time series by the time feature mask to calculate the extracted feature vector time series; and a decoder that inputs the extracted feature vector time series and generates the time waveform D1600 of the extracted signal. Both the encoder and the decoder are, for example, neural networks consisting of one-dimensional convolutional layers. The time feature mask generation neural network may be a neural network consisting of only a fully connected layer, or may be a neural network consisting of multiple convolutional layers, activation functions, and pooling layers stacked with a self-attention layer or skip connections interposed between them. When using a time-frequency mask as in the first embodiment, a time-frequency representation is required, but using the time-frequency representation does not necessarily result in high extraction accuracy. In contrast, a signal extraction model that converts to temporal features within a neural network uses a temporal feature representation that has been trained to achieve high extraction accuracy, which has the advantage of providing higher extraction accuracy than when using a time-frequency representation.

[0155] The learning unit 119c learns the parameters of the signal extraction model and the text-embedded extraction model by minimizing the value of a loss function defined by the distance between the time waveform D1600 of the extracted signal and the time waveform D10 of the target signal.

[0156] The learning unit 119c further calculates a convergence determination function used to determine whether convergence has occurred. For example, the convergence determination function is defined by the magnitude of change in the loss function value from the previous iteration (learning). The convergence determination function may also be defined by the magnitude of change in the signal extraction model parameters from the previous iteration (learning), the magnitude of change in the text embedding extraction model parameters from the previous iteration (learning), or the product of these. If the change is sufficiently small, convergence is determined. For example, convergence is determined if the convergence determination function is smaller than a predetermined threshold.

[0157] (Overview of operation) An overview of the operation of the learning subsystem 110 will be described. The learning subsystem 110 reads a set of triplet data from the training dataset database 160: a time waveform D10 of a target signal, a time waveform D20 of a mixed signal obtained by mixing the time waveform of the target signal with a signal corresponding to noise other than the target signal, and a variable-length onomatopoeia text D30.

[0158] The variable-length onomatopoeia text D30 is input to the phoneme conversion unit 118 and converted into a variable-length phoneme string D31. The variable-length phoneme string D31 is input to the text embedding extraction unit 119 and converted into a text embedding vector D32.

[0159] The time waveform D20 of the mixed signal and the text-embedded vector D32 are input to a signal extraction unit 1600, which generates a time waveform D1600 of the extracted signal.

[0160] The time waveform D10 of the target signal and the time waveform D1600 of the extracted signal are input to the learning unit 119c. The learning unit 119c learns and updates the parameters of the signal extraction model and the text-embedded extraction model. The signal extraction model and the text-embedded extraction model with their parameters updated are stored in the signal extraction model database 1510 and the text-embedded extraction model database 130, respectively. For convenience, the signal extraction model and the text-embedded extraction model with their parameters updated are also referred to as the "trained model."

[0161] (Specific operation) The specific operation of the learning subsystem 110 will now be described. FIG. 17 shows an example of a processing flow of the learning subsystem 110. The learning subsystem 110 executes the processing flow of FIG. 17. The learning subsystem 110 reads a set of triplet data, namely, a time waveform D10 of a target signal, a time waveform D20 of a mixed signal, and a variable-length onomatopoeia text D30, from the training dataset database 160, and starts processing at step 1700 in FIG. 17 and proceeds to step 1701. In step 1701, the learning subsystem 110 calculates a variable-length phoneme string D31 by phoneme conversion processing from the variable-length onomatopoeia text D30 using the phoneme conversion unit 118 (converting the onomatopoeia text D30 into the phoneme string D31).

[0162] Thereafter, the learning subsystem 110 proceeds to step 1702 and determines whether the learning termination condition is met. The learning termination condition is met when either condition 1 or condition 2 described below is met. Condition 1 is met when a predetermined convergence condition is met (for example, when the convergence determination function is smaller than a predetermined threshold value). Condition 2 is met when the counter C1 is greater than the threshold value ThC (C1>ThC). Note that the learning termination condition may be condition 2 alone.

[0163] If the learning end condition is not met, the learning subsystem 110 determines "NO" in step 1702, executes the processes of steps 1703 to 1707 described below in order, and then returns to step 1702.

[0164] Step 1703: The learning subsystem 110 uses the text embedding extraction unit 119 to calculate (extract) the text embedding vector D32 from the phoneme sequence D31 using the latest text embedding extraction model.

[0165] Step 1704: The learning subsystem 110 causes the signal extractor 1600 to generate a time waveform D1600 of an extracted signal from the time waveform D20 of the mixed signal and the text-embedding vector D32 using the latest signal extraction model.

[0166] Step 1705: The learning subsystem 110, using the learning unit 119c, learns each parameter of the signal extraction model and the text-embedded extraction model (each parameter of the neural network (NN)) by minimizing the value of a loss function defined by the distance between the time waveform D1600 of the extracted signal and the time waveform D10 of the target signal (i.e., updates each model).

[0167] Step 1706: The learning subsystem 110, using the learning unit 119c, calculates a convergence condition indicating whether convergence has occurred. The convergence condition is defined, for example, by the magnitude of change in the loss function from the previous iteration (learning). Alternatively, the convergence condition is defined by the magnitude of change in each parameter of the signal extraction model and the text embedding extraction model from the previous iteration (learning). If the change is sufficiently small, it is determined that convergence has occurred (step 1702).

[0168] Step 1707: The learning subsystem 110 increases the current value of the counter C1 by "1".

[0169] If the learning termination condition is met in step 1702, the learning subsystem 110 determines "YES" in step 1702 and proceeds to step 1708, where it saves the signal extraction model and the text-embedded extraction model (each parameter of the neural network (NN)) in each database (the signal extraction model database 1510 and the text-embedded extraction model database 130). Thereafter, the learning subsystem 110 proceeds to step 1795, where it temporarily ends this processing flow.

[0170] <Sound extraction subsystem> (Sound extraction subsystem function) The configuration of the sound extraction subsystem 150 will be explained mainly for each function. Fig. 18 is a block diagram for explaining an example of the configuration of the sound extraction subsystem 150 for each function.

[0171] 18, the sound extraction subsystem 150 includes a phoneme conversion unit 156, a text-embedded extraction unit 157, and a signal extraction unit 1800. The phoneme conversion unit 156, the text-embedded extraction unit 157, and the signal extraction unit 1800 are configured by various programs (not shown) stored in the ROM 202 and / or the storage device 204 of the information processing device 200.

[0172] The phoneme conversion unit 156 outputs a variable-length phoneme string D61 by phoneme conversion processing from the variable-length onomatopoeia text D60 (converts the onomatopoeia text D60 into the phoneme string D61).

[0173] The text embedding extraction unit 157 uses the latest text embedding extraction model to calculate (extract) a text embedding vector D62 (text embedding vector D62) from the phoneme sequence D61.

[0174] The signal extraction unit 1800 uses the latest signal extraction model to generate a time waveform D72 of an extracted signal from the time waveform D50 of the mixed signal and the text-embedding vector D62.

[0175] (Specific operation) A specific operation of the sound extraction subsystem 150 will now be described. Fig. 19 shows an example of a processing flow of the sound extraction subsystem 150. The sound extraction subsystem 150 executes the processing flow of Fig. 19. When the time waveform D50 of the mixed signal and the variable-length onomatopoeia text D60 are input, the sound extraction subsystem 150 starts processing from step 1900 in Fig. 19 and executes the processing of steps 1901 to 1903 described below in order, and then proceeds to step 1995, where the processing flow is temporarily ended.

[0176] Step 1901: The sound extraction subsystem 150 outputs a variable-length phoneme string D61 by phoneme conversion processing from the variable-length onomatopoeia text D60 using the phoneme conversion unit 156 (converts the onomatopoeia text D60 into the phoneme string D61).

[0177] Step 1902: The sound extraction subsystem 150 uses the text embedding extraction unit 157 to calculate (extract) the text embedding vector D62 from the phoneme sequence D61 using the latest text embedding extraction model.

[0178] Step 1903: The sound extraction subsystem 150 uses the latest signal extraction model to generate, by the signal extraction unit 1800, the time waveform D72 of the extracted signal from the time waveform D50 of the mixed signal and the text-embedding vector D62.

[0179] <Effects> As described above, the sound extraction system 1500 according to the fifth embodiment of the present invention can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wants to extract from a mixed signal. Furthermore, the sound extraction system 1500 according to the fifth embodiment can provide an infinite number of texts and specify an infinite number of sound ranges. Therefore, even if the user cannot predefine the range of sounds that the user wants to extract as a certain type of event, the sound extraction system 1500 according to the fifth embodiment can accurately extract a signal corresponding to the sound the user wants to extract from a mixed signal by providing text corresponding to the sound the user wants to extract. Furthermore, unlike the first embodiment, the sound extraction system 1500 according to the fifth embodiment directly inputs the time waveform D50 of the mixed signal to a neural network without using a time-frequency representation, thereby avoiding the reduction in extraction accuracy that occurs when using a time-frequency representation. Furthermore, the sound extraction system 1500 according to the fifth embodiment generates the time waveform D72 of the extracted signal without undergoing phase restoration processing, which has the advantage of eliminating distortion that occurs when undergoing phase restoration processing.

[0180] <<Sixth Embodiment>> A sound extraction system 2000 according to a sixth embodiment of the present invention will be described. Fig. 20 is a block diagram showing a schematic configuration example of the sound extraction system 2000 according to the sixth embodiment of the present invention. As shown in Fig. 20, the sound extraction system 2000 differs from the sound extraction system 1500 according to the fifth embodiment only in the following points.

[0181] The sound extraction system 2000 does not include the learning subsystem 110 of the sound extraction system 1500 according to the fifth embodiment, and instead uses a signal extraction model database 2010 and a text-embedded extraction model database 830, which store signal extraction models and text-embedded extraction models that have been trained in advance based on a correspondence database between general environmental sounds and onomatopoeia. The following explanation will focus on this difference.

[0182] 20, the sound extraction system 2000 includes a sound extraction subsystem 150, a signal extraction model database 2010, and a text-embedded extraction model database 830. When the time waveform of a mixed signal and a variable-length onomatopoeia text are input, the sound extraction subsystem 150 outputs the time waveform of an extracted signal using an existing signal extraction model and text-embedded extraction model. Note that the details of this processing are the same as those in the fifth embodiment except for the use of an existing signal extraction model and text-embedded extraction model, and therefore will not be described here.

[0183] <Effects> As described above, the sound extraction system 2000 according to the sixth embodiment of the present invention, like the fifth embodiment, can accurately extract (extract or emphasize) a signal corresponding to a sound the user wants to extract from a mixed signal. Furthermore, the sound extraction system 2000 according to the sixth embodiment directly inputs the time waveform of the mixed signal to a neural network without using a time-frequency representation, thereby avoiding the reduction in extraction accuracy that accompanies the use of a time-frequency representation. Furthermore, the sound extraction system 2000 according to the sixth embodiment generates the time waveform of the extracted signal without undergoing phase restoration processing, which has the advantage of eliminating distortion that accompanies phase restoration processing. Furthermore, the sound extraction system 2000 according to the sixth embodiment can use a signal extraction model and a text-embedded extraction model that have been trained in advance based on a database of correspondences between general environmental sounds and onomatopoeia. Therefore, unlike the sound extraction system 1500 according to the fifth embodiment, a new learning process by the learning subsystem 110 is not required. This has the advantage of eliminating the need to construct a new training dataset for each site.

[0184] <<Seventh Embodiment>> A sound extraction system 2100 according to a seventh embodiment of the present invention will be described. FIG. 21 is a block diagram showing an example of the schematic configuration of the sound extraction system 2100 according to the seventh embodiment of the present invention. As shown in FIG. 21, in the sound extraction system 2100, the learning subsystem 110 uses a signal extraction model database 2110 and a text-embedded extraction model database 930, which store existing signal extraction models and text-embedded extraction models trained in advance based on a database of correspondences between general environmental sounds and onomatopoeia. The learning subsystem 110 learns using a learning dataset (training dataset) for each site, thereby optimizing the model to suit the site and improving accuracy. The sound extraction system 2100 according to the seventh embodiment differs from the sound extraction system 1500 according to the fifth embodiment only in the above respects. Therefore, the following description will focus on these differences.

[0185] 21, the sound extraction system 2100 has a configuration in which a signal extraction model database 2110 and a text-embedded extraction model database 930 are added to the sound extraction system 1500 according to the fifth embodiment. For convenience, the existing models stored in the signal extraction model database 2110 and the text-embedded extraction model database 930 are also referred to as "initial signal extraction models and initial text-embedded extraction models," and these are also referred to as "initial trained models."

[0186] The learning subsystem 110 learns using a learning dataset (training dataset) for each site, thereby optimizing (updating) the models (existing signal extraction models and text-embedded extraction models) to suit the site, and stores each optimized model in the signal extraction model database 1510 and the text-embedded extraction model database 130, respectively.

[0187] When the sound extraction subsystem 150 receives the time waveform of the mixed signal and the variable-length onomatopoeia text, it outputs the time waveform of the extracted signal using models that are optimized from existing signal extraction models and text-embedded extraction models. Note that the details of this process are the same as those in the fifth embodiment except for the use of a signal extraction model and a text-embedded extraction model that are optimized from existing models, and therefore will not be described here.

[0188] <Effects> As described above, the sound extraction system 2100 according to the seventh embodiment of the present invention, like the fifth embodiment, can accurately extract (extract or emphasize) a signal corresponding to a sound that a user wants to extract from a mixed signal. Furthermore, the sound extraction system 2100 according to the seventh embodiment can avoid the reduction in extraction accuracy that accompanies the use of time-frequency representation by directly inputting the time waveform of the mixed signal to a neural network without going through time-frequency representation. Furthermore, the sound extraction system 2100 according to the seventh embodiment generates the time waveform of the extracted signal without going through phase restoration processing, which has the advantage of not generating distortion that accompanies going through phase restoration processing. Furthermore, the sound extraction system 2100 according to the seventh embodiment has the advantage of improving the accuracy of the model according to the site, while using an existing model, which allows only a small number of training datasets to be newly constructed for each site.

[0189] <<Eighth Embodiment>> A sound extraction system 2200 according to an eighth embodiment of the present invention will be described. Fig. 22 is a block diagram showing a schematic configuration example of the sound extraction system 2200 according to the eighth embodiment of the present invention. As shown in Fig. 22, the sound extraction system 2200 differs from the sound extraction system 1500 according to the fifth embodiment only in that it uses explanatory text (for example, "A clatter sound is followed by a boom," "A shocking sound is followed by a clatter sound," etc.) instead of onomatopoeia as text representing the range of sounds. Therefore, the following description will mainly focus on this difference.

[0190] <Learning Subsystem> (Functions of the learning subsystem) Fig. 23 is a block diagram for explaining a functional example of the configuration of the learning subsystem 110 of the sound extraction system 2200. As shown in Fig. 23, the learning subsystem 110 includes a text-embedded extraction unit 119, a signal extraction unit 1600, and a learning unit 119c.

[0191] (Overview of operation) An overview of the operation of the learning subsystem 110 will be given below. The learning subsystem 110 reads from the training dataset database 160 a set of triplet data: a time waveform D10 of a target signal (a signal corresponding to a sound to be extracted), a time waveform D20 of a mixed signal obtained by mixing the time waveform of the target signal with a signal corresponding to noise other than the target signal (noise other than the sound to be extracted), and a variable-length explanatory text D1100 (an explanatory text corresponding to a sound to be extracted).

[0192] The variable-length explanatory text D1100 is input to the text embedding extraction unit 119, and converted into an embedding vector D1101 of the explanatory text D1100 (text embedding vector D1101).

[0193] The time waveform D20 of the mixed signal and the text-embedded vector D1101 are input to a signal extraction unit 1600, which generates a time waveform D1600 of the extracted signal.

[0194] When the time waveform D10 of the target signal and the time waveform D1600 of the extracted signal are input to the learning unit 119c, the parameters of the signal extraction model and the text-embedded extraction model are learned and updated.

[0195] The signal extraction model and text-embedded extraction model with their parameters updated are stored in the signal extraction model database 1510 and the text-embedded extraction model database 130, respectively.

[0196] (Specific operation) The specific operation of the learning subsystem 110 will be described. FIG. 24 shows an example of a processing flow of the learning subsystem 110. The learning subsystem 110 executes the processing flow of FIG. 24. After reading a set of triplet data, namely, the time waveform D10 of the target signal, the time waveform D20 of the mixed signal, and the variable-length explanatory text D1100, from the training dataset database 160, the learning subsystem 110 starts processing at step 2400 in FIG. 24 and proceeds to step 2401, where it determines whether or not a learning termination condition is met. The learning termination condition is met when either condition 1 or condition 2 described below is met. Condition 1 is met when a predetermined convergence condition is met (for example, when the convergence determination function is smaller than a predetermined threshold value). Condition 2 is met when the counter C1 is greater than the threshold value ThC (C1>ThC). Note that the learning termination condition may be only condition 2.

[0197] If the learning end condition is not met, the learning subsystem 110 determines "NO" in step 2401, executes the processes of steps 2402 to 2406 described below in order, and then returns to step 2401.

[0198] Step 2402: The learning subsystem 110 uses the text embedding extraction unit 119 to calculate (extract) the text embedding vector D1101 from the variable-length explanatory text D1100 using the latest text embedding extraction model.

[0199] Step 2403: The learning subsystem 110 uses the latest signal extraction model by the signal extraction unit 1600 to generate a time waveform D1600 of an extracted signal from the time waveform D20 of the mixed signal and the text-embedding vector D1101.

[0200] Step 2404: The learning subsystem 110, using the learning unit 119c, learns (updates) each parameter of the signal extraction model and the text-embedded extraction model (each parameter of the neural network (NN)) by minimizing the value of a loss function defined by the distance between the time waveform D1600 of the extracted signal and the time waveform D10 of the target signal.

[0201] Step 2405: The learning subsystem 110 calculates a convergence condition that indicates whether convergence has occurred. The convergence condition is defined, for example, by the magnitude of change in the loss function from the previous iteration (learning). Alternatively, the convergence condition is defined by the magnitude of change in each parameter of the signal extraction model and the text embedding extraction model from the previous iteration (learning). If the change is sufficiently small, it is determined that convergence has occurred (step 2401).

[0202] Step 2406: The learning subsystem 110 increases the current value of the counter C1 by "1".

[0203] If the learning termination condition is met in step 2401, the learning subsystem 110 determines "YES" in step 2401 and proceeds to step 2407, where it saves the signal extraction model and the text-embedded extraction model (each parameter of the neural network (NN)) in each database. Thereafter, the learning subsystem 110 proceeds to step 2495, where it temporarily ends this processing flow.

[0204] <Sound extraction subsystem> (Sound extraction subsystem function) The configuration of the sound extraction subsystem 150 will be explained below, mainly for each function. Fig. 25 is a block diagram for explaining an example of the configuration of the sound extraction subsystem 150 for each function.

[0205] As shown in FIG. 25, the sound extraction subsystem 150 includes a text-embedded extraction unit 157 and a signal extraction unit 1800.

[0206] (Overview of operation) 25, variable-length explanatory text D1300 is input to the text embedding extraction unit 157, which converts it into an embedding vector D1301 (text embedding vector D1301) of the explanatory text D1300. The time waveform D50 of the mixed signal and the text embedding vector D1301 are input to the signal extraction unit 1800, which generates a time waveform D72 of the extracted signal.

[0207] (Specific operation) Fig. 26 shows an example of a processing flow of the sound extraction subsystem 150. The sound extraction subsystem 150 executes the processing flow of Fig. 26. When the time waveform D50 of the mixed signal and the variable-length explanatory text D1300 are input, the sound extraction subsystem 150 starts processing from step 2600 in Fig. 26 and executes the processing of steps 2601 and 2602 described below in order, and then proceeds to step 2695, where the processing flow is temporarily ended.

[0208] Step 2601: The sound extraction subsystem 150 uses the latest text embedding extraction model to calculate (extract) the text embedding vector D1301 from the variable-length explanatory text D1300 by the text embedding extraction unit 157.

[0209] Step 2602: The sound extraction subsystem 150 uses the latest signal extraction model to generate, by the signal extraction unit 1800, the time waveform D72 of the extracted signal from the time waveform D50 of the mixed signal and the text-embedding vector D1301.

[0210] <Effects> As described above, the sound extraction system 2200 according to the eighth embodiment of the present invention can accurately extract (extract or emphasize) a signal corresponding to a sound that the user wants to extract from a mixed signal. With this basic configuration, the sound extraction system 2200 according to the eighth embodiment can extract sound even when the range of sounds that the user wants to extract as a certain type of event cannot be defined in advance. The explanatory text is relatively versatile and can be used across a wide range of application sites. Furthermore, the sound extraction system 2200 according to the eighth embodiment can avoid the reduction in extraction accuracy that accompanies the use of time-frequency representation by directly inputting the time waveform of the mixed signal into a neural network without using a time-frequency representation. Furthermore, the sound extraction system 2200 according to the eighth embodiment generates the time waveform of the extracted signal without undergoing phase restoration processing, which has the advantage of eliminating distortion that accompanies phase restoration processing.

[0211] <<Modifications>> The present invention is not limited to the above-described embodiments, and various modifications can be adopted within the scope of the present invention. Furthermore, the above-described embodiments can be combined with each other without departing from the scope of the present invention. Furthermore, within the scope of the present invention, part of the configuration of one embodiment can be replaced with the configuration of another embodiment. Furthermore, within the scope of the present invention, the configuration of one embodiment can be added to the configuration of another embodiment. Furthermore, within the scope of the present invention, part of the configuration of each embodiment can be added to, deleted from, or replaced with another configuration.

[0212] Furthermore, in each of the above embodiments, the text input to the sound extraction subsystem 150 may be input by operating an operating device such as a keyboard. Furthermore, in each of the above embodiments, the text input to the sound extraction subsystem 150 may be input by converting human speech into text using speech recognition technology. Furthermore, in each of the above embodiments, the mixed signal input to the sound extraction subsystem 150 may be input from an acoustic device such as a microphone. [Explanation of symbols]

[0213] 100...sound extraction system, 110...learning subsystem, 120...feature extraction model database, 130...text embedding extraction model database, 140...time-frequency mask generation model database, 150...sound extraction subsystem, 160...training dataset database

Claims

1. A sound extraction system including a sound extraction device that extracts a signal corresponding to a sound to be extracted from a mixed signal including the signal corresponding to the sound to be extracted, The sound extraction device includes: generating a time-frequency mask for extracting a signal corresponding to the sound to be extracted based on the mixed signal and a text that is a description specifying an onomatopoeia or an onomatopoeia following a specific sound, which corresponds to the sound to be extracted; and applying the time-frequency mask to the mixed signal to extract the signal corresponding to the sound to be extracted from the mixed signal. It was configured as follows: Sound extraction system.

2. The sound extraction system according to claim 1, Further comprising a storage device storing a trained model used to generate the time-frequency mask; The sound extraction device includes: generating the time-frequency mask based on the text and the mixed signal corresponding to the target sound using the trained model; It was configured as follows: Sound extraction system.

3. The sound extraction system according to claim 2, the storage device stores, as the trained models, a text embedding extraction model that outputs an embedding vector of the text from data obtained by preprocessing the text corresponding to the sound to be extracted, and a time-frequency mask generation model that generates the time-frequency mask from the embedding vector of the text and a sound feature vector of the mixed signal; The sound extraction device includes: calculating a feature vector of the sound of the mixed signal from the mixed signal; Using the text embedding extraction model, an embedding vector of the text corresponding to the sound to be extracted is calculated from the text; generating the time-frequency mask from the calculated embedding vector of the text and the sound feature vector of the mixed signal using the time-frequency mask generation model; It was configured as follows: Sound extraction system.

4. The sound extraction system according to claim 2, the storage device stores, as the trained models, a feature extraction model that outputs a sound feature vector of the mixed signal from the mixed signal, a text embedding extraction model that outputs an embedding vector of the text from data obtained by preprocessing the text corresponding to the sound to be extracted, and a time-frequency mask generation model that generates the time-frequency mask from the embedding vector of the text and the sound feature vector of the mixed signal; The sound extraction device includes: Using the feature extraction model, a feature vector of the sound of the mixed signal is calculated from the mixed signal; Using the text embedding extraction model, an embedding vector of the text corresponding to the sound to be extracted is calculated from the text; generating the time-frequency mask from the calculated embedding vector of the text and the sound feature vector of the mixed signal using the time-frequency mask generation model; It was configured as follows: Sound extraction system.

5. The sound extraction system according to claim 1, The present invention further includes a learning device that generates a trained model used to generate the time-frequency mask by performing machine learning using a training dataset including a target signal corresponding to the sound to be extracted, a training mixed signal obtained by mixing the target signal with a signal corresponding to noise other than the sound to be extracted, and training text corresponding to the sound to be extracted, The sound extraction device includes: generating the time-frequency mask based on the text and the mixed signal corresponding to the sound to be extracted, using the trained model generated by the training device; It was configured as follows: Sound extraction system.

6. The sound extraction system according to claim 5, The learning device includes, as the trained model: a text embedding extraction model that outputs an embedding vector of the text from data obtained by preprocessing the text corresponding to the sound to be extracted; a time-frequency mask generation model that generates the time-frequency mask from the feature of the mixed signal and the embedding vector of the text; configured to generate The sound extraction device includes: Calculating a feature vector of the sound of the mixed signal from the mixed signal; calculating an embedding vector of the text from the text corresponding to the sound to be extracted using the text embedding extraction model generated by the learning device; generating the time-frequency mask from the calculated embedding vector of the text and the sound feature vector of the mixed signal using the time-frequency mask generation model generated by the learning device; It was configured as follows: Sound extraction system.

7. The sound extraction system according to claim 5, The learning device includes, as the trained model: a feature extraction model that outputs a feature vector of the sound of the mixed signal from the mixed signal; a text embedding extraction model that outputs an embedding vector of the text from data obtained by preprocessing the text corresponding to the sound to be extracted; a time-frequency mask generation model that generates the time-frequency mask from the sound feature vector of the mixed signal and the text embedding vector; configured to generate The sound extraction device includes: calculating a feature vector of the sound of the mixed signal from the mixed signal using the feature extraction model generated by the learning device; calculating an embedding vector of the text from the text corresponding to the sound to be extracted using the text embedding extraction model generated by the learning device; generating the time-frequency mask from the calculated embedding vector of the text and the sound feature vector of the mixed signal using the time-frequency mask generation model generated by the learning device; It was configured as follows: Sound extraction system.

8. The sound extraction system according to claim 1, a learning device that acquires an initial trained model from outside, and updates the initial trained model by performing machine learning using a training dataset including a target signal corresponding to the sound to be extracted, a training mixed signal obtained by mixing the target signal with a signal corresponding to noise other than the sound to be extracted, and a training text corresponding to the sound to be extracted, thereby generating a trained model used to generate the time-frequency mask; The sound extraction device includes: generating the time-frequency mask based on the text and the mixed signal corresponding to the sound to be extracted, using the trained model generated by the training device; It was configured as follows: Sound extraction system.

9. The sound extraction system according to claim 1, The method further includes a storage device storing a trained model used to extract a signal corresponding to the target sound from the mixed signal based on the text corresponding to the target sound and the mixed signal, The sound extraction device includes: extracting a signal corresponding to the sound to be extracted from the mixed signal based on the text corresponding to the sound to be extracted and the mixed signal using the trained model; It was configured as follows: Sound extraction system.

10. The sound extraction system according to claim 9, the storage device stores, as the trained models, a text embedding extraction model that outputs an embedding vector of the text from data obtained by preprocessing the text corresponding to the sound to be extracted, and a signal extraction model that generates a time waveform of a signal corresponding to the sound to be extracted from the embedding vector of the text and a time waveform of the mixed signal; The sound extraction device includes: Using the text embedding extraction model, an embedding vector of the text corresponding to the sound to be extracted is calculated from the text; extracting a signal corresponding to the sound to be extracted from the mixed signal by generating a time waveform of a signal corresponding to the sound to be extracted from the calculated embedding vector of the text and the time waveform of the mixed signal using the signal extraction model; It was configured as follows: Sound extraction system.

11. The sound extraction system according to claim 9, The present invention further includes a learning device that generates the trained model used to extract a signal corresponding to the sound to be extracted from the mixed signal, based on the text corresponding to the sound to be extracted and the mixed signal, by performing machine learning using a training dataset including a target signal corresponding to the sound to be extracted, a training mixed signal obtained by mixing the target signal with a signal corresponding to a noise other than the sound to be extracted, and training text corresponding to the sound to be extracted, extracting a signal corresponding to the sound to be extracted from the mixed signal based on the text corresponding to the sound to be extracted and the mixed signal using the trained model generated by the training device; It was configured as follows: Sound extraction system.

12. The sound extraction system according to claim 9, the learning device further includes a learning device that acquires an initial trained model from outside, and updates the initial trained model by performing machine learning using a training dataset including a target signal corresponding to the sound to be extracted, a training mixed signal obtained by mixing the target signal with a signal corresponding to a noise other than the sound to be extracted, and a training text corresponding to the sound to be extracted, thereby generating the trained model used to extract the signal corresponding to the sound to be extracted from the mixed signal based on the text corresponding to the sound to be extracted and the mixed signal; The sound extraction device includes: extracting a signal corresponding to the sound to be extracted from the mixed signal based on the text corresponding to the sound to be extracted and the mixed signal using the trained model generated by the training device; It was configured as follows: Sound extraction system.

13. In the sound extraction system according to any one of claims 5, 8, 11 and 12, Further comprising a storage device in which the target signal corresponding to the sound to be extracted and a signal corresponding to a noise other than the sound to be extracted are stored, The learning device reading out from the storage device the target signal corresponding to the sound to be extracted and a signal corresponding to a noise other than the sound to be extracted, and mixing the target signal corresponding to the sound to be extracted and the signal corresponding to a noise other than the sound to be extracted to generate the training mixed signal; It was configured as follows: Sound extraction system.

14. A sound extraction method using a sound extraction device that extracts a signal corresponding to a sound to be extracted from a mixed signal including the signal corresponding to the sound to be extracted, The sound extraction device generating a time-frequency mask for extracting a signal corresponding to the sound to be extracted based on the mixed signal and a text that is a description specifying an onomatopoeia or an onomatopoeia following a specific sound, which corresponds to the sound to be extracted; and applying the time-frequency mask to the mixed signal to extract the signal corresponding to the sound to be extracted from the mixed signal. Sound extraction method.

Citation Information

Patent Citations

  • Device and method of adding sound effect

    JP2000081892A

  • Environmental sound retrieval device and environmental sound retrieval method

    JP2014178886A

  • Estimation apparatus, estimation method, and estimation program

    JP2017228164A

  • Determination program, determination device, and determination method

    JP2018156627A

  • Signal processing device, signal processing method and signal processing program

    JP2020134567A