Information processing apparatus, information processing method, and information processing program

The information processing device uses VAE, Tokenizer, and Transformer models to separate multiple speakers from mixed audio data by extracting and reconstructing speaker features, addressing the limitations of existing LSTM-based methods and enabling speaker separation without order information.

JP2026003368APending Publication Date: 2026-01-13DENSO CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024101283
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing speaker separation technologies using LSTM are limited in their ability to separate voices of multiple speakers without requiring information on the order of speech and can only handle a single speaker effectively, while other methods require speaker order information.

Method used

An information processing device employing a combination of VAE, Tokenizer, and Transformer models to extract, generate, and subtract speaker features from mixed audio data, allowing separation of multiple speakers without order information, using a method that includes an encoder for feature extraction, a decoder for reconstruction, and a second model for generating speaker features.

Benefits of technology

Enables the separation of individual speakers from mixed audio data of an unspecified number of speakers in a simple manner, without requiring information on the order of speech, by performing extraction, generation, and subtraction processes on spectrograms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026003368000001_ABST
    Figure 2026003368000001_ABST
Patent Text Reader

Abstract

To separate each speaker from mixed voices of many and unspecified speakers by a simple method without requiring information such as a speaking order.SOLUTION: An information processor 100 includes a first model 122a learned to extract a feature quantity when a spectrogram is input and to reconfigure input data when the feature quantity is input, a second model 122c learned to generate a feature quantity of a certain speaker when a feature quantity extracted from a spectrogram obtained by performing short-time Fourier transformation on voice data of the speaker is input, and a control unit 11 for converting mixed voice data of a plurality of speakers into a mixed spectrogram, extracting a mixed feature quantity from the mixed spectrogram, inputting the mixed feature quantity into the first model to reconfigure a first speaker feature quantity corresponding to the first speaker, and subtracting the first speaker feature quantity from the first spectrogram.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] Conventionally, speaker separation techniques using LSTM (Long Short-Term Memory) are known (for example, Patent Document 1 and Patent Document 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-042152 [Patent Document 2] Japanese Patent Publication No. 2020-013034 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the speaker separation technology using LSTM described in Patent Document 1 can only separate the voice of a single speaker (main voice) from mixed voice data in which the voices of multiple speakers are superimposed, and cannot separate the voices of other speakers from the mixed voice data.Furthermore, the speaker separation technology using LSTM described in Patent Document 2 can separate multiple speakers from mixed voice data in which the voices of multiple speakers are superimposed, but requires information on the order in which the multiple speakers speak in order to separate the multiple speakers.

[0005] In view of the above problems, an object of the present invention is to separate individual speakers from mixed speech of an unspecified number of speakers by a simple method without requiring information such as the order of speech. [Means for solving the problem]

[0006] The present invention employs the following technical solutions to solve the above problems. The reference symbols in parentheses in the claims and this section are merely examples showing the correspondence with the specific solutions described in the embodiments below as one aspect, and do not limit the technical scope of the present invention.

[0007] An information processing device (100, 200) according to one aspect of the present invention is an information processing device comprising at least one memory (12, 22) in which a computer program is recorded, and at least one processor (11, 21) capable of executing the computer program stored in the at least one memory, wherein the at least one memory includes a first model (122a, 222a) trained to include an encoder that extracts features when a spectrogram is input, and a decoder that reconstructs input data when the features are input, and a decoder that is trained to generate features corresponding to a speaker of the at least one speaker when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data containing speech of the at least one speaker are input. and a second model (122c, 222c1-222c4), and the at least one processor receives mixed audio data including speech sounds of a plurality of speakers, performs a transformation process to convert the mixed audio data into a mixed spectrogram using a short-time Fourier transform, performs an extraction process to extract mixed features by inputting the mixed spectrogram to the encoder of the first model, performs a generation process to generate first speaker features corresponding to a first speaker of the plurality of speakers by inputting the mixed features to the second model, performs a reconstruction process to reconstruct a first speaker spectrogram by inputting the first speaker features to the decoder of the first model, and performs a subtraction process to subtract the first speaker spectrogram from the mixed spectrogram.

[0008] With the above configuration, the information processing device can obtain a spectrogram in which the first speaker is excluded from the unspecified number of speakers by subtracting the first speaker spectrogram from the mixed spectrogram corresponding to mixed audio data containing speech sounds from an unspecified number of speakers. This makes it possible to separate speech data from the first speaker from the mixed audio data, even if the mixed audio data is from an unspecified number of speakers. By performing the above extraction process, generation process, reconstruction process, and subtraction process in this order on the spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram, other speakers can be separated. By repeating these processes, speakers can be separated one by one from mixed audio data containing speech sounds from an unspecified number of speakers.

[0009] An information processing method according to one aspect of the present invention uses a first model trained to include an encoder that extracts features when a spectrogram is input, and a decoder that reconstructs input data when the features are input, and a second model trained to generate features corresponding to a speaker among the at least one speaker when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech from at least one speaker are input, the information processing method comprising the steps of: receiving mixed speech data including speech from a plurality of speakers; and The method includes the steps of: performing a conversion process to convert speech data into a mixture spectrogram; performing an extraction process to extract mixture features by inputting the mixture spectrogram to the encoder of the first model; performing a generation process to generate first speaker features corresponding to a first speaker of the plurality of speakers by inputting the mixture features to the second model; performing a reconstruction process to reconstruct a first speaker spectrogram by inputting the first speaker features to the decoder of the first model; and performing a subtraction process to subtract the first speaker spectrogram from the mixture spectrogram.

[0010] An information processing program according to one aspect of the present invention uses a first model trained to include an encoder that extracts features when a spectrogram is input, and a decoder that reconstructs input data when the features are input, and a second model trained to generate features corresponding to a certain speaker among the at least one speaker when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech from at least one speaker are input, the information processing program including the steps of: receiving mixed speech data including speech from a plurality of speakers; and performing a conversion process to convert the mixed speech data into a mixed spectrogram, an extraction process to extract mixed features by inputting the mixed spectrogram to the encoder of the first model, a generation process to generate first speaker features corresponding to a first speaker among the plurality of speakers by inputting the mixed features to the second model, a reconstruction process to reconstruct a first speaker spectrogram by inputting the first speaker features to the decoder of the first model, and a subtraction process to subtract the first speaker spectrogram from the mixed spectrogram. [Effects of the Invention]

[0011] According to the present invention, it is possible to provide an information processing device or the like that can separate each speaker from a mixed voice of an unspecified number of speakers in a simple manner without requiring information such as the order of speech. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram illustrating an example of a configuration of an information processing device according to the first embodiment. [Figure 2] FIG. 2 is a schematic diagram illustrating an example of learning of a VAE model. [Figure 3] FIG. 3 is a schematic diagram illustrating an example of learning of the Tokenizer model. [Figure 4]FIG. 4 is a diagram showing an example of logical functional blocks realized in the control unit of the information processing device to execute information processing according to the first embodiment. [Figure 5] FIG. 5 is a flowchart showing a flow of processing by the control unit according to the first embodiment. [Figure 6] FIG. 6 is a block diagram illustrating an example of a configuration of an information processing device according to the second embodiment. [Figure 7] FIG. 7 is a diagram showing an example of logical functional blocks realized in a control unit of an information processing device for executing information processing according to the second embodiment. [Figure 8] FIG. 8 is a flowchart showing a flow of processing by the control unit according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the embodiment described below shows an example of how the present invention can be implemented, and the present invention is not limited to the specific configuration described below. When implementing the present invention, a specific configuration corresponding to the embodiment may be appropriately adopted.

[0014] (First embodiment) FIG. 1 is a block diagram showing an example of the configuration of an information processing device 100 according to a first embodiment. The information processing device 100 is an information processing device capable of executing various types of information processing, such as a server computer or a personal computer. As will be described later, the information processing device 100 generates a Transformer model 122c by learning training data, and separates multiple speakers using the Transformer model 122c. The Transformer model 122c is a deep learning model that is mainly used in natural language processing. Since the Transformer model 122c is a widely known deep learning model, a specific internal structure thereof will not be described.

[0015] The information processing device 100 includes a control unit 11, a storage unit 12, an audio input unit 13, and an audio output unit 14. These are communicably connected to each other via a bus, for example. However, the information processing device 100 does not necessarily have to include at least one of the audio input unit 13 and the audio output unit 14. The information processing device 100 of this embodiment can operate in a standalone state.

[0016] The control unit 11 has one or more processors as hardware. The control unit 11 according to this embodiment has a CPU (Central Processing Unit) 11a and a GPU (Graphics Processing Unit) 11b. The control unit 11 reads and executes a computer program (not shown) that includes at least one of computer program code and computer program instructions stored in the storage unit 12 (specifically, the auxiliary storage unit 122), thereby performing various information processing, control processing, and the like.

[0017] The storage unit 12 includes a main storage unit 121 and an auxiliary storage unit 122, and includes at least one memory capable of storing desired data. The main storage unit 121 is a temporary storage area, such as an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory), and temporarily stores data required for the control unit 11 to execute arithmetic processing. As will be described later, the main storage unit 121 includes a first memory M01 that stores input spectrograms and a second memory M02 that stores output spectrograms.

[0018] The auxiliary storage unit 122 is a non-volatile storage area such as a large-capacity memory or a hard disk, and stores computer programs and other data necessary for the control unit 11 to execute processing. The auxiliary storage unit 122 also stores a VAE model 122a, a Tokenizer model 122b, and a Transformer model 122c necessary for executing the computer programs. The auxiliary storage unit 122 may also store training data for training the VAE model 122a, training data for training the Tokenizer model 122b, and training data for training the Transformer model 122c. The VAE model 122a, the Tokenizer model 122b, and the Transformer model 122c are machine learning models that have trained predetermined training data. The VAE model 122a, the Tokenizer model 122b, and the Transformer model 122c are expected to be used as program modules that constitute part of artificial intelligence software. The auxiliary storage unit 122 also stores an input audio file 122d of audio data containing the speech of multiple speakers (i.e., audio data in which the speech of multiple speakers is superimposed), and an output audio file 122e of audio data containing the speech of one speaker separated from the multiple speakers.

[0019] The auxiliary storage unit 122 may be an external storage device connected to the information processing device 100. The information processing device 100 may be a multi-computer consisting of multiple computers, or may be a virtual machine virtually constructed by software.

[0020] The voice input unit 13 includes an input interface for connecting to an external device, and receives input of voice data containing speech from multiple speakers. The voice input unit 13 transmits the received input content to the control unit 11. A microphone 15 may be connected to the voice input unit 13 via a wired or wireless connection.

[0021] The audio output unit 14 includes an output interface for connecting to an external device, and outputs audio data generated by the control unit 11. A speaker 16 may be connected to the audio output unit 14 by wire or wirelessly.

[0022] In addition, in this embodiment, the information processing device 100 may be provided with a reading unit that reads portable storage media such as CD (Compact Disk)-ROMs and DVD (Digital Versatile Disc)-ROMs, and may read and execute computer programs from the portable storage media.

[0023] FIG. 2 is a schematic diagram illustrating an example of training of the VAE model 122a. The VAE model 122a is a neural network that acquires features from input data and outputs restored data of the input data from the features. The VAE model 122a is, for example, a variational autoencoder. The VAE model 122a is a machine learning model that has been trained to include an encoder 122a1 that obtains features z (latent variables) from input data and a decoder 122a2 that generates restored data from the features z. Note that during training, the entire VAE model 122a including the encoder 122a1 and the decoder 122a2 is used.

[0024] To train the VAE model 122a, a spectrogram is input to the encoder 122a1, and machine learning of the VAE model 122a is performed to minimize the error between the input data and the restored data. Here, the spectrogram as input data to the VAE model 122a corresponds to a so-called voiceprint, in which the spectra of each audio data segment (i.e., frame) obtained by extracting the frequency components and amplitude components of the audio signal from an audio waveform through a Fourier transform are arranged along the time axis. In this embodiment, the input data to the VAE model 122a is audio data in which the voices of one or more speakers are superimposed. This audio data can be obtained, for example, by synthesizing appropriate short sentences (approximately 20 characters or less) collected from the web. For audio data from one speaker, an appropriate short sentence can be synthesized using a certain voice A. For audio data from two speakers, a mixture of the voice data of voice A and the voice data of voice B can be used. For audio data from three speakers, a mixture of the voice data of voice A, the voice data of voice B, and the voice data of voice C can be used.

[0025] The latent space shown in Fig. 2 indicates the feature z (latent variable) output when each training data (spectrogram) is input to the encoder 122a1 of the trained VAE model 122a. Note that although the latent space shown in Fig. 2 is expressed as a two-dimensional space, it is actually a multidimensional space of three or more dimensions (1024 real dimensions in this embodiment), and the feature of the multidimensional space is projected onto the two-dimensional space.

[0026] FIG. 3 is a schematic diagram illustrating an example of training of the Tokenizer model 122b. The Tokenizer model 122b is a machine learning model that functions as an encoder 122b1 that tokenizes input data into one-dimensional integers so that the Transformer model 122c can process it, and as a decoder 122b2 that restores the tokenized one-dimensional integer data to the original data. The Tokenizer model 122b according to this embodiment is a machine learning model that is trained to tokenize features extracted from a spectrogram and restore the original features from the tokenized features. Each token is assigned an ID, and this ID is used when inputting the tokens to the Transformer model 122c. Because the features are 1024-dimensional vectors of real numbers, the Tokenizer model 122b according to this embodiment converts the vectors into one-dimensional integer values. Because real numbers are uncountably infinite and integers are countably infinite, there is no one-to-one correspondence between the features of a 1024-dimensional vector of real numbers and one-dimensional integer tokens. Therefore, the Tokenizer model 122b according to this embodiment has a many-to-one correspondence between features of vectors with 1024 real dimensions and tokens with one integer dimension. Specifically, features that are close to each other (or adjacent to each other) are considered to be the same, and the same ID is assigned to these features. Note that the Euclidean distance between features may be used to determine whether features are close to each other.

[0027] For training of the Tokenizer model 122b, the feature amount is input to the encoder 122b1, and machine learning of the Tokenizer model 122b is performed so as to reduce the error between the input data and the restored data. The Tokenizer model 122b can also embed an instruction token into the token string of the input data.

[0028] The Transformer model 122c is a machine learning model that uses features (latent variables) as input and identifies and separates features of different speakers using an attention mechanism. In this embodiment, the Transformer model 122c is a model constructed using a neural network, and is, for example, a large-scale language model (LLaMA: Large Language Model) such as Transformer-based LLaMA (Large Language Model Meta AI). The large-scale language model is not limited to LLaMA, and may be GPT (Generative Pretrained Transformer, including GPT-1, GPT-2, GPT-3, GPT-4, and GPT-4o), BERT (Bidirectional Encoder Representations from Transformers), BART (Bidirectional and Auto-regressive Transformer), etc.

[0029] The Transformer model 122c includes an input layer that accepts input of latent variables, an intermediate layer (hidden layer) that extracts features of the latent variables, and an output layer that outputs features of different speakers. For example, the latent variables may be generated by converting spectrograms generated by a short-time Fourier transform (STFT) of speech data from multiple speakers into latent variables using the encoder 122a1 of the VAE model 122a. The latent variables may then be input to the input layer as tokens generated by the Tokenizer model 122b. The intermediate layer includes multiple nodes that extract features of each input value and transfers the extracted features using various parameters to the output layer. The output layer is configured to output features of one of the multiple speakers based on the features output from the intermediate layer. The Transformer model 122c includes an Attenuation layer in the intermediate layer, which can identify and separate the features of one of the multiple speakers.

[0030] The Transformer model 122c can be constructed by a supervised learning method. In supervised learning, machine learning is performed using training data (training data). The training data is composed of a pair of input data for training and output data (correct answer data). In this embodiment, training of the Transformer model 122c is performed using a trained VAE model 122a and a trained Tokenizer model 122b. As shown in FIG. 4, which will be described later, the trained VAE model 122a and the trained Tokenizer model 122b are connected to the Transformer model 122c. In this state, training data is provided, the output result is compared with the correct answer data, and parameters are optimized using, for example, backpropagation so that the output result approaches the correct answer data. The parameters are, for example, weights (coupling coefficients) between neurons.

[0031] Transformer model 122c is a machine learning model that is trained to generate tokenized features corresponding to one of the at least one speaker when tokenized features extracted from a spectrogram obtained by performing a short-time Fourier transform on audio data containing the speech of at least one speaker are input.

[0032] 4 is a diagram showing an example of logical functional blocks realized in the control unit 11 of the information processing device 100 in order to execute information processing according to this embodiment. As shown in Fig. 4, the control unit 11 of the information processing device 100 executes a computer program to realize the functions of an STFT unit 111, a subtraction unit 112, a VAE encoder unit 113, a Tokenizer unit 114, a Transformer unit 115, a De-Tokenizer unit 116, a VAE decoder unit 117, and an ISTFT unit 118.

[0033] The control unit 11 receives mixed audio data consisting of audio waveforms containing speech sounds from multiple speakers via the audio input unit 13. The control unit 11 may also receive mixed audio data by reading it from an input audio file 122d that stores mixed audio data containing speech sounds from multiple speakers. The control unit 11 may also receive audio data containing speech sounds from a single speaker. The input audio data is, for example, data in WAV format, and is 256 x 256 dimensional data.

[0034] The STFT unit 111 performs a conversion process to convert the mixed audio data into a mixed spectrogram using a short-time Fourier transform. Here, the mixed audio data includes audio data of a single speaker's speech in addition to speech from multiple speakers. Here, a spectrogram corresponds to a so-called voiceprint, in which the spectra of each audio data segment (i.e., frame) obtained by extracting the frequency components and amplitude components of an audio signal from an audio waveform using a Fourier transform are arranged along the time axis.

[0035] The VAE encoder unit 113 performs extraction processing to extract mixed features by inputting a mixed spectrogram to the encoder 122a1 of the VAE model 122a. The mixed features are, specifically, real-valued 1024-dimensional feature vectors. The VAE encoder unit 113 can reduce the amount of information by converting 256×256-dimensional audio data into 1024-dimensional features.

[0036] The Tokenizer unit 114 inputs the mixed feature extracted by the VAE encoder unit 113 to the Tokenizer model 122b, thereby tokenizing the mixed feature into integer one-dimensional vector values.

[0037] The Transformer unit 115 performs a generation process of generating tokenized first speaker features corresponding to a first speaker, who is one of multiple speakers, by inputting the tokenized mixed features to the Transformer model 122c. The data output from the Transformer unit 115 is an integer one-dimensional vector value. In this embodiment, the mixed features are tokenized by the Tokenizer unit 114 before being input to the Transformer unit 115. However, the mixed features may be input to the Transformer unit 115 without being tokenized by the Tokenizer unit 114. Tokenizing the mixed features by the Tokenizer unit 114 is preferable for speeding up processing because it reduces the amount of information processed by the Transformer unit 115. Furthermore, tokenization can capture global features of the data, thereby improving data accuracy. It is considered that the Transformer unit 115 separates the first speaker from multiple speakers based on voice frequency, voice speed, voice intonation, etc.

[0038] The De-Tokenizer unit 116 inputs the tokenized first speaker features generated by the Transformer unit 115 to the Tokenizer model 122b, thereby restoring the original first speaker features.

[0039] The VAE decoder unit 117 performs a reconstruction process to reconstruct (or restore) a 256 x 256-dimensional first speaker spectrogram (i.e., the spectrogram of the first speaker) by inputting the real-valued 1024-dimensional first speaker features restored by the De-Tokenizer unit 116 into the decoder 122a2 of the VAE model 122a.

[0040] The ISTFT transform unit 118 performs a transform process using an inverse short-time Fourier transform to transform the spectrogram of the first speaker into voice data of the first speaker. The voice data output by the ISTFT transform unit 118 is written to an output voice file 122e.

[0041] Subtraction unit 112 performs a subtraction process to subtract the first speaker spectrogram restored by VAE decoder unit 117 from the mixed spectrogram. By performing the above extraction process, generation process, reconstruction process, and subtraction process in this order on the spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram, it is possible to separate different speakers. By repeating these processes, it is possible to separate each speaker one by one from mixed audio data containing speech sounds from an unspecified number of multiple speakers.

[0042] Next, a specific operation of the control unit 11 according to this embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the flow of processing by the control unit 11.

[0043] In step S501, the control unit 11 reads out mixed audio data including speech sounds from a plurality of speakers from the input audio file 122d. The control unit 11 may also receive mixed audio data consisting of audio waveforms including speech sounds from a plurality of speakers via the audio input unit 13.

[0044] In step S502, the control unit 11 (STFT transformation unit 111) executes a transformation process for transforming the mixed audio data into a mixed spectrogram using a short-time Fourier transform.

[0045] In step S503, the control unit 11 stores the mixed spectrogram, which is the result of the short-time Fourier transform, in the first memory M01.

[0046] In step S504, the control unit 11 (subtraction unit 112) executes subtraction processing to subtract the output spectrogram stored in the second memory M02 in step S510 (described later) from the mixed spectrogram, which is the input spectrogram stored in the first memory M01 in step S503, and stores (i.e., overwrites) the new mixed spectrogram, which is the result of the subtraction processing, in the first memory M01. Because the initial value of the second memory M02 is empty, if no data is stored in the second memory M02, nothing is subtracted from the mixed spectrogram, which is the input spectrogram.

[0047] In step S505, the control unit 11 (VAE encoder unit 113) inputs the mixed spectrogram, which is the subtraction result in step S504, to the encoder 122a1 of the VAE model 122a, thereby executing extraction processing to extract mixed features.

[0048] In step S506, the control unit 11 (Tokenizer unit 114) inputs the mixed feature extracted in step S505 to the Tokenizer model 122b, thereby tokenizing the mixed feature into an integer one-dimensional vector value.

[0049] In step S507, the control unit 11 (Transformer unit 115) executes a generation process to generate tokenized first speaker features corresponding to a first speaker who is one of the multiple speakers, by inputting the tokenized mixed features to the Transformer model 122c. Note that in this embodiment, the mixed features are tokenized in step S506 and then the generation process is performed in step S507, but the generation process may be performed in step S507 without tokenizing the mixed features.

[0050] In step S508, the control unit 11 (De-tokenizer unit 116) inputs the tokenized first speaker features generated in step S507 into the Tokenizer model 122b, thereby restoring the original first speaker features.

[0051] In step S509, the control unit 11 (VAE decoder unit 117) performs a reconstruction process to reconstruct (or restore) the first speaker spectrogram (i.e., the spectrogram of the first speaker) by inputting the first speaker features restored in step S508 to the decoder 122a2 of the VAE model 122a.

[0052] In step S510, the control unit 11 stores the first speaker spectrogram restored in step S509 as an output spectrogram in the second memory M02.

[0053] In step S511, the control unit 11 (ISTFT transformation unit 118) starts another thread and executes a transformation process for transforming the first speaker spectrogram into the first speaker's voice data using an inverse short-time Fourier transform.

[0054] In step S512, the control unit 11 determines whether the output speech of the speech data generated in step S511 is silent. If the output speech is silent, it is determined that there is no speech to separate, and the control unit 11 ends the process. If the output speech is not silent, the flow proceeds to step S513. Here, to determine whether the output speech is silent, any algorithm or technique for performing speech-to-text transcription (SST) known in the speech recognition field may be used. If text representing the spoken content is not acquired from the output speech using SST, the output speech can be determined to be silent. For example, if there is noise such as environmental sound, incomprehensible text or nothing will be output. Alternatively, whether the output speech is silent may be determined by performing voice detection using a library called voice-activity-detection.

[0055] In step S513, the control unit 11 writes the audio data output in step S511 into the output audio file 122e, and ends the separate thread started in step S511.

[0056] In this embodiment, the processes from step S511 to step S513 are performed in a separate thread. While the inverse short-time Fourier transform is being performed in the separate thread, the subtraction process in step S504 is performed in this thread. However, if the output audio is not silent, the process of separating speakers can be performed in parallel in this thread, thereby realizing a speedup of the entire process.

[0057] As described above, the information processing device 100 according to the present embodiment can obtain a spectrogram in which the first speaker is excluded from the multiple speakers by subtracting the first speaker spectrogram from the mixed spectrogram corresponding to mixed audio data containing speech sounds from multiple speakers. This allows the speech data of the first speaker to be separated from the mixed audio data, even if the mixed audio data is composed of an unspecified number of multiple speakers. Furthermore, by performing the above-described extraction process, generation process, reconstruction process, and subtraction process in this order on the spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram, other speakers can be separated. By repeating these processes, speakers can be separated one by one from mixed audio data containing speech sounds from an unspecified number of multiple speakers.

[0058] (Second embodiment) Next, an information processing device 200 according to a second embodiment will be described. Fig. 6 is a block diagram showing an example of the configuration of the information processing device 200 according to the second embodiment. The same reference numerals are used to designate the same components as those in the information processing device 100 according to the first embodiment, and the description thereof may be omitted.

[0059] In the information processing device 100 according to the first embodiment, the speaker separation process is performed using one Transformer model 122c.

[0060] On the other hand, the information processing device 200 according to the second embodiment uses a plurality of Transformer models 222c1, 222c2, 222c3, and 222c4 to separate a maximum of the number of Transformer models 222c1, 222c2, 222c3, and 222c4 from a plurality of speakers at one time. This enables the processing of separating speakers to be performed faster than the information processing device 100 according to the first embodiment. In this embodiment, the number of Transformer models 222c1, 222c2, 222c3, and 222c4 will be described as four, but the number of Transformer models 222c1, 222c2, 222c3, and 222c4 is not limited to four, and may be three, five, six, or the like, depending on the number of attributes, which will be described later.

[0061] The information processing device 200 according to this embodiment is an information processing device capable of executing various types of information processing, such as a server computer or a personal computer. As will be described later, the information processing device 200 generates Transformer models 222c1, 222c2, 222c3, and 222c4 by learning training data, and separates multiple speakers using the Transformer models 222c1, 222c2, 222c3, and 222c4. The Transformer models 222c1, 222c2, 222c3, and 222c4 are deep learning models that are mainly used in natural language processing. Since the Transformer models 222c1, 222c2, 222c3, and 222c4 are widely known deep learning models, a detailed description of their internal structures will be omitted.

[0062] The information processing device 200 includes a control unit 21, a storage unit 22, an audio input unit 13, and an audio output unit 14. These are connected to each other so that they can communicate with each other, for example, via a bus. However, the information processing device 200 does not necessarily have to include at least one of the audio input unit 13 and the audio output unit 14. The information processing device 200 of this embodiment can operate in a standalone state.

[0063] The control unit 21 has one or more processors as hardware. The control unit 21 according to the present embodiment has a CPU (Central Processing Unit) 21a and a GPU (Graphics Processing Unit) 21b. The control unit 21 reads and executes a computer program (not shown) stored in the auxiliary storage unit 222, the computer program including at least one of computer program code and computer program instructions, to perform various information processing, control processing, and the like.

[0064] The storage unit 22 includes a main storage unit 221 and an auxiliary storage unit 222, and includes at least one memory capable of storing desired data. The main storage unit 221 is a temporary storage area, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), and temporarily stores data necessary for the control unit 11 to execute arithmetic processing. As will be described later, the main storage unit 121 includes a first memory M01 that stores input spectrograms and multiple second memories M02a, M02b, M02c, and M02d that store output spectrograms. Since a maximum of the number of output spectrograms generated is equal to the number of Transformer models 222c1, 222c2, 222c3, and 222c4, the number of second memories M02a, M02b, M02c, and M02d is four in this embodiment.

[0065] The auxiliary storage unit 222 is a nonvolatile storage area such as a large-capacity memory or a hard disk, and stores computer programs and other data necessary for the control unit 11 to execute processing. The auxiliary storage unit 222 also stores a VAE model 222a, Tokenizer models 222b1, 222b2, 222b3, and 222b4, and Transformer models 222c1, 222c2, 222c3, and 222c4 necessary for executing the computer programs. The auxiliary storage unit 222 may also store training data for training the VAE model 222a, training data for training the Tokenizer models 222b1, 222b2, 222b3, and 222b4, and training data for training the Transformer models 222c1, 222c2, 222c3, and 222c4. The VAE model 222a, the Tokenizer models 222b1, 222b2, 222b3, and 222b4, and the Transformer models 222c1, 222c2, 222c3, and 222c4 are machine learning models that have been trained using predetermined training data. The VAE model 222a, the Tokenizer models 222b1, 222b2, 222b3, and 222b4, and the Transformer models 222c1, 222c2, 222c3, and 222c4 are intended for use as program modules that constitute part of artificial intelligence software. The auxiliary storage unit 222 also stores an input audio file 222d of audio data containing speech from multiple speakers and an output audio file 222e of audio data containing speech from one speaker separated from the multiple speakers.

[0066] The auxiliary storage unit 222 may be an external storage device connected to the information processing device 200. The information processing device 200 may be a multi-computer consisting of multiple computers, or may be a virtual machine virtually constructed by software.

[0067] The voice input unit 13 includes an input interface for connecting to an external device, and receives input of voice data including speech sounds from multiple speakers. The voice input unit 13 transmits the received input content to the control unit 21. A microphone 15 may be connected to the voice input unit 13 via a wired or wireless connection.

[0068] The audio output unit 14 includes an output interface for connecting to an external device, and outputs audio data generated by the control unit 21. A speaker 16 may be connected to the audio output unit 14 by wire or wirelessly.

[0069] In addition, in this embodiment, the information processing device 200 may be provided with a reading unit that reads portable storage media such as CD (Compact Disk)-ROMs and DVD (Digital Versatile Disc)-ROMs, and may read and execute computer programs from the portable storage media.

[0070] The information processing device 200 according to this embodiment stores a plurality of Transformer models 222c1, 222c2, 222c3, and 222c4. Each of the plurality of Transformer models 222c1, 222c2, 222c3, and 222c4 is a machine learning model trained to generate tokenized features corresponding to speakers with different attributes when tokenized features extracted from spectrograms obtained by performing a short-time Fourier transform on speech data including speech from at least one speaker are input. Here, the different attributes include, for example, male, female, adult, and child. That is, the Transformer model 222c1 is a machine learning model trained to generate tokenized features corresponding to a single speaker whose attribute is male when tokenized features extracted from spectrograms obtained by performing a short-time Fourier transform on speech data including speech from multiple speakers including male, female, adult, and child are input. Transformer model 222c2 is a machine learning model trained to generate tokenized features corresponding to a single speaker whose attribute is female when, for example, tokenized features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data containing speech from multiple speakers, including males, females, adults, and children, are input. Transformer model 222c3 is a machine learning model trained to generate tokenized features corresponding to a single speaker whose attribute is adult when, for example, tokenized features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data containing speech from multiple speakers, including males, females, adults, and children, are input. Transformer model 222c4 is a machine learning model trained to generate tokenized features corresponding to a single speaker whose attribute is child when, for example, tokenized features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data containing speech from multiple speakers, including males, females, adults, and children, are input.

[0071] In this embodiment, the information processing device 200 inputs tokenized mixed features based on mixed speech data of multiple speakers to each of multiple Transformer models 222c1, 222c2, 222c3, and 222c4, thereby generating multiple tokenized features as tokenized first speaker features. That is, if the mixed speech data contains speech data with a male attribute, the Transformer model 222c1 can generate tokenized features with a male attribute. Also, if the mixed speech data contains speech data with a child attribute, the Transformer model 222c4 can generate tokenized features with a child attribute. On the other hand, if the mixed speech data does not contain speech data with a female attribute, the Transformer model 222c2 cannot generate tokenized features with a female attribute and outputs noise. In this way, if the Transformer model 222c2 outputs noise, the Transformer model 222c2 is considered to have not output anything, and no processing is performed on the output result of the Transformer model 222c2. Furthermore, even if the Transformer model 222c2 outputs noise, it may be possible for the output noise to be female voice data, and therefore it may be determined whether or not the noise is present. While the processing of this noise has been described using the Transformer model 222c2 as an example, similar processing may also be performed on the other Transformer models 222c1, 222c3, and 222c4. As will be described later, the generation of multiple tokenized features is performed using multiple different threads.

[0072] 7 is a diagram showing an example of logical functional blocks realized in the control unit 21 of the information processing device 200 in order to execute information processing according to this embodiment. As shown in Fig. 7, the control unit 21 of the information processing device 200 executes a computer program to realize the functions of an STFT unit 211, a subtraction unit 212, a VAE encoder unit 213, a Tokenizer unit 314, Transformer units 215a-215d, De-Tokenizer units 216a-216d, a duplication determination unit 219, VAE decoder units 217a-217d, and ISTFT units 218a-218d.

[0073] The control unit 21 receives mixed audio data consisting of audio waveforms containing speech sounds from multiple speakers via the audio input unit 13. The control unit 21 may also receive mixed audio data by reading it from an input audio file 222d that stores mixed audio data containing speech sounds from multiple speakers. The control unit 21 may also receive audio data containing speech sounds from a single speaker. The input audio data is, for example, data in WAB format, and is 256 x 256 dimensional data.

[0074] The STFT transform unit 211 performs a transformation process to transform the mixed audio data into a mixed spectrogram using a short-time Fourier transform. Here, the mixed audio data includes audio data of a single speaker's speech in addition to speech from multiple speakers. Here, a spectrogram corresponds to a so-called voiceprint, in which the spectra of each audio data segment (i.e., frame) obtained by extracting the frequency components and amplitude components of an audio signal from an audio waveform using a Fourier transform are arranged along the time axis.

[0075] The VAE encoder unit 213 performs extraction processing to extract mixed features by inputting a mixed spectrogram to an encoder (not shown) of the VAE model 222a. Specifically, the mixed features are real-valued 1024-dimensional feature vectors. The VAE encoder unit 213 can reduce the amount of information by converting 256×256-dimensional audio data into 1024-dimensional features. Furthermore, the VAE encoder unit 213 duplicates the extracted mixed features by the number of Transformer models 222c1, 222c2, 222c3, and 222c4 (four in this embodiment).

[0076] The Tokenizer units 214a-214d input the mixed features extracted by the VAE encoder unit 213 to the Tokenizer models 222b1-222b4, respectively, and tokenize the mixed features into integer one-dimensional vector values. In this embodiment, the same Tokenizer models 222b1-222b4 are used, and therefore the same Tokenizer units 214a-214d are also used. It is also possible to make the Tokenizer models 222b1-222b4 different in accordance with the Transformer models 222c1-222c4, and use different Tokenizer units 214a-214d, respectively.

[0077] The Transformer units 215a-215d each input the tokenized mixed features to the corresponding Transformer models 222c1-222c4, thereby performing a generation process to generate tokenized first speaker features corresponding to a first speaker, who is one of the multiple speakers. The data output from the Transformer units 215a-215d is an integer, one-dimensional vector value. Note that in this embodiment, the mixed features are tokenized by the Tokenizer units 214a-214d before being input to the Transformer units 215a-215d. However, the mixed features may also be input to the Transformer units 215a-215d without being tokenized by the Tokenizer units 214a-214d. Tokenizing the mixed features by the Tokenizer units 214a-214d is preferable for speeding up processing, since it reduces the amount of information processed by the Transformer units 215a-215d. Furthermore, tokenization allows global features of the data to be captured, thereby improving data accuracy. It is believed that the transformer separates the first speaker from the other speakers based on the frequency, speed, and intonation of the voice.

[0078] The De-Tokenizer units 216a-216d restore the original first speaker features by inputting the tokenized first speaker features generated by the corresponding Transformer units 215a-215d to the Tokenizer models 222b1-222b4, respectively. Since the same Tokenizer models 222b1-222b4 are used, the De-Tokenizer units 216a-216d are also all the same. It is also possible to use different Tokenizer models 222b1-222b4 to match the Transformer models 222c1-222c4, and use different De-Tokenizer units 216a-216d for each.

[0079] The duplication determination unit 219 determines whether the tokenized features output from each of the de-tokenizer units 216a-216d overlap. For example, consider a case where the mixed audio data includes an adult male voice. In this case, features may be extracted from both the male Transformer model 222c1 and the adult Transformer model 222c3. Therefore, the duplication determination unit 219 determines whether the extracted (tokenized) features overlap by using the Euclidean distance between the features. Specifically, if the Euclidean distance between the two features being compared is equal to or less than a predetermined threshold, the features are determined to overlap, and one of the features is deleted. Note that the predetermined threshold can be calculated in advance from the training audio data.

[0080] The VAE decoder units 217a-217d each perform a reconstruction process to reconstruct (or restore) a (256 x 256 dimensional) first speaker spectrogram by inputting the first speaker features (1024 real dimensions) restored by the De-Tokenizer units 216a-216d into the decoder (not shown) of the VAE model 222a.

[0081] Each of the ISTFT units 218a-218d performs a conversion process using an inverse short-time Fourier transform to convert the spectrogram of the first speaker into speech data of the first speaker. The speech data output by the ISTFT units 218a-218d is written to an output speech file 222e.

[0082] The subtraction unit 212 performs subtraction processing to subtract the first speaker spectrogram restored by the VAE decoder units 217a-217d from the mixed spectrogram.

[0083] By performing the above extraction process, generation process, reconstruction process, and subtraction process in this order on the spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram, it is possible to separate different speakers. By repeating these processes, it is possible to separate speakers from mixed speech data that contains speech sounds from an unspecified number of multiple speakers.

[0084] Next, a specific operation of the control unit 21 according to this embodiment will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the flow of processing by the control unit 21.

[0085] In step S801, the control unit 21 reads out mixed audio data including speech sounds from a plurality of speakers from the input audio file 222d. The control unit 21 may also receive mixed audio data consisting of audio waveforms including speech sounds from a plurality of speakers via the audio input unit 13.

[0086] In step S802, the control unit 21 (STFT transformation unit 211) executes a transformation process to transform the mixed audio data into a mixed spectrogram using a short-time Fourier transform.

[0087] In step S803, the control unit 21 stores the mixed spectrogram, which is the result of the short-time Fourier transform, in the first memory M01.

[0088] In step S804, the control unit 21 (subtraction unit 212) executes subtraction processing to subtract the output spectrogram stored in the second memories M02a-M02d in step S811 (described later) from the mixed spectrogram, which is the input spectrogram stored in the first memory M01 in step S803, and stores (i.e., overwrites) the new mixed spectrogram, which is the result of the subtraction processing, in the first memory M01. Because the initial values ​​of the second memories M02a-M02d are empty, if no data is stored in the second memories M02a-M02d, nothing is subtracted from the mixed spectrogram, which is the input spectrogram.

[0089] In step S805, the control unit 21 (VAE encoder unit 213) executes extraction processing to extract mixed features by inputting the mixed spectrogram, which is the subtraction result in step S804, to the encoder of the VAE model 222a. Then, the control unit 21 (VAE encoder unit 213) copies the extracted mixed features by the number of Transformer models 222c1-222c4 (four in this embodiment).

[0090] In step S806, the control unit 21 (Tokenizer units 214a-214d) inputs the mixed features extracted in step S805 to the respective Tokenizer models 222b1-222b4, thereby tokenizing the mixed features into integer one-dimensional vector values.

[0091] In step S807, the control unit 21 (Transformer units 215a-215d) executes a generation process to generate tokenized first speaker features corresponding to a first speaker who is one of the multiple speakers, by inputting the tokenized mixed features to each of the Transformer models 222c1-222c4. Note that in this embodiment, the mixed features are tokenized in step S806 and then the generation process is performed in step S807, but the generation process in step S807 may be performed without tokenizing the mixed features.

[0092] In step S808, the control unit 21 (De-Tokenizer units 216a-216d) inputs the tokenized first speaker features generated in step S807 into the respective Tokenizer models 222b1-222b4, thereby restoring the original first speaker features.

[0093] Control unit 21 starts up separate threads equal to the number of Transformer models 222c1-222c4 (four in this embodiment), and executes the processes from step S806 to step S808 for each thread. This allows extraction and restoration of first speaker features to be performed in parallel, thereby speeding up the process.

[0094] In step S809, the control unit 21 (duplication determination unit 219) determines whether the multiple tokenized features output in step S808 overlap. If the features are determined to overlap, they indicate the same speaker, so one of the overlapping features is deleted. Note that the duplication confirmation process in step S809 is executed in this thread.

[0095] In step S810, the control unit 21 (VAE decoder unit 217a-217d) performs a reconstruction process to reconstruct (or restore) the first speaker spectrogram (i.e., the spectrogram of the first speaker) by inputting the first speaker features restored in step S808 into the decoder of the VAE model 222a.

[0096] In step S811, the control unit 21 stores the first speaker spectrogram restored in step S810 as an output spectrogram in the second memories M02a-M02d.

[0097] In step S812, the control unit 21 (ISTFT transform units 218a-218d) executes a conversion process to convert the first speaker spectrogram into the first speaker's voice data using an inverse short-time Fourier transform.

[0098] In step S813, the control unit 21 determines whether the output speech of the speech data generated in step S812 is silent. If the output speech is silent, it is determined that there is no speech to separate, and the control unit 21 ends the process. If the output speech is not silent, the flow proceeds to step S814. Here, to determine whether the output speech is silent, any algorithm or technique for performing speech-to-text transcription (STT) known in the speech recognition field may be used. If no text representing the spoken content is obtained from the output speech using SST, the output speech can be determined to be silent. For example, if the output speech is noise such as environmental sound, incomprehensible text or nothing will be output. Alternatively, whether the output speech is silent may be determined by performing voice detection using a library called voice-activity-detection.

[0099] In step S814, the control unit 21 writes the audio data output in step S813 into the output audio file 222e.

[0100] The control unit 21 executes the processes from step S810 to step S814 for each thread in the separate threads started up in step S806. This allows the conversion process to audio data to be performed in parallel, thereby speeding up the process.

[0101] In step S815, the control unit 21 determines whether or not the execution of the process in step S814 has been completed in all of the launched separate threads. If the execution of the process in step S814 has not been completed in all of the separate threads, the control unit 21 waits for the execution of the process in step S814 to be completed in all of the threads. If the execution of the process in step S814 has been completed in all of the separate threads, the flow returns to step S804, and the control unit 21 (subtraction unit 212) executes the subtraction process. In this way, the subtraction process is executed after the execution of all of the separate threads has been completed.

[0102] As described above, the information processing device 200 according to this embodiment can obtain a spectrogram in which the first speaker is excluded from the unspecified number of speakers by subtracting the first speaker spectrogram from the mixed spectrogram corresponding to mixed audio data containing speech sounds from an unspecified number of speakers. This allows for separation of speech data from the first speaker from the mixed audio data, even when the mixed audio data is composed of an unspecified number of speakers. In this embodiment, parallel processing is performed for the number of Transformer models 222c1-222c4, so that multiple speakers can be separated at once, up to the number of Transformer models 222c1-222c4. Furthermore, by performing the extraction process, generation process, reconstruction process, and subtraction process in this order on the spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram, other speakers can be separated. Repeating these processes allows for separation of speakers from mixed audio data containing speech sounds from an unspecified number of speakers.

[0103] (Other embodiments) The above describes embodiments of the present disclosure, but the present disclosure should not be construed as being limited to the above embodiments, and can be applied to various embodiments and combinations within the scope that does not deviate from the gist of the present disclosure.

[0104] Furthermore, the processing flow described in the above embodiment is also an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged within the scope of the present invention. [Explanation of symbols]

[0105] 100, 200: Information processing device, 11, 21: Control device, 12, 22: Memory unit, 13: Audio input unit, 14: Audio output unit, 15···Microphone, 16···Speaker, 122a, 222a···VAE model, 122b, 222b1-222b4···Tokenizer model, 122c, 222c1-222c4...Transformer model

Claims

1. At least one memory (12, 22) in which a computer program is recorded; an information processing device comprising: at least one processor (11, 21) capable of executing the computer program stored in the at least one memory; The at least one memory includes: a first model (122a, 222a) trained to include an encoder that extracts a feature when a spectrogram is input, and a decoder that reconstructs input data when the feature is input; a second model (122c, 222c1-222c4) that has been trained to generate a feature corresponding to a certain speaker of the at least one speaker when a feature extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech of the at least one speaker is input; The at least one processor: Accepting mixed voice data containing speech sounds from multiple speakers; converting the mixed audio data into a mixed spectrogram using a short-time Fourier transform; performing an extraction process of extracting a mixed feature by inputting the mixed spectrogram to the encoder of the first model; performing a generation process of generating a first speaker feature corresponding to a first speaker among the plurality of speakers by inputting the mixed feature into the second model; performing a reconstruction process of reconstructing a first speaker spectrogram by inputting the first speaker feature to the decoder of the first model; an information processing device that performs a subtraction process of subtracting the first speaker spectrogram from the mixed spectrogram;

2. 2. The information processing device according to claim 1, wherein the at least one processor performs the extraction processing, the generation processing, the reconstruction processing, and the subtraction processing in this order on a spectrogram obtained by subtracting the first speaker spectrogram from the mixed spectrogram.

3. the at least one memory further stores a third model (122b, 222b1-222b4) that is trained to tokenize features extracted from a spectrogram and restore the original features from the tokenized features; The at least one processor: inputting the extracted mixed features into the third model to tokenize the mixed features; The information processing apparatus according to claim 1 , wherein the first speaker feature corresponding to the first speaker among the plurality of speakers is generated by inputting the tokenized mixed feature to the second model.

4. the second model includes a plurality of models trained to generate features corresponding to speakers having different attributes when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech of at least one speaker are input, and The information processing apparatus according to claim 1 , wherein the at least one processor generates a plurality of the features as the first speaker features by inputting the mixture features to each of the plurality of models.

5. The information processing device according to claim 4 , wherein the attributes include male, female, adult, and child.

6. The information processing apparatus according to claim 4 , wherein the at least one processor generates the plurality of features as the first speaker features using a plurality of threads different from each other.

7. The information processing apparatus according to claim 4 , wherein the at least one processor determines whether the plurality of feature amounts overlap.

8. The information processing apparatus according to claim 7 , wherein the at least one processor determines whether the plurality of feature quantities overlap each other by using a Euclidean distance between the plurality of feature quantities.

9. An information processing method using a first model trained to include an encoder that extracts features when a spectrogram is input, and a decoder that reconstructs input data when the features are input, and a second model trained to generate features corresponding to a certain speaker of at least one speaker when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech of the at least one speaker are input, receiving mixed voice data including speech voices of a plurality of speakers; performing a transformation process to convert the mixed audio data into a mixed spectrogram using a short-time Fourier transform; performing an extraction process of extracting a mixture feature by inputting the mixture spectrogram to the encoder of the first model; performing a generation process of generating first speaker features corresponding to a first speaker among the plurality of speakers by inputting the mixed features into the second model; performing a reconstruction process of reconstructing a first speaker spectrogram by inputting the first speaker features to the decoder of the first model; performing a subtraction process of subtracting the first speaker spectrogram from the mixed spectrogram.

10. An information processing program using a first model trained to include an encoder that extracts features when a spectrogram is input, and a decoder that reconstructs input data when the features are input, and a second model trained to generate features corresponding to a certain speaker of at least one speaker when features extracted from a spectrogram obtained by performing a short-time Fourier transform on speech data including speech of the at least one speaker are input, On the computer, receiving mixed voice data including speech voices of a plurality of speakers; performing a transformation process to convert the mixed audio data into a mixed spectrogram using a short-time Fourier transform; performing an extraction process of extracting a mixture feature by inputting the mixture spectrogram to the encoder of the first model; performing a generation process of generating first speaker features corresponding to a first speaker among the plurality of speakers by inputting the mixed features into the second model; performing a reconstruction process of reconstructing a first speaker spectrogram by inputting the first speaker features to the decoder of the first model; performing a subtraction process of subtracting the first speaker spectrogram from the mixed spectrogram.

Citation Information

Patent Citations

  • Voice recognition device and program

    JP2016042152A

  • Voice recognition device and voice recognition method

    JP2020013034A