Speech processing learning method, speech processing learning device, and program
The voice processing learning method addresses the challenges of training target speaker extraction models by using a voice database without speaker labels, enabling effective target speaker extraction and expanding the range of usable voice data.
Patent Information
- Application Number
- JP2023573787
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-17
AI Technical Summary
Existing target speaker extraction technologies face challenges in training models without speaker labels, datasets with unspecified speakers, or datasets where each speaker has only one utterance, leading to difficulties in domain adaptation and requiring costly manual annotation.
A voice processing learning method that includes feature extraction, speaker expression extraction, target speaker extraction, and optimization processes, allowing for the creation of a target speaker extraction model using a voice database without speaker labels or with unspecified speakers, by utilizing paired data with enrollment voices and mixed sounds.
Enables practical and effective target speaker extraction performance even with voice databases lacking speaker labels or containing voices of unspecified speakers, thereby expanding the range of usable voice data and improving robustness against speaker variations.
Smart Images

Figure 0007683746000001 
Figure 0007683746000002 
Figure 0007683746000003
Abstract
Description
Technical Field
[0001] The present invention relates to speech recognition technology, and more particularly to a target speaker extraction technology for extracting only the speech audio of a target speaker from a mixed sound including the speech of other speakers and noise in addition to the speech of the target speaker.
Background Art
[0002] In recent years, the performance of speech recognition has been improved due to the development of deep learning technology. However, even now, an example of a situation where speech recognition is difficult is the mixed sound (overlapping speech) of multiple people. To address this, the following technologies have been devised.
[0003] Blind source separation enables speech recognition by separating a speech that is difficult to recognize as it is in the mixed sound into the speech of each speaker (see, for example, Non-Patent Document 1).
[0004] Target speaker extraction uses the speech pre-registered by the target speaker as auxiliary information and obtains only the speech of the pre-registered speaker from the mixed sound (see, for example, Non-Patent Document 2). It has the advantage of not requiring prior information regarding the number of speakers included in the mixed sound and is a practically useful technology. Since the extracted speech contains only the voice of the target speaker, speech recognition is possible.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] Target speaker extraction is implemented using a neural network. When training the network, paired data including a mixed sound containing the target speaker's voice as input, the target speaker's pre-registered utterances, and the target speaker's voice to be extracted from the mixed sound as output is utilized. That is, for training to create a target speaker extraction model, the voice of an utterance by the target speaker that is different from the utterance to be extracted (enrollment voice) needs to be included in the dataset as the pre-registered utterance of the target speaker to be extracted.
[0007] Therefore, it was not possible to train target speaker extraction for a dataset without speaker labels, a dataset containing the voices of unspecified speakers, etc. Also, it was not possible to train target speaker extraction for a dataset with speaker labels but where each speaker uttered only one utterance.
[0008] As a result, it was difficult to train a target speaker extraction model or perform domain adaptation for voice logs anonymized from the perspective of privacy, voice logs of applications where multiple utterances are often not obtained from the same speaker, etc.
[0009] Furthermore, when performing learning for target speaker extraction using recordings of conference voices or the like, it may be necessary to manually annotate the speaker voices, which can be costly. Also, when there are speakers with similar voices, it has not always been possible to perfectly annotate the speakers.
[0010] In view of the above problems, an object of the present invention is to perform learning for target speaker extraction that is practically usable based on a voice database without speaker labels or a voice database including voices of an unspecified number of speakers.
Means for Solving the Problems
[0011] To solve the above problems, a voice processing learning method according to an aspect of the present invention includes a feature extraction process of extracting a feature amount composed of time-series data that is a fixed-length vector from the same voice as the target voice, which is the utterance voice of a target speaker, received as enrollment voice; a speaker expression extraction process of extracting a speaker expression that is a fixed-length vector from the feature amount extracted by the feature extraction process; a target speaker extraction process of extracting a voice estimated to be the target voice from a mixed voice composed of the target voice, non-target voices that are voices of speakers different from the target speaker, and noise using the speaker expression extracted by the speaker expression extraction process; and an optimization process of calculating a loss function using the voice extracted by the target speaker extraction process and the target voice, and optimizing the target speaker extraction process so that the calculated value becomes minimum.
Effects of the Invention
[0012] According to the present invention, it is possible to perform learning for target speaker extraction that is practically usable based on a voice database without speaker labels or a voice database including voices of an unspecified number of speakers.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present invention will be described in detail. Note that components having the same function are denoted by the same reference numerals, and redundant descriptions are omitted.
[0015] Fig. 1 shows a diagram illustrating a functional configuration example of an audio processing learning apparatus according to an embodiment of the present invention. The audio processing learning apparatus 1 shown in Fig. 1 includes a feature extraction unit 11, a speaker expression extraction unit 13, a target speaker extraction unit 15, and an optimization unit 16. Fig. 2 is a diagram showing an example of a processing flow of an audio processing learning method in the audio processing learning apparatus according to an embodiment of the present invention. By the audio processing learning apparatus 1 performing the processing of each step from step S11 to step S16 illustrated in Fig. 2, the audio processing learning method of the embodiment is realized. One aspect of the audio processing learning apparatus 1 learns using the target audio i, the enrollment audio i', and the mixed audio X that constitute the learning data γ created by the learning data creation apparatus 2 described later, thereby creating a target speaker extraction model Z that is a learned model.
[0016] Hereinafter, while explaining the processing performed by each element of the audio processing learning apparatus 1 using Fig. 1 and Fig. 2, the function of the audio processing learning apparatus 1 and the audio processing learning method performed by the audio processing learning apparatus 1 will be described.
[0017] [Feature Extraction Unit 11] The feature extraction unit 11 performs feature extraction processing (step S11). That is, the feature extraction unit 11 acquires the enrollment audio i' from among the paired data stored in the learning data γ described later, performs feature extraction on this enrollment audio i', and acquires time-series data of a fixed-length vector representation. The feature extraction unit 11 outputs the acquired time-series data to the speaker expression extraction unit 13. As a feature extraction method, for example, known filter bank features for a short-time window or short-time Fourier transform results can be used. The feature extraction unit 11 can also be configured by a neural network. An example of the neural network of the feature extraction unit 11 includes a known convolutional neural network. In the normal method, the target audio and the enrollment audio are utterances of the same speaker, but different utterances are used for each. However, in the present invention, as will be described later, the enrollment audio i' stored in the learning data γ is set to the same audio as the target audio i that is the uttered voice of the target speaker.
[0018] Therefore, the feature extraction unit 11 may be configured to receive the target voice i from the training data γ, set the received target voice i as the enrollment voice i', and perform the above-described feature extraction on this enrollment voice i'.
[0019] [Speaker expression extraction unit 13] The speaker expression extraction unit 13 performs speaker expression extraction processing (step S13). That is, the speaker expression extraction unit 13 extracts a speaker expression, which is a fixed-length vector, from the time-series data received from the feature extraction unit 11 by the feature extraction processing in step S11. The speaker expression extraction unit 13 can be constructed using a known neural network or a known speaker expression extractor such as i-vector. The speaker expression extraction unit 13 outputs the extracted speaker expression to the target speaker extraction unit 15.
[0020] [Target speaker extraction unit 15] The target speaker extraction unit 15 performs target speaker extraction processing (step S15). That is, the target speaker extraction unit 15 uses the speaker expression received from the speaker expression extraction unit 13 by the processing in step S13 to obtain the mixed sound X from among the paired data stored in the training data γ described later, and extracts the utterance that is estimated to be the target speaker from that mixed sound X. The target speaker extraction unit 15 can be configured by a known neural network. As an example of the neural network of the target speaker extraction unit 15, a network structure such as a known ConvTasNet can be mentioned. The target speaker extraction unit 15 outputs the extracted utterance estimated to be the target speaker to the optimization unit 16. Note that the target speaker extraction unit 15 receives the optimized parameters from the optimization unit 16 by the processing in step S16 of the optimization unit 16 described later, and reflects those parameters in the parameters included in the target speaker extraction model Z of the target speaker extraction unit 15 to create the target speaker extraction model Z.
[0021] [Optimization unit 16] The optimization unit 16 performs an optimization process (step S16). That is, the optimization unit 16 calculates a loss function using the speech presumed to be the target speaker received from the target speaker extraction unit 15 by the process of step S15 and the target audio i received from the paired data stored in the learning data γ described later. The optimization unit 16 optimizes the parameters included in the model related to target speaker extraction that the target speaker extraction unit 15 has so that the calculated value of the loss function indicates the minimum value.
[0022] As an example of the loss function used by the optimization unit 16, a known sd-SNR can be cited. Also, as an example of the optimization method, a known Adam method can be cited. The optimization process performed by the optimization unit 16 can be performed on all the parameters included in the model related to target speaker extraction. The optimization process performed by the optimization unit 16 can also be performed only on the remaining parameters while fixing some parameters. The optimization unit 16 outputs the parameters included in the optimized model related to target speaker extraction to 15. Note that the parameters output from the optimization unit 16 to the target speaker extraction unit 15 are reflected in the parameters included in the model related to target speaker extraction that the target speaker extraction unit 15 has, and the target speaker extraction model Z is created.
[0023] By the voice processing learning device 1 performing the processes from step S11 to step S16 described above, the voice processing learning method is realized. As a result, the target speaker extraction model Z shown in FIG. 1, which is a learned model, is created.
[0024] <Learning data creation device> Hereinafter, a method for creating learning data used by the voice processing learning device 1 according to an embodiment of the present invention will be described. FIG. 3 shows a diagram showing a functional configuration example of a learning data creation device according to an embodiment of the present invention.
[0025] The learning data creation device 2 shown in FIG. 3 is a device that creates learning data γ consisting of paired data used in the speech processing learning device 1 in order to learn target speaker extraction. As shown in FIG. 3, the learning data creation device 2 includes a target utterance extraction unit 21, a non-target utterance extraction unit 22, a noise extraction unit 23, a voice mixing unit 24, and a data creation unit 25. FIG. 4 is a diagram showing an example of a processing flow of a learning data creation method in the learning data creation device according to an embodiment of the present invention. By the learning data creation device 2 performing the processing of each step from step S21 to step S25 illustrated in FIG. 4, the learning data creation method according to an embodiment of the present invention is realized.
[0026] Hereinafter, while explaining the processing performed by each element of the learning data creation device 2 using FIGS. 3 and 4, the function of the learning data creation device 2 and the learning data creation method performed by the learning data creation device 2 will be explained.
[0027] [Target Utterance Extraction Unit 21] The target utterance extraction unit 21 performs target utterance extraction processing (step S21). That is, the target utterance extraction unit 21 receives audio data from an audio database α containing a large number of unspecified utterances prepared in advance, and extracts target audio i, which is the utterance audio of the target speaker, from the audio data. The target utterance extraction unit 21 outputs the extracted target audio i to the audio mixing unit 24 and the data creation unit 25.
[0028] [Non-Target Utterance Extraction Unit 22] The non-target utterance extraction unit 22 performs non-target utterance extraction processing (step S22). That is, the non-target utterance extraction unit 22 receives audio data from the above-mentioned audio database α, and extracts non-target audio k, which is the utterance audio of a speaker other than the target speaker, from the audio data. The non-target utterance extraction unit 22 outputs the extracted non-target audio k to the audio mixing unit 24.
[0029] [Noise Extraction Unit 23] The noise extraction unit 23 performs a noise extraction process (step S23). That is, the noise extraction unit 23 receives noise data from a previously prepared noise database β, and extracts noise r, which is a noise signal, from the noise data. The noise extraction unit 23 outputs the extracted noise r to the audio mixing unit 24.
[0030] [Audio mixing unit 24] The audio mixing unit 24 performs an audio mixing process (step S24). That is, the audio mixing unit 24 mixes the target audio i received from the target utterance extraction unit 21, the non-target audio k received from the non-target utterance extraction unit 22, and the noise r received from the noise extraction unit 23 with an arbitrary gain to create a mixed sound X. The audio mixing unit 24 outputs the created mixed sound X to 25.
[0031] [Data creation unit 25] The data creation unit 25 performs a data creation process (step S25). That is, the data creation unit 25 creates learning data γ, which is a data set, based on the target audio i received from the target utterance extraction unit 21 and the mixed sound X received from the audio mixing unit 24. The data creation unit 25 sets the same audio as the target audio i, which is the utterance audio of the target speaker, as the enrollment audio i', and creates paired data of the three audios, namely the target audio i, the enrollment audio i', and the mixed sound X, as the learning data γ.
[0032] By performing the processes of step S21 to step S25 described above by the learning data creation device 2, the learning data creation method is realized, and the learning data used for the speech processing learning device 1 can be created.
[0033] The learning data creation device 2 can create the learning data used for the speech processing learning device 1 without preparing an enrollment audio other than the target audio i. The learning data creation device 2 can construct a learning data set based on a speech database without speaker labels or a speech database α that includes a large number of unspecified utterances.
[0034] It is also conceivable that the above-described voice database α is composed of an unspecified number of utterance data and data with speaker labels. In that case, for the unspecified number of utterance data, pair data can be generated using the above-described learning data creation method, and for the voice with speaker labels, pair data can also be generated by a conventional method using the speaker labels. Alternatively, pair data can also be generated for the voice with speaker labels using the method of the embodiment of the present invention described above.
[0035] In the above description, each process of step S21, step S22, and step S23 has been described as being performed in this order, but the order of these three processes can be processed in any order. Alternatively, two or three of these processes can be processed in parallel.
[0036] <First Modification Example> Hereinafter, a first modification example according to an embodiment of the present invention will be described. FIG. 5 is a diagram showing a functional configuration example of a first modification example of a voice processing learning device according to an embodiment of the present invention. In this modification example, the voice processing learning device 1 has a first data expansion unit 12 in addition to each element of the voice processing learning device 1 shown in FIG. 1 described above. FIG. 6 shows an example of a processing flow of a voice processing learning method in the first modification example of the voice processing learning device according to an embodiment of the present invention. By the voice processing learning device 1 shown in FIG. 5 performing the processing of each step from step S11 to step S16 illustrated in FIG. 6, the voice processing learning method of this modification example is realized.
[0037] Hereinafter, with reference to FIGS. 5 and 6, while explaining the functions and processes different from those of the voice processing learning device 1 in FIG. 1, the functions of the voice processing learning device 1 in FIG. 5 and the voice processing learning method performed by the voice processing learning device 1 will be described.
[0038] [First Data Expansion Unit 12] The first data augmentation unit 12 performs the first data augmentation process (step S12). That is, the first data augmentation unit 12 performs a data augmentation transformation on the time-series data, which is the feature amount extracted by the feature extraction unit 11 in the process of step S11. As a data augmentation method, for example, a method of randomly replacing a part of the time-frequency representation with a predetermined value such as 0 (zero) or the average value can be mentioned. As a specific method, known SpecAugment can be mentioned. As an implementation form, any one or a plurality of data augmentation related to frequency, data augmentation related to time, and time warping in SpecAugment can be adopted. The first data augmentation unit 12 outputs the transformed data to the speaker expression extraction unit 13.
[0039] [Speaker expression extraction unit 13] In this modification example, the process of step S13 performed by the speaker expression extraction unit 13 performs speaker expression extraction using the transformed data received from the first data augmentation unit 12. Other processes are the same as the process of step S13 shown in FIG. 2.
[0040] By the voice processing learning device 1 in FIG. 5 performing the processes from step S11 to step S16 in FIG. 6, a voice processing learning method is realized. As a result, the target speaker extraction model Z shown in FIG. 5, which is a learned model, is created.
[0041] <Second modification example> Hereinafter, a second modification example according to the embodiment of the present invention will be described. FIG. 7 is a diagram showing a functional configuration example of a second modification example of a voice processing learning device according to an embodiment of the present invention. In this modification example, the voice processing learning device 1 has a second data augmentation unit 14 in addition to each element of the voice processing learning device 1 shown in FIG. 5 described above. FIG. 8 shows an example of a processing flow of a voice processing learning method in a second modification example of a voice processing learning device according to an embodiment of the present invention. By the voice processing learning device 1 shown in FIG. 7 performing the processes of each step from step S11 to step S16 illustrated in FIG. 8, the voice processing learning method of this modification example is realized.
[0042] Hereinafter, while explaining functions and processes different from those of the voice processing learning device 1 in FIG. 5 by using FIGS. 7 and 8, the functions of the voice processing learning device 1 in FIG. 7 and the voice processing learning method performed by the voice processing learning device 1 will be described.
[0043] [Second data augmentation unit 14] The second data augmentation unit 14 performs a second data augmentation process (step S14). That is, the second data augmentation unit 14 performs a data augmentation conversion on the fixed-length speaker expression extracted by the speaker expression extraction unit 13 through the process of step S13. As a data augmentation method, for example, a method of replacing some elements of a vector with a predetermined value such as 0 (zero) or an average value using a known dropout method can be mentioned.
[0044] [Target speaker extraction unit 15] In this modification example, the process of step S15 performed by the target speaker extraction unit 15 performs target speaker extraction using the data after conversion received from the second data augmentation unit 14. Other processes are the same as the process of step S15 shown in FIG. 2.
[0045] By the voice processing learning device 1 in FIG. 7 performing the processes from step S11 to step S16 in FIG. 8, a voice processing learning method is realized. As a result, the target speaker extraction model Z shown in FIG. 7, which is a learned model, is created.
[0046] Note that in FIG. 8, both the process of the first data expansion unit 12 (step S12) and the process of the second data expansion unit 14 (step S14) are to be performed. However, the audio signal processing learning device 1 may be configured to perform only one of them. That is, similar to the first modification example, the audio signal processing learning device 1 may be configured to perform only the process of the first data expansion unit 12 (step S12) without performing the process of the second data expansion unit 14 (step S14). Further, the audio signal processing learning device 1 may be configured to perform only the process of the second data expansion unit 14 (step S14) without performing the process of the first data expansion unit 12 (step S12).
[0047] <Performance evaluation result> The performance evaluation results of the target speaker extraction model learned using the learning method for target speaker extraction described above will be described. FIG. 9 shows an example of the performance evaluation results of the target speaker extraction model in the case of performing a target speaker extraction experiment. In FIG. 9, as the target speaker extraction model for performing the target speaker extraction experiment, the case of using a model learned using a conventional speaker label (FIG. 9(a)), the case of using a model that implements the audio signal processing learning method of FIG. 2 (FIG. 9(b)), the case of using a model that implements the audio signal processing learning method of FIG. 6 (FIG. 9(c)), and the case of using a model that implements the audio signal processing learning method of FIG. 8 (FIG. 9(d)) are shown.
[0048] In this experiment, the Spontaneous Japanese Corpus (CSJ) was used. Also, as evaluation metrics for measuring the target speaker extraction performance, two evaluation metrics, the Signal to Distortion Ratio (SDR) and the Character Error Rate (CER), were adopted. A larger SDR value indicates higher extraction performance. Also, a smaller CER value indicates higher extraction performance.
[0049] In Fig. 9, the Japanese Spontaneous Speech Corpus (CSJ) is used in all cases, and the evaluation is performed using the same evaluation set. The conditions with and without using speaker labels use the same pair of mixed sound X and target speech i during learning, and the only difference lies in the method of selecting the enrollment speech i'. Therefore, it is possible to directly compare the performance of each model.
[0050] As shown in Fig. 9(b), when using the model implementing the speech processing learning method of Fig. 2, the signal-to-distortion ratio (SDR) was 15.5 and the character error rate (CER) was 12.3%. Even when compared with the results (Fig. 9(a)) of using a model trained with conventional speaker labels, where the signal-to-distortion ratio (SDR) was 17.3 and the character error rate (CER) was 8.1%, it can be said that a target speaker extraction performance sufficient for practical use was achieved.
[0051] By adopting the speech processing learning method in the speech processing learning apparatus 1 of the present embodiment, the information of the enrollment speech i' is appropriately varied, and it can be seen that the enrollment speech i' sufficiently plays the role of a so-called original enrollment speech. That is, from the results of Fig. 9(b) above, it can be said that practical-level target speaker extraction performance is achieved even for speech data without speaker labels.
[0052] Also, as shown in Fig. 9(c), when using the model implementing the speech processing learning method of Fig. 6, the signal-to-distortion ratio (SDR) was 16.4 and the character error rate (CER) was 11.5%. By comparing with the model of Fig. 9(b), it can be seen that both the SDR and CER have higher target speaker extraction performance.
[0053] Similarly, as shown in Fig. 9(d), when using the model implementing the speech processing learning method of Fig. 8, the signal-to-distortion ratio (SDR) was 17.2 and the character error rate (CER) was 9.7%. By comparing with the model of Fig. 9(b), it can be seen that both the SDR and CER have higher target speaker extraction performance.
[0054] From the results of FIGS. 9(c) and 9(d), it can be seen that the data augmentation method achieves high target speaker extraction performance even for voice data without speaker labels.
[0055] When the target voice i is used as the enrollment voice i', there is no variation in speaker characteristics between the target voice and the enrollment voice. Therefore, compared with the variation in speaker characteristics when the enrollment voice is another utterance of the target speaker as in the conventional method, it may occur that sufficient robustness against the variation cannot be obtained. It can be seen that the adoption of the data augmentation method of the present embodiment contributes to obtaining robustness against this variation.
[0056] In particular, in the case of FIG. 9(d), the signal-to-distortion ratio (SDR) is 17.2, which is only slightly different from the signal-to-distortion ratio (SDR) of 17.3 in FIG. 9(a) with speaker labels. The result of the character error rate (CER) in the case of FIG. 9(d) is 9.7%, which is close to the value of the character error rate (CER) of 8.1% in the result of FIG. 9(a) with speaker labels. Voice resources without speaker labels can be obtained in large quantities, especially in real data, compared with those with voice labels. In this experiment using the same amount of data, the model performance in FIG. 9(d) did not result in a significant inferiority compared to the model performance in FIG. 9(a). Therefore, it is considered that the practical performance can be greatly improved by learning using a large amount of available speaker-labeled data.
[0057] As described above, an embodiment of the present invention and its modification have been described for the voice processing learning device and the voice processing learning method. It is considered that the practical-level target speaker extraction performance is achieved even for voice data without speaker labels by this method. That is, regardless of whether speaker labels are given, practical target speaker extraction can be realized without preparing an enrollment voice.
[0058] Due to the above effects, it becomes possible to learn target speaker extraction using data that could not be utilized conventionally. That is, the range of voice data that can be utilized for learning the target speaker extraction model is expanded.
[0059] The expansion of the data utilization range yields the following two utilities. One is to utilize as learning data data that could not be included in the conventional learning data, thereby increasing the variations in utterances and speakers in the data. Therefore, the robustness regarding the differences between speakers in voice emphasis can be enhanced, and the performance of target speaker extraction can be improved. The other is to achieve domain adaptation by performing re-learning of voice emphasis also for a data set without speaker labels or a data set containing a large amount of data in which each speaker utters only one utterance.
[0060] Note that the various processes described above are not only executed in time series according to the description, but may also be executed in parallel or individually according to the processing ability of the device that executes the processes. Needless to say, appropriate changes can be made without departing from the spirit of the present invention.
[0061] [Program, Recording Medium] The various processes described above can be implemented by causing the recording unit 2020 of the computer 2000 shown in FIG. 10 to read a program for executing each step of the above method and causing the control unit 2010, the input unit 2030, the output unit 2040, the display unit 2050, etc. to operate.
[0062] The program describing this processing content can be recorded on a computer-readable recording medium. As the computer-readable recording medium, for example, any of a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, etc. may be used.
[0063] In addition, the distribution of this program can be carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in the storage device of a server computer and distributed by transferring the program from the server computer to other computers via a network.
[0064] A computer that executes such a program first stores, for example, the program recorded on a portable recording medium or the program transferred from a server computer in its own storage device. Then, when executing the process, this computer reads the program stored in its own recording medium and executes the process according to the read program. As another execution form of this program, the computer may directly read the program from the portable recording medium and execute the process according to the program. Furthermore, each time a program is transferred from the server computer to this computer, the computer may sequentially execute the process according to the received program. Also, the transfer of the program from the server computer to this computer may not be performed, and the above-described process may be executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only by the execution instruction and result acquisition. Note that the program in this embodiment includes information used for processing by an electronic computer and similar to the program (data having the property of defining the processing of the computer but not being a direct instruction to the computer).
[0065] In addition, in this embodiment, the present apparatus is configured by causing a computer to execute a predetermined program, but at least a part of these processing contents may be realized hardware-wise.
Claims
1. A feature extraction process for extracting a feature amount composed of time series data, which is a fixed-length vector, from the same voice as the target voice, which is the utterance voice of the target speaker received as enrollment voice; A speaker expression extraction process for extracting a speaker expression, which is a fixed-length vector, from the feature amount extracted by the feature extraction process; An intended speaker extraction process for extracting a voice estimated to be the target voice from a mixed sound composed of the target voice, a non-target voice that is the voice of a speaker different from the target speaker, and noise, using the speaker expression extracted by the speaker expression extraction process; An optimization process for calculating a loss function using the voice extracted by the intended speaker extraction process and the target voice, and optimizing the intended speaker extraction process so that the calculated value is minimized A speech processing learning method for performing.
2. Performing a first data augmentation process for performing data augmentation on the feature amount extracted by the feature extraction process, The speaker expression extraction process extracts a speaker expression, which is a fixed-length vector, from the feature amount data-augmented by the first data augmentation process. The speech processing learning method according to claim 1.
3. The data augmentation in the first data augmentation process replaces a part of the feature amount with a predetermined value. The speech processing learning method according to claim 2.
4. Performing a second data augmentation process for performing data augmentation on the speaker expression extracted by the speaker expression extraction process, The intended speaker extraction process extracts a voice estimated to be the target voice from the mixed sound using the speaker expression data-augmented by the second data augmentation process. The speech processing learning method according to claim 1.
5. The data augmentation in the second data augmentation process replaces a part of the speaker expression with a predetermined value. The speech processing learning method according to claim 4.
6. A feature extraction unit that extracts a feature amount composed of time series data, which is a fixed-length vector, from the same voice as the target voice, which is the utterance voice of the target speaker received as enrollment voice; A speaker expression extraction unit that extracts a speaker expression, which is a fixed-length vector, from the feature amount extracted by the feature extraction unit; An intended speaker extraction unit that extracts a voice estimated to be the target voice from a mixed sound composed of the target voice, a non-target voice that is the voice of a speaker different from the target speaker, and noise, using the speaker expression extracted by the speaker expression extraction unit; An optimization unit that calculates a loss function using the voice extracted by the target speaker extraction unit and the target voice, and optimizes the target speaker extraction unit so that the calculated value is minimized A voice processing learning device having the same. **Claim 7** A first data augmentation unit that performs data augmentation on the feature amount extracted by the feature extraction unit, A second data augmentation unit that performs data augmentation on the speaker expression extracted by the speaker expression extraction unit, and The speaker expression extraction unit extracts a speaker expression, which is a fixed-length vector, from the feature amount data-augmented by the first data augmentation unit, The target speaker extraction unit extracts a voice estimated as the target voice from the mixed voice using the speaker expression data-augmented by the second data augmentation unit The voice processing learning device according to claim 6. **Claim 8** A program for causing a computer to function as the voice processing learning method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Speaker recognition system
JP1993143094A
Speaker model creation system, recognition system, program and control device
JP2019219574A