Learning device, signal processing device, learning method, and learning program

By focusing on reducing artifact errors in single-channel speech enhancement using a specialized loss function, the learning device improves speech recognition performance, addressing the limitations of single-channel techniques.

JP7713147B2Active Publication Date: 2025-07-25NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023572283
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-07-25
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

Single-channel voice enhancement techniques often deteriorate speech recognition performance despite noise removal, limiting the effectiveness of speech recognition systems, especially when compared to multi-channel methods.

Method used

A learning device and method that uses a voice enhancement unit to generate enhanced signals by preferentially reducing artifact errors in single-channel speech enhancement, utilizing a loss function to minimize artifact errors while maintaining noise errors, thereby improving speech recognition performance.

Benefits of technology

The approach significantly enhances speech recognition accuracy by reducing artifact errors in single-channel systems, leading to improved speech recognition performance compared to conventional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713147000020
    Figure 0007713147000020
  • Figure 0007713147000021
    Figure 0007713147000021
  • Figure 0007713147000022
    Figure 0007713147000022
Patent Text Reader

Abstract

A learning device (10) has: a voice enhancement unit (11) for generating an enhanced signal in which the voice of a speaker is enhanced from an inputted learning observation signal, using a model for generating an enhanced signal in which the voice of a speaker is enhanced; and an updating unit (12) for updating a parameter of the model using, as a loss function for calculating the degree of similarity between a reference signal corresponding to an estimated target objective sound source signal of a learning observation signal and an enhanced signal generated by the model from the learning observation signal, a loss function defined so as to reduce artifact errors with priority from among noise errors and artifact errors included in the enhanced signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, a signal processing device, a learning method, and a learning program.

Background Art

[0002] Constructing a speech recognition system that is robust against acoustic interference such as background noise and reverberation is an issue in speech processing. Here, it has been confirmed that a multi-channel voice enhancement technique (beamformer) using multiple microphones greatly improves speech recognition performance.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] On the other hand, in a single-channel voice enhancement technique using a single microphone, even when using an enhanced signal with noise removed, the speech recognition performance may sometimes deteriorate rather than that of the observed signal with noise, and the effect on improving speech recognition performance is limited.

[0005] In fact, many devices have only a single microphone. Therefore, in order to realize a robust speech recognition system, it is important to develop speech enhancement techniques for both single-channel and multi-channel, along with speech enhancement techniques for multi-channel.

[0006] The present invention has been made in view of the above, and an object thereof is to provide a learning device, a signal processing device, a learning method, and a learning program that enable improvement of speech recognition performance by speech enhancement.

Means for Solving the Problems

[0007] In order to solve the above-described problems and achieve the object, a learning device according to the present invention uses a model that generates an enhanced signal in which a speaker's voice is enhanced, and from an input learning observation signal, an enhanced signal in which the speaker's voice is enhanced. A voice enhancement unit that generates a voice enhancement unit, a reference signal corresponding to an estimated target target sound source signal of the learning observation signal, and a loss function that calculates a similarity between the enhanced signal generated by the model from the learning observation signal, among the noise error and the artifact error included in the enhanced signal, An update unit that updates the parameters of the model using a loss function defined to preferentially reduce the artifact error.

[0008] In order to solve the above-described problems, a voice enhancement unit that generates an enhanced signal in which a speaker's voice is enhanced from an input observation signal using a model that generates an enhanced signal in which the speaker's voice is enhanced, and a voice recognition unit that performs voice recognition on the enhanced signal, and the model is a reference signal corresponding to an estimated target target sound source signal of the learning observation signal, and a loss function that calculates a similarity between the enhanced signal generated by the model from the learning observation signal, among the noise error and the artifact error included in the enhanced signal, It is a model in which the parameters are updated using a loss function defined to preferentially reduce the artifact error.

Effects of the Invention

[0009] According to the present invention, it is possible to improve speech recognition performance by speech enhancement.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0011] Hereinafter, with reference to the drawings, an embodiment of the present invention will be described in detail. Note that the present invention is not limited by this embodiment. Also, in the description of the drawings, the same parts are denoted by the same reference numerals. Hereinafter, for a vector or matrix A, when it is described as “^A”, it is assumed to be the same as “a symbol with ‘^’ written immediately above ‘A’”.

[0012] [Embodiment] In this embodiment, as an example, a learning method for executing learning of a model capable of improving speech recognition performance based on an analysis result of analyzing factors that deteriorate speech recognition performance using a single-channel speech enhancement (SE) enhanced signal, and a signal processing method using the model are proposed. In this embodiment, a signal processing method for a speech signal (observation signal) recorded by a single microphone (single channel) will be described, but the method is applicable not only to a single channel but also to speech signals recorded by a plurality of microphones (multi-channel).

[0013] [Analysis of Enhanced Signal] First, factors that deteriorate speech recognition performance were analyzed for the enhanced signal by single-channel SE.

[0014] Generally, it is often assumed that the processing distortion caused by single-channel SE is the cause of the deterioration of speech recognition performance. However, no systematic and detailed analysis or clarification of such distortion, particularly its impact on speech recognition, has been done so far. It is considered essential to clarify the influence of single-channel SE estimation error on speech recognition in order to improve the SE front-end design.

[0015] Here, focus on the single-channel SE task. y ∈ R T represents the T-length time-domain waveform of the observation signal. The observation signal y is modeled as Equation (1). s ∈ R T represents the sound source signal. n ∈ R T represents the background noise signal.

[0016] [Equation]

[0017] SE aims to reduce the noise signal n from the observation signal y. When the observation signal y is input, the enhanced signal ^s ∈ R Tis estimated as ^s = SE(y). SE(·) represents, for example, the SE process performed by a neural network.

[0018] Subsequently, in order to analyze the influence of the SE estimation error on the speech recognition performance, the SE estimation error decomposition was examined using orthogonal projection. FIG. 1 is a diagram for explaining the signal decomposition of the enhanced signal by orthogonal projection.

[0019] Since the enhanced signal ^s is obtained by performing an estimation process, it is inevitable to include an estimation error. The enhanced signal ^s is decomposed using orthogonal projection as shown in Equation (2).

[0020]

Equation

[0021] In Equation (2), s target represents the estimated target target sound source signal (hereinafter referred to as the target sound source signal), and e noise ∈R T represents the noise error, and e artif ∈R T represents the artifact error (see FIG. 1).

[0022] Specifically, by the error decomposition by orthogonal projection, the estimation error in SE is decomposed into the noise error e noise and the artifact error e artif . These two elements are obtained by projecting the SE error onto the speech / noise subspace spanned by the speech / noise signal and the subspace orthogonal to the speech / noise subspace. Ps ∈ R T×T represents the orthogonal projection matrix on the subspace spanned by the sound source signal (Equation (3)). Similarly, P s,n ∈R T×T represents the orthogonal projection matrix on the subspace spanned by the sound source signal and the noise signal (Equation (4)). Note that L-1 is the number of maximum allowable delays. These matrices are obtained by Equations (5) and (6).

[0023]

Equation

[0024]

Number

[0025]

Number

[0026]

Number

[0027] The decomposition terms of formula (2) are obtained using the projection matrices of formulas (7), (8), and (9).

[0028]

Number

[0029]

Number

[0030]

Number

[0031] Noise error e noise is composed of a linear combination of the speech signal and the noise signal, and is expected to be an observable signal as a real-world signal. These are called natural signals. Similar noise errors e noise naturally appear in the training samples, so the influence of this natural signal on speech recognition performance may be limited.

[0032] On the other hand, the artifact error e artifis composed of signals that cannot be represented by a linear combination of an audio signal and a noise signal (see Fig. 1), and is an artificial / unnatural signal. This unnatural signal is very diverse and may rarely appear in training samples. Therefore, it is hypothesized that speech recognition is more sensitive to artifact error e noise than to noise error e artif .

[0033] As SE evaluation metrics, Signal to Distortion Ratio (SDR) (Equation (10)), Signal to Noise Ratio (SNR) (Equation (11)), and Signal to Artifact Ratio (SAR) (Equation (12)) are used. SDR is derived as shown in Equation (10) by applying Equation (2).

[0034] [Number]

[0035] [Number]

[0036] [Number]

[0037] Next, an experiment was conducted to examine the influence of the error factor of artifact error e artif on speech recognition performance. In the experiment, in order to measure the influence of artifact error e artif and noise error e noise on speech recognition performance, the emphasized signal was changed by varying the magnitude of the error factor, and speech recognition was performed with the changed emphasized signal as the input.

[0038] Specifically, after decomposing the emphasized signal ^s using orthogonal projection, artifact error e artif and noise error e noiseBy increasing or decreasing as in Equation (13), the enhanced signal ^s ω ∈R T was synthesized.

[0039]

Number

[0040] ω noise is a parameter that controls the amount of noise error e noise , and ω nartif is a parameter that controls the amount of artifact error e artif . In this experiment, to obtain various enhanced signals ^s noise with different ratios of noise error e artif and artifact error e ω , the values of ω noise and ω artif were changed. As a result, while controlling the values of SNR and SAR, the same target sound source signal s target could be maintained. By inputting such a modified enhanced signal into the speech recognition system as an evaluation enhanced signal, the influence of each error factor on the speech recognition performance was directly measured.

[0041] Figure 2 is a diagram showing the WER for the evaluation enhanced signal. (a) in Figure 2 is a 3D plot showing the speech recognition results for the evaluation enhanced signal with the ratio of noise error e noise / artifact error e artif changed. (b) in Figure 2 is the corresponding 2D plot obtained by changing only one of the weights of ω noise and ω artif . The baseline (obs.) in (b) of Figure 2 represents the reference WER score of the observed signal, and the square symbols represent the WER scores of the original enhanced signal without change.

[0042] As shown in Figure 2, it can be confirmed that the original enhanced signal actually degrades the speech recognition performance compared to the observed signal. As shown in Figure 2, the artifact error e artifIt has been observed that by reducing [it], a significant improvement in speech recognition performance is possible. On the other hand, the speech recognition performance was not significantly affected even when the noise error e noise was increased or decreased. Based on these results, it was confirmed that among the noise error e noise and the artifact error e artif , the artifact error e artif has a greater impact on the degradation of speech recognition performance.

[0043] Therefore, based on this finding, in this embodiment, a learning method and a signal processing method for improving speech recognition performance are proposed. In this embodiment, as an approach to reducing the influence of the artifact error e artif , a method of reducing the artifact error e artif included in the enhanced signal ^s input to the speech recognition system was considered.

[0044] In the embodiment, when generating the enhanced signal ^s from the observed signal y, the learning of the voice enhancement unit is executed so that an enhanced signal with a more focused reduction in the magnitude of the artifact error e artif included in the enhanced signal ^s can be generated.

[0045] Specifically, in this embodiment, as a loss function for obtaining the similarity between the enhanced signal ^s and the reference signal s, among the noise error e noise included in the enhanced signal ^s and the artifact error e artif which is an unnatural signal, a loss function defined to preferentially reduce the artifact error e artif is used to train the model in the voice enhancement unit.

[0046] For example, in the embodiment, the model in the voice enhancement unit is trained using the loss function L1 defined in Equation (14). The loss function L1 includes the loss function L noise (the first loss function) and the loss function L artif (the second loss function), and the loss function L artif is a weighted function.

[0047]

Number

[0048] The loss function L of Equation (14) noise is a loss function defined to minimize the noise error. For example, as the loss function L noise , SDR (see Equation (15)), Classical_SNR (see Equation (16)), and SNR (see Equation (11)) can be used.

[0049]

Number

[0050]

Number

[0051] Also, the loss function L artif is a loss function defined to minimize the artifact error e artif . For example, as the loss function L artif , SAR (see Equation (11)) can be used. α is the weight added to the loss function L artif , that is, the weight coefficient (hyperparameter) that determines the priority of the artifact error e artif , and it can be appropriately changed according to the network configuration and data such as the observed signal.

[0052] Note that as the loss function, a loss function L1´ (see Equation (17)) in which the sum of the weight added to the loss function L noise and the weight β added to the loss function L artif is 1 may be used. In this case, β, which is the weight coefficient that determines the priority of the artifact error e artif , takes a value between 0 and 1.

[0053]

Number

[0054] Also, in the embodiment, the model in the voice enhancement unit is trained using the loss function L2 defined in Equation (18). The loss function L2 is the target sound source signal s target divided by the sum of the noise error e noise and the artifact error e artif weighted by the weight γ, and is expressed as a logarithmic function with the resulting value as the argument. The weight γ is a weight coefficient (hyperparameter) that determines the priority of the artifact error e artif .

[0055]

Equation

[0056] [Learning Device] Next, the learning device according to the embodiment will be described. FIG. 3 is a diagram schematically showing an example of the configuration of the learning device according to the embodiment.

[0057] The learning device 10 according to the embodiment is realized, for example, by loading a predetermined program into a computer including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and the CPU executing the predetermined program. Further, the learning device 10 has a communication interface for transmitting and receiving various information to and from other devices connected via a network or the like.

[0058] As shown in FIG. 3, the learning device 10 has a voice enhancement unit 11 and an update unit 12. A learning observation signal y t recorded in a single channel is input to the learning device 10. The learning device 10 uses a reference signal s corresponding to the target sound source signal of the learning observation signal y t to train the model used by the voice enhancement unit 11.

[0059] The voice enhancement unit 11 is a learning observation signal y recorded in a single channel tReceives the input. The voice enhancement unit 11 generates an enhanced signal ^s that emphasizes the speaker's voice using a model that generates an enhanced signal that emphasizes the speaker's voice from the input learning observation signal y t from the input learning observation signal y. The model is configured by, for example, a neural network.

[0060] The update unit 12 updates the model parameters based on the loss function L1, loss function L1', or loss function L2 described above as a loss function for calculating the similarity between the reference signal s and the enhanced signal ^s generated by the model from the learning observation signal.

[0061] [Learning Process] Next, the learning method processing procedure executed by the learning device 10 will be described. FIG. 4 is a flowchart showing the processing procedure of the learning method according to the embodiment.

[0062] As shown in FIG. 4, when the learning device 10 receives the input of the learning observation signal y t the voice enhancement unit 11 performs a voice enhancement process of generating an enhanced signal ^s that emphasizes the speaker's voice from the input learning observation signal y t (Step S11).

[0063] Then, the update unit 12 updates the model parameters based on the loss function L1, loss function L1', or loss function L2 (Step S12). The learning device 10 determines whether a predetermined end condition is satisfied (Step S13). The end condition is, for example, when the loss becomes equal to or less than a predetermined threshold, when the number of parameter updates reaches a predetermined number, when the parameter update amount becomes equal to or less than a predetermined threshold, and the like.

[0064] When the predetermined end condition is not satisfied (step S13: No), the learning device 10 returns to step S11. The learning device 10 repeats the voice enhancement process and the parameter update process until the predetermined end condition is satisfied. When the predetermined end condition is satisfied (step S13: Yes), the learning device 10 ends the learning process. The learning device 10 outputs the model (including model parameters) of the voice enhancement unit 11 to the signal processing device 20 (described later).

[0065] [Signal processing device] Next, the signal processing device according to the embodiment will be described. FIG. 5 is a diagram schematically showing an example of the configuration of the signal processing device according to the embodiment.

[0066] The signal processing device 20 according to the embodiment is realized, for example, by a computer including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and the CPU executes a predetermined program. Further, the signal processing device 20 has a communication interface for transmitting and receiving various information to and from other devices connected via a network or the like. As shown in FIG. 5, the signal processing device 20 includes a voice enhancement unit 21 and a voice recognition unit 22. An observation signal y recorded in a single channel is input to the signal processing device 20, and for example, a voice recognition result obtained by converting the voice signal into text is output.

[0067] The voice enhancement unit 21 receives the input of the observation signal y recorded in a single channel. The voice enhancement unit 21 generates an enhanced signal ^s in which the voice of the speaker is enhanced from the observation signal y. The voice enhancement unit 21 performs voice enhancement processing using the model trained by the learning device 10. The model is a model in which the parameters are updated based on the above-described loss function L1, loss function L1', or loss function L2.

[0068] The voice recognition unit 22 performs voice recognition on the enhanced signal ^s. For example, the voice recognition unit 22 outputs a voice recognition result obtained by converting an audio signal into text. For example, the voice recognition unit 22 performs voice recognition processing using a trained deep learning model.

[0069] [Signal Processing Method] Next, the signal processing method executed by the signal processing apparatus 20 will be described. FIG. 6 is a flowchart showing the processing procedure of the signal processing method according to the embodiment.

[0070] As shown in FIG. 6, when the signal processing apparatus 20 receives an input of the observation signal y, the voice enhancement unit 21 performs voice enhancement processing for generating an enhanced signal ^s obtained by enhancing the voice of the speaker from the observation signal y using the model trained by the learning apparatus 10 (step S21). The voice recognition unit 22 performs voice recognition processing on the enhanced signal ^s (step S22) and outputs a voice recognition result.

[0071] [Effects of the Embodiment] In the learning of the conventional voice enhancement unit, a loss function for reducing the entire estimation error is used without considering that the estimation error in the enhanced voice includes the artifact error e artif and the noise error e noise That is, in the learning of the conventional voice enhancement unit, a loss function defined to equally reduce the artifact error e artif and the noise error e noise is used.

[0072] For example, in the learning of the conventional voice enhancement unit, without considering the difference between the artifact error e artif and the noise error e noise , a loss function (see, for example, Equation (11), Equation (15), Equation (16)) that equally weights the artifact error e artif and the noise error e noise is used.

[0073] In contrast, in the present embodiment, as a loss function for obtaining the similarity between the enhanced signal ^s and the reference signal s, the noise error e noise and the artifact error e artif among them, the loss function L1, loss function L1', or loss function L2 defined to preferentially reduce the artifact error e artif is used to train the model in the voice enhancement unit.

[0074] This loss function L1, loss function L1', or loss function L2 induces learning to preferentially reduce the artifact error e artif by weighting the artifact error e noise so that, among the noise error e artif and the artifact error e artif the artifact error e is preferentially reduced.

[0075] Actually, the speech recognition accuracy of the signal processing device 20 was evaluated. As the voice enhancement unit 21, a time-domain noise reduction network (Denoising-TasNet) based on a neural network was adopted. As the speech recognition unit 22, a deep neural network hidden Markov model (DNN-HMM) hybrid ASR (Automatic Speech Recognition) system based on the standard recipe of Kaldi was adopted.

[0076] A dataset of reverberant noise background voice signals was generated from the Wall Street Journal (WSJ0) corpus of the voice source and the CHiME-3 corpus of the noise source, and used as the training set, development set, and evaluation set. The model used by the voice enhancement unit 21 was trained in the learning device 10 using the loss function L2.

[0077] Table 1 is a table showing the SAR and WER for the enhanced signal ^s generated by the voice enhancement unit 21.

[0078]

Table 1

[0079] Table 1 shows the SAR and WER when the enhanced signal ^s is generated using the model trained with the loss function L2 (in this embodiment). Table 1 shows the cases where γ is 2.0 and 3.0. For comparison, Table 1 also shows the SAR and WER when the enhanced signal is generated using the model trained with γ = 1.0, that is, the conventional method that does not weight the artifact error e artif of the conventional method without weighting the artifact error e during training.

[0080] As shown in Table 1, by setting γ to 2.0 and 3.0 and weighting the artifact error e artif of the loss function L2 to train the model, the SAR value can be increased compared to the conventional method. Therefore, by using the model in the embodiment, the ratio of the artifact error e artif included in the enhanced signal ^s can be reduced compared to the model trained by the conventional method.

[0081] As shown in Table 1, by setting γ to 2.0 and 3.0 and weighting the artifact error e artif of the loss function L2 to train the model, the WER can be improved compared to the model trained by the conventional method. Thus, it is demonstrated that the speech recognition performance is improved by using the model trained in the learning device 10 according to the embodiment compared to the model trained by the conventional method.

[0082] Thus, in the embodiment, by adopting the loss function L1, the loss function L1´, or the loss function L2 to train the model of the speech enhancement unit 11, 21, an enhanced signal with a more significantly reduced magnitude of the artifact error e artif can be generated. Therefore, according to the embodiment, it is possible to reduce the artifact error e artif included in the enhanced signal ^s input to the speech recognition system, and improve the speech recognition performance by speech enhancement.

[0083] [System Configuration of Embodiment] Each component of the learning device 10 and the signal processing device 20 is a functional concept, and does not necessarily have to be physically configured as shown in the figure. That is, the specific forms of the distribution and integration of the functions of the learning device 10 and the signal processing device 20 are not limited to those shown in the figure, and all or part of them can be functionally or physically distributed or integrated in any unit according to various loads, usage situations, etc.

[0084] In addition, each process performed in the learning device 10 and the signal processing device 20 may be realized in whole or in any part by a program analyzed and executed by a CPU, a GPU (Graphics Processing Unit), and a CPU, a GPU. Also, each process performed in the learning device 10 and the signal processing device 20 may be realized as hardware by wired logic.

[0085] Also, among the processes described in the embodiment, all or part of the processes described as being automatically performed can be manually performed. Or, all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the above-described and illustrated process procedures, control procedures, specific names, and information including various data and parameters can be appropriately changed unless otherwise specified.

[0086] [Program] FIG. 7 is a diagram showing an example of a computer in which the learning device 10 and the signal processing device 20 are realized by executing a program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0087] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores a boot program such as BIOS (Basic Input Output System). Hard disk drive interface 1030 is connected to hard disk drive 1090. Disk drive interface 1040 is connected to disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into disk drive 1100. Serial port interface 1050 is connected to, for example, mouse 1110 and keyboard 1120. Video adapter 1060 is connected to, for example, display 1130.

[0088] Hard disk drive 1090 stores, for example, OS (Operating System) 1091, application program 1092, program module 1093, and program data 1094. That is, the programs defining the respective processes of learning device 10 and signal processing device 20 are implemented as program module 1093 in which code executable by computer 1000 is described. Program module 1093 is stored in, for example, hard disk drive 1090. For example, program module 1093 for executing processes similar to the functional configurations in learning device 10 and signal processing device 20 is stored in hard disk drive 1090. Note that hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0089] Also, the setting data used in the processes of the above-described embodiments is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 to RAM 1012 and executes them as needed.

[0090] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090. For example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a LAN (Local Area Network) or a WAN (Wide Area Network)). Then, the program module 1093 and the program data 1094 may be read by the CPU 1020 from another computer via the network interface 1070.

[0091] As described above, the embodiments to which the invention made by the present inventor is applied have been described. However, the present invention is not limited by the description and drawings that form part of the disclosure of the present invention according to the present embodiment. That is, all other embodiments, examples, operation techniques, etc. made by those skilled in the art based on the present embodiment are included in the scope of the present invention.

Explanation of Reference Numerals

[0092] 10 Learning device 11, 21 Voice enhancement unit 12 Update unit 20 Signal processing device 22 Voice recognition unit

Claims

1. An audio enhancement unit that generates an enhanced signal with the speaker's voice emphasized from the input learning observation signal using a model that generates an enhanced signal with the speaker's voice emphasized; As a loss function for calculating the similarity between a reference signal corresponding to the estimated target target sound source signal of the learning observation signal and the enhanced signal generated by the model from the learning observation signal, among the noise error and the artifact error included in the enhanced signal, an update unit that updates the parameters of the model using a loss function defined to preferentially reduce the artifact error; A learning device characterized by comprising:

2. The update unit includes, as the loss function, a first loss function defined to reduce the noise error and a second loss function defined to reduce the artifact error, and uses a loss function in which the second loss function is weighted. The learning device according to claim 1.

3. The update unit uses, as the loss function, a loss function expressed by a logarithmic function having, as a real number, a value obtained by dividing the estimated target target sound source signal by the sum of the noise error and the weighted artifact error. The learning device according to claim 1.

4. The learning observation signal is an audio signal recorded by a single microphone. The learning device according to claim 1.

5. An audio enhancement unit that generates an enhanced signal with the speaker's voice emphasized from the input observation signal using a model that generates an enhanced signal with the speaker's voice emphasized; An audio recognition unit that performs audio recognition on the enhanced signal; Comprising: The model is a model whose parameters are updated using a loss function defined to preferentially reduce the artifact error among the noise error and the artifact error included in the enhanced signal, as a loss function for calculating the similarity between a reference signal corresponding to the estimated target target sound source signal of the learning observation signal and the enhanced signal generated by the model from the learning observation signal. A signal processing device characterized by the above.

6. The observation signal and the learning observation signal are audio signals recorded by a single microphone. The signal processing device according to claim 5.

7. A learning method executed by a learning device, A step of generating an emphasized signal that emphasizes the speaker's voice from the input learning observation signal using a model that generates an emphasized signal that emphasizes the speaker's voice; As a loss function for calculating the similarity between a reference signal corresponding to the estimated target target sound source signal of the learning observation signal and the emphasized signal generated by the model from the learning observation signal, among the noise error and the artifact error included in the emphasized signal, a step of updating the parameters of the model using a loss function defined to preferentially reduce the artifact error; A learning method characterized by including the above.

8. A step of generating an emphasized signal that emphasizes the speaker's voice from the input learning observation signal using a model that generates an emphasized signal that emphasizes the speaker's voice; As a loss function for calculating the similarity between a reference signal corresponding to the estimated target target sound source signal of the learning observation signal and the emphasized signal generated by the model from the learning observation signal, among the noise error and the artifact error included in the emphasized signal, a step of updating the parameters of the model using a loss function defined to preferentially reduce the artifact error; A learning program for causing a computer to execute the above.

Citation Information

Patent Citations

  • Signal processing device, signal processing program, signal processing method, and sound collection device

    JP2020012980A

  • Method, device, apparatus, and computer-readable storage medium for speech recognition

    JP2021086154A

  • Speech recognition system and method for using a speech recognition system

    JP2021507312A

  • Using a predictive model to automatically enhance audio having various audio quality issues

    US20210343305A1

  • Voice recognition method, device, and computer-readable storage medium

    WO2021143327A1