Learning device, inference device, learning method, inference method, and program
The speech understanding model addresses the challenge of recognizing fine-grained speech information by integrating speech encoder outputs over time intervals, enhancing the accuracy of systems that rely on speech recognition results.
Patent Information
- Application Number
- PCT/JP2024/001906
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-31
AI Technical Summary
Existing technologies face challenges in accurately recognizing fine-grained non-verbal and para-linguistic information from speech due to differences in configuration between image and speech encoders, and the variable length of speech, making it difficult to output speech information as natural language sentences effectively.
A deep learning model, known as the speech understanding model, integrates information from multiple layers of a speech encoder over time intervals using a voice encoder output integration block and a time information integration block, followed by a linear transformation and a large language model to generate natural language sentences from speech.
The model enables accurate recognition and output of various speech information, including delicate non-verbal and para-verbal information, improving processing accuracy in systems that utilize these recognition results.
Smart Images

Figure JP2024001906_31072025_PF_FP_ABST
Abstract
Description
Learning device, inference device, learning method, inference method, and program
[0001] The present disclosure relates to a learning device, an inference device, a learning method, an inference method, and a program.
[0002] It is known that speech contains three types of information: linguistic information, non-linguistic information, and paralinguistic information (hereinafter, these three types of information are collectively referred to as "speech information") (Non-Patent Document 1). Furthermore, technology for recognizing non-linguistic information and paralinguistic information from speech is known (Non-Patent Document 2).
[0003] On the other hand, in the field of image processing, a technique called image understanding technology is known as a technique for outputting various information contained in an image in natural language (Non-Patent Document 3).
[0004] H. Fujisaki, "Prosody, Models, and Spontaneous Speech," in Computing Prosody, Y. Sagisaka, N. Campbell, and N. Higuchi, Springer, pp.27-42, 1996. A. Ando, S. Kobashikawa, H. Kamiyama, R. Masumura, Y. Ijima and Y. Aono, "SOFT-TARGET TRAINING WITH AMBIGUOUS EMOTIONAL UTTERANCESFOR DNN-BASED SPEECH EMOTION CLASSIFICATION," in Proc. of ICASSP, 2018, pp. 4964-4968.D. Zhu, J. Chen, X. Shen, X. Li, M. Elhoseiny, "MINIGPT-4: ENHANCING VISION-LANGUAGE UNDERSTANDING WITH ADVANCED LARGE LANGUAGE MODELS," in arXiv preprint arXiv:2304.10592, 2023.
[0005] It is believed that by using a speech encoder instead of the image encoder used in image understanding technology, it is possible to realize a technology that outputs speech information contained in speech in natural sentences (hereinafter also referred to as "speech understanding technology"). However, due to differences in the configurations of image encoders and speech encoders and differences in the properties of images and speech, it is difficult to realize speech understanding technology simply by using a speech encoder instead of an image encoder.
[0006] The present disclosure has been made in consideration of the above points, and aims to realize a speech understanding technology.
[0007] A learning device according to one aspect of the present disclosure includes an input unit that inputs learning data including speech, a first sentence related to the speech, and a second sentence corresponding to the first sentence; a speech feature generation unit that generates information representing features of the speech for each predetermined time interval based on a speech feature extractor configured with multiple layers; a first integration unit that generates first integrated information for each time interval based on a first parameter, the first integrated information integrating the information representing the features generated by each of the predetermined multiple layers of the speech feature extractor; a second integration unit that generates second integrated information for each time interval based on a second parameter, the second integrated information integrating the first integrated information in the time direction; a calculation unit that calculates a generation probability of a third sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model; and a learning unit that learns learning target parameters including the first parameter and the second parameter based on the generation probability of the third sentence and the second sentence.
[0008] Speech understanding technology can be realized.
[0009] 1 is a diagram showing an example of a speech understanding model; FIG. 2 is a diagram showing an example of a speech encoder and a speech encoder output integrated block; FIG. 3 is a diagram showing an example of a time information integrated block; FIG. 4 is a diagram showing an example of a hardware configuration of a speech understanding device during learning; FIG. 5 is a diagram showing an example of a functional configuration of a speech understanding device during learning; FIG. 6 is a diagram showing an example of a learning dataset; FIG. 7 is a diagram showing an example of a detailed functional configuration of a model learning unit; FIG. 8 is a flowchart showing an example of a model construction process; FIG. 9 is a flowchart showing an example of a model learning process; FIG. 10 is a diagram showing an example of a functional configuration of a speech understanding device during inference; FIG. 11 is a diagram showing an example of a detailed functional configuration of an output sentence generation unit; FIG. 12 is a flowchart showing an example of an output sentence generation process.
[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0011] <Technology for Recognizing Non-Verbal Information and Paralinguistic Information> It is known that speech contains speech information (i.e., linguistic information, non-linguistic information, and paralinguistic information) (Non-Patent Document 1). Here, linguistic information refers to information about the words spoken by a speaker. Non-linguistic information refers to information that is not linguistic information and cannot be changed at will (e.g., information that represents the speaker's identity, gender, emotions, etc.). Paralinguistic information refers to information that is not linguistic information and can be changed at will (e.g., information that represents intentions, attitudes, etc.).
[0012] Conventional technologies for recognizing non-verbal and paralinguistic information from speech often predefine a finite number of states and then estimate which of those states most closely matches the speech. For example, the technology described in Non-Patent Document 2 uses a statistical model based on deep learning to estimate which emotional state, such as anger, joy, or sadness, the speech most closely matches. However, such conventional technologies are unable to recognize detailed non-verbal and paralinguistic information. For example, they are unable to estimate undefined emotional states such as "irritated" or emotional states spanning multiple emotional states such as "angry and sad." This results in a problem of reduced processing accuracy in downstream systems that use the recognition results of non-verbal and paralinguistic information (e.g., the accuracy of call analysis in contact center systems, the accuracy of dialogue control and analysis in voice dialogue systems, etc.).
[0013] <Image Understanding Technology> In the field of image processing, a technology known as image understanding technology, which outputs various information contained in an image in natural language, is known (Non-Patent Document 3). Note that natural language refers to a sentence written in a natural language (e.g., a language used by humans for communication, such as Japanese, English, or Chinese). Image understanding technology is composed of a deep learning model that combines a large-scale language model that acquires relationships and co-occurrences between words using a large amount of text data with an image encoder that extracts information about objects in the image from the image. When this deep learning model is given an input image and a natural language question about the input image (e.g., a sentence such as "What do you think about the logo in this image?"), it outputs an output sentence corresponding to the question (e.g., a sentence such as "This log is a simple and symbolic logo").
[0014] A deep learning model that realizes image understanding technology can estimate various information contained in an image in natural language by providing a pair of an input image, a natural language question about the input image, and a correct output sentence corresponding to the question. For example, in response to the question "What color is the logo in this image?", a sentence such as "It's pink" can be output as the output sentence.
[0015] <Speech understanding technology> By realizing speech understanding technology that outputs speech information contained in speech in natural sentences, it will be possible to recognize, for example, a variety of speech information including detailed non-linguistic and paralinguistic information, and output that speech information in natural sentences.
[0016] A simple way to realize speech understanding technology would be to use a speech encoder that extracts information representing the characteristics of speech from speech instead of the image encoder used in image understanding technology. However, in practice, it is difficult to realize speech understanding technology using this method. There are two reasons for this.
[0017] The first reason is the difference in the configuration between image encoders and speech encoders. That is, existing speech encoders (e.g., wav2vec2.0 (Reference 1), WavLM (Reference 2), etc.) extract different information at each layer. Therefore, as with image understanding technology, diverse speech information cannot be understood by using only the information ultimately output from the speech encoder. For example, Reference 2 suggests that information closer to physical properties, such as speaker information, is extracted at the lower layers of the speech encoder, while information closer to abstract properties, such as phonemicity, is extracted at the higher layers of the speech encoder. For this reason, for example, while speaking style recognition uses information extracted at the lower layers of the speech encoder, speech recognition must use information extracted at the higher layers of the speech encoder; otherwise, accurate natural-sounding output would be difficult.
[0018] The second reason is the difference in the nature of images and audio. Unlike images, audio has a variable length, so it is thought that time-domain processing is necessary to recognize the audio information contained in audio of any length and output that audio information in natural language.
[0019] <Speech Understanding Model> Therefore, we propose a deep learning model (hereinafter referred to as the "speech understanding model") that can solve the problems caused by the above two causes. This speech understanding model makes it possible to recognize diverse speech information, including detailed non-verbal and paralinguistic information, from speech of any length and output the speech information in natural sentences. In other words, it is possible to realize a speech understanding technology that outputs diverse speech information contained in speech of any length (more specifically, speech information related to at least one of the physical properties and abstract properties of the speech) in natural sentences. This can be expected to improve the processing accuracy of downstream systems that use the recognition results of non-verbal and paralinguistic information (e.g., the accuracy of call analysis in contact center systems, the accuracy of dialogue control and analysis in voice dialogue systems, etc.).
[0020] An example of a speech understanding model 1000 proposed in this embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the speech understanding model 1000.
[0021] As shown in FIG. 1, the speech understanding model 1000 is composed of a speech encoder 1100 , a speech encoder output integration block 1200 , a temporal information integration block 1300 , a linear transformation layer 1400 , and a large-scale language model 1500 .
[0022] The audio encoder 1100 is any existing audio encoder (e.g., wav2vec2.0, WavLM, etc.). The audio encoder 1100 receives audio (hereinafter also referred to as "input audio") as input and outputs information representing the features of the audio. At this time, the audio encoder 1100 receives, for each predetermined time interval, the input audio for that time interval and outputs information representing the features of the audio for that time interval.
[0023] The audio encoder output merging block 1200 merges, in each time interval, the outputs of multiple pre-specified layers from among the outputs of each layer of the audio encoder 1100. Hereinafter, each of the multiple pre-specified layers will be referred to as a "layer to be merged." However, it is assumed that each layer to be merged has the same number of output dimensions. The layer to be merged is specified by a user or the like from among the layers of the audio encoder 1100 that have the same number of output dimensions.
[0024] Hereinafter, the index representing the time interval is defined as t, where t = 1, ..., T. T is an index representing the last time interval of the input speech, and its value may vary depending on the length of the input speech. In addition, hereinafter, the number of layers to be integrated is defined as N, and the output of the nth (where n = 1, ..., N) layer to be integrated in the time interval t is defined as h n (t). For each n = 1, ..., N, h n (t) represents some feature of the input speech (e.g., physical or abstract properties), and each h n Since (t) can be described as a vector with a predetermined number of dimensions, each h n (t) will be called the "first audio feature vector."
[0025] In this case, the speech encoder output integration block 1200 integrates, in each time interval t, each first speech feature vector h n (t) as input, and each first speech feature vector h n (t) is output as a vector. n The vector obtained by integrating (t) is called the "first integrated vector" and is represented by e(t).
[0026] For example, as shown in Figure 2, suppose the speech encoder 1100 is composed of one convolutional layer and N transformer layers, and these N transformer layers are layers to be integrated. In this case, the output of the n-th transformer layer in time interval t is the first speech feature vector h n(t), and the speech encoder output synthesis block 1200 synthesizes these first speech feature vectors h n (t) to create a first integrated vector e(t).
[0027] Each first speech feature vector h n As a method for integrating (t), for example, weighted sum or linear transformation sum can be used. When weighted sum is used, e(t)=α 1 h 1 (t) + ... + α N h N (t) to create a first integrated vector e(t), where α 1 , ..., α N is also called the weighting coefficient, and α 1 +...+α N On the other hand, when using a linear transformation sum, e(t) = ((W 1 h 1 (t) + b 1 ) + ... + (W N h N (t) + b N )) / N to create the first integrated vector e(t), where W 1 , ..., W N , b 1 , ..., b N is also called a linear transformation coefficient and is a parameter to be learned.
[0028] The temporal information integration block 1300 integrates the first integrated vector e(t) in the time direction. That is, the temporal information integration block 1300 receives the first integrated vector e(t) for each time interval t as input and outputs a vector obtained by integrating each first integrated vector e(t) in the time direction. In this case, it is not sufficient to simply integrate the first integrated vector e(t) for all time intervals t; it is necessary to consider which parts of the time period should be emphasized and which parts should not. For example, when recognizing a speaker's emotion or speaker information from speech, it is necessary to ignore time periods representing short pauses in the speech or time periods where breathing occurs, and to emphasize the time periods during which the speaker is speaking. For this reason, the temporal information integration block 1300 integrates each first integrated vector e(t) by a weighted sum. Hereinafter, the vector obtained by integrating each first integrated vector e(t) in the time direction will be referred to as a "second integrated vector" and represented by v.
[0029] For example, as shown in FIG. 3, when first integrated vectors e(1), ..., e(T) are input, the time information integration block 1300 creates a second integrated vector v by integrating these first integrated vectors e(1), ..., e(T) in the time direction.
[0030] The first integrated vectors e(1), ..., e(T) can be integrated using, for example, a self-attentive pooling layer: E = [e(1), ..., e(T)] τ In this case, when using a self-attention pooling layer, the second integrated vector v is created by v = aE, where a = softmax(ReLU(W 1 'E)W 2 '). Also, W 1 ' and W 2 where ' is a training parameter, and τ is a symbol representing transposition. As a result, the first integrated vectors e(1), ..., e(T) are integrated by a weighted sum to obtain a second integrated vector v.
[0031] The time information integration block 1300 may integrate each of the first integrated vectors e(1), ..., e(T) in the time direction using a convolutional neural network. In this case, a matrix V = [v(1), ..., v(K)] consisting of K second integrated vectors v(1), ..., v(K) is obtained by V = conv1D(E). τ Here, the trainable parameters of the one-dimensional convolutional neural network are the training target parameters. K is an integer equal to or greater than 1, determined by the window size of the one-dimensional convolutional neural network and the sequence length T of the first integrated vectors e(1), ..., e(T).
[0032] The linear transformation layer 1400 linearly transforms the second integrated vector v. That is, the linear transformation layer 1400 receives the second integrated vector v as input and outputs a vector obtained by linearly transforming this second integrated vector v. Hereinafter, the vector obtained by linearly transforming the second integrated vector v will be referred to as the "second speech feature vector" and represented by w. The second speech feature vector w is created by w = Wv + b. Here, W and b are also called linear transformation coefficients and are training target parameters. Note that if the number of dimensions of the token embedding space in the large-scale language model 1500 is M, the number of dimensions of the second speech feature vector w is also M. This means that the linear transformation layer 1400 creates a second speech feature vector sequence with a length of 1 and M dimensions.
[0033] The time information integration block 1300 provides the matrix V=[v(1), . . . , v(K)]. τ is output, for example, v(1), ..., v(K) can be linearly transformed into K vectors, which are then used as second speech feature vectors w(1), ..., w(K). This means that a second speech feature vector sequence with a length of K and a dimensionality of M is created.
[0034] The large-scale language model 1500 is any existing large-scale language model (LLM). The large-scale language model 1500 receives an input sentence, which is a natural language question regarding the input speech, and a second speech feature vector w (or a second speech feature vector w(1), ..., w(K)), and outputs an output sentence corresponding to the question. Such an output sentence is generated according to a posterior probability given the input sentence and the second speech feature vector w (or a second speech feature vector w(1), ..., w(K)). Note that the large-scale language model 1500 calculates the posterior probability by processing a vector sequence combining an embedding vector sequence representing the embedded representation of the tokens constituting the input sentence and the second speech feature vector w (or a second speech feature vector w(1), ..., w(K)).
[0035] 1 shows a case where an input sentence of "Please tell me the emotional state of the person in the following speech" is given, and an output sentence of "This person is male and is a little irritated" is generated. Note that, for example, Llama 2 7B (Reference 3) or the like can be used as the large-scale language model 1500. However, a language model other than a large-scale language model may be used as the large-scale language model 1500 as long as it is a language model that can generate an output sentence according to the posterior probability when an input sentence and a second speech feature vector w (or second speech feature vectors w(1), ..., w(K)) are given.
[0036] For simplicity, the following description will be given assuming that a self-attention pooling layer is used in the temporal information integration block 1300, that a second integrated vector v is obtained in the temporal information integration block 1300, and that a second audio feature vector w is created in the linear transformation layer 1400. However, if a convolutional neural network is used in the temporal information integration block 1300, the following embodiments can be similarly applied by replacing the "second integrated vector v" with the "second integrated vector v(1), ..., v(K)" and the "second audio feature vector w" with the "second audio feature vector w(1), ..., w(K)."
[0037] The speech encoder 1100 may be referred to as, for example, a "speech feature extractor" or a "speech coder." The large-scale language model 1500 may be referred to as, for example, a "language model," a "natural language model," or a "natural language processing model." Furthermore, the components constituting the speech understanding model 1000 (the speech encoder 1100, the speech encoder output integration block 1200, the temporal information integration block 1300, the linear transformation layer 1400, and the large-scale language model 1500) may be referred to as, for example, a "module."
[0038] A speech understanding device 10 that realizes speech understanding technology using a speech understanding model 1000 shown in Fig. 1 will be described below. The speech understanding device 10 operates during a "model construction" phase during which the speech understanding model 1000 is constructed, a "learning" phase during which learning parameters for the speech understanding model 1000 are learned, and an "inference" phase during which an output sentence is generated by the speech understanding model 1000 using the learned parameters. During learning, the speech understanding device 10 is provided with a set of training data (hereinafter also referred to as a "training dataset") consisting of pairs of input speech, input sentences that are questions in natural language related to the input speech, and correct output sentences corresponding to the questions. Meanwhile, during inference, the speech understanding device 10 is provided with test data consisting of pairs of input speech and input sentences that are questions in natural language related to the input speech.
[0039] The speech understanding device 10 during model construction may be called, for example, a "model construction device" or a "model creation device." The speech understanding device 10 during learning may be called, for example, a "learning device," a "parameter estimation device," a "parameter optimization device." The speech understanding device 10 during inference may be called, for example, an "inference device," an "estimation device," a "natural sentence generation device," etc.
[0040] For simplicity, the following description will be given assuming that model construction is included in learning, and a case will be described in which the speech understanding device 10 also constructs the speech understanding model 1000 during learning.
[0041] [During Learning] The following describes the speech understanding device 10 during learning.
[0042] <Example of Hardware Configuration of Speech Understanding Device 10 During Learning> An example of the hardware configuration of the speech understanding device 10 during learning will be described with reference to Fig. 4. Fig. 4 is a diagram showing an example of the hardware configuration of the speech understanding device 10 during learning.
[0043] 4, the speech understanding device 10 during learning includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.
[0044] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the speech understanding device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.
[0045] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0046] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).
[0047] 4 is an example, and the hardware configuration of the speech understanding device 10 is not limited to this. For example, the speech understanding device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.
[0048] <Example of functional configuration of the speech understanding device 10 during learning> An example of the functional configuration of the speech understanding device 10 during learning will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the functional configuration of the speech understanding device 10 during learning.
[0049] 5, the speech understanding device 10 during training includes a model construction unit 201 and a model training unit 202. These units are realized, for example, by processing in which one or more programs installed in the speech understanding device 10 are executed by the processor 108 or the like. The speech understanding device 10 during training also includes a trained speech encoder storage unit 203, a trained large-scale language model storage unit 204, a speech understanding model storage unit 205, and a training dataset storage unit 206. Each of these storage units is realized, for example, by a storage area of the auxiliary storage device 107 or the like. However, at least one of these storage units may be realized by a storage area of a storage device (e.g., a storage device included in a database server or the like) communicatively connected to the speech understanding device 10.
[0050] The model construction unit 201 constructs the speech understanding model 1000 shown in FIG. 1 using the trained speech encoder stored in the trained speech encoder storage unit 203 and the trained large-scale language model stored in the trained large-scale language model storage unit 204. That is, the model construction unit 201 constructs the speech understanding model 1000 shown in FIG. 1 using the trained speech encoder as the speech encoder 1100 and the trained large-scale language model as the large-scale language model 1500. At this time, the model construction unit 201 initializes the training target parameters of the speech encoder output integration block 1200, the training target parameters of the time information integration block 1300, and the training target parameters of the linear transformation layer 1400. The training target parameters may be initialized by any method, such as random initialization or sampling from a predetermined distribution. Note that the trained speech encoder refers to a speech encoder whose parameters have been trained. Similarly, the trained large-scale language model refers to a large-scale language model whose parameters have been trained.
[0051] Furthermore, the model construction unit 201 stores the voice understanding model 1000 in the voice understanding model storage unit 205 .
[0052] The model training unit 202 trains the speech understanding model 1000 stored in the speech understanding model storage unit 205 using the training dataset stored in the training dataset storage unit 206. At this time, the model training unit 202 trains the training parameters of the speech encoder output integration block 1200, the training parameters of the temporal information integration block 1300, and the training parameters of the linear transformation layer 1400, while keeping the parameters of the speech encoder 1100 and the large-scale language model 1500 fixed. More specifically, the model training unit 202 uses, as a loss function, a cross entropy between an output sentence generated by the speech understanding model 1000 when input speech and input sentences included in the training data are given and a correct output sentence included in the training data, and trains the training parameters by an existing optimization method so as to minimize the loss function. Note that a detailed example of the functional configuration of the model training unit 202 will be described later.
[0053] The trained speech encoder storage unit 203 stores a trained speech encoder. The trained large-scale language model storage unit 204 stores a trained large-scale language model. The speech understanding model storage unit 205 stores the speech understanding model 1000 constructed by the model construction unit 201. The training dataset storage unit 206 stores a given training dataset.
[0054] <Learning Data Set> An example of the learning data set stored in the learning data set storage unit 206 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the learning data set.
[0055] 6, a training data set is composed of one or more training data, and each training data set includes an input speech, an input sentence, and a correct output sentence. Generally, a training data set is composed of a large number of training data.
[0056] The input speech is speech data input to the speech understanding model 1000. The input sentence is text data representing a question in natural language related to the input speech. The correct output sentence is text data representing a response or answer in natural language that is the correct answer to the question represented by the input sentence. Note that the input speech does not necessarily have to be speech data recording a human voice, but may be speech data recording any sound. The correct output sentence may also be called, for example, "teaching data."
[0057] For example, the training data in the first row of the example shown in FIG. 6 includes an input speech "Speech A," an input sentence "Please transcribe this speech," and a correct output sentence "Please explain what's going on." Similarly, the training data in the second row of the example shown in FIG. 6 includes an input speech "Speech A," an input sentence "Please tell me how this speech is spoken," and a correct output sentence "A woman is speaking quickly and loudly." Similarly, the training data in the third row of the example shown in FIG. 6 includes an input speech "Speech B," an input sentence "Please tell me how this speech is spoken," and a correct output sentence "A man speaks slowly and calmly." Similarly, the training data in the fourth row of the example shown in FIG. 6 includes an input speech "Speech B," an input sentence "Please tell me the gender of the speaker of this speech," and a correct output sentence "It's a man." Similarly, the training data in the fifth row of the example shown in Figure 6 includes the input speech "Speech C," the input sentence "What is the emotion of the speaker of this speech?", and the correct output sentence "This speaker is somewhat irritated."
[0058] In this way, the training data set is composed of training data represented by pairs of input speech, input sentences, and correct output sentences. Note that, as shown in Fig. 6, the training data set may contain multiple training data containing different input sentences and correct output sentences for the same input speech.
[0059] <<Example of Detailed Functional Configuration of Model Learning Unit 202>> An example of a detailed functional configuration of the model learning unit 202 will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of a detailed functional configuration of the model learning unit 202.
[0060] As shown in FIG. 7, the model learning unit 202 includes a learning data input unit 211, an audio encoding unit 212, a first integration unit 213, a second integration unit 214, a linear transformation unit 215, a posterior probability calculation unit 216, a parameter update unit 217, and an end determination unit 218.
[0061] The learning data input unit 211 inputs one piece of learning data from the learning data set stored in the learning data set storage unit 206 .
[0062] The speech encoding unit 212 is realized by the speech encoder 1100 included in the speech understanding model 1000. The speech encoding unit 212 receives input speech included in the training data input by the training data input unit 211, and generates N first speech feature vectors h from N integration target layers in each time interval t (t=1, ..., T). n (t) (n=1, . . . , N) are output respectively.
[0063] The first integration unit 213 is realized by the speech encoder output integration block 1200 included in the speech understanding model 1000. The first integration unit 213 integrates the first speech feature vector h n (t) (n=1, . . . , N) as inputs, and each of these first speech feature vectors h n (t) (n=1, . . . , N) to output a first integrated vector e(t).
[0064] The second integration unit 214 is realized by the time information integration block 1300 included in the speech understanding model 1000. The second integration unit 214 receives the first integrated vector e(t) for each time interval t as input, and outputs a second integrated vector v obtained by integrating each of the first integrated vectors e(t) (t=1, ..., T) in the time direction.
[0065] The linear transformation unit 215 is realized by the linear transformation layer 1400 included in the speech understanding model 1000. The linear transformation unit 215 receives the second integrated vector v as input and outputs a second speech feature vector w obtained by linearly transforming the second integrated vector v.
[0066] The posterior probability calculation unit 216 is realized by the large-scale language model 1500 included in the speech understanding model 1000. The posterior probability calculation unit 216 receives an input sentence included in the training data input by the training data input unit 211 and a second speech feature vector w, and calculates the posterior probability of an output sentence when the input sentence and the second speech feature vector w are given.
[0067] More specifically, the i-th token constituting the output sentence is expressed as s i(where s 1 is a token representing the beginning of a sentence.) When the input sentence and the second speech feature vector w are given, the token s 1 The posterior probability of generating p(s 1 ), the input sentence and the second speech feature vectors w and s 1 , ..., s i-1 Given a token s i The posterior probability of generating p(s i ) (where i≧2). In this case, the posterior probability calculation unit 216 calculates the posterior probability p(s i ) is calculated. I is the number of tokens included in the output sentence (i.e., the length of the output sentence), and may be set to, for example, the length of the correct output sentence included in the training data input by the training data input unit 211. Here, the posterior probability p(s i ) is expressed as an M-dimensional vector in which, for example, when the number of dimensions of the token embedding space is M, the probability that the m-th type of token is generated is in the m-th element, and the sum of the values of all elements is 1.
[0068] A token is a basic processing unit when a language model such as a large-scale language model processes a character string. A typical example of a token is a word, but a token is not limited to a word and may be, for example, a character, a morpheme, a subword, a certain coherent character string, etc.
[0069] The parameter update unit 217 uses the posterior probability calculated by the posterior probability calculation unit 216 and the correct output sentence included in the learning data input by the learning data input unit 211 to learn the learning target parameters of the speech understanding model 1000 using an existing optimization method.
[0070] More specifically, the i-th token constituting the correct output sentence is denoted by s i ' (where s 1 ' is a token that indicates the beginning of a sentence, s I ' is a token that indicates the end of a sentence. i The probability that ' is generated is p(s i ') Probability p(s iFor example, when the number of dimensions of the token embedding space is M, the token s i , I. In this case, the parameter update unit 217 calculates -p(s i ') logp(s i ) (i.e., cross entropy) as a loss function, and the training parameters are updated using an existing optimization method so as to minimize the loss function. Note that the optimization method that can be used when updating the training parameters is not limited to a specific method, but for example, an online optimization method based on stochastic gradient descent can be used.
[0071] The termination determination unit 218 determines whether to terminate the update of the training parameters. At this time, the termination determination unit 218 determines to terminate the update of the training parameters if a predetermined termination condition is met, and determines not to terminate the update of the training parameters if a predetermined termination condition is not met. As a result, the training parameters of the audio encoder output integrated block 1200, the training parameters of the temporal information integrated block 1300, and the training parameters of the linear transformation layer 1400 are repeatedly updated until the predetermined termination condition is met. Here, examples of the predetermined termination condition include the training parameters being updated a predetermined number of times or more, the number of epochs being a predetermined number of epochs or more, the value of the loss function being less than a predetermined value, the loss function converging, etc.
[0072] <Model Building Process> The model building process will be described below with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the model building process.
[0073] The model construction unit 201 constructs the speech understanding model 1000 shown in Fig. 1 using the trained speech encoder stored in the trained speech encoder storage unit 203 and the trained large-scale language model stored in the trained large-scale language model storage unit 204 (step S101). At this time, the model construction unit 201 initializes the training parameters of the speech encoder output integration block 1200, the training parameters of the time information integration block 1300, and the training parameters of the linear transformation layer 1400 by any method. In this way, the untrained speech understanding model 1000 is constructed.
[0074] Then, the model construction unit 201 stores the voice understanding model 1000 constructed in the above step S101 in the voice understanding model storage unit 205 (step S102).
[0075] <Model Learning Process> The model learning process will be described below with reference to Fig. 9. Fig. 9 is a flowchart showing an example of the model learning process.
[0076] The training data input unit 211 of the model training unit 202 inputs one piece of training data from the training data set stored in the training data set storage unit 206 (step S201). The training data input unit 211 inputs, for example, one piece of training data that has not yet been input for the current number of epochs from among the training data that make up the training data set. The epoch number starts from 0 and is incremented by 1 each time all the training data that make up the training data set are input.
[0077] The speech encoding unit 212 of the model training unit 202 receives the input speech included in the training data input in step S201 and generates N first speech feature vectors h from the N integration target layers in each time interval t (t=1, . . . , T). n (t) (n=1, . . . , N) are output (step S202).
[0078] The first integration unit 213 of the model learning unit 202 integrates, for each time interval t (t=1, . . . , T), a first speech feature vector h n(t) (n=1, . . . , N) as inputs, and each of these first speech feature vectors h n (t) (n=1, . . . , N) are integrated and a first integrated vector e(t) is output (step S203).
[0079] The second integration unit 214 of the model learning unit 202 receives the first integrated vector e(t) for each time interval t as input, and outputs a second integrated vector v obtained by integrating each of these first integrated vectors e(t) (t = 1, ..., T) in the time direction (step S204).
[0080] The linear transformation unit 215 of the model learning unit 202 receives the second integrated vector v as input, and outputs the second speech feature vector w obtained by linearly transforming the second integrated vector v (step S205).
[0081] The posterior probability calculation unit 216 of the model learning unit 202 receives as input the input sentence included in the learning data input in step S201 above and the second speech feature vector w output in step S205 above, and calculates the posterior probability of the output sentence when the input sentence and the second speech feature vector w are given (step S206).
[0082] The parameter update unit 217 of the model learning unit 202 learns the learning target parameters of the speech understanding model 1000 by an existing optimization method using the posterior probability calculated in the above step S206 and the correct output sentence included in the learning data input in the above step S201 (step S207). That is, the parameter update unit 217 calculates the cross entropy (specifically, -p(s i ') logp(s i ) as a loss function, and the training parameters are updated using an existing optimization method so as to minimize the loss function.
[0083] The termination determination unit 218 of the model learning unit 202 determines whether to terminate the update of the learning parameter (step S208). That is, the termination determination unit 218 determines to terminate the update of the learning parameter if a predetermined termination condition is met, and determines not to terminate the update of the learning parameter if a predetermined termination condition is not met.
[0084] If it is determined in step S208 that the update of the learning target parameters is not to be terminated, the model learning unit 202 returns to step S201, whereby steps S201 to S207 are repeatedly executed until a predetermined termination condition is met.
[0085] On the other hand, if it is determined in step S208 that the updating of the learning parameters is to be terminated, the model training unit 202 terminates the model training process, thereby training the learning parameters and obtaining the trained speech understanding model 1000.
[0086] [During Inference] The following describes the speech understanding device 10 during inference. The following mainly describes differences from during learning, and omits descriptions of points that may be the same as during learning, as appropriate.
[0087] <Example of Hardware Configuration of Speech Understanding Device 10 During Inference> The hardware configuration of the speech understanding device 10 during inference may be the same as that during learning, and therefore a description thereof will be omitted.
[0088] <Example of Functional Configuration of Speech Understanding Device 10 During Inference> An example of the functional configuration of the speech understanding device 10 during inference will be described with reference to Fig. 10. Fig. 10 is a diagram showing an example of the functional configuration of the speech understanding device 10 during inference.
[0089] 10 , the speech understanding device 10 at the time of inference has an output sentence generation unit 207. The output sentence generation unit 207 is realized, for example, by a process in which one or more programs installed in the speech understanding device 10 are executed by the processor 108 or the like. The speech understanding device 10 at the time of inference also has a trained speech understanding model storage unit 208 and a test data storage unit 209. Each of these storage units is realized, for example, by a storage area of the auxiliary storage device 107 or the like. However, at least one of these storage units may be realized by a storage area of a storage device (for example, a storage device included in a database server or the like) communicatively connected to the speech understanding device 10.
[0090] The output sentence generation unit 207 uses the test data stored in the test data storage unit 209 and the trained speech understanding model 1000 stored in the trained speech understanding model storage unit 208 to generate and output an output sentence corresponding to a question expressed by an input sentence included in the test data (i.e., text data representing a natural language response or answer to the question). Here, the test data refers to data represented by a pair of an input speech and an input sentence that is a natural language question related to the input speech. The trained speech understanding model 1000 refers to the speech understanding model 1000 whose learning target parameters have been trained. A detailed example of the functional configuration of the output sentence generation unit 207 will be described later.
[0091] The trained speech understanding model storage unit 208 stores the trained speech understanding model 1000. The test data storage unit 209 stores the given test data.
[0092] <<Detailed Functional Configuration Example of Output Sentence Generation Unit 207>> A detailed functional configuration example of the output sentence generation unit 207 will be described with reference to Fig. 11. Fig. 11 is a diagram showing an example of the detailed functional configuration of the output sentence generation unit 207.
[0093] As shown in FIG. 11, the output sentence generation unit 207 includes a test data input unit 221, a speech encoding unit 222, a first integration unit 223, a second integration unit 224, a linear conversion unit 225, a generation unit 226, and an output unit 227.
[0094] The test data input unit 221 inputs one piece of test data stored in the test data storage unit 209 .
[0095] The speech encoding unit 222 is realized by the speech encoder 1100 included in the trained speech understanding model 1000. The speech encoding unit 222 receives input speech included in the test data input by the test data input unit 221, and generates N first speech feature vectors h from N integration target layers in each time interval t (t=1, ..., T). n (t) (n=1, . . . , N) are output respectively.
[0096] The first integration unit 223 is realized by the speech encoder output integration block 1200 included in the trained speech understanding model 1000. The first integration unit 223 integrates the first speech feature vector h n (t) (n=1, . . . , N) as inputs, and each of these first speech feature vectors h n (t) (n=1, . . . , N) to output a first integrated vector e(t).
[0097] The second integration unit 224 is realized by the time information integration block 1300 included in the trained speech understanding model 1000. The second integration unit 224 receives the first integrated vector e(t) for each time interval t as input, and outputs a second integrated vector v obtained by integrating each of the first integrated vectors e(t) (t=1, ..., T) in the time direction.
[0098] The linear transformation unit 225 is realized by the linear transformation layer 1400 included in the trained speech understanding model 1000. The linear transformation unit 225 receives the second integrated vector v as input, and outputs a second speech feature vector w obtained by linearly transforming the second integrated vector v.
[0099] The generation unit 226 is realized by the large-scale language model 1500 included in the trained speech understanding model 1000. The generation unit 226 receives an input of an input sentence included in the test data input by the test data input unit 221 and a second speech feature vector w, and generates an output sentence when the input sentence and the second speech feature vector w are given.
[0100] More specifically, the i-th token constituting the output sentence is expressed as s i Furthermore, when the input sentence and the second speech feature vector w are given, a token s 1 The posterior probability of generating p(s 1 ), the input sentence and the second speech feature vectors w and s 1 , ..., s i-1 Given a token s i The posterior probability of generating p(s i ) (where i≧2). In this case, the generating unit 226 continues to generate the posterior probability p(s i ) according to the token s i The output sentence is generated by generating
[0101] The output unit 227 outputs the output sentence generated by the generation unit 226 to a predetermined output destination. Here, examples of the predetermined output destination include a storage area such as the auxiliary storage device 107, the display device 102 such as a display, other devices or equipment connected in a communicable manner, etc.
[0102] <Output Sentence Generation Processing> The output sentence generation processing will be described below with reference to Fig. 12. Fig. 12 is a flowchart showing an example of the output sentence generation processing.
[0103] The test data input unit 221 of the output statement generation unit 207 inputs one piece of test data stored in the test data storage unit 209 (step S301).
[0104] The speech encoding unit 222 of the output sentence generation unit 207 receives the input speech included in the test data input in step S301, and generates N first speech feature vectors h from N integration target layers in each time interval t (t=1, . . . , T). n (t) (n=1, . . . , N) are output (step S302).
[0105] The first integration unit 223 of the output sentence generation unit 207 integrates, in each time interval t (t=1, . . . , T), a first speech feature vector h n (t) (n=1, . . . , N) as inputs, and each of these first speech feature vectors h n (t) (n=1, . . . , N) are integrated and a first integrated vector e(t) is output (step S303).
[0106] The second integration unit 224 of the output sentence generation unit 207 receives the first integrated vector e(t) for each time interval t as input, and outputs a second integrated vector v obtained by integrating each of these first integrated vectors e(t) (t = 1, ..., T) in the time direction (step S304).
[0107] The linear transformation unit 225 of the output sentence generation unit 207 receives the second integrated vector v as input, and outputs a second speech feature vector w obtained by linearly transforming the second integrated vector v (step S305).
[0108] The generation unit 226 of the output sentence generation unit 207 receives as input the input sentence included in the test data input in step S301 above and the second speech feature vector w output in step S305 above, and generates an output sentence when the input sentence and the second speech feature vector w are given (step S306).
[0109] The output unit 227 of the output sentence generation unit 207 outputs the output sentence generated in step S306 to a predetermined output destination (step S307), thereby obtaining an output sentence that is a response or answer to the question related to the input voice.
[0110] <Summary> As described above, the speech understanding device 10 according to this embodiment can realize speech understanding technology using the speech understanding model 1000 in which the speech encoder output integration block 1200 and the time information integration block 1300 are present between the speech encoder 1100 and the large-scale language model 1500. For this reason, by using the speech understanding device 10 according to this embodiment, it is possible to expect, for example, improvement in processing accuracy in a downstream system that uses the recognition results of non-linguistic information and paralinguistic information.
[0111] The following supplementary note is further disclosed regarding the above embodiment: (Supplementary Note 1) A learning device including: a memory; and at least one processor connected to the memory, wherein the processor receives training data including speech, a first sentence related to the speech, and a second sentence corresponding to the first sentence; generates information representing features of the speech for each predetermined time interval based on a speech feature extractor configured with multiple layers; generates first integrated information for each time interval based on first parameters, the information representing the features generated in each of the predetermined multiple layers of the speech feature extractor; generates second integrated information for each time interval based on second parameters, the first integrated information being integrated in the time direction; calculates a generation probability of a third sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model; and learns training target parameters including the first parameter and the second parameter based on the generation probability of the third sentence and the second sentence. and at least one processor connected to the memory, wherein the processor: inputs test data including speech and a first sentence related to the speech; generates information representing features of the speech for each predetermined time interval based on a speech feature extractor configured with multiple layers; generates first integrated information by integrating the information representing the features generated in each of the predetermined multiple layers of the speech feature extractor for each time interval based on trained first parameters; generates second integrated information by integrating the first integrated information for each time interval in the time direction based on trained second parameters; and generates a second sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model. (Supplementary Note 3) The learning device according to Supplementary Note 1, wherein the processor: linearly transforms the second integrated information based on third parameters; and calculates a generation probability of the third sentence based on the first sentence, the second integrated information after the linear transformation, and the language model.(Supplementary Note 4) The learning device according to Supplementary Note 1 or 3, wherein the first parameters are weights used in a weighted sum or linear transformation coefficients used in a linear transformation sum, and the processor generates first integrated information by integrating the features using the weighted sum or the linear transformation sum. (Supplementary Note 5) The learning device according to Supplementary Note 1 or 3, wherein the second parameters are weights of a self-attention pooling layer or parameters of a one-dimensional convolutional neural network, and the processor generates second integrated information by integrating the first integrated information in a time direction using the self-attention pooling layer or the one-dimensional convolutional neural network. (Supplementary Note 6) A non-transitory storage medium storing a program executable by a computer to execute a learning process, wherein the learning process includes: inputting learning data including speech, a first sentence related to the speech, and a second sentence corresponding to the first sentence; generating information representing features of the speech for each predetermined time interval based on a speech feature extractor configured with multiple layers; generating first integrated information for each time interval based on first parameters, the first integrated information integrating the information representing the features generated in each of the predetermined multiple layers of the speech feature extractor; generating second integrated information for each time interval based on second parameters, the first integrated information being integrated in the time direction; calculating a generation probability of a third sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model; and learning learning target parameters including the first parameter and the second parameter based on the generation probability of the third sentence and the second sentence.(Supplementary Note 7) A non-transitory storage medium storing a program executable by a computer to execute an inference process, wherein the inference process inputs test data including speech and a first sentence related to the speech, generates information representing features of the speech for each predetermined time interval based on a speech feature extractor composed of multiple layers, generates first integrated information for each time interval based on trained first parameters by integrating the information representing the features generated in each of the multiple predetermined layers of the speech feature extractor, generates second integrated information for each time interval by integrating the first integrated information in the time direction based on trained second parameters, and generates a second sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model.
[0112] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0113] [References] Reference 1: Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, Michael Auli, "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations," arXiv preprint arXiv:2006.11477, 2020. Reference 2: S. Chen et al., "WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing," in IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505-1518, Oct. 2022, doi: 10.1109 / JSTSP.2022.3188113. Reference 3: H. Touvron et al., "Llama 2: Open Foundation and Fine-Tuned Chat Models," arXiv preprint arXiv:2307.09288, 2023.
[0114] 10 Speech understanding device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Model construction unit 202 Model learning unit 203 Trained speech encoder storage unit 204 Trained large-scale language model storage unit 205 Speech understanding model storage unit 206 Training dataset storage unit 207 Output sentence generation unit 208 Trained speech understanding model storage unit 209 Test data storage unit 211 Training data input unit 212 Speech encoding unit 213 First integration unit 214 Second integration unit 215 Linear transformation unit 216 Posterior probability calculation unit 217 Parameter update unit 218 End determination unit 221 Test data input unit 222 Audio encoding unit 223 First integration unit 224 Second integration unit 225 Linear conversion unit 226 Generation unit 227 Output unit
Claims
1. An input unit that inputs learning data including voice, a first sentence related to the voice, and a second sentence corresponding to the first sentence; a voice feature generation unit that generates information representing features of the voice for each predetermined time interval based on a voice feature extractor composed of a plurality of layers; a first integration unit that generates first integrated information obtained by integrating information representing the features respectively generated in a predetermined plurality of layers of the voice feature extractor for each time interval based on a first parameter; a second integration unit that generates second integrated information obtained by integrating the first integrated information for each time interval in the time direction based on a second parameter; a calculation unit that calculates a generation probability of a third sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model; and a learning unit that learns learning target parameters including the first parameter and the second parameter based on the generation probability of the third sentence and the second sentence. A learning device having the above components.
2. An input unit that inputs test data including voice and a first sentence related to the voice; a voice feature generation unit that generates information representing features of the voice for each predetermined time interval based on a voice feature extractor composed of a plurality of layers; a first integration unit that generates first integrated information obtained by integrating information representing the features respectively generated in a predetermined plurality of layers of the voice feature extractor for each time interval based on a learned first parameter; a second integration unit that generates second integrated information obtained by integrating the first integrated information for each time interval in the time direction based on a learned second parameter; and a sentence generation unit that generates a second sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model. An inference device having the above components.
3. A linear transformation unit that linearly transforms the second integrated information based on a third parameter, and the calculation unit calculates the generation probability of the third sentence based on the first sentence, the second integrated information after linear transformation, and the language model. The learning device according to claim 1.
4. The first parameter is a weight used for a weighted sum or a linear transformation coefficient used for a linear transformation sum, and the first integration unit generates first integrated information obtained by integrating the features by the weighted sum or the linear transformation sum. The learning device according to claim 1 or 3.
5. The second parameter is a weight of the self-attention pooling layer or a parameter of the one-dimensional convolutional neural network, and the second integration unit generates second integrated information obtained by integrating the first integrated information in the time direction by the self-attention pooling layer or the one-dimensional convolutional neural network. The learning device according to claim 1 or 3.
6. An input procedure for inputting learning data including audio, a first sentence related to the audio, and a second sentence corresponding to the first sentence; An audio feature generation procedure for generating information representing features of the audio for each predetermined time interval based on an audio feature extractor composed of a plurality of layers; A first integration procedure for generating first integrated information obtained by integrating information representing the features respectively generated in a plurality of predetermined layers of the audio feature extractor for each of the time intervals based on a first parameter; A second integration procedure for generating second integrated information obtained by integrating the first integrated information for each of the time intervals in the time direction based on a second parameter; A calculation procedure for calculating a generation probability of a third sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model; A learning procedure for learning learning target parameters including the first parameter and the second parameter based on the generation probability of the third sentence and the second sentence. A learning method executed by a computer.
7. An input procedure for inputting test data including audio and a first sentence related to the audio; An audio feature generation procedure for generating information representing features of the audio for each predetermined time interval based on an audio feature extractor composed of a plurality of layers; A first integration procedure for generating first integrated information obtained by integrating information representing the features respectively generated in a plurality of predetermined layers of the audio feature extractor for each of the time intervals based on a learned first parameter; A second integration procedure for generating second integrated information obtained by integrating the first integrated information for each of the time intervals in the time direction based on a learned second parameter; A sentence generation procedure for generating a second sentence corresponding to the first sentence based on the first sentence, the second integrated information, and a language model. An inference method executed by a computer.
8. A program for causing a computer to function as the learning device according to claim 1 or the inference device according to claim 2.
Citation Information
Patent Citations
Machine learning device, machine learning method, machine learning program and inference device
JP2023117248A
Recognition device, learning device, method for same, and program
WO2021166207A1