Speech understanding device, speech understanding method, and program
The speech understanding device addresses the challenge of generating accurate speech information by training with correct answer data and speech factor phrases, improving output sentence accuracy and enhancing applications in speech understanding systems.
Patent Information
- Application Number
- PCT/JP2024/021042
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-10
- Publication Date
- 2025-12-18
AI Technical Summary
Conventional speech understanding technologies struggle to generate output sentences with accurate speech information due to training methods that treat syntax and speech information errors similarly, leading to a high probability of generating incorrect speech information.
A speech understanding device that uses correct answer data including speech factor phrases and delimiters to train a deep learning model, performing Chain-of-Thought prompting to explicitly learn and generate output sentences with accurate speech information.
The device generates output sentences with improved accuracy in speech information, enhancing applications like call analysis, customer voice extraction, and human-like dialogue systems by reducing estimation errors in speech factors.
Smart Images

Figure JP2024021042_18122025_PF_FP_ABST
Abstract
Description
Speech understanding device, speech understanding method, and program
[0001] The present invention relates to a speech understanding device, a speech understanding method, and a program.
[0002] Speech is considered to include three types of information: linguistic information, non-linguistic information, and paralinguistic information (see, for example, Non-Patent Document 1). In this specification, these three types of information are collectively referred to as speech information.
[0003] In recent years, a speech understanding technology has been proposed that infers speech information in a free-form format (see Non-Patent Document 2).
[0004] H. Fujisaki, "Prosody, Models, and Spontaneous Speech", in Computing Prosody, Y. Sagisaka, N. Campbell, and N. Higuchi, Springer, pp.27-42, 1996.M.Wang,W. Han, I. Shafran, Z.Wu, C.-C. Chiu, Y. Cao, N. Chen, Y. Zhang, H. Soltau, PK Rubenstein, L. Zilka, D. Yu, G. Pundak, N. Siddhartha, J. Schalkwyk, and Y. Wu, "SLM: Bridge the thin gap between speech and text foundation models", in Proc. ASRU, 2023.
[0005] In speech understanding technology, it is important to generate output sentences that contain correct speech information. However, conventional technology trains to generate correct output sentences, which can increase the possibility of generating output sentences that contain incorrect speech information.
[0006] An embodiment of the present invention has been made in consideration of the above-mentioned problems, and enables a speech understanding device that understands speech using input speech, input sentences, and a speech understanding model to generate output sentences that include more accurate speech information.
[0007] In order to solve the above problems, a speech understanding device according to an embodiment of the present invention is a speech understanding device that understands speech using input speech, an input sentence, and a speech understanding model, and includes a correct answer data generation unit that generates correct answer data including speech factor phrases, delimiters, and correct answer output sentences that represent speech information, and a learning unit that learns the speech understanding model using learning data including the input speech, the input sentence, and the correct answer data.
[0008] According to an embodiment of the present invention, a speech understanding device that understands speech using input speech, an input sentence, and a speech understanding model can generate an output sentence that includes more accurate speech information.
[0009] FIG. 1 is a diagram for explaining an example of the configuration of a speech understanding model. FIG. 2 is a diagram for explaining correct answer data according to the present embodiment. FIG. 3 is a diagram for explaining a method for generating speech factor phrases according to the present embodiment. FIG. 4 is a diagram for explaining an example of a prompt sentence for speech factor estimation according to the present embodiment. FIG. 5 is a diagram for explaining an example of a configuration of a speech understanding device according to Example 1. FIG. 6 is a flowchart for explaining an example of a process during learning according to Example 1. FIG. 7 is a flowchart for explaining an example of a process during inference according to Example 1. FIG. 8 is a diagram for explaining an example of an output sentence generation process according to Example 2. FIG. 9 is a flowchart for explaining an example of a process during inference according to Example 2. FIG. 10 is a diagram for explaining an example of a hardware configuration of a computer.
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.
[0011] <Overview> (Background) Speech is considered to include three types of information: linguistic information, non-linguistic information, and paralinguistic information (see, for example, Non-Patent Document 1). Linguistic information is information about spoken words, non-linguistic information is information that is not linguistic but cannot be changed at will (e.g., speaker identity, gender, emotions, etc.), and paralinguistic information is information that is not linguistic but can be changed at will (e.g., intention, attitude, etc.). In this embodiment, these three types of information are collectively referred to as speech information.
[0012] Understanding speech information is important in, for example, call analysis and customer voice extraction in contact centers, dialogue control for human-like voice dialogue systems, and human-controllable speech synthesis.
[0013] In recent years, a speech understanding technology has been proposed that infers speech information from free-form descriptions (see, for example, Non-Patent Document 2). This technology, as shown in Fig. 1, is configured with a deep learning-based speech understanding model 10 that combines a large-scale language model decoder 11 with a speech encoder 12 that acquires sound information and associations for each subsection using a large amount of speech data. The large-scale language model decoder 11 is a trained large language model (LLM) that acquires inter-word relationships and co-occurrences using a large amount of text data.
[0014] When an input speech and an input sentence representing a question are given to this speech understanding model 10, the speech understanding model 10 can generate an output sentence corresponding to the question. For example, when the input sentence "Please tell me the speaking style of the speaker included in this speech" is given to the speech understanding model 10, the speech understanding model 10 outputs the output sentence "The man is speaking in a rather fast tone, at a loud volume and in a low voice."
[0015] This speech understanding model 10 is obtained by optimizing the entire model using, for example, a set of input speech, input sentence, and corresponding correct output sentence. Furthermore, by providing various sets of input speech, input sentence, and correct output sentence, it becomes possible to estimate various pieces of information contained in speech in natural language. For example, if the input sentence "Please transcribe this speech" is provided for the previous example, the sentence "The transcription result of the speech is 'It might rain today'" can be output.
[0016] (Problem) In speech understanding technology, it is important to generate output sentences containing correct speech information, but this can be difficult with conventional technologies. This is because conventional technologies are trained to generate correct output sentences. A correct output sentence contains both speech information-related words and syntax-related words. For example, in the sentence "A man is speaking at a high volume and in a low voice, with a somewhat rapid tone," "male," "speaks at a high volume," and "low voice" correspond to speech information-related words, while the rest correspond to syntax-related words. When a speech understanding model is trained to generate an output sentence as close as possible to this correct output sentence, both an "output sentence with correct syntax but a single incorrect speech information" and an "output sentence with a single incorrect syntax but correct speech information" are treated as the same single-word error. In other words, conventional training methods have difficulty explicitly correcting only errors in speech information, resulting in a high probability of generating output sentences containing incorrect speech information.
[0017] (Outline of this embodiment) Therefore, the speech understanding device according to this embodiment learns a speech understanding model using correct answer data 200, which is a correct answer output sentence with speech factor phrases including speech factor phrases, delimiters, and the original correct answer output sentence, as shown in FIG. 2, instead of the original correct answer output sentence.
[0018] Here, a speech factor phrase is a short phrase that expresses only speech information. A speech factor phrase is a phrase that has a fixed format and includes words that represent speech factors (categories of each speech information). For example, speech factors for gender are "male" and "female," and speech factors for volume are "loud," "normal," and "low," respectively. Speech factors for pitch are "high pitch," "normal pitch," and "low pitch," and speech factors for speaking speed are "fast speaking speed," "normal speaking speed," and "slow speaking speed." A speech factor phrase is a phrase in which speech factors are connected with periods, such as "male, loud, low pitch, fast speaking speed."
[0019] 3 is a diagram illustrating a method for generating speech factor phrases according to this embodiment. Speech factor phrases can be generated by, for example, estimating speech factors from an original correct output sentence using a large-scale language model such as a Generative Pretrained Transformer (GPT) (Reference 1) (step S1) and substituting the speech factors into a template sentence (step S2). Alternatively, speech factor phrases can be generated by, for example, manually labeling speech factors (step S3) and substituting the speech factors into a template sentence (step S2).
[0020] An example of a prompt sentence 400 for estimating speech factors using the GPT is shown in Fig. 4. For example, the original correct output sentence is substituted for [correct output sentence] in the prompt sentence 400 shown in Fig. 4, and the resulting sentence is input to the GPT. As a result, estimation results of speech factors such as "gender: male, voice volume: loud, voice pitch: low, speaking rate: fast" can be obtained, as shown in Fig. 3.
[0021] By substituting this estimation result into a template sentence of a speech factor phrase such as "gender," "voice volume," "voice pitch," and "speaking rate," speech factor phrases can be generated.
[0022] By training the deep learning model using the correct answer data 200 (correct output sentence with speech factor phrases) instead of the original correct sentence, the deep learning model first estimates the speech factor phrases and then estimates the original correct output sentence. This is because the large-scale language model decoder of the deep learning model is an autoregressive model (a model that predicts the next word based on past words), and so the original correct output sentence is estimated taking into account the speech factor phrase estimated immediately before.
[0023] In other words, after explicitly estimating only the speech information as an intermediate step, the output sentence (including words related to the speech information and words related to the syntax) is estimated based on the speech information. Because there is a step of estimating only the speech information, it is possible to explicitly learn only the speech information, and it is expected that estimation errors of the speech information will be reduced.
[0024] This embodiment can be considered to perform complex inference (inference of speech information and syntax) via intermediate inference (inference of speech information only), and this inference method is called Chain-of-Thought prompting (Reference 2). This Chain-of-Thought prompting has been confirmed to be effective in the field of natural language processing. This embodiment can also be said to train a speech understanding model that always performs Chain-of-Thought for speech understanding.
[0025] <Specific Description of the Present Embodiment> As the speech understanding model of the present embodiment, for example, a speech understanding model 10 configured with a large-scale language model decoder 11, a speech encoder 12, and a bridge network 13 as shown in Fig. 1 can be used. In this case, the large-scale language model decoder 11 uses an autoregressive model. Furthermore, the large-scale language model decoder 11 and the speech encoder 12 are assumed to have been individually trained in advance. The other blocks are initialized with random numbers.
[0026] (During learning) 1) The speech understanding device generates correct answer data 200, which is a correct answer output sentence with a speech factor phrase. When generating from an original correct answer output sentence, the speech understanding device estimates speech factors for the original correct answer output sentence using the GPT and a prompt sentence 400 to be explained in FIG. 4. If speech factors corresponding to the correct answer output sentence are manually provided, the speech understanding device uses the manually provided speech factors as is. Furthermore, the speech understanding device combines speech factors with a template sentence to generate a speech factor phrase. Note that the template sentence is assumed to have been manually created in advance, for example. Furthermore, the speech understanding device inserts a delimiter between the speech factor phrase and the original correct answer output sentence to generate the correct answer data 200.
[0027] 2) The speech understanding device uses a set of input speech, input sentence, and correct answer data 200 (correct output sentence with speech factor phrases) as a training dataset and trains the speech understanding model (updates model parameters). At this time, the parameters of the large-scale language model decoder 11 and the speech encoder 12 are fixed, and only the parameters of the bridge network 13 are updated. However, an adapter (Reference 3), which is a deep learning block with trainable parameters that is inserted into part of the decoder, may be inserted into the large-scale language model decoder 11, and parameters of the adapter may be updated in addition to the bridge network 13. The parameter update is performed, for example, by online optimization based on stochastic gradient descent. Furthermore, the cross entropy of words for the correct output sentence with speech factor phrases is used as the loss function.
[0028] (During speech understanding) An input speech and an input sentence are input to a trained speech understanding model and forward propagated to obtain an output sentence corresponding to the input sentence. Since the trained speech understanding model is trained with the correct answer data 200 (correct output sentence with speech factor phrase), the sentence is generated in the order of speech factor phrase, delimiter, and output sentence.
[0029] Next, an example of this embodiment will be described.
[0030] [Example 1] Fig. 5 is a diagram illustrating an example of the configuration of a speech understanding device according to Example 1. The speech understanding device 500 is, for example, an information processing device having a computer configuration, or a system including multiple computers. The speech understanding device 500 realizes each of the functional components, such as a speech factor estimation unit 501, a correct answer data generation unit 502, a learning unit 503, and an output sentence generation unit 504, by, for example, executing a predetermined program on the computer included in the speech understanding device 500. Note that at least a part of the above functional components may be realized by hardware.
[0031] The speech factor estimation unit 501 executes a speech factor estimation process to estimate speech factors based on the correct output sentence 513 of the original training data 510. For example, the speech factor estimation unit 501 substitutes the correct output sentence 513 into the prompt sentence 400 described in FIG. 4 , inputs the result into the GPT, and outputs the return value as the speech factor estimation result 505. Specifically, the speech factor estimation unit 501 specifies a GPT model (gpt-4-0613) via the OpenAI API (https: / / openai.com / blog / openai-api), substitutes the correct output sentence 513 into the prompt sentence 400, inputs the result into the API, and sets the return value as the speech factor estimation result. Note that the prompt sentence is not limited to the prompt sentence 400 shown in FIG. 4 . For example, a different prompt sentence that prompts the user to select a different speech factor (e.g., emotion, speaker age, voice quality, etc.) may be used. Furthermore, a language model other than the GPT may be used to estimate the speech factors included in the correct sentence.
[0032] The correct data generation unit 502 executes a correct data generation process to generate correct data 200 (correct output sentence with speech factor phrases) including speech factor phrases expressing speech information, delimiters, and correct output sentences. For example, the correct data generation unit 502 substitutes the speech factor estimation results 505 into template sentences to generate speech factor phrases, and generates the correct data 200, which is a correct output sentence with speech factor phrases.
[0033] The template sentence may be, for example, a template sentence of a speech factor phrase such as "[gender], [voice volume], [voice pitch], and [speaking rate]" as described above. The correct data generation unit 502 creates a speech factor phrase by substituting words from the speech factor estimation result 505 into this template sentence. Thereafter, the correct data generation unit 502 combines phrases so that the speech factor phrase, delimiter, and correct output sentence are arranged as shown in FIG. 2, for example, to create correct data 200 (correct output sentence with speech factor phrase). The delimiter may be, for example, "<|cap|>", and this character string is treated as one word. However, other delimiters may be used.
[0034] The learning unit 503 executes a learning process to learn a speech understanding model using learning data including an input speech 511, an input sentence 512, and the correct answer data 200.
[0035] For example, the learning unit 503 initializes the speech understanding model, updates the parameters of the speech understanding model using learning data consisting of a set of input speech, input sentence, and correct answer data 200 (correct answer output sentence with speech factor phrase), and learns the speech understanding model.
[0036] In this embodiment, initialization of a speech understanding model refers to, for example, constructing a speech understanding model 10 as shown in FIG. 1. In this case, a trained speech encoder 520 is used as the speech encoder 12 in FIG. 1, and a trained large-scale language model 530 is used as the large-scale language model decoder 11. However, as described above, an adapter (Reference 3), which is a deep learning block with trainable parameters that is inserted into part of the decoder, may be inserted, and the parameters of the adapter may be updated in addition to the bridge network 13. Furthermore, for example, a Time and Layer-wise Transformer (Reference 4) may be used for the bridge network 13. However, other bridge networks, such as a bridge network that uses the final layer of the speech encoder as input and is composed of a fully connected layer and average pooling, may also be used. Note that the parameters of the bridge network 13 are initialized using random numbers.
[0037] In this embodiment, model training refers to an operation of updating model parameters based on the stochastic gradient descent method, performing a predetermined number of parameter updates (e.g., 100,000 times) to obtain final model parameters. The loss function required for updating the model parameters is the cross entropy between the correct answer data 200 (correct answer output sentences with speech factor phrases) and the output sentences of the speech understanding model. Furthermore, when updating the parameters, only the parameters of the large-scale language model decoder adapter and bridge network are updated, and all other parameters are not updated. The speech understanding model having parameters obtained by these operations is referred to as the trained speech understanding model 540.
[0038] The output sentence generation unit 504 executes an output sentence generation process to generate an output sentence corresponding to the input speech and the input sentence using the input speech, the input sentence, and the trained speech understanding model. Here, the input speech and the input sentence are the input speech and the input sentence that are the targets of speech understanding. The output sentence is a sentence that indicates the result of speech understanding.
[0039] For example, the output sentence generation unit 504 estimates an output sentence corresponding to the input speech and the input sentence by inputting the input speech and the input sentence into the trained speech understanding model 540 and performing forward propagation. Since the trained speech understanding model 540 is trained with the correct answer data 200, which is a correct output sentence with speech factor phrases, the output of the trained speech understanding model is in the order of speech factor phrases, delimiters, and estimated output sentence. When only the output sentence is needed (when speech factor phrases are not needed), the output sentence generation unit 504 treats only the sentence after the delimiters as the final output. To estimate the output sentence, for example, a greedy search is used. That is, the output sentence generation unit 504 estimates the occurrence probability of the next word based on the word string contained in the output sentence up to now, and considers the word with the highest probability to be the next word. The output sentence generation unit 504 repeats this process and sets the word string as the final output sentence.
[0040] <Processing Flow> Next, the processing flow of the speech understanding method according to the first embodiment will be described.
[0041] (During Learning) Fig. 6 is a flowchart showing an example of processing during learning according to Example 1. This processing shows an example of processing during learning executed by the speech understanding device 500 according to this embodiment.
[0042] In step S601, the speech factor estimation unit 501 estimates speech factors based on the correct output sentence 513 included in the original training data 510. For example, the speech factor estimation unit 501 substitutes the correct output sentence 513 into the prompt sentence 400 described in FIG. 4 , inputs the result into a large-scale master ledger model such as GPT, and outputs the return value as the speech factor estimation result 505.
[0043] In step S602, the correct data generation unit 502 generates correct data 200 (correct output sentence with speech factor phrase) including speech factor phrases, delimiters, and correct output sentences. For example, the correct data generation unit 502 generates speech factor phrases by substituting the speech factor estimation results 505 into a template sentence, and adds delimiters and the original correct output sentence to generate correct data 200 as shown in FIG. 2.
[0044] In step S603, the learning unit 503 uses the learning data including the input speech 511, the input sentence 512, and the correct answer data 200 to learn the speech understanding model.
[0045] For example, the speech understanding device 100 repeatedly executes the process of FIG. 6 using the original training data 510 and updates the parameters of the speech understanding model a predetermined number of times, thereby obtaining a trained speech understanding model 540.
[0046] (Inference Time) Fig. 7 is a flowchart showing an example of processing at the time of inference according to the embodiment 1. This processing shows an example of processing at the time of inference executed by the speech understanding device 500 according to the embodiment 1 described with reference to Fig. 5 .
[0047] In step S701, an input speech and an input sentence to be understood are input to the speech understanding device 500.
[0048] In step S702, the output sentence generation unit 504 inputs the input speech and the input sentence into the trained speech understanding model 540, and generates (estimates) an output sentence corresponding to the input speech and the input sentence.
[0049] In step S703, the output sentence generation unit 504 outputs an output sentence corresponding to the input speech and the input sentence.
[0050] 6 and 7, the speech understanding device 500 that understands speech using input speech, input sentence, and a speech understanding model can generate an output sentence that includes more accurate speech information.
[0051] [Example 2] When generating an output sentence, the sentence can be generated by next word prediction based on sampling search in order to use a variety of phrases and expressions like a human. Sampling search is a method for predicting the next word by randomly selecting the next word based on the occurrence probability of the next word, rather than the word with the highest probability (Reference 4). Note that sampling search is also called, for example, random search or random sampling search.
[0052] However, when performing a sampling search, there is a problem in that there is an increased possibility that a speech factor with a low probability (i.e., incorrect for the input speech) will be output.
[0053] In order to solve the above problems, a speech understanding device 500 according to a second embodiment has, for example, the functional configuration shown in Fig. 8. The speech understanding device 500 according to the second embodiment shown in Fig. 8 has an output sentence generation unit 801 (hereinafter referred to as the output sentence generation unit 801) that uses a combination of a greedy search and a sampling search, instead of the output sentence generation unit 504 described in Fig. 5.
[0054] The output sentence generation unit 801 is realized, for example, by a program executed on a computer included in the speech understanding device 500. The output sentence generation unit 801 searches for words to be used in the output sentence using a combination of a greedy search and a sampling search.
[0055] For example, the output sentence generation unit 801 estimates an output sentence corresponding to the input speech and the input sentence by inputting the input speech and the input sentence to the trained speech understanding model 540 and propagating them forward.
[0056] 9, the output sentence generation unit 801 predicts the next word by greedy search in a portion (speech factor phrase) 901 of an output sentence 900 up until the appearance of a delimiter, and predicts the next word by sampling search in a portion 902 after the appearance of the delimiter. Also, when only the output sentence is needed (when the speech factor phrase is not needed), the output sentence generation unit 801 treats only the sentence after the delimiter as the final output.
[0057] In Example 2, a greedy search is performed on the speech factor phrases, so that speech factors with low probability are not output. Furthermore, since the trained speech understanding model 540 is trained using the correct answer data 200, which are correct output sentences with speech factor phrases, output sentences are generated taking into account the speech factor phrases. Therefore, the output sentence generation unit 801 according to Example 2 can generate output sentences with various wordings while taking into account information on speech factors (= speech factors with the highest probability) obtained by the greedy search. In other words, the output sentence generation unit 801 can generate output sentences with various wordings without outputting incorrect speech factors compared to Example 1.
[0058] Among the functional components of the speech understanding device 500 according to the second embodiment, the functional components other than the output sentence generating unit 801 may be the same as the functional components of the speech understanding device 500 according to the first embodiment.
[0059] <Processing Flow> The learning process according to the second embodiment may be similar to the learning process according to the first embodiment described with reference to FIG.
[0060] (Processing at the Time of Inference) Fig. 10 is a flowchart showing an example of processing at the time of inference according to Example 2. This processing shows an example of processing at the time of inference executed by the speech understanding device 500 according to Example 2 described with reference to Fig. 8 .
[0061] In step S1001 , an input speech and an input sentence to be understood are input to the speech understanding device 500 .
[0062] In step S1002, the output sentence generation unit 801 inputs the input speech and the input sentence into the trained speech understanding model 540, and generates (estimates) an output sentence corresponding to the input speech and the input sentence using a combination of greedy search and sampling search.
[0063] In step S1003, the output sentence generation unit 504 outputs an output sentence corresponding to the input speech and the input sentence.
[0064] By the process of FIG. 7, the speech understanding device 500 can generate output sentences with a variety of phrases without outputting incorrect speech factors.
[0065] 5 and 8 are merely examples. For example, the speech factor estimation unit 501, the correct answer data generation unit 502, the learning unit 503, the output sentence generation unit 504, 801, etc. may be distributed across multiple servers, etc. Furthermore, the speech understanding device 500 may include the trained speech understanding model 540, etc., within the device.
[0066] <Hardware Configuration> The speech understanding device 500 according to this embodiment has, for example, the hardware configuration of a computer 1100 as shown in Fig. 11. Alternatively, the speech understanding device 500 is realized by a plurality of computers 1100. Note that the computer is not limited to a physical machine, and may be, for example, a virtual machine on a cloud.
[0067] Fig. 11 is a diagram showing an example of the hardware configuration of a computer. In the example of Fig. 11, a computer 1100 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, and an output device 1008, all of which are interconnected via a bus B. The computer 1100 may further include a GPU (Graphics Processing Unit) or the like.
[0068] A program for implementing processing on the computer 1100 is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.
[0069] The memory device 1003 reads and stores the program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the speech understanding device 500 in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a communication network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, a mouse, buttons, and / or a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.
[0070] The CPU 1004 may be another processor such as a DSP (Digital Signal Processor), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0071] <Effects of the embodiment> According to the present embodiment, the speech understanding device 500 that understands speech using an input speech, an input sentence, and a speech understanding model can generate an output sentence that includes correct speech information.
[0072] Furthermore, according to this embodiment, it is possible to train a speech understanding model that can easily generate output sentences containing correct speech information. As a result, the accuracy of speech understanding is improved, and the accuracy of downstream systems, such as call analysis and customer voice extraction in contact centers, dialogue control for human-like voice dialogue systems, and speech synthesis that can be controlled by humans, is improved.
[0073] Summary of Embodiments This specification discloses at least the following speech understanding devices, speech understanding methods, and programs. (Item 1) A speech understanding device that understands speech using input speech, an input sentence, and a speech understanding model, comprising: a correct data generation unit that generates correct data including speech factor phrases, delimiters, and a correct output sentence that express speech information; and a learning unit that learns the speech understanding model using learning data including the input speech, the input sentence, and the correct data. (Item 2) The speech understanding device according to Item 1, wherein the speech information includes information on gender, volume, pitch, and / or speaking rate, and the speech factor phrases include one or more speech factors that express the speech information. (Item 3) The speech understanding device according to Item 2, comprising a speech factor estimation unit that estimates the speech factors based on the correct output sentence. (Item 4) The speech understanding device according to Item 3, wherein the speech factor estimation unit estimates the speech factors using a large-scale language model. (Clause 5) A speech understanding device according to any one of clauses 1 to 4, comprising an output sentence generation unit that generates an output sentence corresponding to the input speech and the input sentence using input speech, an input sentence, and the trained speech understanding model. (Clause 6) A speech understanding device according to clause 5, wherein the output sentence generation unit searches for words to be used in the output sentence using a combination of greedy search and sampling search. (Clause 7) A speech understanding method, in which a computer that understands speech using input speech, an input sentence, and a speech understanding model executes the following processes: generating correct answer data including speech factor phrases that express speech information, delimiters, and a correct output sentence; and training the speech understanding model using training data including the input speech, the input sentence, and the correct answer data. (Clause 8) A program, or a storage medium storing a program, that causes a computer to execute the speech understanding method according to clause 7.
[0074] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0075] (References) Reference 1: OpenAI, "GPT-4 Technical Report", 2023, arXiv:2303.08774. Reference 2: J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. ichter, F. Xia, E. Chi, QV Le, and D. Zhou, "Chain-of-thought prompting elicits reasoning in large language models," in Advances in NeurIPS, vol. 35, 2022, pp. 24 824-24 837. Reference 3: EJ Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L.Wang, and W. Chen, "LoRA: Low-rank adaptation of large language models," in Proc. ICLR, 2022. Reference 4: A. Holtz, J. Buys, L. Du, M. Forbes, and Y. Choi, "The curious case of neural text degeneration," in Proc. ICLR, 2020.
[0076] 200 Correct answer data 500 Speech understanding device 501 Speech factor estimation unit 502 Correct answer data generation unit 503 Learning unit 504 Output sentence generation unit 530 Large-scale language model 801 Output sentence generation unit (output sentence generation unit using a combination of greedy search and sampling search) 1100 Computer
Claims
1. A speech understanding device that understands speech using input speech, an input sentence, and a speech understanding model, comprising: a correct answer data generation unit that generates correct answer data including speech factor phrases, delimiters, and correct answer output sentences that represent speech information; and a learning unit that learns the speech understanding model using learning data including the input speech, the input sentence, and the correct answer data.
2. The speech understanding device according to claim 1, wherein the speech information includes information on gender, volume, pitch, and / or speaking rate, and the speech factor phrases include one or more speech factors that express the speech information.
3. The speech understanding device according to claim 2, further comprising a speech factor estimation unit that estimates the speech factors based on the correct output sentence.
4. The speech understanding device according to claim 3, wherein the speech factor estimation unit estimates the speech factors using a large-scale language model.
5. A speech understanding device according to any one of claims 1 to 4, comprising an output sentence generation unit that generates an output sentence corresponding to the input speech and the input sentence using an input speech, an input sentence, and the trained speech understanding model.
6. The speech understanding device according to claim 5, wherein the output sentence generation unit searches for words to be used in the output sentence using a combination of greedy search and sampling search.
7. A speech understanding method in which a computer that understands speech using input speech, an input sentence, and a speech understanding model performs the following processes: generating correct answer data including speech factor phrases, delimiters, and correct output sentences that represent speech information; and training the speech understanding model using training data including the input speech, the input sentence, and the correct answer data.
8. A program for causing a computer to execute the speech understanding method according to claim 7.
Citation Information
Patent Citations
Program, information processing device, and information processing method
JP2023135405A
Image processing device, image processing method, and program
WO2023084833A1