Voice recognition apparatus and voice recognition method

The voice recognition system addresses the need for multiple microphones by using two models to process voice data in pronunciation sections, enhancing accuracy through integrated results based on content consistency and grammatical analysis.

WO2026069580A1PCT designated stage Publication Date: 2026-04-02NTT DOCOMO INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Conventional voice recognition technologies require multiple microphones or different microphone settings to achieve accurate results, limiting their applicability and increasing equipment needs.

Method used

A voice recognition system that utilizes two voice recognition models with different characteristics to process voice data divided into pronunciation sections, integrating results based on content consistency, grammatical accuracy, and registered words using a large-scale language model.

Benefits of technology

Achieves highly accurate voice recognition with a simple configuration by combining results from multiple models, reducing equipment requirements and improving recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024034674_02042026_PF_FP_ABST
    Figure JP2024034674_02042026_PF_FP_ABST
Patent Text Reader

Abstract

This first acquisition unit acquires a first voice recognition result obtained by inputting voice data representing an utterance divided into a plurality of pronunciation sections to a first voice recognition model, and a second voice recognition result obtained by inputting the voice data to a second voice recognition model. A second acquisition unit acquires an integrated voice recognition result created on the basis of the first voice recognition result and the second voice recognition result. An output unit outputs the integrated voice recognition result. The integrated voice recognition result is created on the basis of the consistency of the content of the utterance over the plurality of pronunciation sections.
Need to check novelty before this filing date? Find Prior Art

Description

Voice Recognition Device and Voice Recognition Method

[0001] The present disclosure relates to a voice recognition device and a voice recognition method.

[0002] Conventionally, voice recognition technology that outputs text indicating the content of voice based on voice data obtained by recording voice uttered by a user has become widespread, and efforts have been made to improve the accuracy of voice recognition. For example, the voice recognition device described in Patent Document 1 below includes a voice recognition unit that performs voice recognition on each of a plurality of voice data obtained by inputting the speaker's uttered voice under different recording conditions, and a recognition result selection unit that compares a plurality of voice recognition results obtained by voice recognition in the voice recognition unit and selects the optimal one. As a result, in relation to the surrounding situation in which the utterance was made, a voice recognition result based on voice data recorded under the recording conditions most suitable for voice recognition is selected.

[0003] International Publication No. 2011 / 121978

[0004] In the above-described conventional technology, it is necessary to record voice data under a plurality of recording conditions. To record voice data under a plurality of recording conditions, for example, a large amount of equipment is required such as preparing a plurality of microphones of different types or preparing the same type of microphone with different settings. Therefore, the conventional technology has a problem that the applicable range is limited.

[0005] An object of the present disclosure is to obtain a highly accurate voice recognition result with a simple configuration.

[0006] The voice recognition device according to the present disclosure includes a first acquisition unit that acquires a first voice recognition result obtained by inputting voice data representing an utterance divided into a plurality of pronunciation sections to a first voice recognition model and a second voice recognition result obtained by inputting the voice data to a second voice recognition model, a second acquisition unit that acquires an integrated voice recognition result created based on the first voice recognition result and the second voice recognition result, and an output unit that outputs the integrated voice recognition result, and the integrated voice recognition result is created based on the consistency of the content of the utterance over the plurality of pronunciation sections.

[0007] The speech recognition method according to this disclosure involves a computer that inputs speech data representing utterances divided into multiple pronunciation segments into a first speech recognition model and obtains a first speech recognition result obtained by inputting the speech data into a second speech recognition model, obtains an integrated speech recognition result created based on the first speech recognition result and the second speech recognition result, outputs the integrated speech recognition result, and the integrated speech recognition result is created based on the consistency of the content of the utterances across the multiple pronunciation segments.

[0008] According to this disclosure, highly accurate speech recognition results can be obtained with a simple configuration.

[0009] This is a schematic diagram showing the configuration of the speech recognition system 1. This is a block diagram showing the configuration of the user terminal 20. This is a block diagram showing the configuration of the speech recognition device 10. This is a schematic diagram showing an example of the speech recognition flow in the speech recognition device 10. This is a schematic diagram showing an example of a prompt PT to be input to the large-scale language model 30. This is a flowchart showing the operation of the processing unit 103 of the speech recognition device 10. This is a schematic diagram showing another example of the speech recognition flow in the speech recognition device 10. This is a schematic diagram showing another example of the speech recognition flow in the speech recognition device 10. This is a block diagram showing the configuration of the speech recognition device 10 according to a third modified example.

[0010] [Embodiment] Figure 1 is a schematic diagram showing the configuration of the speech recognition system 1. The speech recognition system 1 according to this embodiment includes a speech recognition device 10 and a user terminal 20. In this embodiment, the case in which the speech recognition system 1 creates a speech recognition result corresponding to the speech data to be recognized DO (see Figure 4) transmitted from the user terminal 20 and provides the speech recognition result to the user terminal 20 will be described. The speech data to be recognized DO is an example of speech data. The speech recognition device 10 may, for example, provide the user terminal 20 with a speech recognition result of the speech data to be recognized DO acquired from a device other than the user terminal 20. The speech recognition device 10 may also, for example, provide the user terminal 20 with a terminal other than the user terminal 20 a speech recognition result of the speech data to be recognized DO transmitted from the user terminal 20.

[0011] The user terminal 20 is an information processing terminal. More specifically, the user terminal 20 may be a personal computer, a smartphone, a tablet, or a wearable device (e.g., a smartwatch, smart glasses, or smart ring). The user terminal 20 is held by user U.

[0012] The voice recognition device 10 and each user terminal 20 are interconnected by a network N. Network N may be a WAN (Wide Area Network) or a LAN (Local Area Network).

[0013] Furthermore, an information processing device (not shown in the diagram) that provides a Large Language Model (LLM) 30 is connected to the network N. As will be described later, the speech recognition device 10 uses the Large Language Model 30 to obtain speech recognition results to be provided to the user terminal 20.

[0014] [User Terminal 20] Figure 2 is a block diagram showing the configuration of the user terminal 20. Figure 2 illustrates the configuration when the user terminal 20 is a personal computer. The user terminal 20 comprises a display device 201, an input device 202, a microphone 203, a speaker 204, a communication device 206, a storage device 208, a processing device 209, and a bus 220 that connects these devices to each other.

[0015] The display device 201 is a device that displays information to the outside (for example, various display panels such as liquid crystal display panels or organic EL display panels). The input device 202 is a device that receives input from the outside (for example, a keyboard, mouse, microphone, switch, button, or sensor). The display device 201 and the input device 202 may be configured as an integrated unit (for example, a touch panel).

[0016] Microphone 203 picks up sound from the user terminal 20 and generates audio data. The audio data generated by microphone 203 can be the audio data DO to be recognized. The sound picked up by microphone 203 is mainly the speech of user U. Speaker 204 plays back the audio data and outputs sound corresponding to the audio data. The audio data played back by speaker 204 is, for example, the speech of the other party when user U makes a call with another user U using a voice call application, or the playback sound when user U watches a video or listens to music. Microphone 203 and speaker 204 may be separate from the user terminal 20 and not be included in the user terminal 20. In this case, the user terminal 20 and microphone 203 and speaker 204 are connected by an interface (not shown).

[0017] The communication device 206 has an interface that can connect to the network N and communicates with other devices connected to the network N using wireless or wired communication. The communication device 206 may also be an interface for short-range wireless communication with other user terminals 20 (wearable terminals, etc.) owned by user U.

[0018] The storage device 208 is a recording medium that can be read by the processing device 209. The storage device 208 includes, for example, non-volatile memory and volatile memory. Non-volatile memory includes, for example, ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), and EEPROM (Electrically Erasable Programmable Read Only Memory). Volatile memory includes, for example, RAM (Random Access Memory). The storage device 208 stores program PG2. Program PG2 is a program for operating the user terminal 20.

[0019] The processing unit 209 includes one or more CPUs (Central Processing Units). One or more CPUs are examples of one or more processors. Each processor and CPU is an example of a computer.

[0020] The processing unit 209 reads the program PG2 from the storage device 208. By executing the program PG2, the processing unit 209 functions as an audio data transmission unit 211 and an audio recognition result output unit 212. At least one of the audio data transmission unit 211 and the audio recognition result output unit 212 may be composed of circuits such as a DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), PLD (Programmable Logic Device), and FPGA (Field Programmable Gate Array).

[0021] The voice data transmission unit 211 transmits the recognition target voice data DO, which is the utterance H of user U, to the voice recognition device 10. In this embodiment, the recognition target voice data DO represents an utterance H divided into multiple pronunciation sections HT (HT1, HT2). A single utterance H may contain multiple sentences. In this embodiment, if there is a silent section of a predetermined time or longer in the recognition target voice data DO, that section is considered to be the boundary of a single utterance H. The portion of the utterance H separated by a silent section is hereinafter referred to as the "pronunciation section HT".

[0022] The voice data DO to be recognized is, for example, one voice data specified by user U from among multiple voice data stored in the storage device 208. Furthermore, the voice data DO to be recognized may be voice data generated by the microphone 203, or voice data acquired by user U from another information processing device via the network N. Alternatively, voice data generated by collecting user U's speech H in real time with the microphone 203 may also be the voice data DO to be recognized.

[0023] The speech recognition result output unit 212 outputs the integrated speech recognition result KT transmitted from the speech recognition device 10. As will be described later, the integrated speech recognition result KT is text data. Therefore, the speech recognition result output unit 212 displays text indicating the integrated speech recognition result KT on, for example, the display device 201. Alternatively, the speech recognition result output unit 212 may print the text information indicating the integrated speech recognition result KT using, for example, a printer (not shown).

[0024] [Voice Recognition Device 10] Figure 3 is a block diagram showing the configuration of the voice recognition device 10. The voice recognition device 10 includes a communication device 101, a storage device 102, a processing device 103, and a bus 120 that connects these devices to each other.

[0025] The communication device 101 has an interface that can be connected to the network N and communicates with other devices connected to the network N using wireless or wired communication.

[0026] The storage device 102 is a recording medium that can be read by the processing device 103. The storage device 102 includes, for example, non-volatile memory and volatile memory. The non-volatile memory is, for example, ROM, EPROM, and EEPROM. The volatile memory is, for example, RAM. The storage device 102 stores program PG1, first speech recognition model M1, and second speech recognition model M2.

[0027] Program PG1 is a program for operating the speech recognition device 10. The first speech recognition model M1 and the second speech recognition model M2 receive the speech data DO to be recognized as input, perform speech recognition processing on the speech data DO, and output text data which is the speech recognition result. "Speech recognition model" may be replaced with "speech recognition engine". The speech recognition result output from the first speech recognition model M1 is called the "first speech recognition result K1", and the speech recognition result output from the second speech recognition model M2 is called the "second speech recognition result K2".

[0028] In this embodiment, the first speech recognition model M1 and the second speech recognition model M2 each have different characteristics. For example, the first speech recognition model M1 has high general accuracy, while the second speech recognition model M2 has particularly high accuracy in noisy environments. Therefore, when the same audio data is input to the first speech recognition model M1 and the second speech recognition model M2, different speech recognition results may be obtained. That is, the first speech recognition result and the second speech recognition result for the same audio data DO may differ.

[0029] Furthermore, word registration may be performed in at least one of the first speech recognition model M1 and the second speech recognition model M2. Word registration involves registering words that are highly likely to be included in the speech data DO to be recognized as "registered words" in the speech recognition model. Generally, words registered as registered words have a higher probability of being correctly recognized compared to words that are not registered as registered words. Examples of registered words include proper nouns such as company names, service names, or personal names, or specialized terms such as medical terms.

[0030] The processing unit 103 includes one or more CPUs. One or more CPUs are examples of one or more processors. Each of the processors and CPUs is an example of a computer. By executing the program PG1, the processing unit 103 functions as a voice data acquisition unit 110, a first voice recognition result acquisition unit 111, a second voice recognition result acquisition unit 112, and an output unit 113. At least some of the functions of these components may be configured by circuits such as a DSP, ASIC, PLD, and FPGA.

[0031] The voice data acquisition unit 110 acquires the voice data DO to be recognized. In this embodiment, the voice data acquisition unit 110 acquires the voice data DO to be recognized from the user terminal 20.

[0032] The first speech recognition result acquisition unit 111 acquires a first speech recognition result K1 obtained by inputting recognition target speech data DO representing an utterance H divided into multiple pronunciation intervals HT to the first speech recognition model M1, and a second speech recognition result K2 obtained by inputting recognition target speech data DO to the second speech recognition model M2. The first speech recognition result acquisition unit 111 is an example of the first acquisition unit. As described above, the first speech recognition model M1 and the second speech recognition model M2 each have different characteristics. Therefore, even when the same recognition target speech data DO is input, the first speech recognition result K1 and the second speech recognition result K2 may differ.

[0033] In this embodiment, the utterance H represented by the recognition target speech data DO includes multiple pronunciation intervals HT. The first speech recognition result acquisition unit 111 divides the recognition target speech data DO into silent intervals to create speech data for each pronunciation interval (hereinafter sometimes referred to as "unit speech data DT"). The first speech recognition result acquisition unit 111 then inputs the multiple unit speech data DT corresponding to the multiple pronunciation intervals HT to the first speech recognition model M1 and the second speech recognition model M2, respectively. For example, if there is recognition target speech data DO that represents an utterance H including a first pronunciation interval HT1 and a second pronunciation interval HT2, a first unit speech data DT1 representing the first pronunciation interval HT1 and a second unit speech data DT2 representing the second pronunciation interval HT2 are created.

[0034] The first speech recognition model M1 outputs the first unit speech recognition result K1-1 for the first unit speech data DT1 and the first unit speech recognition result K1-2 for the second unit speech data DT2. The second speech recognition model M2 outputs the second unit speech recognition result K2-1 for the first unit speech data DT1 and the second unit speech recognition result K2-2 for the second unit speech data DT2. This ensures that each speech recognition model considers the range of a single pronunciation interval HT to be the same, preventing problems such as comparing speech recognition results from different pronunciation intervals HT when comparing speech recognition results.

[0035] Furthermore, the method of dividing the audio data DO to be recognized into pronunciation intervals HT is not limited to dividing it into silent intervals. Also, the unit for dividing the audio data DO to be recognized is not limited to pronunciation intervals HT; the audio data DO to be recognized can be divided into any unit as long as the comparison of speech recognition results can be performed accurately (for example, relatively short audio data DO to be recognized may not need to be divided at all).

[0036] Figure 4 is a schematic diagram showing an example of the speech recognition flow in the speech recognition device 10. For example, suppose user U makes an utterance H which is divided into a first utterance section HT1 and a second utterance section HT2. The first utterance section HT1 is "Next, I would like to talk about YY, a new product from XX Company." The second utterance section HT2 is "YY is scheduled to be released in August." XX Company is a proper noun indicating the name of a company, and YY is a proper noun indicating the name of a new product. The speech data acquisition unit 110 acquires the speech data DO to be recognized, which represents the utterance H, and divides the speech data DO to be recognized into a first unit speech data DT1 corresponding to the first utterance section HT1 and a second unit speech data DT2 corresponding to the second utterance section HT2.

[0037] The first speech recognition result acquisition unit 111 inputs the first unit speech data DT1 and the second unit speech data DT2 to the first speech recognition model M1 and the second speech recognition model M2, respectively. Here, the first speech recognition model M1 has "XX Company" and "YY" registered as words. On the other hand, the second speech recognition model M2 has no registered words.

[0038] The first speech recognition result K1 output from the first speech recognition model M1 includes the first unit speech recognition result K1-1, which is the speech recognition result of the first pronunciation section HT1 (first unit speech data DT1), and the first unit speech recognition result K1-2, which is the speech recognition result of the second pronunciation section HT2 (second unit speech data DT2). Similarly, the second speech recognition result K2 output from the second speech recognition model M2 includes the second unit speech recognition result K2-1, which is the speech recognition result of the first pronunciation section HT1 (first unit speech data DT1), and the second unit speech recognition result K2-2, which is the speech recognition result of the second pronunciation section HT2 (second unit speech data DT2).

[0039] The first unit speech recognition result K1-1 is, "Next, we will discuss YY, a new product from XX Company." The first unit speech recognition result K1-2 is, "YY is scheduled to be released in August." The second unit speech recognition result K2-1 is, "Next, we will discuss the editor's new product, Beloved Wife." The second unit speech recognition result K2-2 is, "YY is scheduled to be released in August."

[0040] Of the first speech recognition results K1, the first unit speech recognition result K1-1 matches the utterance content of the first pronunciation section HT1, but the first unit speech recognition result K1-2 does not match the utterance content of the second pronunciation section HT2. Also, of the second speech recognition results K2, the second unit speech recognition result K2-1 does not match the utterance content of the first pronunciation section HT1, but the second unit speech recognition result K2-2 matches the utterance content of the second pronunciation section HT2. In other words, both the first speech recognition result K1 and the second speech recognition result K2 contain some speech recognition errors.

[0041] Returning to the explanation in Figure 3, the second speech recognition result acquisition unit 112 acquires an integrated speech recognition result KT created based on the first speech recognition result K1 and the second speech recognition result K2. The second speech recognition result acquisition unit 112 is an example of a second acquisition unit.

[0042] In this embodiment, the second speech recognition result acquisition unit 112 acquires the integrated speech recognition result KT using the large-scale language model 30. More specifically, the second speech recognition result acquisition unit 112 receives a prompt PT (see Figure 5), a first speech recognition result K1, and a second speech recognition result K2 as input to the large-scale language model 30. The large-scale language model 30 creates and outputs the integrated speech recognition result KT based on the prompt PT, the first speech recognition result K1, and the second speech recognition result K2. The second speech recognition result acquisition unit 112 acquires the integrated speech recognition result KT output from the large-scale language model 30.

[0043] In this embodiment, the large language model 30 compares the first speech recognition result K1 and the second speech recognition result K2 for each pronunciation section HT, and creates an integrated speech recognition result KT by selecting one more appropriate speech recognition result. Specifically, the large language model 30 selects the one that is estimated to be more appropriate between the first unit speech recognition result K1-1 and the second unit speech recognition result K2-1, which are the speech recognition results for the first pronunciation section HT1 (the first unit speech data DT1). Also, the large language model 30 selects the one that is estimated to be more appropriate between the first unit speech recognition result K1-2 and the second unit speech recognition result K2-2, which are the speech recognition results for the second pronunciation section HT2 (the second unit speech data DT2). Then, the integrated speech recognition result KT is created by connecting the selected speech recognitions in the order of the pronunciation sections HT.

[0044] In other words, the utterance H includes the first pronunciation section HT1 and the second pronunciation section HT2. The integrated speech recognition result KT selects either the part corresponding to the first pronunciation section HT1 in the first speech recognition result K1 (the first unit speech recognition result K1-1) or the part corresponding to the first pronunciation section HT1 in the second speech recognition result K2 (the second unit speech recognition result K2-1) as the speech recognition result for the first pronunciation section HT1. Also, the integrated speech recognition result KT is created by selecting either the part corresponding to the second pronunciation section HT2 in the first speech recognition result K1 (the first unit speech recognition result K1-2) or the part corresponding to the second pronunciation section HT2 in the second speech recognition result K2 (the second unit speech recognition result K2-2) as the speech recognition result for the second pronunciation section HT2.

[0045] Here, in this embodiment, the large language model 30 creates the integrated speech recognition result KT in consideration of the following [1] to [3].

[0046] [1] Consistency of the content of speech H across multiple pronunciation intervals HT The consistency of the content of speech H across multiple pronunciation intervals HT (hereinafter sometimes simply referred to as "consistency") means whether the content of speech H between the preceding and subsequent pronunciation intervals HT is consistent. The consistency of the content of speech H across multiple pronunciation intervals HT may be paraphrased as the context of speech H. For example, consider the case where the first pronunciation interval HT1 is "When is the new product scheduled to be released?" For the second pronunciation interval HT2 following the first pronunciation interval HT1, assume that the first speech recognition result K1 is "It is scheduled to be released in August." and the second speech recognition result K2 is "It is scheduled to disappear in August." Although both the first speech recognition result K1 and the second speech recognition result K2 are appropriate as individual sentences, the first speech recognition result K1 is judged to have more consistent content of speech H in the preceding and subsequent pronunciation intervals HT.

[0047] [2] Grammatical accuracy of the speech recognition result The grammatical accuracy of the speech recognition result means whether the speech recognition result forms an appropriate sentence when viewed as a single sentence. For example, when the first speech recognition result K1 is "YY is scheduled to be released in August." and the second speech recognition result K2 is "YY does not have a release schedule in August.", the first speech recognition result K1 is judged to be more grammatically accurate.

[0048] [3] Registered words As described above, a registered word is a word registered as a word that is likely to be included in the recognition target speech data DO. Registered words are often unknown words (words that generally do not exist), and if registered words are not considered, the appropriateness of the sentence may be unduly reduced. Therefore, the large language model 30 evaluates [1] and [2] above after considering registered words.

[0049] Thus, in this embodiment, the integrated speech recognition result KT is created based on the consistency of the content of speech H across multiple pronunciation intervals HT. In addition, the integrated speech recognition result KT may be created based on at least one of the grammatical accuracy between the first speech recognition result K1 and the second speech recognition result K2 and the pre-registered words (registered words) related to the content of speech H, in addition to the consistency of the content of speech H across multiple pronunciation intervals HT.

[0050] Figure 5 is a schematic diagram showing an example of a prompt PT to be input to the large-scale language model 30. The prompt PT shown in Figure 5 includes an instruction P1, constraints P2, and example sentences P3. In this embodiment, the prompt PT is written in natural language. The instruction P1 contains a sentence that instructs the system to select the more appropriate of two speech recognition results. The constraints P2 contain the processing rules. Among these, "context" refers to the consistency of the content of the utterance H across multiple pronunciation intervals HT. Example sentences P3 contain examples of output sentences (examples of more appropriate sentences) for two types of input sentences.

[0051] In other words, the prompt PT given to the large-scale language model 30 instructs it to select the more appropriate speech recognition result from the first speech recognition result K1 and the second speech recognition result K2 for each speech segment HT, based on the consistency of the content of the utterance H across multiple speech segments HT.

[0052] Given the first speech recognition result K1 shown in Figure 4, the second speech recognition result K2, and the prompt PT shown in Figure 5, the large-scale language model 30 outputs the integrated speech recognition result KT shown in Figure 4. The integrated speech recognition result KT includes the first unit recognition result KT-1 corresponding to the first pronunciation interval HT1 and the second unit recognition result KT-2 corresponding to the second pronunciation interval HT2. The first unit recognition result KT-1 is selected as the first unit speech recognition result K1-1. The second unit speech recognition result KT-2 is selected as the second unit speech recognition result K2-2. The integrated speech recognition result KT matches the content of user U's utterance H, and it can be said that a correct speech recognition result has been obtained.

[0053] We will now examine in more detail the method for selecting the integrated speech recognition result KT. Hereafter, the correct speech recognition result will be referred to as the "correct answer." Comparing the first unit speech recognition result K1-1 and the second unit speech recognition result K2-1 corresponding to the first pronunciation segment HT1, the first unit speech recognition result K1-1 contains the registered words "XX Company" and "YY," and is therefore more likely to be the correct answer than the second unit speech recognition result K2-1. Thus, the large-scale language model 30 selects the first unit speech recognition result K1-1 as a candidate for the speech recognition result of the first pronunciation segment HT1.

[0054] Furthermore, for example, the second unit speech recognition result K2-1, "Next, let's talk about the editor's new product, his beloved wife," sounds unnatural in the phrase "the editor's new product," and also unnatural in the phrase "(the new product is) his beloved wife." Therefore, from the standpoint of grammatical accuracy, it can be said that the second unit speech recognition result K2-1 is highly likely to be inappropriate as the speech recognition result for the first pronunciation segment HT1.

[0055] Furthermore, comparing the first unit speech recognition result K1-2 and the second unit speech recognition result K2-2 corresponding to the second pronunciation segment HT2, both contain the registered word "YY". Therefore, from the perspective of registered words, it is not possible to determine which is correct. On the other hand, considering the relationship with the first unit speech recognition result K1-1, which is a candidate for the speech recognition result of the first pronunciation segment HT1, the phrase "new product" contained in the first unit speech recognition result K1-1 has a higher correlation with "scheduled for release" contained in the second unit speech recognition result K2-2. In other words, considering the consistency of the content of utterance H across the first pronunciation segment HT1 and the second pronunciation segment HT2, the second unit speech recognition result K2-2 is more likely to be correct. Therefore, the large-scale language model 30 selects the second unit speech recognition result K2-2 as a candidate for the speech recognition result of the second pronunciation segment HT2.

[0056] Returning to the explanation in Figure 3, the output unit 113 outputs the integrated speech recognition result KT. In this embodiment, the output unit 113 transmits the integrated speech recognition result KT to the user terminal 20. The speech recognition result output unit 212 of the user terminal 20, upon receiving the integrated speech recognition result KT, outputs the integrated speech recognition result KT to, for example, the display device 201. If the speech recognition device 10 has a display device, the output unit 113 may display the integrated speech recognition result KT on the display device. Also, if a printer is connected to the speech recognition device 10, the output unit 113 may cause the printer to print the integrated speech recognition result KT.

[0057] [Flowchart] Figure 6 is a flowchart showing the operation of the processing unit 103 of the speech recognition device 10. The processing unit 103 functions as a speech data acquisition unit 110 and acquires the speech data DO to be recognized from the user terminal 20 (step S100). The processing unit 103 divides the speech data DO to be recognized into multiple unit speech data DT for each pronunciation interval HT (step S102).

[0058] The processing unit 103 functions as the first speech recognition result acquisition unit 111 and inputs multiple unit speech data DT to the first speech recognition model M1 and the second speech recognition model M2 (indicated as "multiple speech recognition models" in the figure) (step S104). The processing unit 103 functions as the first speech recognition result acquisition unit 111 and acquires multiple speech recognition results (first speech recognition result K1 and second speech recognition result K2) corresponding to the multiple unit speech data DT from each of the first speech recognition model M1 and the second speech recognition model M2 (step S106).

[0059] The processing unit 103 functions as a second speech recognition result acquisition unit 112 and inputs the first speech recognition result K1, the second speech recognition result K2, and the prompt PT to the large-scale language model 30 (step S108). The processing unit 103 functions as a second speech recognition result acquisition unit 112 and acquires the integrated speech recognition result KT output from the large-scale language model 30 (step S110). The processing unit 103 functions as an output unit 113 and outputs the integrated speech recognition result KT (step S112), ending the processing of this flowchart.

[0060] [Summary of Embodiments] As described above, the speech recognition system 1 according to the embodiment considers the consistency of the content of the utterance H across multiple pronunciation intervals HT, i.e., the context, when creating an integrated speech recognition result KT based on the first speech recognition result K1 output from the first speech recognition model M1 and the second speech recognition result K2 output from the second speech recognition model M2. This improves the likelihood of obtaining a correct speech recognition result. Furthermore, by creating an integrated speech recognition result KT using the speech recognition results of multiple speech recognition models, it is possible to obtain a speech recognition result with fewer errors than if the speech recognition result of a single speech recognition model were used as is.

[0061] Furthermore, in the speech recognition system 1, the integrated speech recognition result KT is created by comparing the first speech recognition result K1 and the second speech recognition result K2 for each pronunciation interval HT and selecting the more appropriate one. Therefore, a more accurate speech recognition result (integrated speech recognition result KT) can be obtained than, for example, adopting either the first speech recognition result K1 or the second speech recognition result K2 as the speech recognition result at once. In addition, compared to, for example, dividing the first speech recognition result K1 and the second speech recognition result K2 into intervals smaller than the pronunciation interval HT (e.g., at the word level) and comparing them, the processing load in creating the integrated speech recognition result KT can be reduced.

[0062] Furthermore, in the speech recognition system 1, the integrated speech recognition result KT is created by the large-scale language model 30. This reduces the cost of training (time cost and human cost) compared to, for example, creating a dedicated learning model.

[0063] Furthermore, in the speech recognition system 1, the integrated speech recognition result KT is created based on the consistency of the content of the utterance H across multiple pronunciation intervals HT, as well as the grammatical accuracy of the first speech recognition result K1 and the second speech recognition result K2, and registered words related to the content of the utterance H. This makes it possible to obtain a more accurate speech recognition result.

[0064] [Regarding Modifications] The following are examples of modifications in the above-described embodiment. Two or more modifications can be arbitrarily selected from the following examples and combined as appropriate, provided they do not contradict each other.

[0065] [First Modification] In the embodiment described above, the large-scale language model 30 created an integrated speech recognition result KT by selecting either the first speech recognition result K1 or the second speech recognition result K2 as the appropriate speech recognition result for each of the multiple pronunciation intervals HT. However, the large-scale language model 30 may also create an integrated speech recognition result KT by generating an appropriate speech recognition result based on the first speech recognition result K1 and the second speech recognition result K2.

[0066] Figure 7 is a schematic diagram showing another example of the speech recognition flow in the speech recognition device 10. In Figure 7, the utterance H of user U is the same as that shown in Figure 4. Assume that the recognition target speech data DO (first unit speech data DT1 and second unit speech data DT2) representing user U's utterance H is input to the first speech recognition model M1 and the second speech recognition model M2, respectively, and as a result, the first speech recognition result K1A and the second speech recognition result K2A are obtained.

[0067] The first speech recognition result K1A includes the first unit speech recognition result K1A-1, which is the speech recognition result of the first pronunciation section HT1 (first unit speech data DT1), and the first unit speech recognition result K1A-2, which is the speech recognition result of the second pronunciation section HT2 (second unit speech data DT2). The second speech recognition result K2A includes the second unit speech recognition result K2A-1, which is the speech recognition result of the first pronunciation section HT1 (first unit speech data DT1), and the second unit speech recognition result K2A-2, which is the speech recognition result of the second pronunciation section HT2 (second unit speech data DT2).

[0068] The first unit speech recognition result K1A-1 is, "Um, next, we have YY, which is a product submitted by XX company." The first unit speech recognition result K1A-2 is, "YY is scheduled to be released in August." The second unit speech recognition result K2A-1 is, "Next, we have Aizai, the editor's new product." The second unit speech recognition result K2A-2 is, "YY is scheduled to be released in August." The first unit speech recognition results K1A-1 and K1A-2, and the second unit speech recognition results K2A-1 and K2A-2 all contain some speech recognition errors.

[0069] In such cases, the large-scale language model 30 generates a speech recognition result for the first pronunciation segment HT1 that is more likely to be correct, based on the first unit speech recognition results K1A-1 and K1A-2. The large-scale language model 30 also generates a speech recognition result for the second pronunciation segment HT2 that is more likely to be correct, based on the second unit speech recognition results K2A-1 and K2A-2. Specifically, for example, by swapping some words in the first unit speech recognition results K1A-1 and K1A-2, or changing some words to other words (for example, homophones), the model generates a speech recognition result for the first pronunciation segment HT1 that is more likely to be correct.

[0070] Thus, in the first modified example, instead of selecting either the first speech recognition result K1 or the second speech recognition result K2 as the appropriate speech recognition result, an appropriate speech recognition result is generated based on both the first speech recognition result K1 and the second speech recognition result K2. This increases the likelihood of obtaining a correct integrated speech recognition result KT, even if, for example, either the first speech recognition result K1 or the second speech recognition result K2 contains errors.

[0071] [Second Modification] The above-described embodiment explained the case of speech recognition of Japanese utterances H. However, the speech recognition system 1 according to this embodiment can also be applied to other languages ​​such as English, Chinese, and Korean.

[0072] Figure 8 is a schematic diagram showing another example of the speech recognition flow in the speech recognition device 10. For example, suppose user U makes an utterance HB in English which is divided into a first pronunciation section HTB1 and a second pronunciation section HTB2. The first pronunciation section HTB1 is "Next, we will discuss XX company's new product, YY." The second pronunciation section HTB2 is "YY is scheduled to be released in August." XX is a proper noun representing the name of a company, and YY is a proper noun representing the name of a new product. The voice data acquisition unit 110 acquires the recognition target voice data DOB representing the utterance HB, and divides the recognition target voice data DOB into a first unit voice data DTB1 corresponding to the first pronunciation section HTB1 and a second unit voice data DTB2 corresponding to the second pronunciation section HTB2.

[0073] The first speech recognition result acquisition unit 111 inputs the first unit speech data DTB1 and the second unit speech data DTB2 to the first speech recognition model M1 and the second speech recognition model M2, respectively. Here, the first speech recognition model M1 has "XX" and "YY" registered as registered words. On the other hand, the second speech recognition model M2 has no registered words.

[0074] The first speech recognition result K1B output from the first speech recognition model M1 includes the first unit speech recognition result K1B-1, which is the speech recognition result of the first pronunciation section HTB1 (first unit speech data DTB1), and the first unit speech recognition result K1B-2, which is the speech recognition result of the second pronunciation section HTB2 (second unit speech data DTB2). Similarly, the second speech recognition result K2 output from the second speech recognition model M2 includes the second unit speech recognition result K2B-1, which is the speech recognition result of the first pronunciation section HTB1 (first unit speech data DTB1), and the second unit speech recognition result K2B-2, which is the speech recognition result of the second pronunciation section HTB2 (second unit speech data DTB2).

[0075] The first unit speech recognition result K1B-1 is "Next, we will discuss XX company's new product, YY." The first unit speech recognition result K1B-2 is "YY is scheduled to be really since August." The second unit speech recognition result K2B-1 is "Next, we will discuss eggs company's new product, why why." The second unit speech recognition result K2B-2 is "YY is scheduled to be released in August."

[0076] Of the first speech recognition results K1B, the first unit speech recognition result K1B-1 matches the utterance content of the first pronunciation section HTB1, but the first unit speech recognition result K1B-2 does not match the utterance content of the second pronunciation section HTB2. Similarly, of the second speech recognition results K2B, the second unit speech recognition result K2B-1 does not match the utterance content of the first pronunciation section HTB1, but the second unit speech recognition result K2B-2 matches the utterance content of the second pronunciation section HTB2. In other words, both the first speech recognition result K1B and the second speech recognition result K2B contain some speech recognition errors.

[0077] The second speech recognition result acquisition unit 112 receives the prompt PT, the first speech recognition result K1B, and the second speech recognition result K2B as input to the large-scale language model 30 capable of handling English. The large-scale language model 30 compares the first speech recognition result K1B and the second speech recognition result K2B for each pronunciation section HTB and creates an integrated speech recognition result KTB by selecting the most appropriate speech recognition result. In the example in Figure 8, the first unit speech recognition result K1B-1 is selected as the speech recognition result for the first pronunciation section HTB1, and the second unit speech recognition result K2B-2 is selected as the speech recognition result for the second pronunciation section HTB2.

[0078] Thus, according to this second modification, the accuracy of speech recognition can be improved not only in Japanese but also in various other languages.

[0079] [Third Modification] In the embodiment described above, the second speech recognition result acquisition unit 112 acquires the integrated speech recognition result KT by inputting the first speech recognition result K1 and the second speech recognition result K2 into the large-scale language model 30. In the third modification, the integrated speech recognition result KT is acquired using a dedicated learning model LM instead of the large-scale language model 30.

[0080] Figure 9 is a block diagram showing the configuration of the speech recognition device 10 according to the third modified example. In the third modified example, the learning model LM is stored in the storage device 102 of the speech recognition device 10. The learning model LM compares the first speech recognition result K1 and the second speech recognition result K2 for each pronunciation interval HT and selects the speech recognition result that is more likely to be correct. Then, it arranges the selected speech recognition results in order of pronunciation interval HT to create an integrated speech recognition result KT.

[0081] The learning model LM is provided with a set of multiple (two in this embodiment) learning speech recognition results for the same learning audio data (input learning data) and correct answer information indicating which of the multiple learning speech recognition results is correct (output learning data). For example, a human judges which of the multiple learning speech recognition results is correct and labels the correct learning speech recognition result. Furthermore, to enable contextual learning, the content of the utterances H before and after the learning audio data is also provided to the learning model LM as relevant information. Additionally, if there are registered words, these registered words are also provided to the learning model LM as relevant information.

[0082] More specifically, the learning model LM is given a pair of input and output training data, such as the following: Input training data: {Registered words: {XX Company, YY}, Introductory sentence: That concludes this topic., Concluding sentence: It is scheduled to be released in August., First speech recognition result: Next, we will discuss YY, XX Company's new product., Second speech recognition result: Next, we will discuss Ambiguous, the editor's new product.}. Output training data: {Next, we will discuss YY, XX Company's new product.}

[0083] Thus, in the third modified example, the second speech recognition result acquisition unit 112 inputs the first speech recognition result K1 and the second speech recognition result K2 to a learning model LM that has been machine-learned using learning data which includes a plurality of learning speech recognition results and correct answer information indicating which of the plurality of learning speech recognition results is correct, and acquires the integrated speech recognition result KT output from the learning model LM.

[0084] According to the third modification, the integrated speech recognition result KT can be obtained without using the large-scale language model 30. This is effective, for example, when the speech data DO to be recognized cannot be transmitted to a device other than the speech recognition device 10 due to contractual obligations with the user U, or when using the large-scale language model 30 would incur significant costs. Furthermore, the learning model LM can be trained specifically for the first speech recognition model M1 and the second speech recognition model M2, thereby improving the accuracy of speech recognition.

[0085] Furthermore, the learning model LM is not limited to being stored in the speech recognition device 10; like the large-scale language model 30, it may also be stored in an information processing device other than the speech recognition device 10 that is connected to the network N.

[0086] [Fourth Modification] In a typical speech recognition model, the acoustic score, linguistic score, and confidence level of the speech recognition result are output along with the speech recognition result. These values ​​may be input to the large-scale linguistic model 30 along with the first speech recognition result K1 and the second speech recognition result K2.

[0087] The acoustic score is the acoustic likelihood (a value representing the likelihood from the perspective of the acoustic model) for the speech recognition result. The linguistic score is the linguistic likelihood (a value representing the likelihood from the perspective of the linguistic model) for the speech recognition result. The confidence level is the likelihood of the current speech recognition result compared to other candidate speech recognition results, and is a value evaluated relatively on a range of 0 to 1.0.

[0088] The first speech recognition result acquisition unit 111 may acquire the acoustic score, language score, and confidence level of the first speech recognition result K1 from the first speech recognition model M1, and may acquire the acoustic score, language score, and confidence level of the second speech recognition result K2 from the second speech recognition model M2. If at least one of the acoustic score, language score, and confidence level is acquired, the second speech recognition result acquisition unit 112 may input the acquired acoustic score, language score, and confidence level, along with the first speech recognition result K1 and the second speech recognition result K2, to the large-scale language model 30.

[0089] The large-scale language model 30 creates an integrated speech recognition result KT by considering at least one of the input acoustic score, language score, and confidence level. Specifically, for example, the speech recognition result with the higher acoustic score, language score, and confidence level is included in the integrated speech recognition result KT.

[0090] Furthermore, as in the third modified example, when creating a dedicated learning model LM, the learning model LM can be made to perform learning that takes these values ​​into account by including the acoustic score, linguistic score, and confidence level of each speech recognition result in the training data.

[0091] As described above, in the fourth modified example, the first speech recognition model M1 selects the first speech recognition result K1 from among a plurality of first speech recognition result candidates, and outputs the first speech recognition result K1 along with a first acoustic score indicating the acoustic validity of the first speech recognition result K1, a first linguistic score indicating the linguistic validity of the first speech recognition result K1, and a first confidence score indicating the likelihood of the first speech recognition result K1 among the plurality of first speech recognition result candidates. The second speech recognition model M2 selects the second speech recognition result K2 from among a plurality of second speech recognition result candidates, and outputs the second speech recognition result K2 along with a second acoustic score indicating the acoustic validity of the second speech recognition result K2, a second linguistic score indicating the linguistic validity of the second speech recognition result K2, and a second confidence score indicating the likelihood of the second speech recognition result K2 among the plurality of second speech recognition result candidates. The integrated speech recognition result KT is created based on the consistency of the content of the utterance H across multiple pronunciation segments HT, as well as at least one of the following: the first acoustic score, the first language score, the first confidence level, the second acoustic score, the second language score, and the second confidence level.

[0092] According to the fourth modification, the accuracy of the integrated speech recognition result KT can be improved based on existing information such as acoustic score, language score, and confidence level.

[0093] [Fifth Variation] User U's utterance H may differ in its manner of speaking (words used, sentence endings, etc.) depending on the situation in which the utterance H is made. For example, comparing "conversation between friends" and "discussion in a workplace meeting," it is thought that the way of speaking will be relatively informal in "conversation between friends" and formal in "discussion in a workplace meeting." Therefore, the second speech recognition result acquisition unit 112 may input information indicating the situation of the utterance H along with the recognition target speech data DO to the large-scale language model 30 or the learning model LM.

[0094] When using the large-scale language model 30, for example, the prompt PT can be set to specify a situation such as "conversation between friends" or "discussion in a workplace meeting." In the case of the learning model LM, for example, by specifying a situation such as "business meeting," "small talk," or "sports" during the learning stage, the model learns how to speak in that situation.

[0095] According to the fifth modification, the large-scale language model 30 or the learning model LM creates speech recognition results by taking into account the situation of the utterance H, thereby improving the accuracy of speech recognition.

[0096] [Sixth Modification] In the embodiment described above, the first speech recognition model M1 and the second speech recognition model M2 were stored in the speech recognition device 10. However, the embodiment is not limited to this, and at least one of the first speech recognition model M1 and the second speech recognition model M2 may be stored in an information processing device other than the user terminal 20 connected to the network N. In this case, the first speech recognition result acquisition unit 111 transmits the recognition target speech data DO (unit speech data DT) to the information processing device via the network N and acquires the speech recognition result returned via the network N.

[0097] According to the sixth modification, the speech recognition device 10 does not need to maintain a speech recognition model, thereby reducing the processing load on the speech recognition device 10. Furthermore, various speech recognition models, each with different characteristics, can be used, and it is expected that the speech recognition accuracy of the integrated speech recognition result KT will improve.

[0098] [Seventh Modification] In the embodiment described above, the large-scale language model 30 was stored in an information processing device other than the speech recognition device 10 connected to the network N. However, the speech recognition device 10 may also store the large-scale language model 30.

[0099] [Other] (1) In the embodiments described above, ROM and RAM were given as examples of storage devices 102 and 207, but storage devices 102 and 207 may be flexible disks, magneto-optical disks (e.g., compact disks, digital multipurpose disks, Blu-ray® disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROMs (Compact Disc-ROMs), registers, removable disks, hard disks, floppy® disks, magnetic strips, databases, servers, or other suitable storage media.

[0100] (2) In the embodiments described above, the information, signals, etc. may be represented using any of the various different techniques. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0101] (3) In the embodiments described above, the input and output information may be stored in a specific location (e.g., memory) or managed using a management table. The input and output information may be overwritten, updated, or appended to. The output information may be deleted. The input information may be transmitted to other devices.

[0102] (4) In the embodiments described above, the determination may be made by a value represented by one bit (0 or 1), by a Boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).

[0103] (5) The processing procedures, sequences, flowcharts, etc., exemplified in the embodiments described above may be rearranged in order, as long as there is no inconsistency. For example, in the methods described herein, various step elements are presented using an exemplary order and are not limited to the specific order presented.

[0104] (6) Each function illustrated in Figures 2 and 3 is implemented by any combination of at least one of hardware and software. Furthermore, the method of implementing each function block is not particularly limited. That is, each function block may be implemented using one device that is physically or logically coupled, or it may be implemented using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired, wireless, etc.). A function block may also be implemented by combining the one or more devices with software.

[0105] (7) The programs illustrated in the embodiments described above should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether they are called software, firmware, middleware, microcode, hardware description languages ​​or by any other name.

[0106] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0107] (8) In each of the above-mentioned forms, the terms “system” and “network” shall be used interchangeably.

[0108] (9) The information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values ​​from a given value, or other corresponding information.

[0109] (10) In the embodiments described above, the portable device may be a Mobile Station (MS). A Mobile Station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or several other appropriate terms. In this disclosure, terms such as “mobile station,” “user terminal,” “user equipment (UE),” and “terminal” may be used interchangeably.

[0110] (11) In the embodiments described above, the terms “connected,” “coupled,” or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” with each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, “connection” may be read as “access.” As used in this disclosure, two elements may be considered to be “connected” or “coupled” with each other using at least one of one or more wires, cables, and printed electrical connections, and, in some non-limiting and non-exclusive examples, electromagnetic energy having wavelengths in the radio frequency domain, microwave domain, and optical (both visible and invisible) domain.

[0111] (12) In the embodiments described above, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on".

[0112] (13) The terms “determinating” and “deciding” as used in this disclosure may encompass a wide variety of actions. “Determinating” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, search, inquiry (for example, searching in a table, database or another data structure), and confirming. Furthermore, "judgment" and "decision" may include considering something as a "judgment" or "decision" based on actions such as receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, and access (e.g., accessing data in memory). Additionally, "judgment" and "decision" may include considering something as a "judgment" or "decision" based on actions such as resolving, selecting, choosing, establishing, and comparing. In short, "judgment" and "decision" may include considering something as a "judgment" or "decision" based on some action. Furthermore, "judgment (decision)" may be reinterpreted as "assuming," "expecting," or "considering."

[0113] (14) Where the terms “include,” “including,” and variations thereof are used in the embodiments described above, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to be exclusive OR.

[0114] (15) Where articles are added in translation, for example, a, an, and the in English, the disclosure may include the fact that the noun following these articles is plural.

[0115] (16) In this disclosure, the term “A and B are different” may mean “A and B are different from each other.” The term may also mean “A and B are each different from C.” Terms such as “separate” and “combine” may be interpreted in the same way as “different.”

[0116] (17) Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of certain information (e.g., notification that "it is X") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).

[0117] 1...Speech recognition system, 10...Speech recognition device, 20...User terminal, 30...Large-scale language model, 101...Communication device, 102...Storage device, 103...Processing device, 110...Speech data acquisition unit, 111...First speech recognition result acquisition unit, 112...Second speech recognition result acquisition unit, 113...Output unit, M1...First speech recognition model, M2...Second speech recognition model, N...Network, U...User.

Claims

1. A speech recognition device comprising: a first acquisition unit that acquires a first speech recognition result obtained by inputting speech data representing utterances divided into multiple pronunciation segments to a first speech recognition model, and a second speech recognition result obtained by inputting the speech data to a second speech recognition model; a second acquisition unit that acquires an integrated speech recognition result created based on the first speech recognition result and the second speech recognition result; and an output unit that outputs the integrated speech recognition result, wherein the integrated speech recognition result is created based on the consistency of the content of the utterances across the multiple pronunciation segments.

2. The speech recognition device according to claim 1, wherein the utterance includes a first pronunciation section and a second pronunciation section, and the integrated speech recognition result is created by selecting either the portion of the first speech recognition result corresponding to the first pronunciation section and the portion of the second speech recognition result corresponding to the first pronunciation section as the speech recognition result for the first pronunciation section, and selecting either the portion of the first speech recognition result corresponding to the second pronunciation section and the portion of the second speech recognition result corresponding to the second pronunciation section as the speech recognition result for the second pronunciation section.

3. The speech recognition device according to claim 1, wherein the second acquisition unit inputs a prompt, the first speech recognition result, and the second speech recognition result to a large-scale language model, acquires the integrated speech recognition result output from the large-scale language model, and the prompt instructs the system to select a more appropriate speech recognition result from the first speech recognition result and the second speech recognition result for each of the multiple speech segments based on the consistency of the content of the utterances across the segments.

4. The speech recognition device according to claim 1, wherein the second acquisition unit inputs the first speech recognition result and the second speech recognition result to a learning model that has been machine-learned using learning data including a plurality of learning speech recognition results and correct answer information indicating which of the plurality of learning speech recognition results is correct, and acquires the integrated speech recognition result output from the learning model.

5. The speech recognition device according to claim 1, wherein the integrated speech recognition result is created based on at least one of the following: the consistency of the content of the utterance across the plurality of pronunciation segments, the grammatical accuracy of the first speech recognition result and the second speech recognition result, and pre-registered words related to the content of the utterance.

6. The first speech recognition model selects the first speech recognition result from among a plurality of candidate first speech recognition results, and outputs the first speech recognition result along with a first acoustic score indicating the acoustic validity of the first speech recognition result, a first language score indicating the linguistic validity of the first speech recognition result, and a first confidence score indicating the likelihood of the first speech recognition result among the plurality of candidate first speech recognition results; The second speech recognition model selects the second speech recognition result from among a plurality of candidate second speech recognition results, and outputs the second speech recognition result along with a second acoustic score indicating the acoustic validity of the second speech recognition result, a second language score indicating the linguistic validity of the second speech recognition result, and a second confidence score indicating the likelihood of the second speech recognition result among the plurality of candidate second speech recognition results; The speech recognition device according to claim 1, wherein the integrated speech recognition result is created based on the consistency of the content of the utterance across the plurality of pronunciation segments, as well as at least one of the first acoustic score, the first language score, the first confidence level, the second acoustic score, the second language score, and the second confidence level.

7. A speech recognition method comprising: obtaining a first speech recognition result obtained by inputting speech data representing utterances divided into multiple pronunciation segments into a first speech recognition model using a computer; obtaining a second speech recognition result obtained by inputting the same speech data into a second speech recognition model; obtaining an integrated speech recognition result created based on the first speech recognition result and the second speech recognition result; outputting the integrated speech recognition result, wherein the integrated speech recognition result is created based on the consistency of the content of the utterances across the multiple pronunciation segments.

Citation Information

Patent Citations

  • Voice-recognition system, device, method and program

    WO2011121978A1

  • Language processing device, language processing method, learning method, and program

    WO2023248456A1