Information processing apparatus, information processing method, and computer program
By using a voice recognition model and a large-scale language model to correct transcription errors with context-aware proofreading, the method improves transcription accuracy and naturalness, addressing the limitations of conventional and end-to-end models.
Patent Information
- Application Number
- JP2024112965
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-07-12
AI Technical Summary
Conventional speech recognition engines struggle to produce accurate transcriptions that consider the context of multiple utterances, leading to transcription errors, and end-to-end models require large amounts of high-quality data to improve accuracy.
A method involving a voice recognition model followed by a large-scale language model to correct transcription errors by outputting proofreading information that matches transcription data, with specific instructions to ensure similar pronunciation and context.
Enhances transcription accuracy by reducing proofreading omissions and improving the naturalness of transcriptions, especially for technical terms and homonyms.
Smart Images

Figure 2026011949000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a computer program that perform processing related to conversion of audio data into text data. [Background technology]
[0002] In recent years, transcription tools that convert the contents of audio data into text have been attracting attention, and are expected to be used in a variety of fields, such as creating minutes at meetings and recording phone calls at call centers.
[0003] Conventional speech recognition engines for transcribing speech data combine an acoustic model that converts speech features into a phoneme sequence, a pronunciation model that converts the phoneme sequence into word candidates, and a language model that predicts the sequence of word candidates and generates text. However, these conventional speech recognition engines focus on the problem of transcribing each utterance independently, and are unable to transcribe while taking into account the context of multiple utterances contained in the speech data. As a result, they are unable to produce natural-sounding transcriptions that take into account the context of the speech data, and transcription accuracy has been lacking.
[0004] In contrast to this, in recent years, so-called end-to-end transcription has been attracting attention, which uses a speech recognition engine that generates text from speech features using a single model, for example, a transformer-based model (for example, Patent Document 1). An end-to-end speech recognition engine that uses a transformer-based model makes it possible to produce more natural transcription that takes into account long contexts. Known examples of speech recognition engines that use a transformer-based speech recognition model include "Whisper" by OpenAI and "Wav2Vec" by Meta.
[0005] Although end-to-end speech recognition engines using transformer-based models significantly improve transcription accuracy compared to conventional speech recognition engines, the inventors' verification has shown that transcription using currently available end-to-end speech recognition models still contains many transcription errors. However, improving the transcription accuracy of end-to-end speech recognition models requires a huge amount of high-quality text and speech data sets, and it is not easy to improve the transcription accuracy of end-to-end speech recognition models (e.g., Patent Document 2). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Publication No. 2022-010410 [Patent Document 2] Japanese Patent Application Publication No. 2024-010464 Summary of the Invention [Problem to be solved by the invention]
[0007] The present invention has been made in consideration of the above-mentioned problems of the conventional technology, and in one aspect, it is an object of the present invention to provide an information processing device that enables more accurate transcription of voice data into text data. [Means for solving the problem]
[0008] In the course of their extensive research efforts to solve the above-mentioned problems, the inventors came up with the idea that, as an attempt distinct from attempts to improve the speech recognition accuracy of speech recognition models installed in speech recognition engines, it might be possible to correct transcription errors contained in the transcription data output by a speech recognition model with high accuracy by processing the transcription data output by the speech recognition model with a large-scale language model.
[0009] However, when large-scale language models are instructed to proofread transcription data and then directly output the proofread transcription data, the output text contains many obvious proofreading errors, and the results are not necessarily satisfactory in terms of improving transcription accuracy. This is thought to be because large-scale language models have difficulty remembering and continuing to apply the tasks initially instructed to them to long input data.
[0010] Therefore, in the course of further research efforts, the inventors discovered that, rather than providing transcription data to a large-scale language model and directly outputting proofread transcription data, it is possible to reduce proofreading omissions by the large-scale language model and obtain highly accurate transcription data by outputting proofreading information that matches the proofreading points contained in the transcription data, or more preferably, the proofreading points contained in the transcription data that contain transcription errors, with proofreading suggestions for those proofreading points.
[0011] That is, in one aspect, the present invention provides: a means for inputting voice data into a voice recognition model that outputs transcription data in response to input of voice data, and obtaining transcription data of the voice data that is output in response to the input; a means for inputting an input sentence containing an instruction for outputting proofreading information associating a pre-proofreading text containing transcription errors and the proofreading text into a large-scale language model for the acquired transcription data, and obtaining a first proofreading proposal including proofreading information associating a pre-proofreading text and the proofreading text that is output in response to the input; A means for proofreading the transcription data based on the obtained first proofreading proposal and generating proofread transcription data; and means for outputting the generated proofread transcription data; The above-mentioned problems are solved by providing an information processing device comprising:
[0012] In a preferred embodiment, the input sentence to be input to the large-scale language model may further include instructions for outputting the pre-proofreading text and the proofreading text as proofreading information only if the pronunciation of the pre-proofreading text and the proofreading text is the same or similar. Such instructions may be, for example, instructions to output the pre-proofreading text and the proofreading text as proofreading information only if the pronunciation of the pre-proofreading text and the proofreading text is the same or similar.
[0013] The similarity in pronunciation between a pre-proofreading text and its proofread text may be determined independently by the large-scale language model, or a definition or condition may be set. For example, a definition or condition may be set that the pronunciations of the pre-proofreading text and the proofreading text are similar if the number of different pronunciations between the pre-proofreading text and the proofreading text is two or less, more preferably one or less, or if the number of different pronunciations between the pre-proofreading text and the proofreading text is 33% or less, more preferably 25% or less, and even more preferably 20% or less of the total pronunciations of the pre-proofreading text or the proofreading text. When transcription data is provided to a large-scale language model and an attempt is made to output a pre-proofreading text and a proofreading text containing transcription errors, the meaning of the sentence may be prioritized, resulting in a proofreading text whose pronunciation is significantly different from the pre-proofreading text. If the pronunciations of the pre-proofreading text and the proofreading text differ significantly, this may be undesirable, as it may go beyond the scope of proofreading a transcription error. If the input sentence input to the large-scale language model includes the above instruction, proofreading information including a more appropriate proofreading text may be output as a proofreading proposal for proofreading the transcription result.
[0014] In a preferred embodiment, when executing the means for obtaining the first proofreading proposal, a word list that associates words with their pronunciations may be referenced. That is, in a preferred embodiment, the input sentence input to the large-scale language model may further include an instruction to output the proofreading information by referring to a word list that associates words with their pronunciations. Such an instruction may be, for example, an instruction to output the proofreading information by referring to a word list that associates words with their pronunciations.
[0015] The large-scale language model can execute a task according to an instruction to output the proofreading information by referring to a word list that associates words with their pronunciations. However, more preferably, more specific instructions are provided. For example, the instruction to output proofreading information by referring to a word list may be an instruction to output proofreading information including text and words included in the transcription data as pre-proofreading text and proofreading text, respectively, only if the pronunciation of the text is the same as or similar to the pronunciation of a word included in the word list. Even more preferably, the instruction may be an instruction to output proofreading information including text and words included in the transcription data, respectively, associated with pre-proofreading text and proofreading text, if the pronunciation of the text is the same as or similar to the pronunciation of a word included in the word list and the meaning of a sentence including the text included in the transcription data is not unnatural when the text is replaced with the word. In addition, in a sentence including the text included in the transcription data, the instruction to consider whether the meaning of the sentence obtained by replacing the text with the word is not unnatural may be explicitly stated or implicit in the input sentence. Furthermore, when the meaning of the sentence is not unnatural, the meaning of the sentence is considered to be natural.
[0016] Technical terms, new words, personal names, homonyms, and the like used in specific fields or specific groups are difficult for speech recognition engines or large-scale language models to understand, and it is not easy to transcribe them correctly. On the other hand, when an information processing device according to one aspect of the present invention has the above-described configuration using a word list, by referring to a pre-recorded word list, it becomes possible to more accurately read and write technical terms, new words, personal names, homonyms, and the like that have not been learned by a large-scale language model or that are difficult for a large-scale language model to understand or output, or to proofread transcription data for specific terms as desired by the user.
[0017] In addition to the words and their pronunciations, the word list may contain other information, such as the part of speech of the word, the context in which the word is used, the field in which the word is used, etc. If the word list contains such other information, it goes without saying that the input sentence input to the large-scale language model may include instructions for outputting the proofreading information by referring to the other information.
[0018] In a preferred embodiment, the information processing device according to an aspect of the present invention may further include means for acquiring a second proofreading suggestion by selecting proofreading information from the first proofreading suggestion that has the same or similar pronunciation as the unproofread text and the proofread text. In this case, the information processing device according to an aspect of the present invention proofreads the transcription data based on the acquired second proofreading suggestion, thereby generating proofread transcription data.
[0019] When transcription data is provided to a large-scale language model and an attempt is made to output a pre-proofreading text and its proofread text containing transcription errors, the proofread text may have a pronunciation significantly different from that of the pre-proofreading text, which may exceed the scope of a transcription error. By issuing an instruction to output the pre-proofreading text and the proofread text as proofreading information only if the pre-proofreading text and the proofread text have the same or similar pronunciation in an input sentence input to a large-scale language model, the output of proofread text with a pronunciation significantly different from that of the pre-proofreading text is suppressed, but it is not easy to completely eliminate such output. In response to this, an information processing device according to one aspect of the present invention further includes means for acquiring a second proofreading proposal by selecting proofreading information included in the first proofreading proposal that has the same or similar pronunciation as the pre-proofreading text. When inspecting the proofreading proposal (first proofreading proposal) output by the large-scale language model, proofreading information with the same or similar pronunciation as the pre-proofreading text and the proofread text can be selected, thereby achieving more accurate transcription.
[0020] There are no particular restrictions on the specific implementation method of the above means for obtaining a second proofreading proposal by selecting proofreading information from the first proofreading proposal that has the same or similar pronunciation as the pre-proofreading text and the proofread text, and the means may be implemented using a rule-based algorithm, a model such as a large-scale language model or AI, or a combination thereof.
[0021] When selecting using a rule-based algorithm, for example, if the number of sounds with different readings in the pre-proofreading text and the proofreading text is two or less, more preferably one or less, or if the number of sounds with different readings is 33% or less of the total number of sounds that make up the reading of the pre-proofreading text or the proofreading text, more preferably 25% or less, and even more preferably 20% or less, it can be determined that the readings of the pre-proofreading text and the proofreading text are similar and extracted.
[0022] On the other hand, when selection is performed using a large-scale language model, for example, an input sentence including instructions for selecting and outputting only the proofreading information contained in the first proofreading draft, for which the readings of the unproofreaded text and the proofread text are the same or similar, can be input to the large-scale language model, and the output can be obtained as the second proofreading draft. Here, the definition or condition for the same or similar readings can be given in more detail for the input sentence input to the large-scale language model. For example, a definition or condition can be set such that the readings of the two texts are similar when the number of sounds whose readings differ between the reading of the unproofreaded text and the reading of the proofread text is two or less, more preferably one or less, or when the number of sounds whose readings differ is 33% or less, more preferably 25% or less, and even more preferably 20% or less of the number of sounds constituting the reading of the unproofreaded text or the proofread text.
[0023] The pre-proofreading text and the proofread text may be output in any unit. For example, the pre-proofreading text and the proofread text may be output in sentence or utterance units, or in word or phrase units. Furthermore, the proofreading information may include other elements in addition to the pre-proofreading text and the proofread text. Examples of other elements included in the proofreading information include, for example, information that can identify the proofreading point, i.e., the location of the pre-proofreading text, in the transcription data. Such information may be any information as long as it can identify the location of the pre-proofreading text in the transcription data. For example, a sentence or utterance containing the pre-proofreading text, and, if necessary, one or more utterances or sentences before or after it, may be output. Furthermore, the location of the pre-proofreading text in the transcription data may be identified by the number of lines, characters, bytes, paragraph number, or the order or number of the sentence or utterance in the entire transcription data.
[0024] In a preferred embodiment, an information processing device according to an aspect of the present invention executes the means for dividing the transcription data into predetermined units and acquiring, for each of the predetermined units, a first proofreading proposal including proofreading information associating a pre-proofreading text with the proofread text. That is, in a preferred embodiment, the means for acquiring a first proofreading proposal divides the acquired transcription data into predetermined units and, for each of the transcription data divided into the predetermined units, inputs an input sentence including an instruction to output a first proofreading proposal including pre-proofreading text containing transcription errors and proofreading information associating the proofread text to a large-scale language model, thereby acquiring a first proofreading proposal including proofreading information associating a pre-proofreading text that is output in response to the input with the proofread text.
[0025] There are no particular restrictions on the specified unit, but it is preferable that the unit contains at least two or more utterances or sentences, and more preferably, it is a unit containing 2 to 1000 utterances or sentences, preferably, a unit containing 5 to 500 utterances or sentences, and more preferably, a unit containing 10 to 300 utterances or sentences.
[0026] A sentence refers to a unit separated by a period. On the other hand, an utterance refers to a speech section included in audio data. Audio data consists of speech sections and silent sections (or non-speech sections), and each speech section is separated by a silent section. Those skilled in the art can easily recognize speech sections included in audio data. Alternatively, a speech recognition model that can distinguish and output speech sections and silent sections of audio data, such as "Whisper" by OpenAI, may be used. It is also possible to detect speech sections included in audio data using voice activity detection technology. In this case, audio data can be processed using speech activity detection before being transcribed using a speech recognition model.
[0027] According to the findings of the inventors, if the input transcription data is too short, the context cannot be understood, resulting in oversight of proofreading sections, i.e., missing parts of the transcription containing errors. On the other hand, many large-scale language models are not good at understanding long tasks or at continuing to apply the tasks initially instructed to long input data without forgetting them. Therefore, if the input transcription data is too long, missing parts of the transcription become noticeable. In response to this, by dividing the transcription data into predetermined units, inputting the transcription data for each predetermined unit, and outputting a first proofreading proposal containing proofreading information that associates the pre-proofreading text with the proofreading text, it becomes possible to proofread the transcription data more thoroughly.
[0028] In another aspect, the present invention also provides, for example, an information processing method realized by the information processing device according to one aspect of the present invention, and a computer program for causing a computer to function as the information processing device. [Effects of the Invention]
[0029] According to the present invention, it is possible to transcribe voice data into text data with high accuracy. [Brief explanation of the drawings]
[0030] [Figure 1] 1 is a diagram illustrating an example of a configuration of an information processing system including an information processing device according to an aspect of the present invention. [Figure 2] 1 is a diagram illustrating an example of a configuration of an embodiment of an information processing device according to an aspect of the present invention. [Figure 3] 1 is a diagram showing an example of a flow of information processing by an information processing system configured to include an information processing device according to one aspect of the present invention. [Figure 4] FIG. 10 is a diagram showing an example of transcription data input to a large-scale language model, and proofreading information that associates pre-proofreading text with proofreading text and is output by the large-scale language model in accordance with the transcription data. [Figure 5] FIG. 10 is a diagram showing an example of transcription data input to a large-scale language model, and proofreading information that associates pre-proofreading text with proofreading text and is output by the large-scale language model in accordance with the transcription data. [Figure 6] 1 is a diagram showing an example of a flow of information processing by an information processing system configured to include an information processing device according to one aspect of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the following embodiments are merely examples of embodiments of the present invention and are not intended to limit the technical scope of the present invention.
[0032] [1. Embodiment 1] [1.1. Configuration of First Embodiment] First, the configuration of the information processing device 100 according to the first embodiment will be described.
[0033] The information processing device 100 is an information processing device that transcribes voice data and outputs the transcription data.
[0034] "Audio data" is data that includes audio. The audio data may be basically any type as long as it can be input to a speech recognition model that outputs transcription data in response to input of audio data, as described below. For example, the audio data may be provided as a file in an audio data format. The audio data format file may be, for example, an MP3, M4A, or WAV format file, but is not limited to these. The audio data may also be provided as a file in a format that includes data other than audio, such as video. The file format that includes data other than audio together with audio data may be, for example, an MP4 or WebM format file, but is not limited to these.
[0035] On the other hand, "transcription data" is data containing text obtained by transcribing (converting) the voice included in the voice data. The format of the transcription data output by the information processing device 100 may basically be any format, but as an example, it may be a file in TXT or CSV format.
[0036] As shown in FIG. 1, an information processing device 100 is communicably connected via a network N to an external information processing device 200, an external information processing device 300, an external storage 400, and a user terminal 500, which is a terminal device used by a user.
[0037] The information processing device 200 is an information processing device equipped with a voice recognition model, and is, for example, a cloud server. When the information processing device 200 receives a request to transcribe voice data from the information processing device 100, the information processing device 200 transcribes the voice data using the voice recognition model and converts the voice data into text data (transcription data).
[0038] The information processing device 300 is an information processing device equipped with a large-scale language model, such as a cloud server. When the information processing device 300 receives a request to proofread transcription data from the information processing device 100, the information processing device 300 inputs an input sentence including an instruction to cause the large-scale language model to output a proofreading proposal corresponding to the transcription data, and causes the large-scale language model to generate a proofreading proposal corresponding to the transcription data.
[0039] The external storage 400 is a storage device, such as an online storage, in which various data such as voice data provided from the user terminal 500, transcription data generated in the information processing device 200, proofreading proposals generated in the information processing device 300, and proofread transcription data are stored as needed.
[0040] 1, the information processing devices 100, 200, and 300 are depicted as if they were a single physical information processing device, but the information processing device 100 does not necessarily have to be a single physical information processing device. For example, the information processing device 100 may be a distributed information processing device, such as a distributed server. This also applies to FIG. 3, which will be described later.
[0041] 1, external storage 400 is depicted as being physically one storage device, but it does not necessarily have to be one physical storage device. For example, it may be a distributed storage device, such as a distributed storage. This also applies to FIG. 3, which will be described later.
[0042] Next, the configuration of the information processing device 100 will be described in more detail.
[0043] As shown in FIG. 2, the information processing device 100 conceptually includes a communication unit 110, a storage unit 120, and a processing unit .
[0044] <Communication Unit 110> The communication unit 110 is a means for transmitting and / or receiving various types of information to and from various devices via the network N. The communication unit 110 may be configured to include, for example, a network interface card (NIC).
[0045] <Storage section 120> The storage unit 120 is a means for storing various types of information, and may be configured to include, for example, memories including RAM (Random Access Memory) such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory), flash memory such as NOR flash memory and NAND flash memory, and / or a hard disk.
[0046] <Processing Unit 130> The processing unit 130 is a means for executing various processes, and may include a processor such as a central processing unit (CPU), a micro processing unit (MPU), or a graphics processing unit (GPU).
[0047] Functionally, the processing unit 130 is configured to include a transcription request acquisition unit 131, a transcription data acquisition unit 132, a proofreading proposal acquisition unit 133, a transcription data proofreading unit 134, and a proofread transcription data output unit 135. Each unit will be described below.
[0048] <Transcription request acquisition unit 131> The transcription request acquisition unit 131 is a means for acquiring a transcription request for audio data.
[0049] The transcription request acquisition unit 131 receives a transcription request for audio data from the user terminal 500 via the communication unit 110, for example, to acquire the transcription request for audio data.
[0050] The transcription request is accompanied by information identifying the audio data to be transcribed. The information identifying the audio data may be, for example, a file path of the audio data. For example, the user terminal 500 may upload the audio data to the external storage 400, and provide the file path of the audio data uploaded to the external storage 400 to the transcription request acquisition unit 131 of the information processing device 100 as information identifying the audio data.
[0051] <Transcription Data Acquisition Unit 132> The transcription data acquisition unit 132 is a means for acquiring transcription data of the voice data specified in the transcription request. For example, the transcription data acquisition unit 132 directly or indirectly inputs the voice data specified in the transcription request into a voice recognition model that outputs transcription data in response to input of voice data, and acquires the transcription data that is output in response to the input.
[0052] In this example, a speech recognition model that outputs transcription data in response to input of speech data is provided in an information processing device 200, which is an external information processing device separate from the information processing device 100. The transcription data acquisition unit 132, for example, transmits a transcription request for speech data identified in the transcription request acquired by the transcription request acquisition unit 131 to the information processing device 200. Upon receiving the transcription request from the transcription data acquisition unit 132, the information processing device 200 inputs the speech data identified in the transcription request into a speech recognition model that outputs transcription data in response to input of speech data, acquires text data output in response to the input as transcription data, and transmits this to the information processing device 100. The transcription data acquisition unit 132 of the information processing device 100 receives the transcription data via the communication unit 110, thereby acquiring the transcription data of the speech data.
[0053] In one embodiment, the information processing device 200 may be configured to store transcription data generated by a speech recognition model in the external storage 400. In this case, the transcription data acquisition unit 132 of the information processing device 100 may be configured to acquire the transcription data by reading the transcription data stored in the external storage 400. In any case, it is sufficient that the information processing device 100 is able to acquire the transcription data generated by a speech recognition model.
[0054] As described above, the information processing device 200 having a speech recognition model that outputs transcription data in response to input of speech data may be a single physical information processing device, or may be a distributed information processing device, for example, a distributed server. In this example, the speech recognition model is provided in the information processing device 200, which is an external information processing device, but in another aspect, the information processing device 100 may be configured to have the speech recognition model.
[0055] The speech recognition model that outputs transcription data in response to input speech data may be essentially any type, as long as it is a trained model that can transcribe speech contained in input speech data and generate text data corresponding to the content of the speech data, and may be, for example, a transformer-based speech recognition model. Examples of such speech recognition models include, but are not limited to, "Whisper" provided by OpenAI and "Speech-to-Text" provided by Alphabet.
[0056] Furthermore, a speech recognition model that outputs transcription data in response to input of speech data may be fine-tuned to generate transcription data, etc. For example, the model may be fine-tuned to understand technical terms according to a field, to understand terms used in a specific organization such as a company, to understand a specific language, and / or to understand the habits of speakers in a specific organization such as a company.
[0057] The voice data may be pre-processed by an appropriate method before being transcribed using a voice recognition model. For example, the file format of the voice data may be converted to another file format, or noise contained in the voice data may be removed. The pre-processing of the voice data may be performed by the information processing device 100 or by another information processing device, for example, the information processing device 200.
[0058] <Proofreading proposal acquisition unit 133> The proofreading proposal acquisition unit 133 is a means for acquiring a proofreading proposal (first proofreading proposal) for proofreading transcription data. For example, for the transcription data acquired by the transcription data acquisition unit 132, the proofreading proposal acquisition unit 133 directly or indirectly inputs an input sentence including an instruction for outputting a proofreading proposal corresponding to the transcription data into a large-scale language model, and acquires a proofreading proposal (first proofreading proposal) for the transcription data that is output in response to the input.
[0059] In this example, the large-scale language model is provided in an information processing device 300, which is an external information processing device separate from the information processing device 100. That is, the proofreading proposal acquisition unit 133 transmits an input sentence including an instruction for generating a proofreading proposal (first proofreading proposal) corresponding to the transcription data acquired by the transcription data acquisition unit 132 to the information processing device 300 via the communication unit 110. As a result, the input sentence including an instruction for generating a proofreading proposal (first proofreading proposal) for proofreading the transcription data acquired by the transcription data acquisition unit 132 is input to the large-scale language model provided in the information processing device 300, causing the large-scale language model to generate a proofreading proposal. The information processing device 300 transmits the proofreading proposal (first proofreading proposal) generated by the large-scale language model to the information processing device 100, and the proofreading proposal acquisition unit 133 of the information processing device 100 receives the proofreading proposal via the communication unit 110.
[0060] In one embodiment, the information processing device 300 may be configured to store a proofreading proposal (first proofreading proposal) generated by a large-scale language model in the external storage 400. In this case, the external storage 400 may be configured to store, for example, transcription data and the proofreading proposal (first proofreading proposal) generated by the large-scale language model in association with each other. In this way, the proofreading proposal acquisition unit 133 of the information processing device 100 can acquire a proofreading proposal corresponding to the transcription data by reading out a proofreading proposal stored in the external storage 400 in association with the transcription data. In any case, it is sufficient that the information processing device 100 can acquire a proofreading proposal (first proofreading proposal) generated by a large-scale language model and identify the correspondence between each proofreading proposal and the transcription data.
[0061] As described above, the information processing device 300, which includes a large-scale language model that outputs a proofreading proposal (first proofreading proposal) according to transcription data, may be a single physical information processing device, or may be a distributed information processing device, for example, a distributed server. In this example, the large-scale language model is provided in the information processing device 300, which is an external information processing device, but in another aspect, the information processing device 100 may be configured to include the large-scale language model.
[0062] The large-scale language model that generates a proofreading proposal (first proofreading proposal) based on the transcription data may be, for example, a Transformer-based large-scale language model. Examples of the Transformer-based large-scale language model include, but are not limited to, "GPT" provided by OpenAI, Inc., "Gemini" provided by Google DeepMind, Inc., "Claude" provided by Anthropic, Inc., and "tsuzumi" provided by Nippon Telegraph and Telephone Corporation.
[0063] Additionally, the large-scale language model may be fine-tuned as appropriate, for example, to understand technical terms specific to a field, to understand terms used within a specific organization such as a company, to understand a specific language, and / or to understand the habits of speakers within a specific organization such as a company.
[0064] An input sentence including instructions for outputting a proofreading proposal (first proofreading proposal) based on the transcription data is also referred to as a prompt including instructions for outputting a (first proofreading proposal) based on the transcription data. When a large-scale language model receives an input of an input sentence or a prompt including a predetermined instruction, the large-scale language model executes a task instructed by the instruction included in the input sentence or the prompt.
[0065] In a preferred embodiment, a proofreading proposal (first proofreading proposal) based on transcription data may include proofreading information that associates pre-proofreading text, which includes transcription errors and is included in the transcription data, with proofread text of the pre-proofreading text. The proofreading proposal (first proofreading proposal) may include one or more pieces of proofreading information. For example, if there are two proofreading points, a total of two pieces of proofreading information corresponding to the two proofreading points may be included.
[0066] Unproofread text is text that contains transcription errors in the transcription data, and may be, for example, text in which the meaning of a sentence or utterance containing the text is unclear or unnatural as is. The transcription data output by the speech recognition model may contain a large amount of text whose meaning is unclear or unnatural, for example, because the characters representing the text are incorrect even if the sound is correct. A definition of unproofread text may be explicitly provided for the input sentence to be input to the large-scale language model.
[0067] On the other hand, the proofread text is text that has been proofread to remove transcription errors contained in the pre-proofread text. The proofread text may be, for example, text in which transcription errors contained in the pre-proofread text have been corrected so that the pronunciation of the pre-proofread text and the proofread text is the same or similar. Furthermore, the proofread text may be text in which the meaning of a sentence or utterance containing the pre-proofread text included in the transcription data is not unnatural when the pre-proofread text is replaced with the proofread text. In this example, the proofread text may have a pronunciation similar to that of the pre-proofread text. However, in a preferred embodiment, the proofread text may only include text that has the same pronunciation as the pre-proofread text. A definition of the proofread text may be explicitly provided for an input sentence to be input to a large-scale language model. The same or similar pronunciation has been described above, so a detailed description will be omitted.
[0068] According to the findings of the present inventors, if transcription data is input into a large-scale language model and the transcription data is directly proofread, i.e., if the large-scale language model is directly output as proofread transcription data, many transcription errors will remain unproofread. In contrast, by giving the large-scale language model a more specific task of outputting a proofreading proposal (first proofreading proposal) that includes proofreading information that matches pre-proofreading text containing transcription errors with the proofread text, it becomes possible to perform proofreading with fewer omissions.
[0069] The input sentence for generating a proofreading suggestion based on the transcription data may further include other information. Examples of other information that may be included in the input sentence include various types of information such as the situation or context in which the audio data was recorded, examples of input and output, points to note when performing a task, etc. Examples of the situation or context in which the audio data was recorded may include, but are not limited to, information about the field to which the content of the audio data relates, such as law, medicine, economics, sports, entertainment, etc., the type of speaker, such as elementary school student, high school student, university student, or working adult, and the language of the audio data, such as Japanese, English, or Chinese.
[0070] An input sentence including instructions for generating proofreading information corresponding to the transcription data can be generated, for example, by using a predetermined prompt template. Specifically, the transcription data to be proofread is entered into the predetermined prompt template, and / or a file path specifying the transcription data is entered into the predetermined prompt template, and this is used as the input sentence. The generation of the input sentence including instructions for generating a proofreading proposal corresponding to the transcription data may be configured to be performed by the information processing device 100 or an external information processing device, for example, the information processing device 300. An AI agent may also be used to generate the input sentence. There is no particular limitation on the storage location of the prompt template; it may be stored in the storage unit 120 of the information processing device 100 or in an external storage device, for example, the external storage 400.
[0071] <Transcription Data Proofreading Department 134> The transcription data proofreading unit 134 proofreads the transcription data based on the first proofreading plan acquired by the proofreading plan acquisition unit 133.
[0072] For example, if the first proofreading proposal acquired by the proofreading proposal acquisition unit 133 includes proofreading information that associates pre-proofreading text, which includes a transcription error and is included in the transcription data, with proofread text of the pre-proofreading text, the transcription data proofreading unit 134 executes a process of replacing each pre-proofreading text included in the transcription data with proofread text corresponding to the pre-proofreading text, thereby generating proofread transcription data.
[0073] <Proofread Transcription Data Output Unit 135> The proofread transcription data output unit 135 is a means for outputting the proofread transcription data generated by the transcription data proofreading unit 134.
[0074] For example, the proofread transcription data output unit 135 may transmit the proofread transcription data to the user terminal 500 via the communication unit 110. Alternatively, the proofread transcription data may be stored in the external storage 400 in association with the original audio data.
[0075] [1.2. Example of information processing according to embodiment 1] Next, a description will be given of an example of information processing by the information processing device 100. A flow of an example of information processing by the information processing device 100 is shown in FIG.
[0076] The user operates the user terminal 500 to transmit voice data and a request for transcription of the voice data to the information processing device 100 (step S1). For example, a predetermined display screen provided by the voice recognition system is displayed on the screen of the user terminal 500, and the user uploads the voice data to be transcribed to the external storage 400 via the display screen, and then presses a transcription execution button displayed on the display screen to transmit a transcription request for the voice data to the information processing device 100.
[0077] When the information processing device 100 receives a transcription request for the voice data from the user terminal 500, it transmits the transcription request for the voice data to the information processing device 200 that includes a voice recognition model (step S2). The transcription request may be a transcription request that specifies the voice data to be transcribed, and may be, for example, a request that specifies the file path of the voice data to be transcribed.
[0078] When the information processing device 200, which is equipped with a voice recognition model, receives a transcription request for the voice data from the information processing device 100, it inputs the voice data into the voice recognition model. The voice recognition model transcribes the voice included in the voice data and generates text data. This generates transcription data (step S3).
[0079] The information processing device 200 transmits the generated transcription data (text data) to the information processing device 100 (step S4).
[0080] In this example, the speech recognition model provided in the information processing device 200 is generated based on the speech data specified by the information processing device 100. sesame is located in the Democratic Republic of the Congo in central Africa. year is. sesame teeth MiragonThe following description will be given assuming that the information processing device 200 generates transcription data including the text "The city was destroyed by volcanic lava, and most of the streets, especially the center of town, were buried." and transmits this to the information processing device 100 (Figure 4).
[0081] When the information processing device 100 receives transcription data from the information processing device 200 equipped with a speech recognition model, the information processing device 100 transmits the transcription data and a request for proofreading the transcription data to the information processing device 300 equipped with a large-scale language model in order to proofread the transcription data (step S5).
[0082] The proofreading request for the transcription data sent to the information processing device 300 includes an input sentence including an instruction to output proofreading information corresponding to the transcription data. In this example, the input sentence including an instruction to output a proofreading proposal corresponding to the transcription data is a prompt including an instruction to output a first proofreading proposal including proofreading information that associates the pre-proofreading text included in the transcription data with the proofread text of the pre-proofreading text, specific examples of input and output, and points to note when performing the task.
[0083] Furthermore, in this example, an input sentence including an instruction for generating a proofreading suggestion based on the transcription data is generated by using a prompt template stored in the storage unit 120 of the information processing device 100. Specifically, the information processing device 100 writes the file path of the transcription data to be proofread in the prompt template stored in the storage unit 120, and generates the input sentence. Note that the generation of the input sentence for generating a proofreading suggestion based on the transcription data does not necessarily have to be performed by the information processing device 100, and may be performed by, for example, the information processing device 300. In this case, the proofreading request sent to the information processing device 300 only needs to be accompanied by information identifying the transcription data, for example, the file path of the transcription data.
[0084] When the information processing device 300 equipped with a large-scale language model receives from the information processing device 100 a proofreading request for the transcription data, which includes an input sentence including an instruction to output a proofreading suggestion based on the transcription data, the information processing device 300 inputs the input sentence into the large-scale language model and causes the large-scale language model to generate a proofreading suggestion based on the transcription data (step S6).
[0085] In this example, the output of the large-scale language model is shown in Figure 4, and the " sesame teeth Miragon Volcano For the sentence "The city was destroyed by lava, and most of the streets, especially the center of town, were buried.", the proofreading information including the combination of "sesame" as the pre-proofreading text and "sesame" as the post-proofreading text is output. sesame teeth Miragon Volcano For the sentence "The city was destroyed by lava, burying most of the streets, especially the center of town," proofreading information is output that includes a combination of "Miragon volcano" as the pre-proofreading text and "Niragongo volcano" as the proofreading text. In this way, the output of the large-scale language model includes proofreading results that associate the pre-proofreading text with the proofreading text. In this example, as information identifying the proofreading part, a sentence including the pre-proofreading text is output in association with the pre-proofreading text and the proofreading text.
[0086] The information processing device 300 transmits the proofreading information generated by the large-scale language model to the information processing device 100 as a first proofreading proposal (step S7).
[0087] When the information processing device 100 receives the first proofreading proposal from the information processing device 300 equipped with the large-scale language model, the information processing device 100 proofreads the transcription data based on the received first proofreading proposal (step S8).
[0088] For example, in this example, the first proofreading proposal includes a pre-proofreading text, a proofread text, and a sentence including the pre-proofreading text as information for identifying the proofreading portion, in association with each other. Therefore, the information processing device 100 generates proofread transcription data by replacing the pre-proofreading text with the proofread text in the sentence included in the transcription data. Specifically, the proofread transcription data includes " sesame teeth Nyiragongo The city was destroyed by volcanic lava, burying most of the streets, especially the town center."
[0089] The information processing device 100 outputs the generated proofread transcription data to the user terminal 500 (step S9), thereby enabling the user to obtain the transcription data of the uploaded voice data.
[0090] [1.3. Example of information processing according to embodiment 1] Next, an example of information processing in which the information processing device 100 uses a word list that associates words with their readings will be described.
[0091] The steps of the user terminal 500 sending a transcription request for audio data to the information processing device 100 (step S1); the information processing device 100 receiving the transcription request for audio data sending the transcription request for the audio data to the information processing device 200 equipped with a voice recognition model (step S2); the information processing device 200 equipped with the voice recognition model transcribing the voice contained in the audio data and generating transcription data (step S3); and the step of the information processing device 200 sending the generated transcription data to the information processing device 100 (step S4) are as described above, so further explanation will be omitted.
[0092] On the other hand, in this example, the speech recognition model recognizes "Tokyo KnowledgeThe Special Investigation Unit is an investigative agency that investigates political corruption, tax evasion, economic crimes, etc. Through this investigation, it was discovered that Politician A frequently meets with Politician B. Knowledge The following explanation will be given assuming that transcription data including the text "was obtained." is generated, and that a word list is stored in external storage 400, and that the word "chiken" and its pronunciation "chiken" are registered in association with each other in the word list (Figure 5).
[0093] When the information processing device 100 receives transcription data from the information processing device 200 having a speech recognition model, the information processing device 100 sends a transcription data proofreading request for the transcription data to the information processing device 300 having a large-scale language model in order to proofread the transcription data (step S5).
[0094] The request for proofreading the transcription data sent to the information processing device 300 includes an input sentence including an instruction to output a proofreading proposal corresponding to the transcription data.
[0095] In this example, the input sentence including an instruction to output a proofreading suggestion based on the transcription data is a prompt including an instruction to output a proofreading suggestion based on the transcription data by referencing the word list stored in external storage 400. In addition to an instruction to output the file path of the transcription data and proofreading information that associates the pre-proofreading text included in the transcription data with the proofreading text, the prompt further includes the file path of the word list stored in external storage 400 and an instruction to output the proofreading information by referencing the word list.
[0096] In this example, the instruction to output the proofreading information by referring to the word list is an instruction to output proofreading information in which the text and the word are associated as pre-proofreading text and proofreading text, respectively, when the reading of the text included in the transcription data is the same as or similar to the reading of a word included in the word list, the meaning of a sentence including the text included in the transcription data is unnatural, and the meaning of a sentence including the text included in the transcription data is not unnatural even if the text is replaced with the word.
[0097] In this example, the text "Tokyo" included in the transcription data output by the speech recognition model Knowledge The Special Investigation Unit is an investigative agency that investigates political corruption, tax evasion, economic crimes, etc., and the text "Kenken" has the same pronunciation as the word "Chiken" in the word list. Knowledge The sentence "The Special Investigation Unit is an investigative agency that investigates political corruption, tax evasion, economic crimes, etc." has an unnatural meaning as it is, and the sentence "Tokyo District prosecutor's office The Special Investigation Unit is an investigative agency that investigates political corruption, tax evasion, economic crimes, etc.,'' is not an unnatural meaning.
[0098] On the other hand, the transcription data output by the speech recognition model included the line, "This investigation has revealed that politician A frequently meets with politician B." Knowledge The sentence "I got it." contains the text "Knowledge" which has the same reading as the word "Chiken" in the word list. However, "This investigation has revealed that politician A frequently meets with B." Knowledge The meaning of the sentence "I got it." is not unnatural.
[0099] Therefore, in this example, the large-scale language model is used to KnowledgeRegarding the sentence "The Special Investigation Unit is an investigative agency that investigates political corruption, tax evasion, economic crimes, etc.", the character string "knowledge" is included in the pre-proofreading text as follows: District prosecutor's office " will be output as proofread text. On the other hand, "This investigation has revealed that politician A frequently meets with B. Knowledge For the sentence "I got it.", neither the pre-proofreading text nor the proofreading text is output.
[0100] By providing the above-described predetermined instructions, the large-scale language model functions as: a means for determining whether a sentence included in the transcription data contains text (character strings) whose pronunciation is identical or similar to that of a word included in the word list; a means for determining, if the means determines that a sentence included in the transcription data contains text (character strings) whose pronunciation is identical or similar to that of a word included in the word list, whether the meaning of the sentence including the text (character strings) is unnatural and whether the meaning of the sentence obtained by replacing the text (character strings) included in the sentence with the word is unnatural; and a means for outputting, if the means determines that the meaning of the sentence including the text (character strings) is unnatural and that the meaning of the sentence obtained by replacing the text (character strings) included in the sentence with the word is not unnatural, the text (character strings) as pre-proofreading text, the word as proofreading text, and, if necessary, the sentence including the text as information identifying the proofreading portion. A trained model other than a large-scale language model may be used as long as it can achieve the above-described means.
[0101] The subsequent steps include a step in which the information processing device 300 transmits the first proofreading proposal generated by the large-scale language model to the information processing device 100 (step S7); a step in which the information processing device 100 receives the first proofreading proposal from the information processing device 300 equipped with the large-scale language model and proofreads the transcription data based on the received first proofreading proposal (step S8); and a step in which the information processing device 100 outputs the proofread transcription data to the user terminal 500 (step S9). These steps are as described above, and therefore will not be described again.
[0102] [2. Embodiment 2] As another embodiment of the information processing device according to one aspect of the present invention, an embodiment will be described in which a proofreading proposal proofreading unit 136 is further added to the information processing device 100. Fig. 6 shows an example of the flow of information processing by an information processing system configured to include the information processing device 100 according to this embodiment.
[0103] The proofreading proposal proofreading unit 136 is a means that may be included in the processing unit 130 and that acquires a second proofreading proposal by proofreading the first proofreading proposal acquired by the proofreading proposal acquisition unit 133. In this specification, a proofreading proposal obtained by proofreading the first proofreading proposal is referred to as a second proofreading proposal. If there is no part to be proofread, the contents of the first proofreading proposal and the second proofreading proposal may be identical.
[0104] In this example, the second proofreading proposal is generated by an external information processing device 600 separate from the information processing device 100. That is, the information processing device 100, for example, transmits to the information processing device 600 a proofreading request for the first proofreading proposal acquired by the proofreading proposal acquisition unit 133, and, as necessary, the file path of the first proofreading proposal and / or the first proofreading proposal. Upon receiving the proofreading request for the first proofreading proposal, the information processing device 600 proofreads the first proofreading proposal and generates a second proofreading proposal. Then, the information processing device 600 transmits the generated second proofreading proposal to the information processing device 100. The proofreading proposal proofreading unit 136 of the information processing device 100 receives the second proofreading proposal via the communication unit 110, and thereby acquires the second proofreading proposal.
[0105] In addition, in this example, the transcription data proofreading unit 134 proofreads the transcription data based on the second proofreading plan acquired by the proofreading plan proofreading unit 136.
[0106] In this example, the generation of the second proofreading plan is performed by the information processing device 600, but may be performed by the information processing device 100. Also, it may be performed by an external information processing device (e.g., information processing device 200 or 300) separate from the information processing device 600. Also, in FIG. 6, the information processing device 600 is depicted as if it were a single physical information processing device, but the information processing device 600 does not necessarily have to be a single physical information processing device. For example, it may be a distributed information processing device, such as a distributed server.
[0107] The second proofreading proposal may be, for example, proofreading information selected from the proofreading information included in the first proofreading proposal acquired by the proofreading proposal acquisition unit 133, in which the readings of the pre-proofreading text and the proofread text are the same or similar. The first proofreading proposal acquired by the proofreading proposal acquisition unit 133 is output by a large-scale language model, and the first proofreading proposal output by the large-scale language model may include proofreading information in which the readings of the pre-proofreading text and the proofread text are neither the same nor similar. By selecting proofreading information in which the readings of the pre-proofreading text and the proofread text are the same or similar from the first proofreading proposal output by the large-scale language model, it is possible to obtain an appropriate proofreading result while changing the readings of the transcription result as little as possible.
[0108] Furthermore, the second proofreading proposal may be selected from the proofreading information included in the first proofreading proposal acquired by the proofreading proposal acquisition unit 133, such that the pronunciation of the pre-proofreading text and the proofread text are the same or similar, and further, in a sentence including the pre-proofreading text included in the transcription data, the meaning of the sentence obtained by replacing the pre-proofreading text with the proofread text is not unnatural and / or natural. The first proofreading proposal acquired by the proofreading proposal acquisition unit 133 is output by a large-scale language model, which is not good at performing long tasks. In particular, when a relatively large amount of transcription data is provided, the proofreading proposal output by the large-scale language model may include a combination in which the meaning of the sentence becomes unnatural or unnatural when the pre-proofreading text is replaced with the proofread text. By extracting only the proofreading information from the proofreading proposals output by the large-scale language model that does not make the meaning of the sentence unnatural when the pre-proofreading text is replaced with the proofread text, it is possible to more appropriately proofread the transcription results.
[0109] There are no particular limitations on the specific method for generating the second proofreading proposal, and a person skilled in the art can adopt any appropriate method. For example, the method may be realized using a rule-based algorithm, a model such as a large-scale language model, or AI, or a combination of these.
[0110] When a large-scale language model is used to generate the second proofreading draft, the large-scale language model may be the same as or different from the large-scale language model used to generate the first proofreading draft. For example, the large-scale language model used to generate the second proofreading draft may be a Transformer-based large-scale language model, similar to the large-scale language model used to generate the first proofreading draft. Examples of Transformer-based large-scale language models include, but are not limited to, "GPT" provided by OpenAI, Inc., "Gemini" provided by Google DeepMind, Inc., "Claude" provided by Anthropic, Inc., and "tsuzumi" provided by Nippon Telegraph and Telephone Corporation. It goes without saying that the large-scale language model used to generate the second proofreading draft may also be fine-tuned as appropriate.
[0111] Although the embodiments of the present invention have been described in detail above, these embodiments are merely examples, and the present invention can be embodied in other forms with various modifications and improvements based on the knowledge of those skilled in the art. Furthermore, the terms "means" and "unit" described above can be interpreted as appropriate. For example, the term "acquisition unit" can be interpreted as "acquisition means." [Industrial Applicability]
[0112] The present invention makes it possible to accurately transcribe voice data into text data, thereby preventing forgetting to record, typos, and mishearing during daily meetings, business negotiations, interviews, etc., and making it possible to record and share highly accurate information. [Explanation of symbols]
[0113] 100 Information processing device 110 Communications Department 120 Storage section 130 Processing section 131 Transcription Request Acquisition Unit 132 Transcription Data Acquisition Department 133 Proofreading proposal acquisition department 134 Transcription Data Proofreading Department 135 Proofread Transcription Data Output Unit 200 Information processing device (information processing device equipped with a voice recognition model) 300 Information processing device (information processing device equipped with a large-scale language model) 400 External Storage 500 Terminal Equipment 600 Information Processing Devices N Network
Claims
1. a means for inputting voice data into a voice recognition model that outputs transcription data in response to input of voice data, and obtaining transcription data of the voice data that is output in response to the input; a means for inputting an input sentence containing an instruction for outputting proofreading information associating a pre-proofreading text containing transcription errors and the proofreading text into a large-scale language model for the acquired transcription data, and obtaining a first proofreading proposal including proofreading information associating a pre-proofreading text and the proofreading text that is output in response to the input; A means for proofreading the transcription data based on the obtained first proofreading proposal to generate proofread transcription data; and means for outputting the generated proofread transcription data; An information processing device comprising:
2. The input sentence further includes: The instruction to output the pre-proofreading text and the proofreading text as proofreading information only when the readings of the pre-proofreading text and the proofreading text are the same or similar, 2. The information processing apparatus according to claim 1, wherein:
3. The input sentence further includes: and an instruction to output the proofreading information by referring to a word list that associates words with the pronunciation of the words.
3. The information processing apparatus according to claim 2, wherein:
4. The instruction to output the proofreading information by referring to a word list includes: and if the reading of the text included in the transcription data is the same as or similar to the reading of a word included in the word list, and if replacing the text with the word in a sentence including the text included in the transcription data does not make the meaning unnatural, the program includes instructions to output the text and the word as a pre-proofreading text and a proofreading text, respectively.
4. The information processing apparatus according to claim 3,
5. a means for acquiring a second proofreading plan by selecting proofreading information from the proofreading information included in the first proofreading plan, the proofreading information being the same as or similar to the unproofreading text; The transcription data is proofread based on the obtained second proofreading proposal to generate proofread transcription data. The information processing device according to claim 1 .
6. The means for obtaining a second proofreading proposal utilizes a large-scale language model to obtain the second proofreading proposal.
6. The information processing apparatus according to claim 5,
7. The means for obtaining a second proofreading proposal comprises: An input sentence including an instruction to select and output only the proofreading information included in the first proofreading proposal, the proofreading information being the same or similar in pronunciation to the pre-proofreading text, is input to a large-scale language model, and the output is obtained as a second proofreading proposal.
7. The information processing apparatus according to claim 6,
8. The means for obtaining a first proofreading proposal comprises: The acquired transcription data is divided into predetermined units, and for each of the transcription data divided into the predetermined units, an input sentence including an instruction to output proofreading information associating pre-proofreading text containing transcription errors with the proofreading text is input to a large-scale language model, and a first proofreading proposal is obtained, the first proofreading proposal including proofreading information associating pre-proofreading text and the proofreading text that is output in response to the input.
5. The information processing apparatus according to claim 1, wherein the information processing apparatus is a computer.
9. 9. The information processing apparatus according to claim 8, wherein the predetermined unit includes at least two utterances or sentences.
10. A computer-implemented information processing method, comprising: inputting speech data into a speech recognition model that outputs transcription data in response to input of speech data, and obtaining transcription data of the speech data that is output in response to the input; a step of inputting an input sentence including an instruction for outputting a proofreading proposal including proofreading information that associates a pre-proofreading text containing transcription errors and the proofreading text with each other into a large-scale language model for the acquired transcription data, and acquiring a first proofreading proposal that includes proofreading information that associates a pre-proofreading text and the proofreading text that are output in response to the input; proofreading the transcription data based on the obtained first proofreading proposal to generate proofread transcription data; and outputting the generated proofread transcription data; An information processing method comprising:
11. inputting speech data into a speech recognition model that outputs transcription data in response to input of speech data, and obtaining transcription data of the speech data that is output in response to the input; a step of inputting an input sentence including an instruction for outputting pre-proofreading text containing transcription errors and proofreading information associating the proofreading text with the acquired transcription data into a large-scale language model, and acquiring a first proofreading proposal including proofreading information associating the pre-proofreading text and the proofreading text that are output in response to the input; proofreading the transcription data based on the obtained first proofreading proposal to generate proofread transcription data; and outputting the generated proofread transcription data; A computer program that causes a computer to execute the following.
Citation Information
Patent Citations
Speech recognition device, speech recognition training device, speech recognition method, speech recognition training method, and program
JP2022010410A
Voice recognition accuracy improvement device, and method for improving voice recognition accuracy
JP2024010464A