Information processing device, information processing method, and recording medium
The information processing device generates training text from provisional text with error correction and selection methods, addressing the inefficiencies of manual transcription and error-prone provisional text, thereby improving speech recognition accuracy and data quantity at a lower cost.
Patent Information
- Application Number
- PCT/JP2024/024086
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2026-01-08
AI Technical Summary
Existing speech recognition technologies face challenges in improving accuracy when using provisional text generated from speech signals, as these texts often contain errors, which negatively impact the learning process and require manual transcription, making it costly and inefficient to generate large amounts of high-quality training data.
An information processing device and method that generates training text based on provisional text through instructions, correcting errors and generating multiple natural-looking texts, allowing for efficient creation of training data without manual intervention, and selecting suitable texts for training using similarity and evaluation criteria.
The solution enables the generation of high-quality training text with fewer errors, increasing the amount of training data available for speech recognition at a lower cost and enhancing the robustness of speech recognition models against estimation errors.
Smart Images

Figure JP2024024086_08012026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to the technical fields of an information processing device, an information processing method, and a recording medium.
[0002] As a speech recognition technology, for example, a technology has been proposed in which provisional text corresponding to a speech signal (for which no corresponding text exists) is automatically estimated by speech recognition, the pair of the speech signal and the provisional text is added to training data, and a speech recognition model is repeatedly trained to improve the accuracy of the provisional text (see Non-Patent Document 1).
[0003] Xu, Q., Likhomanenko, T., Kahn, J., Hannun, A., Synnaeve, G., Collobert, R. (2020) Iterative Pseudo-Labeling for Speech Recognition. Proc. Interspeech 2020, 1006-1010, doi: 10.21437 / Interspeech.2020-1800
[0004] An object of this disclosure is to provide an information processing device, an information processing method, and a recording medium that aim to improve the technology related to the prior art documents mentioned above.
[0005] One aspect of the information processing device includes an instruction means for generating instructions to generate training text based on provisional text that is the result of speech recognition of a speech signal, and a generation means for generating the training text in accordance with the instructions and outputting the training text.
[0006] One aspect of the information processing method is an information processing method executed by a computer, which includes generating instructions for generating training text based on provisional text that is the result of speech recognition of a speech signal, generating the training text in accordance with the instructions, and outputting the training text.
[0007] One aspect of the recording medium has recorded thereon a computer program for causing a computer to execute an information processing method including generating instructions for generating training text based on provisional text that is the result of speech recognition of a speech signal, generating the training text in accordance with the instructions, and outputting the training text.
[0008] FIG. 1 is a block diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 2 is a flowchart showing an example of the operation of the information processing device according to an embodiment. FIG. 3 is a block diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 4 is a block diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 5 is a flowchart showing an example of the operation of the information processing device according to an embodiment. FIG. 6 is a block diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 7 is a block diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 8 is a block diagram showing an example of the configuration of an information processing device according to an embodiment.
[0009] Hereinafter, an information processing device, an information processing method, and a recording medium according to an embodiment will be described with reference to the drawings. [1: First Embodiment]
[0010] A first embodiment of an information processing device, an information processing method, and a recording medium will be described with reference to Fig. 1 and Fig. 2. In the following, the first embodiment of the information processing device, the information processing method, and the recording medium will be described using an information processing device 10.
[0011] 1, the information processing device 10 includes an instruction unit 11 and a generation unit 12. The operation performed by the information processing device 10 will be described with reference to the flowchart of FIG.
[0012] As shown in FIG. 2 , the instruction unit 11 generates an instruction to generate a training text (step S11). The training text is a text generated based on a provisional text. The provisional text is a result of speech recognition of a speech signal. The instruction unit 11 generates the instruction based on the provisional text. For example, the instruction unit 11 may generate an instruction such as "Replace some words in the provisional text to create natural-looking text." Alternatively, the instruction unit 11 may generate an instruction such as "Replace some words in the provisional text to create error-free text." Alternatively, the instruction unit 11 may generate an instruction such as "Replace some words in the provisional text with words that have similar pronunciation to create meaningful text."
[0013] The generation unit 12 generates a training text based on the provisional text in accordance with the instructions generated by the instruction unit 11 (step S12). The generation unit 12 outputs the generated training text (step S13). Speech recognition of the speech signal may be performed by a predetermined speech recognition means. The training text may be used to train the predetermined speech recognition means.
[0014] Alternatively, the information processing device 10 may generate training text data from the speech signal. Specifically, the information processing device 10 may generate training text based on the results of speech recognition of the speech signal. More specifically, the information processing device 10 may generate training text based on the results of speech recognition of the speech signal in accordance with instructions generated based on the results of speech recognition of the speech signal.
[0015] In this way, the information processing device 10 performs an information processing method that includes generating instructions for generating a training text based on provisional text that is the result of speech recognition of a speech signal, generating the training text in accordance with the instructions, and outputting the training text.
[0016] The information processing device 10 described above may be realized by a computer reading a computer program recorded on a recording medium. In this case, the computer program may cause the computer to execute an information processing method including generating instructions for generating training text based on provisional text that is a result of speech recognition of a speech signal, generating the training text in accordance with the instructions, and outputting the training text. [Technical Effect]
[0017] The information processing device 10 according to the present disclosure generates training text based on the results of speech recognition of a speech signal in accordance with instructions generated based on the results of speech recognition of the speech signal. This makes it possible to generate training text with few errors and high quality, even if there is an error in the speech recognition of the speech signal. [2: Second Embodiment]
[0018] A second embodiment of the information processing device, information processing method, and recording medium will be described with reference to FIGS. 3 to 6. Hereinafter, the second embodiment of the information processing device, information processing method, and recording medium will be described using an information processing device 20. Note that, for the second embodiment, descriptions that overlap with the description of the first embodiment will be omitted as appropriate. [2-1: Configuration of the information processing device 20]
[0019] The configuration of the information processing device 20 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of the information processing device 20.
[0020] 3, the information processing device 20 includes a calculation device 21, a storage device 22, and a communication device 23. The information processing device 20 may further include an input device 24 and an output device 25. However, the information processing device 20 does not necessarily include at least one of the input device 24 and the output device 25. The calculation device 21, the storage device 22, the communication device 23, the input device 24, and the output device 25 may be connected via a data bus 26.
[0021] The arithmetic device 21 includes at least one processor (i.e., one processor or multiple processors) as hardware. The processor may include, for example, a processor conforming to a von Neumann computer architecture. The processor conforming to the von Neumann computer architecture may include at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor may include, for example, a processor conforming to a non-von Neumann computer architecture. The processor conforming to the non-von Neumann computer architecture may include at least one of an FPGA (Field Programmable Gate Array) and an ASIC (Application Specific Circuit).
[0022] The arithmetic device 21 reads a computer program 221 including at least one of computer program code and computer program instructions. For example, the arithmetic device 21 may read the computer program 221 stored in the storage device 22. For example, the arithmetic device 21 may read the computer program 221 stored in a computer-readable, non-transitory recording medium using a recording medium reading device (not shown) included in the information processing device 20. The computer program 221 read from the recording medium may be stored in the storage device 22. The arithmetic device 21 may acquire (i.e., download or read) the computer program 221 from a device (not shown) located outside the information processing device 20 via the communication device 23 (or another communication device). The downloaded computer program 221 may be stored in the storage device 22.
[0023] The arithmetic device 21 executes the loaded computer program 221. As a result, logical functional blocks for executing information processing to be performed by the information processing device 20 are realized within the arithmetic device 21. In other words, the arithmetic device 21, together with the storage device 22 or the like in which the computer program 221 is recorded (in other words, together with the storage device 22 and the computer program 221 recorded in the storage device 22 or the like), can function as a controller or computer for realizing logical functional blocks for executing processing to be performed by the information processing device 20. In other words, the at least one processor included in the arithmetic device 21, the memory (recording medium) included in the storage device 22 or the like, and the computer program 221 are configured so that the information processing device 20 performs the information processing to be performed by the information processing device 20.
[0024] A computational model that can be constructed by machine learning may be implemented in the computational device 21 by the computational device executing the computer program 221. An example of a computational model that can be constructed by machine learning is a computational model including a neural network (so-called artificial intelligence (AI)). In this case, learning of the computational model may include learning of parameters of the neural network (e.g., at least one of a weight and a bias). The computational device 21 may perform information processing using the computational model. In other words, the operation of performing information processing may include the operation of performing information processing using the computational model. Note that a computational model that has been constructed by offline machine learning using training data may be implemented in the computational device 21. Furthermore, the computational model implemented in the computational device 21 may be updated by online machine learning on the computational device 21. Alternatively, the calculation device 21 may perform information processing using a calculation model implemented in a device external to the calculation device 21 (i.e., a device provided outside the information processing device 20) in addition to or instead of the calculation model implemented in the calculation device 21.
[0025] The recording medium for recording the computer program 221 executed by the arithmetic device 21 may be at least one of a CD-ROM, CD-R, CD-RW, flexible disk, MO, DVD-ROM, DVD-RAM, DVD-R, DVD+R, DVD-RW, DVD+RW, Blu-ray (registered trademark), or other optical disk, a magnetic medium such as a magnetic tape, a magneto-optical disk, a semiconductor memory such as a USB memory, or any other medium capable of storing a program. The recording medium may include a device capable of recording a computer program (for example, a general-purpose device or a dedicated device in which the computer program 221 is implemented in a state in which it can be executed in at least one of the forms of software and firmware). Furthermore, each process or function included in the computer program 221 may be realized by a logical processing block realized within the arithmetic device 21 when the arithmetic device 21 (i.e., processor) executes the computer program 221, or may be realized by hardware such as a predetermined gate array (FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit)) provided in the arithmetic device 21, or may be realized in a form that mixes logical processing blocks and partial hardware modules that realize some elements of the hardware.
[0026] The storage device 22 includes at least one memory capable of storing desired data. In other words, the storage device 22 includes at least one memory containing desired data. For example, the storage device 22 may store a computer program 221 executed by the arithmetic device 21. In this case, the storage device 22 (memory) may be used as the above-mentioned recording medium for recording the computer program 221 executed by the arithmetic device 21. The storage device 22 may temporarily store data used by the arithmetic device 21 when the arithmetic device 21 is executing the computer program 221. The storage device 22 may also store data to be stored long-term by the information processing device 20. The storage device 22 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. In other words, the storage device 22 may include a non-temporary recording medium.
[0027] The storage device 22 may store a speech recognition model 222 (described later). The speech recognition model 222 stored in the storage device 22 may be parameters of a neural network (e.g., at least one of a weight and a bias). The storage device 22 may also store training data 223 (described later).
[0028] The communication device 23 may be capable of communicating with devices external to the information processing device 20. The communication device 23 may perform wired communication or wireless communication.
[0029] The input device 24 is a device capable of accepting information input to the information processing device 20 from outside. The input device 24 may include an operation device (e.g., a keyboard, a mouse, a touch panel, etc.) that can be operated by a user of the information processing device 20. The input device 24 may include a recording medium reading device that can read information recorded on a recording medium that is detachable from the information processing device 20, such as a USB (Universal Serial Bus) memory. Note that when information is input to the information processing device 20 via the communication device 23 (in other words, when the information processing device 20 acquires information via the communication device 23), the communication device 23 may function as an input device.
[0030] The output device 25 is a device capable of outputting information to the outside of the information processing device 20. The output device 25 may output visual information such as text or images, auditory information such as sound, or tactile information such as vibration, as the information. The output device 25 may include, for example, at least one of a display, a speaker, a printer, and a vibration motor. The output device 25 may be capable of outputting information to a recording medium detachable from the information processing device 20, such as a USB memory. Note that when the information processing device 20 outputs information via the communication device 23, the communication device 23 may function as the output device.
[0031] FIG. 3 shows an example of logical functional blocks realized in the arithmetic device 21 to execute information processing. As shown in FIG. 3, a speech recognition unit 213, a training data generation unit 214, and a training unit 215 are realized in the arithmetic device 21. The training data generation unit 214 may include a provisional text evaluation unit 2141, an instruction unit 2142, and a text generation unit 2143. The "instruction unit 2142" is a component corresponding to the "instruction unit 11" in the first embodiment described above, and the "text generation unit 2143" is a component corresponding to the "generation unit 12" in the first embodiment described above. [2-2: Functions Realized by the Information Processing Device 20]
[0032] The information processing device 20 performs voice recognition, generates learning data that can be used for learning voice recognition operations, and performs learning of voice recognition operations. Figures 4 and 5 are block diagrams showing an example of the flow of information processing executed by the information processing device 20.
[0033] 4 shows an example of the flow of information processing executed by the speech recognition unit 213, the training data generation unit 214, and the training unit 215 implemented in the calculation device 21, using, for example, a speech recognition model 222 and training data 223 stored in the storage device 22. The training data 223 may include speech data 2231, generated training data 2232, and training data 2233.
[0034] The audio data 2231 may include an audio signal for which no corresponding text exists, or in other words, the audio data 2231 includes an audio signal for which no text is associated.
[0035] The generated training data 2232 may include training data that is a pair of a speech signal and a training text generated within the information processing device 20. The speech signal corresponds to the training text generated within the information processing device 20. In other words, the generated training data 2232 includes training data in which the speech signal is associated with the training text generated within the information processing device 20. In other words, the generated training data 2232 includes training data in which the speech signal is associated with the training text generated by the training data generation unit 214.
[0036] The training data 2233 may include training data that is a pair of a voice signal and a text. The voice signal and the text correspond to each other. In other words, the training data 2233 includes training data in which the voice signal and the text are associated with each other. The training data included in the training data 2233 may be training data generated outside the information processing device 20. The training data included in the training data 2233 may be, for example, training data generated manually.
[0037] The speech recognition unit 213 recognizes a speech signal included in the speech data 2231 using a speech recognition model. The speech recognition model may be a computational model implemented in the computing device 21. The speech recognition model may include a neural network. Parameters (e.g., at least one of weights and biases) of the neural network included in the speech recognition model may be stored in the speech recognition model 222. The speech recognition model may be a mechanism that, when a speech signal is input, outputs a character string corresponding to the speech signal.
[0038] The speech recognition unit 213 outputs paired data that is a pair of a speech signal and provisional text that is the recognition result of the speech signal. The provisional text may be a character string. The provisional text may also be a character string that represents an intermediate result of speech recognition. The provisional text may also be a predetermined number of character strings that have a high speech recognition score. The provisional text may also have information that indicates the speech recognition score of speech recognition.
[0039] The training data generation unit 214 receives paired data of a voice signal and provisional text and outputs training data that is a paired voice signal and training text corresponding to the voice signal. The training data generation unit 214 converts the provisional text, which is the voice recognition result of the voice recognition unit 213, into training text suitable for training voice recognition operations. For example, if the provisional text corresponding to the voice signal of an utterance of "It's good weather today" is "It's good electricity today," the training data generation unit 214 may convert "It's good electricity today" to "It's good weather today" and output the converted text as training text. The training data generation unit 214 generates training data that is a paired voice signal and training text from the paired data of the voice signal and provisional text. The generated training data 2232 may be training data output by the training data generation unit 214.
[0040] The learning unit 215 uses at least one of the generated training data 2232 and the training data 2233 to train the speech recognition operation of the speech recognition unit 213. In other words, the learning unit 215 trains the speech recognition model used by the speech recognition unit 213. The speech recognition model 222 may store the results of learning by the learning unit 215.
[0041] 5 shows an example of the flow of information processing executed by the provisional text evaluation unit 2141, instruction unit 2142, and text generation unit 2143, which are realized in the training data generation unit 214. As shown in FIG. 5, pair data of a speech signal and provisional text, which is the recognition result of the speech signal, is input to the provisional text evaluation unit 2141 and instruction unit 2142.
[0042] The provisional text evaluation unit 2141 may estimate a speech recognition error portion contained in the provisional text. For example, in the above-described case, the provisional text evaluation unit 2141 may estimate "electricity" contained in the provisional text "Today is good electricity" as a speech recognition error portion. The provisional text evaluation unit 2141 may estimate a speech recognition error related to the provisional text and evaluate the provisional text.
[0043] The instructing unit 2142 may generate an instruction regarding the operation of the text generating unit 2143 according to a pair of the speech signal and the provisional text. The instructing unit 2142 may generate an instruction according to an operation result of the provisional text evaluating unit 2141. For example, the instructing unit 2142 may generate an instruction to "replace some words in the provisional text to create natural-looking text." The instructing unit 2142 may also generate an instruction to "replace some words in the provisional text to create error-free text." The instructing unit 2142 may also generate an instruction to "replace some words in the provisional text with words that have similar pronunciation to create meaningful text." For example, if the "electricity" included in the provisional text "Today is good electricity" described above is estimated to be an error in speech recognition, the instructing unit 2142 may further generate an instruction to "change the "electricity" in the provisional text to another word when replacing some words in the provisional text." The instruction unit 2142 may generate a more specific instruction such as "change 'electricity' in the provisional text to 'weather'." The instruction unit 2142 generates and outputs instructions based on the operation result of the provisional text evaluation unit 2141. The instructions generated by the instruction unit 2142 are input to the text generation unit 2143.
[0044] The text generation unit 2143 generates training text based on the provisional text in accordance with the instructions generated by the instruction unit 2142. Alternatively, the text generation unit 2143 converts the provisional text into training text suitable for training. Alternatively, the training data generation unit 214 corrects the provisional text to a correct transcription. The text generation unit 2143 may tolerate a mismatch between the speech signal and the training text. The training text may be a string of characters that humans perceive as natural. The training text may be a string of characters that exactly corresponds to the speech signal. On the other hand, the training text may be a string of characters that does not exactly correspond to the speech signal. In other words, the training text is not limited to a string of characters that exactly corresponds to the speech signal. Furthermore, the text generation unit 2143 may generate one or more training texts from one provisional text. In other words, the text generation unit 2143 may generate multiple training texts from one provisional text. The text generation unit 2143 may generate multiple natural texts similar to the provisional text as training texts. That is, the text generation unit 2143 may perform data expansion. In order to generate natural text, the text generation unit 2143 may generate training text using large language models (LLM). In this case, the instruction unit 2142 may generate a prompt for the large language model as an instruction. The large language model may generate training text based on the provisional text in accordance with the prompt. The text generation unit 2143 outputs training data that is a pair of a speech signal and the generated training text. [2-3: Information Processing Method Executed by Information Processing Device 20]
[0045] An information processing method executed by the information processing device 20 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the flow of processing in the information processing method executed by the information processing device 20.
[0046] 6, the speech recognition unit 213 recognizes the speech signal included in the speech data and outputs provisional text (step S21). The speech recognition unit 213 outputs set data in which the provisional text is added to the speech data (step S22).
[0047] The provisional text evaluation unit 2141 evaluates the provisional text (step S23). The instruction unit 2142 generates instructions using the evaluation results of the provisional text (step S24). The text generation unit 2143 generates training text based on the provisional text in accordance with the instructions (step S25). The training data generation unit 214 outputs training data in which the training text is added to the speech signal (step S26).
[0048] The learning unit 215 uses the training data to train the speech recognition model (step S27). The learning unit 215 determines whether to continue learning (step S28). If learning is to continue (step S28: Yes), the process returns to step S21 and executes the processes from step S21 to step S28. If learning is to end (step S28: No), the information processing method ends. As illustrated in FIGS. 4 and 6, the information processing device 20 may perform repeated learning. The learning unit 215 may determine to end learning when the number of repetitions reaches a predetermined number. The learning unit 215 may also determine to end learning when the error rate when using the evaluation data falls below a predetermined value. Note that the evaluation data may be data different from the training data. The learning unit 215 may also determine to end learning when the difference between the error rate when using the evaluation data after the (N-1)th learning and before the Nth learning and the error rate when using the evaluation data after the Nth learning falls below a predetermined value. [2-4: Technical Effects]
[0049] Training data for speech recognition includes a speech signal and text corresponding to the speech signal. Training data for speech recognition is often prepared by obtaining a speech signal and manually transcribing the text corresponding to the speech signal. Manually transcribed text is of relatively high quality. On the other hand, manually transcribed text is relatively expensive because it requires time and effort. For this reason, it is not easy to prepare a large amount of speech signals for which corresponding text exists. In contrast, speech signals for which no corresponding text exists are relatively inexpensive and relatively easy to obtain in large quantities. However, text that is not manually transcribed is often of relatively low quality.
[0050] When the text contained in the training data is free of errors, optimal learning is possible. When the text contained in the training data contains errors, the learning effect may decrease. In other words, errors in the text have a negative impact on learning. Even when the text contained in the training data contains errors, if there are few errors, the improvement effect of the speech recognition function due to learning will be greater. On the other hand, when the text contained in the training data contains many errors, the negative impact of the text errors will exceed the improvement effect of the speech recognition function due to learning. When the text contained in the training data contains many errors, repeated learning may not improve the learning effect.
[0051] The information processing device 20 according to this disclosure generates training text in accordance with instructions generated by the instruction unit 2142, and is therefore able to generate training data including training text suitable for speech recognition training. The information processing device 20 can generate natural text from unnatural text in accordance with instructions generated by the instruction unit 2142. The information processing device 20 allows for mismatches between the speech signal and the training text, and is therefore able to generate training text in which mistakes in speech and the like have been corrected. Because the training data generation unit 214 generates training text in accordance with instructions, the information processing device 20 can prepare training data that is expected to have a learning effect, even if the recognition result of the speech recognition unit 213 contains an error.
[0052] In this way, the information processing device 20 can increase the amount of training data for speech recognition without manual intervention. In other words, the information processing device 20 can increase the amount of training data for speech recognition at low cost.
[0053] Furthermore, the information processing device 20 according to this disclosure uses not only speech signals for which corresponding text exists but also speech signals for which no corresponding text exists, and therefore can train a speech recognition model using a large amount of training data. The information processing device 20 using a speech recognition model trained using such training data can perform automatic estimation that is robust against estimation errors. [3: Third Embodiment]
[0054] A third embodiment of an information processing device, an information processing method, and a recording medium will be described with reference to Figures 7 and 8. Hereinafter, the third embodiment of an information processing device, an information processing method, and a recording medium will be described using an information processing device 30. Note that, for the third embodiment, descriptions that overlap with the descriptions of the first and second embodiments will be omitted as appropriate. Note that, in the drawings, parts common to the first and second embodiments are designated by the same reference numerals.
[0055] 7, the arithmetic device 21 included in the information processing device 30 includes, as logical functional blocks, a speech recognition unit 213, a training data generation unit 314, and a training unit 215. The training data generation unit 314 may include a text evaluation unit 3144 in addition to a provisional text evaluation unit 2141, an instruction unit 2142, and a text generation unit 3143. [3-1: Functions Realized by the Information Processing Device 30]
[0056] An information processing method executed by the information processing device 30 will be described with reference to Fig. 8. Fig. 8 is a block diagram showing an example of the flow of information processing executed by the information processing device 30.
[0057] The text generation unit 3143 may generate multiple training texts based on one provisional text. The text generation unit 3143 may generate multiple natural texts similar to the provisional text as training texts. That is, the text generation unit 3143 may perform data augmentation. To generate natural texts, the text generation unit 3143 may generate training texts using large language models (LLMs). This allows the text generation unit 3143 to avoid adding unnatural text to the training data.
[0058] The text evaluation unit 3144 selects one or more training texts from a plurality of training texts. The text evaluation unit 3144 may select a training text suitable for training a speech recognition model from a plurality of training texts. The text evaluation unit 3144 may determine whether a training text is suitable for training based on the similarity between the items indicated in the training text and the items indicated by the corresponding speech signal. The text evaluation unit 3144 may determine whether a training text is suitable for training based on the similarity of character strings. The similarity of character strings is the similarity between the training text T and text Ts, which is the recognition result of the speech signal S. If the training text and the speech signal are completely different, the similarity of character strings will be low. The similarity of character strings may also be referred to as an edit distance. The text evaluation unit 3144 may also determine whether a training text is suitable for training based on an evaluation value of speech recognition. The evaluation value of speech recognition is the evaluation value when a speech recognizer outputs text T as the recognition result of the speech signal S. If the training text and the speech signal are completely different, the evaluation value of the speech recognition will be low. The evaluation value of the speech recognition may also be referred to as a recognition score or an output probability. The text evaluation unit 3144 may determine that a training text that has a relatively high similarity to the items indicated by the corresponding speech signal is suitable for learning.
[0059] Furthermore, the text evaluation unit 3144 may determine whether a training text is suitable for learning based on the degree of similarity between the training text and the corresponding provisional text. The text evaluation unit 3144 may determine that a training text that has a relatively high degree of similarity with the corresponding provisional text is suitable for learning.
[0060] Furthermore, the text evaluation unit 3144 may select one or more training texts based on the degree of similarity between each of the multiple training texts. Learning using multiple dissimilar training texts is often more efficient than learning using multiple similar training texts. Therefore, the text evaluation unit 3144 may thin out some of the similar training texts from the multiple training texts. For example, the text evaluation unit 3144 may determine whether training text data is suitable for learning based on the overlap between multiple training text data. The text evaluation unit 3144 may exclude training text data with a relatively large amount of overlap from the candidates for registration.
[0061] That is, the text evaluation unit 3144 may select one or more training texts based on at least one of the similarity between the items indicated by the audio signal and the items indicated by the training text, the similarity between the provisional text and the training text, and the similarity between each of the multiple training texts. [3-2: Technical Effects]
[0062] The information processing device 30 according to this disclosure evaluates the generated training texts and selects training texts suitable for training, thereby efficiently constructing a speech recognition model capable of performing speech recognition with high accuracy. [4: Fourth Embodiment]
[0063] A fourth embodiment of an information processing device, an information processing method, and a recording medium will be described with reference to Figures 9 and 10. Below, the fourth embodiment of an information processing device, an information processing method, and a recording medium will be described using an information processing device 40. Note that, for the fourth embodiment, descriptions that overlap with the descriptions of the first to third embodiments will be omitted as appropriate. Note that, in the drawings, parts common to the first to third embodiments are indicated by the same reference numerals.
[0064] 9, the arithmetic device 21 included in the information processing device 40 includes, as logical functional blocks, a speech recognition unit 213, a training data generation unit 414, and a training unit 215. The training data generation unit 414 may include a case example extraction unit 4145 in addition to a provisional text evaluation unit 2141, an instruction unit 4142, a text generation unit 2143, and a text evaluation unit 2144. Also, as shown in FIG. 9, the storage device 22 included in the information processing device 40 may realize a provisional text holding unit 424 and a case example holding unit 425 in addition to a computer program 221, a speech recognition model 222, and training data 223.
[0065] The voice data 2231 included in the training data 223 may include a plurality of voice signals. The voice data 2231 may include a plurality of different voice signals. For example, the voice data 2231 may include a first voice signal and a second voice signal.
[0066] 4 and 6, the arithmetic device 21 may repeatedly perform the speech recognition operation of the speech recognition unit 213, the learning data generation operation of the learning data generation unit 414, and the learning operation using the learning data of the learning unit 215. In the repetition, the speech recognition unit 213 may sequentially perform speech recognition on each of the multiple speech signals included in the speech data 2231. For example, if the speech data 2231 includes a first speech signal and a second speech signal, the speech recognition unit 213 may first perform speech recognition on the first speech signal, and then perform speech recognition on the second speech signal. [4-1: Functions Realized by the Information Processing Device 40]
[0067] An information processing method executed by the information processing device 40 will be described with reference to Fig. 10. Fig. 10 is a block diagram showing an example of the flow of information processing executed by the information processing device 30. Note that in Fig. 10, the speech recognition model 222 and the training data 223 are omitted from the illustration so as not to make the drawing difficult to understand.
[0068] In the Nth (N is a positive integer) speech recognition operation, the speech recognition unit 213 may perform speech recognition on each of the multiple speech signals included in the speech data 2231. Similarly, in the N+1th speech recognition operation, the speech recognition unit 213 may perform speech recognition on each of the multiple speech signals included in the speech data 2231.
[0069] The speech recognition unit 213 may store the provisional text, which is the Nth speech recognition result, in the provisional text storage unit 424. The speech recognition unit 213 may store the provisional text, which is the Nth speech recognition result, and the corresponding speech signal in the provisional text storage unit 424 in association with each other.
[0070] The training data generation unit 414 may generate training text based on the Nth speech recognition result in the Nth training data generation operation. Similarly, the training data generation unit 414 may generate training text based on the N+1th speech recognition result in the N+1th training data generation operation.
[0071] The learning unit 215 may perform learning using learning data including the Nth learning data generation result in the Nth learning operation. Similarly, the learning unit 215 may perform learning using learning data including the Nth learning data generation result in the N+1th learning operation.
[0072] The example extraction unit 4145 compares the N+1th speech recognition result with the Nth speech recognition result. The example extraction unit 4145 may compare the provisional text stored in the provisional text storage unit 424 with the provisional text output by the speech recognition unit 213. The example extraction unit 4145 may compare the N+1th speech recognition result of the first speech signal with the Nth speech recognition result of the first speech signal. Similarly, the example extraction unit 4145 may compare the N+1th speech recognition result of the second speech signal with the Nth speech recognition result of the second speech signal. If the N+1th speech recognition result and the Nth speech recognition result differ, the example extraction unit 4145 stores in the example storage unit 425 association information (referred to as an "example") that associates information indicating the Nth speech recognition result with information indicating an instruction generated the Nth time. The instruction generated the Nth time may be an instruction generated based on the provisional text output the Nth time. The example extractor 4145 may store an example in the example holder 425 when the (N+1)th speech recognition result and the Nth speech recognition result differ by a predetermined amount or more.
[0073] In other words, the example extraction unit 4145 compares the provisional text output by the speech recognition unit 213 for the Nth time, which has been trained using training data including training text generated in accordance with the Nth instruction (referred to as "Nth-th training data"), with the provisional text output by the speech recognition unit 213 for the Nth time. The provisional text output by the speech recognition unit 213 for the Nth time may be stored in the provisional text storage unit 424. In other words, the example extraction unit 4145 stores the example in the example storage unit 425 when the provisional text output by the speech recognition unit 213 for the Nth time, which has been trained using the Nth-th training data, differs from the provisional text output by the speech recognition unit 213 for the Nth time.
[0074] In other words, the example extraction unit 4145 compares the N+1th provisional text data output by the speech recognition unit 213 using a speech recognition model (N+1th speech recognition model) trained by the learning unit 215 for the N+1th speech recognition operation, with the Nth provisional text data output by the speech recognition unit 213 using a speech recognition model (Nth speech recognition model) trained by the learning unit 215 for the Nth speech recognition operation. The example extraction unit 4145 may compare the N+1th provisional text data output by the speech recognition unit 213 using the N+1th speech recognition model when a predetermined speech signal is input, with the Nth provisional text data output by the speech recognition unit 213 using the Nth speech recognition model when a predetermined speech signal is input.
[0075] The instructing unit 4142 may generate instructions using the examples. The instructing unit 4142 may extract examples to be stored in the example holding unit 425 based on the provisional text output by the speech recognition unit 213. The instructing unit 4142 may, for example, extract examples similar to the provisional text output by the speech recognition unit 213. The instructing unit 4142 may generate instructions using the extracted examples. For example, the instructing unit 4142 may generate instructions to perform the same conversion on the provisional text as the conversion in the extracted examples. The text generating unit 3143 generates learning text in accordance with the generated instructions.
[0076] For example, if a provisional text that contained an error the Nth time contains no error the N+1th time, the Nth time instruction can be referred to as an example of an instruction to eliminate the error. Also, if a provisional text that was unnatural the Nth time becomes natural the N+1th time, the Nth time instruction can be referred to as an example of an instruction to make the text natural. [4-2: Technical Effects]
[0077] The information processing device 40 according to this disclosure utilizes information obtained during the repeated learning process, and is therefore able to generate training data that improves the learning effect of speech recognition operations. Since a speech recognition model trained using training data including training text generated according to past instructions (for generating training text) is used, it is considered that the provisional text, which is the speech recognition result, has changed. Therefore, the relevant past instructions can be considered useful instructions. Because the information processing device 40 utilizes such useful instructions, it is able to generate training text suitable for learning. [5: Supplementary Note]
[0078] Some or all of the above embodiments may be described as, but are not limited to, the following supplementary notes. [Supplementary Note 1] An information processing device comprising: instruction means for generating an instruction to generate a training text based on a provisional text that is a result of speech recognition of a speech signal; and generation means for generating the training text in accordance with the instruction and outputting the training text. [Supplementary Note 2] The information processing device according to Supplementary Note 1, wherein the instruction means generates the instruction based on the speech signal and the provisional text. [Supplementary Note 3] The information processing device according to Supplementary Note 1 or 2, wherein the information processing device comprises provisional text evaluation means for estimating an error in the speech recognition related to the provisional text and evaluating the provisional text, and the instruction means generates the instruction using a result of evaluation of the provisional text. [Supplementary Note 4] The information processing device according to Supplementary Note 1 or 2, wherein the generation means generates a plurality of the training texts based on one provisional text, and the information processing device comprises selection means for selecting one or more of the training texts from the plurality of training texts. [Supplementary Note 5] The information processing device according to Supplementary Note 4, wherein the selection means selects the one or more training texts based on at least one of a similarity between an item indicated by the speech signal and an item indicated by the training text, a similarity between the provisional text and the training text, and a similarity between each of the plurality of training texts. [Supplementary Note 6] The information processing device according to Supplementary Note 2, further comprising: a learning means that trains a speech recognition operation of a speech recognition means that outputs the provisional text as a result of speech recognition of the speech signal when the speech signal is input, by machine learning using a pair of the speech signal and the training text corresponding to the speech signal. [Supplementary Note 7] The information processing device according to Supplementary Note 2, further comprising: a speech recognition means that outputs the provisional text as a result of speech recognition of the speech signal when the speech signal is input, and the instruction means generates the instruction based on the speech signal input to the speech recognition means and the provisional text output by the speech recognition means.[Supplementary Note 8] The information processing device according to Supplementary Note 7, further comprising: storage means for storing association information that associates information indicating the provisional text output for the Nth time with information indicating an Nth instruction generated based on the provisional text output for the Nth time, when the provisional text output for the Nth time (N is a positive integer) by the speech recognition means to which a predetermined speech signal has been input differs from the provisional text output for the Nth time, and the instruction means generates the instruction using the association information. [Supplementary Note 9] The information processing device according to Supplementary Note 8, wherein the storage means stores the association information when the provisional text output by the speech recognition means trained using an Nth training text generated in accordance with the Nth instruction differs from the provisional text output for the Nth time. [Supplementary Note 10] The information processing device according to Supplementary Note 6, wherein, when the (N+1)th provisional text data output by the speech recognition means having been trained for the Nth (N is a positive integer)+1th time of the speech recognition operation by the learning means and when a predetermined speech signal is input thereto differs from the Nth provisional text data output by the speech recognition means having been trained for the Nth time of the speech recognition operation by the learning means and when the predetermined speech signal is input thereto, the information processing device comprises: storage means for storing association information that associates information indicating the provisional text outputted for the Nth time with information indicating an Nth instruction generated based on the provisional text outputted for the Nth time, and the instruction means generates the instruction using the association information. [Supplementary Note 11] The information processing device according to Supplementary Note 1, wherein the generation means generates the training text using Large Language Models (LLM), and the instruction means generates a prompt of the large language model as the instruction. [Supplementary Note 12] The information processing device according to Supplementary Note 1 or 2, wherein the generating means generates the learning text based on the provisional text in accordance with the instruction.[Supplementary Note 13] An information processing method executed by a computer, comprising: generating instructions for generating training text based on provisional text that is a result of speech recognition of a speech signal; generating the training text in accordance with the instructions, and outputting the training text. [Supplementary Note 14] A recording medium having recorded thereon a computer program for causing a computer to execute an information processing method that includes: generating instructions for generating training text based on provisional text that is a result of speech recognition of a speech signal; generating the training text in accordance with the instructions, and outputting the training text.
[0079] Furthermore, some or all of the configurations described in Supplementary Notes 2 to 11, which are dependent on Supplementary Notes 1, and Supplementary Notes 12 and 13, may be dependent in the same manner as Supplementary Notes 2 to 11. Furthermore, not limited to Supplementary Notes 1, 12, and 13, some or all of the configurations described as Supplements may be dependent on various hardware, software, various recording means for recording software, or systems, within the scope of each of the above-mentioned embodiments.
[0080] This disclosure may be modified as appropriate within the scope that does not contradict the gist or idea of the invention that can be read from the claims and the entire specification, and information processing devices, information processing methods, and programs that involve such modifications are also included in the technical idea of this disclosure.
[0081] 10, 20, 30, 40 Information processing device 11, 2142, 4142 Instruction unit 12 Generation unit 213 Speech recognition unit 214, 314, 414 Training data generation unit 2141 Provisional text evaluation unit 2143, 3143 Text generation unit 215 Training unit 222 Speech recognition model 223 Training data 2231 Speech data 2232 Generated training data 2233 Training data 3144 Text evaluation unit 4145 Example extraction unit 424 Provisional text storage unit 425 Example storage unit
Claims
1. An information processing device comprising: an instruction means for generating instructions to generate training text based on provisional text that is the result of speech recognition of a speech signal; and a generation means for generating the training text in accordance with the instructions and outputting the training text.
2. The information processing device according to claim 1, wherein the instruction means generates the instruction based on the voice signal and the provisional text.
3. An information processing device according to claim 1 or 2, further comprising a provisional text evaluation means for estimating an error in the speech recognition relating to the provisional text and evaluating the provisional text, and wherein the instruction means generates the instruction using the evaluation result of the provisional text.
4. An information processing device according to claim 1 or 2, wherein the generating means generates a plurality of the training texts based on one of the provisional texts, and the information processing device is provided with a selecting means for selecting one or more of the training texts from the plurality of the training texts.
5. The information processing device according to claim 4, wherein the selection means selects the one or more training texts based on at least one of the similarity between the items indicated by the audio signal and the items indicated by the training text, the similarity between the provisional text and the training text, and the similarity between each of the plurality of training texts.
6. The information processing device according to claim 2, further comprising a learning means for learning the speech recognition operation of the speech recognition means, which outputs the provisional text that is the result of speech recognition of the speech signal when the speech signal is input, through machine learning using a pair of the speech signal and the training text corresponding to the speech signal.
7. An information processing device as described in claim 2, further comprising a voice recognition means for outputting the provisional text, which is the result of voice recognition of the voice signal when the voice signal is input, and wherein the instruction means generates the instruction based on the voice signal input to the voice recognition means and the provisional text output by the voice recognition means.
8. An information processing device as described in claim 7, further comprising: a storage means for storing correspondence information that associates information indicating the provisional text outputted the Nth time with information indicating an Nth instruction generated based on the provisional text outputted the Nth time, when the provisional text outputted the Nth time (N is a positive integer) by the speech recognition means to which a predetermined speech signal is input differs from the provisional text outputted the Nth time; and the instruction means generates the instruction using the correspondence information.
9. The information processing device according to claim 8, wherein the storage means stores the correspondence information when the provisional text output by the speech recognition means trained using the Nth training text generated in accordance with the Nth instruction differs from the provisional text output the Nth time.
10. The information processing device according to claim 6, further comprising: storage means for storing association information that associates information indicating the provisional text outputted the Nth time with information indicating an Nth instruction generated based on the provisional text outputted the Nth time, and wherein, when the N+1th provisional text data outputted by the speech recognition means when a predetermined speech signal is inputted, which is the speech recognition means that has learned the speech recognition operation the Nth time (N is a positive integer) by the learning means, differs from the Nth provisional text data outputted by the speech recognition means when the predetermined speech signal is inputted, and wherein the instruction means generates the instruction using the association information.
11. The information processing device according to claim 1, wherein the generating means generates the training text using large language models (LLMs), and the instructing means generates a prompt for the large language model as the instruction.
12. An information processing method executed by a computer, comprising: generating instructions for generating training text based on provisional text that is the result of speech recognition of a speech signal; generating the training text in accordance with the instructions; and outputting the training text.
13. A recording medium having recorded thereon a computer program for causing a computer to execute an information processing method including: generating instructions for generating training text based on provisional text that is the result of speech recognition of a speech signal; generating the training text in accordance with the instructions; and outputting the training text.
Citation Information
Patent Citations
Correction device, correction method and correction program
JP2018156418A
System and program
JP2023131648A
Voice recognition device and method for generating learned model
JP2023137523A
Medical record creation support device and medical record creation support method
JP7501883B1