Information processing device, transcription system, and computer program
The information processing device enhances transcription accuracy by using user proofreading history to adapt large-scale language models, addressing the limitations of existing systems in handling unique vocabularies and improving user experience.
Patent Information
- Application Number
- JP2025099073
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-06-12
AI Technical Summary
Existing transcription systems face challenges in accurately proofreading technical terms, new words, personal names, and homonyms due to limitations in large-scale language models, and creating user-specific vocabulary lists is cumbersome and impractical.
An information processing device that uses a large-scale language model to proofread transcription errors, leveraging user proofreading history information to generate prompts that adapt to individual user vocabularies, reducing the need for manual list creation.
The system efficiently accommodates user-specific vocabulary, improving transcription accuracy and user experience by dynamically learning and correcting errors based on past proofreading history.
Smart Images

Figure 0007759637000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, and more particularly to an information processing device for transcribing audio data, a transcription system including the same, and a computer program. [Background technology]
[0002] In recent years, transcription tools that convert voice data into text have been attracting attention, and various transcription tools have been developed and are being used in a variety of fields, such as creating minutes at meetings and recording telephone calls at call centers (e.g., Patent Documents 1 and 2).
[0003] Typical transcription tools use a speech recognition model to convert speech data into text data and output it as a transcription result. A speech recognition model converts speech data into text data and outputs it. Examples of speech recognition models include a combination of three models: an acoustic model that converts speech features into a phoneme sequence, a pronunciation model that converts the phoneme sequence into word candidates, and a language model that predicts the sequence of word candidates and generates text. Other well-known models include transformer-based speech recognition models such as OpenAI's "Whisper" and Meta's "Wav2Vec." In recent years, transformer-based speech recognition models have made remarkable progress, enabling more natural transcriptions that take long contexts into account. However, based on our testing, when transcription is performed using currently available speech recognition models, the transcription data still contains many transcription errors, and performance improvements are needed. However, improving the performance of transformer-based speech recognition models requires improvements in various aspects, such as data, model structure, and training methods, making it difficult to achieve such improvements.
[0004] In response to this, the present inventors have discovered a method that is distinct from attempts to improve the performance of a speech recognition model itself by combining a speech recognition model with a large-scale language model and having the large-scale language model proofread the transcription data output by the speech recognition model. This method makes it possible to efficiently proofread transcription errors contained in the transcription data, thereby obtaining more suitable transcription results, and the inventors have published this finding in Patent Document 3. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 7645561 [Patent Document 2] Japanese Patent Application Publication No. 2025-023364 [Patent Document 3] Patent No. 7654294 Summary of the Invention [Problem to be solved by the invention]
[0006] On the other hand, it is difficult for a large-scale language model to properly proofread transcription data that includes terms that the large-scale language model has not learned or that the large-scale language model has difficulty understanding and outputting, such as technical terms, new words, personal names, homonyms, etc. In order to have the large-scale language model properly proofread these terms, there is a method, as disclosed by the present inventors in Patent Document 3, in which a word list is prepared that includes words that are thought to be lacking in the large-scale language model, and the large-scale language model proofreads the transcription data while referring to the word list.
[0007] However, it is not easy and time-consuming to prepare a word list by listing vocabulary that is expected to be lacking in a large-scale language model. Furthermore, there are many different users of transcription systems, and each user has their own unique vocabulary. It is nearly impossible to predict such vocabulary in advance and register it in a word list.
[0008] The present invention has been made in consideration of the above-mentioned problems of the conventional technology, and in one aspect, it is an object of the present invention to provide an information processing device for transcribing voice data that can flexibly accommodate user-specific vocabulary without imposing an excessive burden on the system provider and / or user. [Means for solving the problem]
[0009] In one aspect, the present invention provides a transcription data acquisition means for inputting speech data into a speech recognition model and acquiring transcription data of the speech data; prompt generation means for generating a prompt including an instruction for causing a large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data; a proofreading information acquisition means for inputting the generated prompt into a large-scale language model and acquiring proofreading information output from the large-scale language model; and a first proofreading means for generating first proofread data in which transcription errors contained in the transcription data have been corrected based on the acquired proofreading information; An information processing device comprising: a second calibration means for accepting user calibration of the first calibration data; a user proofreading history information storage means for comparing second proofread data obtained by reflecting the user proofreading on the first proofread data with the transcription data, and for words proofread between the first proofread data and the transcription data, storing in a storage unit user proofreading history information that associates the word before proofreading with the word after proofreading; Furthermore, The above problem is solved by providing an information processing device characterized in that the prompt generation means refers to the memory unit and generates a prompt that includes some or all of the user proofreading history information stored in the memory unit, in addition to instructions to cause a large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data.
[0010] As described above, the information processing device according to one aspect of the present invention includes a second proofreading unit that accepts user proofreading of first proofread data output from the large-scale language model. The second proofreading unit then compares the second proofread data, which reflects the accepted user proofreading in the first proofread data, with the transcription data output from the speech recognition model to identify and store the proofread words in both data as user proofreading history information. This information processing device then uses the results to generate prompts for proofreading the transcription data in the large-scale language model. The information processing device according to one aspect of the present invention provides the user with the convenience of proofreading the transcription data themselves. Furthermore, user proofreading history information, which includes correspondences between pre-proofreading words and post-proofreading words and reflects the user's satisfactory proofreading content, is automatically stored in the storage unit upon user proofreading. This eliminates the user's need to go to the trouble of constructing a word list. Furthermore, the user proofreading history information generated based on the results of user proofreading of actual transcription data enables efficient construction of a word list customized to each user's vocabulary and focused on vocabulary that is lacking in the large-scale language model. Therefore, according to the information processing device relating to one aspect of the present invention, it is possible to flexibly respond to user-specific vocabulary without causing excessive effort to the user and / or system provider, thereby dramatically improving the user experience of the transcription tool.
[0011] In a preferred aspect, the prompt generation means may refer to the storage unit to generate a prompt including proofread words included in the user proofreading history information stored in the storage unit. That is, in a preferred aspect, some or all of the user proofreading history information included in the prompt generated by the prompt generation means may be words included as proofread words in the user proofreading history information stored in the storage unit.
[0012] The words included as proofread words in the user proofreading history information stored in the storage unit are words that have been used by the user in the past and that the large-scale language model has mistakenly proofread or overlooked transcription errors in. Such words are likely to be words that are insufficient in the user's vocabulary and / or the large-scale language model, and by providing the large-scale language model with prompts that include these words, it may be possible to flexibly accommodate the user's vocabulary.
[0013] In a preferred aspect, the prompt generation means may refer to the storage unit to generate a prompt that includes, in addition to words included as proofread words in the user proofreading history information stored in the storage unit, corresponding words before proofreading. That is, in a preferred aspect, some or all of the user proofreading history information included in the prompt generated by the prompt generation means may be words included as proofread words in the user proofreading history information stored in the storage unit and words included as pre-proofread words for those words.
[0014] When an information processing device according to one aspect of the present invention has the above-described configuration, the large-scale language model can know what words each word has been proofread in the past, and therefore may be able to properly proofread even words that are not in its vocabulary or are difficult to understand.
[0015] On the other hand, in a preferred aspect, the prompt generation means may refer to the storage unit to generate a prompt including, among words included as proofread words in the user proofreading history information stored in the storage unit, words whose corresponding pre-proofreading words are included in the transcription data to be proofread. That is, in a preferred aspect, some or all of the user proofreading history information included in the prompt generated by the prompt generation means, among words included as proofread words in the user proofreading history information stored in the storage unit, words whose corresponding pre-proofreading words are included in the transcription data to be proofread. Note that, here, the "corresponding pre-proofreading words" refer to words that are associated with the words included as proofread words in the user proofreading history information stored in the storage unit and that are included as pre-proofreading words.
[0016] When an information processing device according to one aspect of the present invention has the above-described configuration, the large-scale language model is provided with a concentrated set of proofread words related to words contained in the transcription data to be proofread, thereby reducing the risk that the large-scale language model will overlook words provided as user proofreading history information.
[0017] In a preferred embodiment, the user proofreading history information storage means may store, as user proofreading history information, in addition to the pre-proofreading words and the proofreading words of the words proofread between the second proofreading data and the transcription data, pre-proofreading sentences including the pre-proofreading words and / or proofreading sentences including the proofreading words, in association with each other. In this case, the user proofreading history information stored in the storage unit includes, in addition to the pre-proofreading words and the proofreading words of the words proofread between the second proofreading data and the transcription data, pre-proofreading sentences including the pre-proofreading words and / or proofreading sentences including the proofreading words, in association with each other.
[0018] Furthermore, when an information processing device according to one aspect of the present invention has the above-described configuration, the prompt generation means can generate prompts that include, in addition to words included as proofread words in the user proofreading history information stored in the storage unit, pre-proofreading sentences including the corresponding pre-proofreading words and / or proofreading sentences including the proofreading words. That is, in a preferred embodiment, some or all of the user proofreading history information included in the prompts generated by the prompt generation means words included as proofreading words in the user proofreading history information, as well as pre-proofreading sentences including the corresponding pre-proofreading words and / or proofreading sentences including the proofreading words. Note that, here, the pre-proofreading sentences including the corresponding pre-proofreading words and / or proofreading sentences including the proofreading words refer to sentences that are associated with the words included as proofreading words in the user proofreading history information stored in the storage unit and are included as pre-proofreading sentences including the pre-proofreading words of the words and / or proofreading sentences including the words (proofreading words).
[0019] When an information processing device according to one aspect of the present invention has the above configuration, the large-scale language model can know, for a word included in transcription data, not only what word the word has been proofread in the past, but also what word the word was proofread in and in what context. Therefore, even when there are multiple proofreading candidates, such as in the case of homonyms, it may be possible to perform appropriate proofreading according to the context.
[0020] In a preferred embodiment, the information processing device according to an aspect of the present invention may further include an estimation means for estimating the pre-proofreading and post-proofreading pronunciations of words proofread between the second proofreading data and the transcription data. In this case, the user proofreading history information storage means may store, in addition to the pre-proofreading and post-proofreading pronunciations of words proofread between the second proofreading data and the transcription data, the pre-proofreading and / or post-proofreading pronunciations estimated by the estimation means in association with each other. The pre-proofreading and post-proofreading pronunciations of words may be used to improve the functionality of the information processing device according to an aspect of the present invention by an appropriate method.
[0021] For example, when an information processing device according to one aspect of the present invention includes the estimation means, the information processing device according to one aspect of the present invention may further include a determination means for determining whether the edit distance between the pronunciation of a word before proofreading and the pronunciation of a word after proofreading, estimated by the estimation means, is within a predetermined range. If the determination means determines that the edit distance is within the predetermined range, the user proofreading history information storage means may store the user proofreading history information in the storage unit. Proofreading to a word with a large edit distance, in other words, a word with a significantly different pronunciation, is likely to be beyond the scope of transcription, i.e., a forceful change that does not reflect the content of the speech data. Even if such proofreading was performed by a user, it is considered that such proofreading content should not be used as reference by a large-scale language model that performs the task of proofreading transcription errors. When the information processing device according to one aspect of the present invention includes the above configuration, it is possible to use the edit distance as a measure to determine whether the user's proofreading is excessive, exceeding the scope of proofreading of the transcription result, and to exclude excessive proofreading that exceeds the scope of proofreading of the transcription result from the user proofreading history information stored in the storage unit.
[0022] Furthermore, when the information processing device according to one aspect of the present invention includes the estimation means, the prompt generation means may generate a prompt that includes, in addition to a word included as a proofread word in the user proofreading history information stored in the storage unit, the reading of the word. Furthermore, when the prompt generated by the prompt generation means includes, in addition to the proofread word, a word before proofreading, the prompt may be generated that further includes the reading of the word before proofreading.
[0023] When an information processing device according to one aspect of the present invention has the above configuration, it is possible to know not only what word a certain word has been proofread in the past, but also what pronunciation the word has and what pronunciation it has been proofread into. Therefore, for example, even if the transcription data to be proofread contains a word whose meaning is unclear in context but the word itself is not included in the user proofreading history information, it may be possible to properly proofread the word by referring to the proofreading history of words with similar pronunciations.
[0024] It is preferable that the user proofreading history information included in the prompt generated by the prompt generation means is included in the prompt as a word list and / or user proofreading history information. When certain information is included in the prompt as a word list and / or user proofreading history information, it means that the information is included in the prompt with the explicit indication that it is a word list and / or user proofreading history information to be referenced.
[0025] In another aspect, the present invention also provides, for example, a transcription system including the information processing device according to one aspect of the present invention, and a computer program for causing a computer to function as the information processing device. [Effects of the Invention]
[0026] According to the present invention, an information processing device for transcribing speech data that flexibly accommodates a user's unique vocabulary without requiring excessive effort on the part of the user can be provided, thereby dramatically improving the user experience of the transcription tool. [Brief explanation of the drawings]
[0027] [Figure 1] 1 is a diagram showing an example of the configuration of a transcription system including an information processing device according to one aspect of the present invention. [Figure 2] 10 is a diagram showing an example of user calibration history information stored in the external storage 400. FIG. [Figure 3] 10 is a diagram showing an example of user calibration history information stored in the external storage 400. FIG. [Figure 4] FIG. 1 is a diagram showing an example of the flow of information processing by a transcription system including an information processing device according to one aspect of the present invention. [Figure 5] FIG. 1 is a diagram showing an example of the flow of information processing by a transcription system including an information processing device according to one aspect of the present invention. [Figure 6] FIG. 1 is a diagram showing an example of the flow of information processing by a transcription system including an information processing device according to one aspect of the present invention. [Figure 7] FIG. 10 is a diagram illustrating an example of transcription data to be proofread. [Figure 8] FIG. 10 is a diagram showing an example of a prompt template stored in external storage 400. [Figure 9] 10 is a diagram showing an example of user proofreading history information included in a prompt generated by a prompt generating unit 133. FIG. [Figure 10] FIG. 10 is a diagram illustrating an example of proofreading information output by a large-scale language model. [Figure 11] FIG. 10 is a diagram showing an example of transcription data (first proofread data) proofread based on proofreading information output by a large-scale language model. [Figure 12] FIG. 10 is a diagram showing an example of a display screen for user calibration of first calibration data and information written thereon. [Figure 13] FIG. 10 is a diagram showing an example of a display screen for user calibration of first calibration data and information written thereon. [Figure 14] 10 is a diagram showing an example of calibration history information stored in the external storage 400. FIG. [Figure 15] 10 is a diagram showing an example of calibration history information stored in the external storage 400. FIG. [Figure 16] 10 is a diagram showing an example of proofreading history information included in a prompt generated by a prompt generating unit 133. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0028] Hereinafter, embodiments of the present invention will be described in detail by way of example with reference to the drawings. However, the following embodiments are merely examples of embodiments of the present invention and are not intended to limit the technical scope of the present invention.
[0029] <1. Transcription system configuration> Fig. 1 shows the overall configuration of a transcription system including an information processing device 100 according to one embodiment of the present invention. As shown in Fig. 1, the transcription system includes the information processing device 100, an external information processing device 200, an external information processing device 300, and an external storage 400, and is communicatively connected to a user terminal 500, which is a terminal device used by a user, via a network N. The information processing device 100 cooperates with the external information processing device 200, the external information processing device 300, and the external storage 400 to perform transcription processing of audio data.
[0030] The "audio data" to be transcribed is data containing audio. The audio data may be in any format as long as it contains audio and can be input to a speech recognition model, but may be, for example, an audio data file. The audio data file may be, for example, an MP3, M4A, or WAV file, but is not limited to these. The audio data may also be provided as data including data other than audio, such as video. The file format of data including audio data and other data other than audio may be, for example, an MP4 or WebM file, but is not limited to these.
[0031] On the other hand, "transcription" means converting the voice contained in the voice data into text, in other words, converting the voice data into text data. "Transcription data" is data containing text obtained by converting the voice contained in the voice data into text. The format of the transcription data output by the information processing device 100 may basically be any format, but as an example, it may be a file in TXT or CSV format, for example.
[0032] The information processing device 200 is an information processing device equipped with a voice recognition model, and is, for example, a cloud server distributed over a network. When the information processing device 200 receives a request to transcribe voice data from the information processing device 100, the information processing device 200 inputs the voice data into the voice recognition model and performs transcription processing of the voice data. The information processing device 200 transmits the transcription data output by the voice recognition model to the information processing device 100, and the information processing device 100 receives the transcription data.
[0033] The speech recognition model is a trained model that can output transcription data in response to input speech data, and can basically be any speech recognition model as long as it is capable of outputting transcription data in response to input speech data. For example, it can be a transformer-based speech recognition model, and specifically, it can be, for example, Whisper (OpenAI), Wav2Vec2.0 (Meta), HuBERT (Meta), Conformer (Google), etc., but is not limited to these.
[0034] The speech recognition model may be fine-tuned as appropriate to generate transcription data, etc. For example, the model may be fine-tuned to understand technical terms appropriate to a particular field, to understand terms used within a specific organization such as a company, to understand a specific language, and / or to understand the habits of speakers within a specific organization such as a company.
[0035] Furthermore, the voice data may be preprocessed by an appropriate method before being transcribed using a voice recognition model. For example, the file format of the voice data may be converted into another file format, or noise contained in the voice data may be removed. The preprocessing of the voice data may be performed by the information processing device 100 or by another information processing device, for example, the information processing device 200.
[0036] The information processing device 300 is an information processing device equipped with a large-scale language model, and is, for example, a cloud server distributed across a network. When the information processing device 300 receives a proofreading request for transcription data from the information processing device 100, it inputs a prompt specified in the proofreading request into the large-scale language model and performs a proofreading process for the transcription data. The prompt includes an instruction to output proofreading information for correcting transcription errors contained in the transcription data to be proofread. Upon receiving the prompt including the instruction, the large-scale language model outputs the proofreading information in accordance with the instruction described in the prompt. The information processing device 300 transmits the proofreading information output by the large-scale language model to the information processing device 100. Upon receiving the proofreading information, the information processing device 100 performs a process of generating first proofread data in which transcription errors contained in the transcription data to be proofread have been corrected based on the proofreading information.
[0037] The large-scale language model may be, for example, a Transformer-based large-scale language model, such as, but not limited to, "GPT" provided by OpenAI, Inc., "Gemini" provided by Google DeepMind, Inc., "Claude" provided by Anthropic, Inc., or "tsuzumi" provided by Nippon Telegraph and Telephone Corporation.
[0038] The large-scale language model may be fine-tuned as appropriate, for example, to understand technical terms specific to a field, to understand the terms used within a specific organization such as a company, to understand a specific language, and / or to understand the habits of speakers within a specific organization such as a company.
[0039] Furthermore, the transcription data may be preprocessed by an appropriate method before proofreading using a large-scale language model, for example, by dividing and / or summarizing the transcription data into predetermined units. The preprocessing of the transcription data may be performed by the information processing device 100 or by another information processing device, for example, the information processing device 300.
[0040] The external storage 400 is a storage device that stores appropriate information, such as an online storage. In this example, the external storage 400 stores a user proofreading history information database, which stores user proofreading history information. In this example, the user proofreading history information database stores, as user proofreading history information, words before proofreading, words after proofreading, readings of the words before proofreading, readings of the words after proofreading, sentences before proofreading that include the words before proofreading, and sentences after proofreading that include the words after proofreading, in association with each other. Examples of user proofreading history information stored in the user proofreading history information database are shown in FIGS. 2 and 3. As shown in FIGS. 2 and 3, in this example, the user proofreading history information is stored in two tables, and each entry is associated with a proofreading ID.
[0041] The external storage 400 also has a voice data database, a transcription data database, a first proofreading data database, and a second proofreading data database, and each database stores voice data provided from the user terminal 500, transcription data generated in the information processing device 200, and first proofreading data and second proofreading data generated in the information processing device 100 in a manner that allows them to be associated with each other.
[0042] In FIG. 1, the information processing devices 100, 200, and 300 are each depicted as if they were a single physical information processing device, but this does not necessarily have to be a single physical information processing device. For example, they may be information processing devices distributed across a network, so-called cloud servers. In addition, in FIG. 1, the external storage 400 is depicted as if it were a single physical storage device, but this does not necessarily have to be a single physical storage device. For example, they may be storage devices distributed across a network, so-called cloud storage.
[0043] 2. Configuration of Information Processing Device 100 The information processing device 100 will be described in more detail below.
[0044] The information processing device 100 may be, for example, a terminal device such as a personal computer, a smartphone, or a tablet terminal, a server device, a cloud server, or the like.
[0045] The information processing device 100 conceptually includes a communication unit 110, a storage unit 120, and a processing unit .
[0046] The communication unit 110 is a means for transmitting and / or receiving various data between various devices via an appropriate network such as wireless LAN, wired LAN, or mobile communications such as 3G, 4G, or 5G, and may be configured to include, for example, a NIC (Network Interface Card).
[0047] The storage unit 120 is a means for storing various data including programs for starting and operating the information processing device 100, and is configured to include, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc. The storage unit 120 has a non-volatile area and a volatile area, and the volatile area is configured to include, for example, a RAM, etc., and provides a working area for the CPU, etc. On the other hand, the non-volatile area is configured to include, for example, a ROM, an HDD, an SSD, etc.
[0048] The control unit 130 includes a processor such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit). The control unit 101 reads various programs from the storage unit 102, executes the programs, and controls the operation of the information processing device 100. In this way, various functions of the information processing device 100 are realized.
[0049] In terms of functional concept, the control unit 130 includes a transcription request acquisition unit 131, a transcription data acquisition unit 132, a prompt generation unit 133, a proofreading information acquisition unit 134, a first proofreading unit 135, a second proofreading unit 136, and a user proofreading history information storage unit 137.
[0050] <Transcription request acquisition unit 131> The transcription request acquisition unit 131 is a means for acquiring a transcription request for audio data.
[0051] The transcription request acquisition unit 131 receives a transcription request for voice data from the user terminal 500 via the communication unit 110, for example, and thereby acquires the transcription request for voice data.
[0052] A request for transcription of audio data typically includes information identifying the audio data to be transcribed and / or the audio data to be transcribed. The information identifying the audio data to be transcribed may be, for example, a file path of the audio data. In this example, the audio data to be transcribed has been uploaded to the external storage 400 via an operation from the user terminal, and therefore the request for transcription of audio data includes the file path of the audio data uploaded to the external storage 400.
[0053] <Transcription Data Acquisition Unit 132> The transcription data acquisition unit 132 is a means for inputting voice data into a voice recognition model and acquiring transcription data of the voice data output from the voice recognition model in response to the input.
[0054] When the transcription request acquisition unit 131 receives a transcription request for audio data, the transcription data acquisition unit 132 transmits a transcription request for audio data to the information processing device 200 based on the transcription request. The transcription request for audio data transmitted to the information processing device 200 is accompanied by audio data to be transcribed and / or information identifying the audio data to be transcribed (in this example, the file path of the audio data), similar to the transcription request for audio data received by the information processing device 100 from the user terminal 500.
[0055] When the information processing device 200 receives a transcription request for audio data, the information processing device 200 inputs the audio data attached to the transcription request or the audio data acquired based on information identifying the audio data attached to the transcription request into a speech recognition model. In this example, the transcription request for audio data received by the information processing device 200 is accompanied by a file path of the audio data as information identifying the audio data to be transcribed, so the information processing device 200 acquires the audio data from the external storage 400 and inputs it into the speech recognition model.
[0056] When voice data is input to the voice recognition model, the voice recognition model converts the voice data into text data and outputs transcription data corresponding to the input voice data. The information processing device 200 transmits the transcription data output by the voice recognition model to the information processing device 100. The information processing device 100 receives the transcription data via the communication unit 110. In this manner, the transcription process is achieved in which voice data is input to the voice recognition model and transcription data of the voice data is acquired. The transcription data acquired by the transcription data acquisition unit 132 is stored in an appropriate storage unit. In this example, it is stored in a transcription data database in the external storage 400.
[0057] <Prompt Generation Unit 133> The prompt generation unit 133 is a means for generating a prompt including an instruction to cause the large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data.
[0058] A prompt including an instruction to cause a large-scale language model to output proofreading information for correcting transcription errors contained in transcription data can be generated using, for example, a predetermined prompt template. The prompt template is a template for generating a prompt, and contains pre-defined routines of instructions to cause a large-scale language model to output proofreading information for correcting transcription errors contained in transcription data. A prompt can be generated by writing, for example, the transcription data to be proofread, information identifying the transcription data to be proofread (e.g., the file path of the transcription data), and information that is particularly desired to be provided to the large-scale language model when proofreading the transcription data (e.g., user proofreading history information, described below) in the prompt template. The prompt template may be stored in the storage unit 120 of the information processing device 100 or in an external storage device. In this example, the prompt template is stored in the external storage 400, and the information processing device 100 appropriately obtains the prompt template from the external storage 400 and uses it to generate a prompt.
[0059] The prompt generation means references an appropriate storage unit in which the user proofreading history information is stored, and generates a prompt that includes some or all of the user proofreading history information stored in the storage unit. That is, the prompt generated by the prompt generation means may further include some or all of the user proofreading history information. As described above, in this example, the user proofreading history information is stored in the user proofreading history information database of the external storage 400, so the prompt generation means references the user proofreading history information database of the external storage 400, and generates a prompt that includes some or all of the user proofreading history information stored in the database.
[0060] Some or all of the user proofreading history information included in the prompt may be, for example, proofread words included in the user proofreading history information stored in the user proofreading history information database of the external storage 400, i.e., words included as proofread words in the user proofreading history information. The proofread words stored in the user proofreading history information database are words that have been used by users in the past and are also words that the large-scale language model has mistakenly proofread or overlooked in proofreading a transcription error. In this way, by providing the large-scale language model with prompts that include words that are likely to be the user's unique vocabulary and / or vocabulary that the large-scale language model lacks, it may be possible to flexibly accommodate the user's unique vocabulary.
[0061] There is no particular limit to the number of proofread words that can be included in a prompt, i.e., the number of words included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400, but it may be, for example, 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, 20 or more, or more. When setting a limit on the number of proofread words that can be included in a prompt, for example, proofread words that have been proofread more frequently may be preferentially included based on the number of times they have been proofread. In this case, the number of times each word has been proofread may be recorded in the user proofreading history information database of the external storage 400 in association with the word before proofreading and / or the word after proofreading.
[0062] In a preferred aspect, the prompt generation unit preferably generates a prompt including words that are included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400, and whose corresponding pre-proofreading words are included in the transcription data to be proofread. In other words, in a preferred aspect, the prompt generated by the prompt generation means does not include all of the words that are included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400, but includes only the words that are included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400, and whose corresponding pre-proofreading words are included in the transcription data to be proofread. Large-scale language models generally have difficulty remembering to continue applying tasks that were initially instructed to them for long input data. Therefore, by narrowing down the user proofreading history information to be included in the prompt to only that related to the transcription data to be proofread, it is possible to prevent the large-scale language model from overlooking something.
[0063] On the other hand, the prompt generated by the prompt generation unit may include, as part or all of the user proofreading history information, other information included in the user proofreading history information in association with a word included as a proofread word in the user proofreading history information stored in the user proofreading history information database of the external storage 400. For example, the prompt generated by the prompt generation unit may include a word included as a pre-proofread word for the word in the user proofreading history information stored in the user proofreading history information database of the external storage 400, and in this case, it is preferable that the pre-proofread word be included in the prompt in association with the corresponding proofread word.
[0064] In a preferred embodiment, the prompt generation unit preferably generates prompts that include, in addition to words included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400, pre-proofread sentences including the corresponding pre-proofread words and / or proofread sentences including the proofread words, in association with each other. In this case, the user proofreading history information storage unit 137, which will be described later, is equipped with means for storing, in association with each other, pre-proofread sentences including the pre-proofread words and / or proofread sentences including the proofread words, in addition to the pre-proofread words and proofread words of the words proofread between the second proofreading data in which the user proofreading is reflected in the first proofreading data and the transcription data output by the speech recognition model.
[0065] In this example, the user proofreading history information database of external storage 400 stores user proofreading history information in correspondence with each other, such as pre-proofreading words, proofreading words, readings of pre-proofreading words, readings of proofreading words, pre-proofreading sentences containing pre-proofreading words, and proofreading sentences containing proofreading words.Therefore, by referring to the user proofreading history information database of external storage 400, the prompt generation unit can obtain, in addition to the words included as proofreading words in the user proofreading history information stored in the user proofreading history information database of external storage 400, pre-proofreading sentences containing the corresponding pre-proofreading words and / or sentences stored in the user proofreading history information database as proofreading sentences containing proofreading words, and generate prompts that include this user proofreading history information.
[0066] In addition, it is preferable that the prompts generated by the prompt generation means include, with explicit indication that they are user proofreading history information, the proofread words, the words before proofreading, the sentences containing the proofread words, the sentences containing the words before proofreading, the readings of the words after proofreading, the readings of the words before proofreading, etc. It is also preferable that the proofread words and the words before proofreading are included with explicit indication that they are pre-proofreading words and proofread words, and this also applies to sentences containing proofread words and sentences containing pre-proofreading words.
[0067] On the other hand, the proofreading information for proofreading transcription errors contained in the transcription data may be, for example, text data proofread for the transcription errors contained in the transcription data (i.e., proofread data), but in a preferred embodiment, it may be proofreading information that associates pre-proofreading text containing transcription errors contained in the transcription data with proofread text of the pre-proofreading text. As disclosed in Patent Document 3, according to the findings of the present inventors, by outputting proofreading information that associates pre-proofreading text containing transcription errors contained in the transcription data with proofread text of the pre-proofreading text, rather than directly outputting proofread data proofread for errors contained in the transcription data, it is possible to reduce proofreading omissions in large-scale language models and enable better proofreading of transcription data.
[0068] In a preferred aspect, the instructions for causing the large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data may further include instructions for outputting the pre-proofreading text and the proofread text as proofreading information only if the pronunciations of the pre-proofreading text and the proofread text are identical or similar. Such instructions may be, for example, instructions to output the pre-proofreading text and the proofread text as proofreading information only if the pronunciations of the pre-proofreading text and the proofread text are identical or similar. The similarity in pronunciation between the pre-proofreading text and the proofread text may be determined independently by the large-scale language model, or a definition or criterion for pronunciation similarity may be provided.
[0069] The pre-proofreading text and the proofread text may be output in any unit. For example, the pre-proofreading text and the proofread text may be output sentence by sentence, utterance by utterance, or word by word or phrase by utterance. Furthermore, the proofreading information may include other elements in addition to the pre-proofreading text and the proofread text. Examples of other elements included in the proofreading information include, for example, information that can identify the proofreading point, i.e., the location of the pre-proofreading text, in the transcription data. Such information may be any information as long as it can identify the location of the pre-proofreading text in the transcription data. For example, a sentence or utterance containing the pre-proofreading text, and, if necessary, one or more utterances or sentences before or after it, may be output. Furthermore, the location of the pre-proofreading text in the transcription data may be identified by the number of lines, characters, bytes, paragraph number, or the order or number of the sentence or utterance in the entire transcription data.
[0070] <Calibration information acquisition unit 134> The proofreading information acquisition unit 134 is a means for inputting the prompt generated by the prompt generation unit 133 into a large-scale language model and acquiring proofreading information for proofreading the transcription data output from the large-scale language model in accordance with the input.
[0071] The proofreading information acquisition unit 134 transmits a proofreading request for the transcription data to the information processing device 300. The proofreading request for the transcription data transmitted to the information processing device 300 is accompanied by a prompt including information that identifies the transcription data to be proofread and / or the transcription data to be proofread.
[0072] When the information processing device 300 receives a request to proofread transcription data, it inputs the transcription data attached to the proofreading request or a prompt including information identifying the transcription data attached to the transcription request into a large-scale language model.
[0073] When a prompt is input to the large-scale language model, the large-scale language model performs a proofreading process on the transcription data in accordance with instructions included in the prompt, and outputs proofreading information corresponding to the input transcription data. The information processing device 300 transmits the proofreading information output by the large-scale language model to the information processing device 100. The information processing device 100 receives the proofreading information via the communication unit 110. In this way, the process of inputting the generated prompt to the large-scale language model and acquiring the proofreading information output from the large-scale language model is achieved. The acquired proofreading information is transmitted to the first proofreading unit 135.
[0074] <1st proofreading section 135> The first proofreading unit 135 is a means for receiving the proofreading information acquired by the proofreading information acquisition unit 134 and generating first proofreading data in which transcription errors contained in the transcription data have been corrected based on the received proofreading information.
[0075] There are no particular limitations on the specific method by which the first proofreading unit 135 proofreads the transcription data based on the proofreading information and generates first proofread data in which the transcription errors contained in the transcription data have been proofread; for example, if the proofreading information acquired by the proofreading information acquisition unit 134 includes pre-proofreading text containing transcription errors in association with proofread text of the pre-proofreading text, a process can be performed in which each of the pre-proofreading texts contained in the transcription data to be proofread is replaced with the corresponding proofread text.
[0076] The first calibration data generated by the first calibration unit 135 is stored in an appropriate storage unit. In this example, it is stored in a first calibration data database in the external storage 400.
[0077] <Second proofing section 136> The second calibration unit 136 is a means for receiving user calibration of the first calibration data generated by the first calibration unit 135.
[0078] There are no particular limitations on the specific method by which the second calibration unit 136 accepts user calibration of the first calibration data. For example, a display screen (calibration screen) including the first calibration data may be displayed on an appropriate display unit, and the user calibration may be accepted on the display screen. Alternatively, the second calibration unit 136 may accept input of second calibration data as user calibration. In this example, a display screen displaying the first calibration data in an editable format is displayed on a display unit (not shown) of the user terminal 500, and the user calibration is accepted on the display screen. More specifically, the user is prompted to edit the first calibration data on the display screen, and the edited calibration data is acquired as second calibration data. That is, in this example, the second calibration unit 136 is configured to accept user calibration via the display screen displayed on the display unit of the user terminal 500 and acquire second calibration data reflecting the user calibration. Meanwhile, the second calibration unit 136 may be calibrated to generate second calibration data based on the accepted user calibration, as necessary.
[0079] The second calibration unit 136 stores the second calibration data in a second calibration data database in an appropriate storage unit, in this example, the external storage 400 .
[0080] The second proofread data is proofread data in which a user further proofreads first proofread data, which is obtained by proofreading transcription data output by a speech recognition model using a large-scale language model, and can be used as a final transcription result in a transcription system including an information processing device according to an aspect of the present invention. Therefore, the information processing device according to an aspect of the present invention may be configured to further include an output unit that outputs the second proofread data in response to a distribution request from, for example, a user terminal 500, as necessary. In this case, if the second proofread data does not exist, the first proofread data may be output instead of the second proofread data. This is because the first proofread data that has not been accepted for user proofreading by the second proofreading unit 136 can be considered to be proofread data that does not require user proofreading and can be used for output as is.
[0081] <Proofreading history information storage unit 137> The proofreading history information storage unit 137 is a means for storing, in the storage unit, user proofreading history information that associates the word before proofreading with the word after proofreading, for a word proofread between the second proofreading data that reflects the user proofreading received by the second proofreading unit 136 and the transcription data generated by the speech recognition model. In this example, the user proofreading history information storage unit 137 conceptually includes a proofreading part extraction unit 137A, an estimation unit 137B, a writing unit 137C, and a determination unit 137D.
[0082] The proofreading part extraction unit 137A, for example, compares the transcription data stored in the external storage 400 with the second proofreading data, identifies words that have been proofread between the second proofreading data that reflects the user proofreading received by the second proofreading unit 136 and the transcription data generated by the speech recognition model, and extracts the word before and after proofreading.
[0083] There are no particular limitations on the specific method by which the proofreading portion extraction unit 137A identifies and extracts the words proofread between the second proofread data reflecting the user proofreading received by the second proofreading unit 136 and the transcription data generated by the speech recognition model, as well as the words before and after proofreading, but for example, the words before and after proofreading may be extracted by analyzing the second proofread data and the transcription data using a morphological analysis tool and a difference algorithm. Also, the proofreading portions may be extracted using a large-scale language model.
[0084] In this example, the pre-proofreading words and post-proofreading words extracted by the proofreading portion extraction unit 137A are transmitted to the estimation unit 137B, which estimates their readings. On the other hand, the extracted pre-proofreading words and post-proofreading words may be transmitted directly to the writing unit 137C. In this case, the writing unit 137C may store the pre-proofreading words and post-proofreading words extracted by the proofreading portion extraction unit 137A in the external storage 400 before the estimation unit 137B estimates the reading of each word. In this case, the estimation unit 137B may be configured to read out and use each word stored in the external storage 400.
[0085] The estimation unit 137B is a means for estimating the reading of a given word extracted by the proofreading part extraction unit 137A, and in this example, it functions as a means for estimating the reading of a word before and after proofreading of a word that has been proofread between the second proofreading data and the transcription data.
[0086] The pronunciation of a word can be estimated using, for example, a large-scale language model. Specifically, for example, the estimation unit 137B generates a prompt including an instruction to estimate the pronunciation of the word before and after proofreading extracted by the proofreading portion extraction unit 137A based on a predetermined prompt template. Next, a request for estimating the pronunciation of the word accompanied by the generated prompt is transmitted to the information processing device 300. The information processing device 300 inputs the prompt included in the received estimation request into the large-scale language model, acquires the pronunciation of the word before and after proofreading output from the large-scale language model in response to the input of the prompt, and transmits this to the information processing device 100. This enables the estimation unit 137B to estimate the pronunciation of the word before and after proofreading extracted by the proofreading portion extraction unit 137A. The large-scale language model that can be used to estimate the pronunciation of a word is the same as the large-scale language model described in the description of the “proofreading information acquisition unit 134.”
[0087] The estimation unit 137B transmits the estimated readings of the pre-proofreading words and the proofreading words to the writing unit 137C along with the pre-proofreading words and the proofreading words extracted by the proofreading portion extraction unit 137A. Note that if the proofreading portion extraction unit 137A has already transmitted the pre-proofreading words and the proofreading words to the writing unit 137C and these words have already been stored in the external storage 400, the estimation unit 137B does not necessarily need to transmit the pre-proofreading words and the proofreading words to the writing unit 137C as long as the writing unit 137C can associate each reading with the reading of the word.
[0088] The writing unit 137C is a means for recording user proofreading history information that associates the pre-proofreading words extracted by the proofreading portion extraction unit 137A with the proofreading words in an appropriate storage unit. When the information processing device according to one aspect of the present invention includes means for estimating the pronunciation of each word as in the present example, the writing unit 137C functions as a means for storing, in an appropriate storage unit, in addition to the pre-proofreading words and the proofreading words extracted by the proofreading portion extraction unit 137A, user proofreading history information that associates the pronunciation of the pre-proofreading words with the pronunciation of the proofreading words. In the present example, the storage unit that stores the user proofreading history information is the external storage 400, but may be any storage unit and / or storage device as long as it is accessible by the information processing device as appropriate.
[0089] Meanwhile, in a preferred embodiment, the user proofreading history information may further include pre-proofreading sentences containing the pre-proofreading words and / or proofreading sentences containing the proofreading words, in association with each other. In this case, the prompt generation unit 133 can generate prompts that include not only the pre-proofreading words and the proofreading words, but also the pre-proofreading sentences and / or the proofreading sentences containing these words. This allows the large-scale language model to obtain information not only about which word each word was proofread into, but also about the context in which the word was proofread into, thereby enabling proofreading that takes context into greater consideration.
[0090] The pre-proofreading sentences containing pre-proofreading words and the proofreading sentences containing proofreading words may be all or part of the transcription data and second proofreading data generated by the speech recognition model, respectively. These sentences may basically be in any unit, but may be, for example, sentences of 1 to 5 sentences (e.g., 1 sentence, 2 sentences, 3 sentences, 4 sentences, 5 sentences), utterances of 1 to 5 utterances (e.g., 1 utterance, 2 utterances, 3 utterances, 4 utterances, 5 utterances), or characters of 10 to 400 characters, 15 to 300 characters, 20 to 200 characters, or 25 to 100 characters. Much of the transcription data generated by the speech recognition model is divided into utterances, and the second proofreading data, which is the proofreading data, can also be easily handled utterance by utterance. Therefore, it is most convenient to handle the pre-proofreading sentences and the proofreading sentences in utterance units, particularly, one utterance at a time.
[0091] A sentence unit may be a unit separated by a period. Meanwhile, an utterance unit refers to a speech interval contained in audio data. Audio data consists of speech intervals and silent intervals (or non-speech intervals), and each speech interval is separated by a silent interval. Those skilled in the art can easily recognize speech units contained in audio data. Alternatively, a speech recognition model capable of distinguishing and outputting speech intervals and silent intervals in audio data, such as "Whisper" by OpenAI, may be used. It is also possible to detect speech intervals contained in audio data by using voice activity detection technology. In this case, audio data may be processed by speech activity detection before being transcribed using a speech recognition model.
[0092] In this example, the user proofreading history information storage unit 137 further includes a determination unit 137D. The determination unit 137D determines whether the edit distance between the pronunciation of a word before proofreading and the pronunciation of the word after proofreading is within a predetermined range, and when the determination unit 137D determines that the edit distance is within the predetermined range, the writing unit 137C is configured to record the user proofreading history information in the storage unit.
[0093] The edit distance, sometimes called the Levenshtein distance, is a measure of how similar two strings are. It is defined as the minimum number of steps required to transform one string into another by inserting, deleting, or substituting a single character. The larger the Levenshtein distance, the more dissimilar the two strings are considered to be. While there are no specific limitations on the range, the normalized edit distance may be within the following ranges: 0.50 or less, 0.45 or less, 0.40 or less, 0.35 or less, 0.30 or less, 0.25 or less, or 0.20 or less. The normalized edit distance is calculated by dividing the edit distance by the number of characters in the string (word) with the longer reading.
[0094] When an information processing device according to one aspect of the present invention has the above configuration, it is possible to use the edit distance as a measure to determine whether a user's proofreading is excessive and exceeds the range of proofreading for the transcription result, and to exclude excessive proofreading that exceeds the range of proofreading for the transcription result from the user proofreading history information recorded in the storage unit. Proofreading for words that are pronounced too differently is a forceful proofreading that goes beyond the scope of transcription and is likely to be unfaithful to the content of the speech data. Such user proofreading, even if it was done by a user, is not appropriate for correcting transcription errors. By excluding such proofreading from the user proofreading history information, it may be possible to provide appropriate user proofreading history information for a large-scale language model.
[0095] <3. An example of information processing> Next, a description will be given of an example of information processing by the information processing device 100. The flow of an example of information processing by the information processing device 100 is shown in FIGS.
[0096] The user operates the user terminal 500 to transmit voice data and a request for transcription of the voice data to the information processing device 100 (step S1). For example, a predetermined display screen provided by the transcription system according to this embodiment is displayed on the screen of the user terminal 500, and the voice data to be transcribed is uploaded to the external storage 400 via the display screen. At the same time, the user presses a transcription execution button displayed on the display screen, whereby a request for transcription of the voice data is transmitted to the information processing device 100.
[0097] This transcription request is a transcription request that specifies the audio data to be transcribed, and includes information that specifies the audio data to be transcribed. The information that specifies the audio data to be transcribed may basically be any information as long as the information processing device 100 that received the transcription request can identify and access the audio data to be transcribed, but it may be, for example, a file path of the audio data to be transcribed. In this example, the transcription request for audio data sent to the information processing device 100 includes the file path of the audio data uploaded to the external storage 400.
[0098] When the information processing device 100 receives a request to transcribe the voice data from the user terminal 500, the information processing device 100 transmits the request to transcribe the voice data to the information processing device 200 that includes a voice recognition model (step S2).
[0099] This transcription request, like the transcription request sent from the user terminal 500 to the information processing device 100, is a transcription request that specifies the audio data to be transcribed, and includes information that specifies the audio data to be transcribed (in this example, the file path of the audio data).
[0100] When the information processing device 200 including the voice recognition model receives the voice data transcription request from the information processing device 100, it acquires the voice data from the external storage 400 based on the file path of the voice data attached to the voice data transcription request, and inputs the acquired voice data to the voice recognition model. This executes the voice data transcription process (step S3). The voice recognition model converts the voice included in the input voice data into text data and outputs it as transcription data.
[0101] The information processing device 200 transmits the output transcription data (text data) to the information processing device 100 (step S4). Upon receiving the transcription data from the information processing device 200, the information processing device 100 stores the transcription data in the external storage 400.
[0102] In this example, the following description will be given on the assumption that the speech recognition model provided in the information processing device 200 outputs transcription data including the text shown in Fig. 7. In the transcription data shown in Fig. 7, utterances included in the audio data are separated by line feed characters. That is, in the transcription data shown in Fig. 7, each line corresponds to one utterance.
[0103] Next, the information processing device 100 generates a prompt for proofreading the transcription data received from the information processing device 200 using a large-scale language model (step S5).
[0104] Specifically, prompt generation unit 133 of information processing device 100 reads out a prompt template stored in external storage 400 and writes necessary information therein to generate a prompt.
[0105] In this example, an example of a prompt template stored in external storage 400 is shown in Fig. 8. As shown in Fig. 8, in this example, the prompt template used to generate a prompt by prompt generation unit 133 includes an instruction to output proofreading information for proofreading transcription errors contained in the transcription data to the large-scale language model, instructing the large-scale language model to output proofreading information that associates pre-proofreading text containing transcription errors contained in the transcription data with proofread text for the pre-proofreading text.
[0106] Meanwhile, as shown in Figure 8, the prompt template has an area indicated by "{sentence}" and an area indicated by "{word_list}." In Figure 8, the text data of the transcription data to be proofread is inserted into the area indicated by "{sentence}," and some or all of the proofreading history information stored in external storage 400 is inserted into the area indicated by "{word_list}." By inserting this information into the prompt template, a prompt is generated that causes the large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data.
[0107] The process of generating user proofreading history information to be inserted into the area indicated by "{word_list}" will be described.
[0108] In this example, an example of user proofreading history information stored in the user proofreading history information database of the external storage 400 is shown in Figures 2 and 3. As shown in Figures 2 and 3, the user proofreading history information database of the external storage 400 stores user proofreading history information for the words "Advance Shipping Nox," "Interfering Agent," and "Red Stock" included in the transcription data.
[0109] Specifically, for the word "Advance Shipping Nokis" (the word before proofreading), the word "Advance Shipping Notice" is stored in correspondence with it as the proofread word for that word, and the reading of the word before proofreading is "Advance Shipping Nokis," and the reading of the word after proofreading is "Advance Shipping Notice," as well as a pre-proofreading sentence containing the pre-proofreading word, "By using Advance Shipping Nokis, stores can find out inventory information before the products arrive," and a proofreading sentence containing the proofreading word, "By using Advance Shipping Notice, stores can find out inventory information before the products arrive." are stored.
[0110] Furthermore, for the word "interfering agent" (the word before proofreading), the word "buffering material" is stored in correspondence with it as the proofread word for that word, and the reading of the word before proofreading is "kanshozai" and the reading of the word after proofreading is "kanshozai," as well as a pre-proofreading sentence containing the pre-proofreading word, "I think it is important to use an interfering agent to pack safely," and a proofreading sentence containing the proofreading word, "I think it is important to use buffering material to pack safely." are stored.
[0111] In addition, for the word "red stock" (the word before proofreading), the word "dead stock" is stored in correspondence with that word as the proofread word, and the reading of the word before proofreading is "RED STOCK" and the reading of the word after proofreading is "DEDD STOCK", as well as a pre-proofreading sentence containing the pre-proofreading word, "By reducing red stock through appropriate inventory management, it is possible to reduce management costs," and a proofreading sentence containing the proofreading word, "By reducing dead stock through appropriate inventory management, it is possible to reduce management costs." are stored.
[0112] The prompt generation unit 133 of the information processing device 100 refers to the user proofreading history information database in the external storage 400, and extracts words whose corresponding pre-proofreading words are included in the transcription data to be proofread, from among the words included as proofread words in the user proofreading history information stored in the user proofreading history information database in the external storage 400. Then, for the extracted word (proofreading word), a prompt is generated that includes the proofreading word, the reading of the proofreading word, and the corresponding sentences before and after proofreading (i.e., the pre-proofreading sentence including the corresponding pre-proofreading word, and the proofreading sentence including the proofreading word), in association with each other.
[0113] In this example, of the proofread words included in the user proofreading history information stored in the user proofreading history information database of the external storage 400, the words whose corresponding pre-proofreading words are included in the transcription data to be proofread are "advance shipping notice" and "dead stock." Therefore, the prompt generation unit 133 generates, as user proofreading history information, prompts that include these proofreading words, the pronunciation of the proofreading words, and the corresponding sentences before and after proofreading (i.e., the pre-proofreading sentences including the corresponding pre-proofreading words and the proofreading sentences including the proofreading words), in association with each other.
[0114] In the prompt generated in this manner, for example, the user proofreading history information shown in Fig. 9 is inserted into the area corresponding to the area indicated by "{word_list}" in the prompt template shown in Fig. 8. As shown in Fig. 9, the prompt generated by the prompt generation unit 133 includes, as user proofreading history information, the proofread words, the pronunciation of the proofread words, and the corresponding sentences before and after proofreading (i.e., the pre-proofreading sentence including the corresponding pre-proofreading word and the proofreading sentence including the proofreading word) for "advance shipping notice" and "dead stock," which are words included as proofread words in the user proofreading history information stored in the user proofreading history information database of the external storage 400 and whose corresponding pre-proofreading words are included in the transcription data to be proofread. In addition, in the user proofreading history information shown in Fig. 9, the sentences before and after proofreading are connected using "→" so that it is clear which is the pre-proofreading sentence and which is the proofreading sentence.
[0115] After generating the prompt, the information processing device 100 transmits a proofreading request for the transcription data to the information processing device 300 equipped with a large-scale language model (step S6). This proofreading request includes the prompt generated by the prompt generation unit 133.
[0116] When the information processing device 300 equipped with a large-scale language model receives a proofreading request for transcription data from the information processing device 100, it inputs the above-mentioned prompt included in the proofreading request into the large-scale language model, causing the large-scale language model to generate proofreading information corresponding to the transcription data (step S7).
[0117] In this example, an example of proofreading information output from the large-scale language model is shown in Figure 10. As shown in Figure 10, in this example, the proofreading information is generated by outputting the sentence "By utilizing Advanced Shipping Nokiss," included in the transcription data, as the pre-proofreading text: Advance Shipping Nokis By utilizing ", the proofread text will be " Advance Shipping NoticeIn addition, the sentence "I want to reduce red stocks," which is included in the transcription data, is translated as " Red Stock I want to reduce the number of ", and after proofreading, I want to use " Deadstock In this way, the output of a large-scale language model contains proofreading information that matches the pre-proofreading text with the proofreading text.
[0118] On the other hand, in this example, the proofreading information is as follows: for the sentence "Is it possible to incorporate the delivery inspection method?" included in the transcription data, the text before proofreading is " Inspected product Is it possible to incorporate this method? ", and after proofreading, " Delivery inspection Is it possible to incorporate the following method? However, the correct translation is: No inspection "Is it possible to incorporate techniques like this?", which means that the large-scale language model incorrectly corrected transcription errors.
[0119] The information processing device 300 transmits the proofreading information generated by the large-scale language model to the information processing device 100 (step S8). The proofreading information may be pre-processed as needed before being transmitted to the information processing device 100. For example, the proofreading information may be pre-processed by deleting unnecessary information, selecting only information identified by a predetermined identifier, or the like.
[0120] When the information processing device 100 receives the proofreading information from the information processing device 300 that includes the large-scale language model, the information processing device 100 proofreads the transcription data based on the received proofreading information (step S9).
[0121] Specifically, the information processing device 100 generates proofread transcription data (first proofread data) by replacing pre-proofread text with proofread text in the sentence included in the transcription data.
[0122] An example of the first proofread data generated in this way is shown in FIG. 11. As shown in FIG. 11, in the first proofread data, the " Red Stock I want to reduce the number of Deadstock I want to reduce it, but,” and the sentence has been corrected. Advance Shipping Nokis By utilizing Advance Shipping Notice By utilizing ", " has been replaced with ", and the sentence has been corrected. Inspected product Is it possible to adopt the following method? Delivery inspection However, the correct sentence is "no inspection," and this revised sentence still contains transcription errors that the large-scale language model has not fully proofread and / or has incorrectly proofread.
[0123] In this way, the information processing device 100 generates proofread transcription data (first proofread data). The information processing device 100 uploads the generated first proofread data to a first proofread data database in the external storage 400 (step S10). The user can access the external storage 400 to appropriately obtain the transcription data of the voice data uploaded to the external storage 400.
[0124] Next, a description will be given of information processing when a user performs user calibration of the first calibration data. First, the user terminal 500 transmits a calibration request for the first calibration data to the information processing device 100 (step S11).
[0125] When the information processing device 100 receives a proofreading request for the first proofreading data from the user terminal 500, it generates a display screen (proofreading screen) for accepting user proofreading of the first proofreading data (step S12) and transmits the generated display screen to the user terminal 500 (step S13).
[0126] The proofreading request for the first proofreading data is accompanied by information identifying the first proofreading data to be subject to user proofreading, and the information processing device 100 identifies the first proofreading data to be subject to proofreading based on the information, and transmits a display screen for proofreading the identified first proofreading data to the user terminal 500.
[0127] An example of the display screen sent to the user terminal 500 and the information displayed thereon is shown in FIG. 12. In the display screen shown in FIG. 12, the first calibration data is displayed in an editable manner, and the 30th line reads: Delivery inspection Would it be possible to adopt this technique?" is displayed as an editable sentence.
[0128] The user can calibrate the first calibration data on the user calibration screen by using an appropriate input means. Specifically, on the display screen shown in FIG. 12, the " Delivery inspection Is it possible to adopt this method? " "Delivery inspection" in the sentence was changed to "No inspection" and No inspection "Is it possible to adopt this method?" Furthermore, for example, on the display screen shown in FIG. 12, the user can correct "deadstock" in the sentence "I would like to reduce deadstock," on line 21, to "deadstock," thereby correcting the sentence to "I would like to reduce deadstock." On the other hand, for the sentence "By utilizing advance shipping notices," on line 33, the user can approve the proofreading by the large-scale language model by leaving the word "advance shipping notice" as is without correcting it. Thus, in one aspect, the user proofreading accepted by the second proofreading means may include user approval.
[0129] An example of the display screen after correction is shown in Fig. 13. When the user proofreading is completed as shown in Fig. 13, by pressing the "Save" button, the user proofreading is transmitted to the information processing device 100 (step S14). In this way, the information processing device 100 accepts the user proofreading on the display screen displayed on the user terminal 500.
[0130] When the information processing device 100 receives the user's proofreading, it records second proofreading data that reflects the user's proofreading and includes the sentence "Is it possible to adopt a no-inspection method?" in the external storage 400 (step S15).
[0131] Next, the information processing device 100 refers to the external storage 400, compares the transcription data generated by the speech recognition model with second proofread data that reflects user proofreading on the first proofread data, and extracts the pre-proofread and post-proofread words of words that have been proofread between the two (step S16). In this example, the information processing device 100 uses a morphological analysis tool and a difference algorithm to analyze the transcription data and the second proofread data, thereby extracting the pre-proofread and post-proofread words of words that have been proofread between the two. Specifically, the transcription data and the second proofread data are each segmented into words using a morphological analysis tool such as MeCab, SudachiPy, or Janome, and changes are identified using an appropriate difference algorithm such as difflib. As a result, in this example, the following changes are extracted as the pre-proofreading word "red stock" and the post-proofreading word "dead stock" in the sentence on line 21, the pre-proofreading word "delivery inspection" and the post-proofreading word "no inspection" in the sentence on line 30, and the pre-proofreading word "advance shipping notice" and the post-proofreading word "advance shipping notice" in the sentence on line 33.
[0132] Next, the information processing device 100 estimates the pronunciation of the word before and after proofreading. In this example, the information processing device 100 estimates the pronunciation of the word using a large-scale language model. Specifically, the information processing device 100 first generates a prompt for causing the large-scale language model to estimate the pronunciation of the word, that is, a prompt including the word whose pronunciation is to be estimated and an instruction to estimate the pronunciation of the word (step S17). After generating the prompt, the information processing device 100 transmits a request for estimating the pronunciation of the word including the prompt to the information processing device 300 equipped with the large-scale language model (step S18).
[0133] The information processing device 300, which has received a request to estimate the pronunciation of a word, inputs the prompt included in the received request to estimate the pronunciation of a word into the large-scale language model (step S19). Then, the information processing device 300 acquires information output from the large-scale language model in response to the input and transmits this to the information processing device 100 (step S20). This allows the information processing device 100 to acquire the pronunciation of the specified word. In this example, the following description will be given assuming that the pronunciations of "red stock" and "defective inventory" are output as "red stock" and "furyozaiko," "delivery inspection product" and "no inspection product" are output as "noukenpin" and "noukenpin," and "advance shipping nokiss" and "advance shipping notice" are output as "advance shipping nokiss" and "advance shipping notice," respectively.
[0134] Next, the information processing device 100 calculates the edit distance between the pronunciation of the word before proofreading and the pronunciation of the word after proofreading.
[0135] For example, the readings of "red stock" and "furyozai-ko" are "red stock" and "furyozai-ko," respectively. In "red stock," "re" is replaced with "fu," "tsu" with "ri," "do" with "yo," "su" with "u," "to" with "za," "tsu" with "i," and "ku" with "ko," thereby converting "red stock" to "furyozai-ko." Therefore, the edit distance between the readings of "red stock" and "furyozai-ko" is calculated to be 7. And, since the edit distance divided by the number of characters in both (7) is "1.00," the normalized edit distance between the readings of "red stock" and "furyozai-ko" is calculated to be "1.00."
[0136] Also, the reading of "nouhinkenpin" is "nouhinkenpin," and the reading of "noukenpin" is also "noukenpin." Therefore, the edit distance between the readings of "nouhinkenpin" and "nokenpin" is calculated to be 0. And, since the edit distance divided by the number of characters in both (6) is "0.00," the normalized edit distance between the readings of "nouhinkenpin" and "nokenpin" is calculated to be "0.00."
[0137] Additionally, the readings of "Advance Shipping Nokis" and "Advance Shipping Notice" are "adobansushippingunokisu" and "adobansushippinguno-tisu," respectively. By replacing "ki" with "-" in "advance shipping nokis" and inserting "te" and "i" between "-" and "su," "advance shipping nokis" is converted to "advance shipping no-tisu." Therefore, the edit distance between the readings of "Advance Shipping Nokis" and "Advance Shipping Notice" is calculated to be 3. Furthermore, the edit distance divided by the longer number of characters (15) is "0.20," so the normalized edit distance between the readings of "Advance Shipping Nokis" and "Advance Shipping Notice" is calculated to be "0.20."
[0138] The information processing device 100 determines whether the normalized edit distance calculated by the above procedure is within a predetermined range, and if it is determined to be within the predetermined range, stores user proofreading history information including the pre-proofreading words and the proofreading words in the external storage 400 (steps S21 and S22). Note that in this example, the following description will be given assuming that "within the predetermined range" is defined as a normalized edit distance of "0.3 or less."
[0139] For example, the information processing device 100 determines whether the normalized edit distance between the pronunciation of the pre-proofreading word "inspection" and the proofreading word "no inspection" is "0.3 or less." As described above, the normalized edit distance between the pronunciations of "inspection" and "no inspection" is 0, which is less than 0.3. Therefore, in this example, the normalized edit distance between the two is determined to be within a predetermined range. Therefore, the proofreading of the pre-proofreading word "inspection" to the proofreading word "no inspection" is further stored in the external storage 400 as user proofreading history information that corresponds to a pre-proofreading sentence including the pre-proofreading word: "By the way, I would like to streamline the inspection work before shipping. Is it possible to adopt the in-progression inspection method? Yes." and a proofreading sentence including the proofreading word: "By the way, I would like to streamline the inspection work before shipping. Is it possible to adopt the no-inspection method? Yes." As described above, in this example, three utterances, including an utterance including the proofread word and the utterances before and after it, are stored as the sentences before and after proofreading. However, the sentences before and after proofreading do not necessarily have to be stored in units of three utterances, and may be stored in any appropriate units.
[0140] The information processing device 100 also determines whether the normalized edit distance between the pronunciation of the pre-proofreading word "Advance Shipping Nokis" and the proofreading word "Advance Shipping Notice" is "0.3 or less." As described above, the normalized edit distance between "Advance Shipping Nokis" and "Advance Shipping Notice" is "0.2," which is less than 0.3. Therefore, in this example, the normalized edit distance between the two words is determined to be within a predetermined range. Therefore, the proofreading of the pre-proofreading word "Advance Shipping Nokis" to the proofreading word "Advance Shipping Notice" is stored in the external storage 400 as user proofreading history information that further includes, in association with the pre-proofreading sentence containing the pre-proofreading word, "I think that by utilizing SMC labels and Advance Shipping Nokis, it will be possible to omit inspection work at the store," and the proofreading sentence containing the proofreading word, "I think that by utilizing SMC labels and Advance Shipping Notice, it will be possible to omit inspection work at the store." The proofreading of "Advance Shipping Nokis" to "Advance Shipping Notice" is a proofreading using a large-scale language model, and is not something the user himself / herself entered on the display screen. However, this proofreading is a correct proofreading approved by the user on the display screen, and may be stored as user proofreading history information.
[0141] Meanwhile, the information processing device 100 determines whether the normalized edit distance between the pronunciation of the pre-proofreading word "red stock" and the proofreading word "deadstock" is "0.3 or less." As described above, the normalized edit distance between "red stock" and "deadstock" is "1.0," which exceeds 0.3. Therefore, in this example, it is determined that the normalized edit distance between the two is not within the predetermined range. Therefore, the proofreading of the pre-proofreading word "red stock" to the proofreading word "deadstock" is not stored in the external storage 400 as user proofreading history information.
[0142] An example of the user calibration history information stored in the user calibration history information database of the external storage 400 updated as described above is shown in FIGS.
[0143] 14 and 15, an entry containing the pre-proofreading words "delivery inspection" and the proofreading words "no inspection" in association with each other has been added to the user proofreading history information database stored in external storage 400, in relation to the proofreading of "delivery inspection" to "no inspection" (the entry with "proofreading id" of "4" in FIG. 14). Also, as shown in FIG. 15, an entry containing the pre-proofreading sentence containing the pre-proofreading words in association with the proofreading sentence containing the proofreading words has been added to the user proofreading history information database stored in external storage 400, in relation to the entry regarding the proofreading of "delivery inspection" to "no inspection" (the entry with "proofreading id" of "4" and "example sentence id" of "1" in FIG. 15).
[0144] Furthermore, with regard to the proofreading of "Advance Shipping Nokis" to "Advance Shipping Notice," a new entry has been added to the user proofreading history information database stored in external storage 400, which is associated with an entry that includes the pre-proofreading words "Advance Shipping Nokis" and the proofreading words "Advance Shipping Notice" in association with each other, and which includes a pre-proofreading sentence that includes the pre-proofreading words and a proofreading sentence that includes the proofreading words in association with each other (in Figure 15, the entry with "proofreading id" as "1" and "example sentence id" as "2").
[0145] On the other hand, with regard to the correction of "red stock" to "dead stock," no new entry is added that includes the word "red stock" before the correction and the word "dead stock" after the correction in association with each other.
[0146] When the prompt generation unit 133 performs the prompt generation process (step S5) described above to proofread transcription data including "Advance Shipping Nokis," "Red Stock," and "Delivery Inspection" after the new entries shown in FIGS. 14 and 15 have been added to the user proofreading history information database of the external storage 400, the generated prompt includes, for example, the user proofreading history information shown in FIG. 16 as user proofreading history information. As shown in FIG. 16, the prompt generated after the new entries shown in FIGS. 14 and 15 have been added to the user proofreading history information database of the external storage 400 includes, as user proofreading history information for the word "Delivery Inspection" included in the transcription data, the pre-proofreading word "Delivery Inspection" and the corresponding post-proofreading word "No Inspection," as well as a pre-proofreading sentence including the pre-proofreading word and a post-proofreading sentence including the post-proofreading word. Furthermore, as user proofreading history information for the word "Advance Shipping Nokis" included in the transcription data, examples of pre-proofreading sentences including the pre-proofreading word and post-proofreading sentences including the post-proofreading word have been added. By referencing the external storage 400 to which the new entries have been added, it is possible to generate prompts that include word lists that newly reflect the user's unique vocabulary. By referencing the word lists, the large-scale language model can flexibly accommodate the user's unique vocabulary.
[0147] The present invention has been described in detail above based on the embodiments, but these embodiments are merely examples, and the present invention can be embodied in other forms with various modifications and improvements based on the knowledge of those skilled in the art. Furthermore, the terms "means" and "unit" described above can be interpreted as appropriate. For example, the term "acquisition unit" can be interpreted as "acquisition means." [Industrial Applicability]
[0148] According to an information processing device relating to one aspect of the present invention, it is possible to build a transcription system that can handle audio data that contains a large amount of user-specific vocabulary, technical terms, etc., without imposing an excessive burden on the system provider and / or user, and to provide excellent transcription and user experience.
Claims
1. a transcription data acquisition means for inputting speech data into a speech recognition model and acquiring transcription data of the speech data; prompt generation means for generating a prompt including an instruction for causing a large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data; a proofreading information acquisition means for inputting the generated prompt into a large-scale language model and acquiring proofreading information output from the large-scale language model; and a first proofreading means for generating first proofread data in which transcription errors contained in the transcription data have been corrected based on the acquired proofreading information; An information processing device comprising: a second calibration means for receiving a user's calibration of the first calibration data via a display screen that displays the first calibration data in an editable manner, thereby acquiring second calibration data in which the user's calibration is reflected in the first calibration data; an extraction means for comparing the second proofread data with the transcription data to extract proofread words between the two; and a user proofreading history information storage means for storing, in a storage unit, user proofreading history information for each extracted word, which associates the word before proofreading with the word after proofreading; Furthermore, The prompt generation means refers to the storage unit, and in addition to issuing an instruction to the large-scale language model to output proofreading information for proofreading transcription errors contained in the transcription data, generates a prompt including a word that corresponds to a word before proofreading among words included as proofread words in the user proofreading history information stored in the storage unit and that is included in the transcription data to be proofread.
1. An information processing device comprising:
2. The user proofreading history information storage means further stores, as user proofreading history information, pre-proofreading sentences including pre-proofreading words and / or proofreading sentences including proofreading words in association with each other.
2. The information processing apparatus according to claim 1, wherein:
3. The prompt generation means generates a prompt including a pre-proofreading sentence including the corresponding pre-proofreading word and / or a proofreading sentence including the proofreading word for a word included as a proofread word in the user proofreading history information stored in the storage unit and corresponding pre-proofreading word included in the transcription data to be proofread.
3. The information processing apparatus according to claim 2, wherein:
4. The system further includes an estimation means for estimating a word pronunciation before and after proofreading of a word proofread between the second proofread data and the transcription data, The user proofreading history information storage means further stores, as user proofreading history information, the pronunciation of the word before proofreading and / or the pronunciation of the word after proofreading estimated by the estimation means in association with each other.
4. The information processing apparatus according to claim 3,
5. The prompt generation means generates a prompt including the reading of a word for a word that is included as a proofread word in the user proofreading history information stored in the storage unit and whose corresponding pre-proofread word is included in the transcription data to be proofread.
5. The information processing apparatus according to claim 4,
6. The system further includes a determination unit that determines whether or not the edit distance between the pronunciation of the word before proofreading and the pronunciation of the word after proofreading estimated by the estimation unit is within a predetermined range, When the determining means determines that the edit distance is within a predetermined range, the user proofreading history information storing means stores the user proofreading history information in the storage unit.
5. The information processing apparatus according to claim 4,
7. A transcription system for audio data comprising an information processing device described in any one of claims 1 to 6.
8. A computer program for causing a computer to function as the information processing device according to any one of claims 1 to 6.
Citation Information
Patent Citations
Information processing device, information processing method, and computer program
JP7654294B1
Recorded data transcription system and program for recorded data transcription system
JP2025023364A
PROGRAM, INFORMATION PROCESSING METHOD AND INFORMATION PROCESSING SYSTEM
JP7645561B1
JPP7654294B