Processing device, processing method, and processing program
The processing device addresses the challenge of supporting daily conversations for individuals with aphasia by using voice recognition and language model-driven text modification, enhancing communication effectiveness.
Patent Information
- Application Number
- JP2023184069
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-05-13
AI Technical Summary
The complexity and difficulty in dispatching long-term communication support for individuals with aphasia make it challenging to provide effective daily conversation support.
A processing device equipped with an input unit for voice data, a voice recognition unit for converting voice data into text, and a correction unit using a language model to modify the text, facilitating communication by outputting modified text to a user interface.
Enables effective daily dialogue support for individuals with aphasia by improving understanding and expression through real-time text modification, allowing interlocutors to engage smoothly despite potential communication barriers.
Smart Images

Figure 2025073359000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a processing device, a processing method, and a processing program. [Background technology]
[0002] Aphasia is one type of language disorder. Aphasia refers to a state in which words cannot be used properly, for example, due to brain damage. Receptive aphasia (Wernicke's aphasia), a type of aphasia, is a state in which a person can speak but cannot understand what others are saying well, and when asked "How are you?" they will respond with "I want to drink some meat," resulting in incoherent speech. Amnesic aphasia is a state in which a person can understand what they hear and speak, but cannot remember the names of things. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Ministry of Health, Labor and Welfare, "Communication Support", [online], [searched September 21, 2023], Internet<https: / / www.mhlw.go.jp / bunya / shougaihoken / sanka / shien.html> Summary of the Invention [Problem to be solved by the invention]
[0004] The government provides communication support to support communication between people with disabilities and others who have communication difficulties. For example, communication support staff are dispatched to people with aphasia to help them understand and express themselves in conversation.
[0005] However, the procedures for actually dispatching communication support staff for people with aphasia are complicated, and it is difficult to dispatch them for long periods of time.
[0006] For this reason, there has been a demand for tools that can support everyday communication with people with aphasia.
[0007] The present invention has been made in view of the above, and has an object to provide a processing device, a processing method, and a processing program that can support daily dialogue with people with aphasia. [Means for solving the problem]
[0008] In order to solve the above-mentioned problems and achieve the object, the processing device of the present invention is characterized by having an input unit that accepts input of voice data uttered by a first user who is an aphasic user, a voice recognition unit that performs voice recognition on the voice data and converts the voice data into text, and a correction unit that uses a language model to correct the text converted by the voice recognition unit and output the corrected first corrected text to a user interface used by a second user who is the interlocutor of the first user. Effect of the Invention
[0009] According to the present invention, it is possible to support daily conversations with people with aphasia. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram for explaining an outline of the process according to the first embodiment. [Diagram 2] FIG. 2 is a diagram for explaining an output example of the first corrected text. [Diagram 3] FIG. 3 is a diagram for explaining an output example of the first corrected text. [Figure 4] FIG. 4 is a diagram for explaining an output example of the first corrected text. [Diagram 5] FIG. 5 is a diagram illustrating an example of a configuration of the processing system according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a system setting prompt given to the language model shown in FIG. [Figure 7] FIG. 7 is a diagram showing an example of the input and output of the language model shown in FIG. [Figure 8] FIG. 8 is a diagram showing a processing procedure of the processing method according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of a configuration of a processing system according to the second embodiment. [Figure 10] FIG. 10 is a diagram showing a processing procedure of the processing method according to the second embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a configuration of a processing system according to the third embodiment. [Figure 12] FIG. 12 is a diagram for explaining the revision history stored in the revision history database (DB) shown in FIG. [Figure 13] FIG. 13 is a diagram illustrating an example of re-learning data. [Figure 14] FIG. 14 is a diagram showing a processing procedure of the processing method according to the third embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of a configuration of a processing system according to the fourth embodiment. [Figure 16] FIG. 16 is a diagram for explaining the revision history stored in the revision history DB shown in FIG. [Figure 17] FIG. 17 is a diagram illustrating the flow of a prompt creation process performed by the prompt correction unit. [Figure 18] FIG. 18 is a diagram for explaining a specific example of prompt correction. [Figure 19] FIG. 19 is a diagram for explaining a specific example of prompt correction. [Figure 20] FIG. 20 is a diagram showing a processing procedure of the processing method according to the fourth embodiment. [Figure 21] FIG. 21 is a diagram illustrating an example of a computer in which each component device of the processing system is realized by executing a program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are indicated by the same reference numerals.
[0012] [Embodiment 1] [Outline of the first embodiment] An overview of the process in the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram for explaining the overview of the process in the first embodiment.
[0013] As shown in Fig. 1, in the first embodiment, speech uttered by a user U1 (first user) who has aphasia is collected by a microphone 20. A processing device 10 converts the speech data of the user U1 into text by speech recognition (Fig. 1(1)), and corrects the converted text using a language model (Fig. 1(2)).
[0014] The processing device 10 transmits the corrected text (hereinafter, referred to as the first corrected text) to an output UI (user interface) 30 (described later) (e.g., a smartphone 30A) used by a user U2 (a second user) who is an interlocutor of the user U1. The smartphone 30A outputs the first corrected text so that the user U2 can recognize it ((3) in FIG. 1).
[0015] 2 to 4 are diagrams for explaining an example of output of the first corrected text. For example, the smartphone 30A renders a subtitle W1 including the first corrected text "evised_res:"Today is hard and it's raining outside.", "reason":"The grammar is unnatural" on the screen of the AR (Augmented Reality) glasses worn by the user U2 on his head ((3) in FIG. 1, (1) in FIG. 2). Also, a subtitle W2 including the first corrected text may be displayed on the screen of the VR (Virtual Reality) goggles 30C worn by the user U2 on his head ((1) in FIG. 3).
[0016] The smartphone 30A may render subtitles W3 including the first corrected text on the camera image showing the user U1 ((1) in FIG. 4), or may display subtitles W4 including the first corrected text on the screen ((2) in FIG. 4). Also, audio including the first corrected text may be output from the output UI 30.
[0017] In this way, even if user U2 does not understand the content of the speech of user U1, who is aphasic, user U2 can check the content corrected by the language model on the spot, and can smoothly proceed with the dialogue with the aphasic person.
[0018] [Processing system] The configuration of the processing system according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the configuration of the processing system according to the first embodiment.
[0019] As shown in FIG. 5, the processing system 100 according to the first embodiment includes a microphone 20, a processing device 10, and an output UI 30.
[0020] The microphone 20 collects speech uttered by the user U1 who has aphasia, and transmits the speech data of the user U1 to the processing device 10. The microphone 20 may be a microphone built into the processing device 10.
[0021] The processing device 10 is a device that corrects the voice data of the user U1, and is, for example, a PC (Personal Computer), a notebook PC, a tablet terminal, a smartphone, etc. The processing device 10 is realized, for example, by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and the CPU executes the predetermined program. The processing device 10 also has a communication interface for transmitting and receiving various information to and from other devices (for example, a microphone 20, an output UI 30) connected via a network, etc. The processing device 10 has a voice data input unit 11 (input unit), a voice recognition unit 12, and a correction unit 13.
[0022] The voice data input unit 11 receives input of voice data uttered by the user U1 from the microphone 20 and outputs it to the voice recognition unit 12.
[0023] The voice recognition unit 12 performs voice recognition on the voice data uttered by the user U1, converts the voice data of the user U1 into text, and outputs it to the correction unit 13. The voice recognition unit 12 performs text conversion of the voice data using existing voice recognition technology.
[0024] The correction unit 13 uses the language model 131 to correct the text converted by the speech recognition unit 12, and outputs the corrected first corrected text to the output UI 30 used by the user U2.
[0025] The language model 131 is a language model (machine learning model) that can be used via APIs (Application Programming Interfaces) of various businesses. The language model 131 is a natural language processing model trained using a large amount of text data, that is, a large-scale language model (LLM: Large-Language-Model), and generates sentence data (text) with a natural context. The language model 131 is, for example, OpenAI's GPT (Generative Pre-trained Transformer).
[0026] Then, when a prompt corresponding to a case of aphasia is given and text corresponding to the speech data of user U1 is input, language model 131 generates a first corrected text by correcting the input text with text that occurs in a natural context.
[0027] The language model 131 is instructed to carry out a conversation support task for a person with aphasia using a system setting prompt or a context prompt. Fig. 6 is a diagram showing an example of a system setting prompt given to the language model 131 shown in Fig. 5.
[0028] The system setting prompt in Figure 6 instructs the user to take on the task of correcting sentences using a language model (Block B1), to correct the sentences step by step (Block B2), and includes a specific example of correction corresponding to aphasia cases (Block B3). In Block B2, the user is instructed to carry out the following steps in the correction process: Step 1, which searches for an incorrect part of the input text as a correction part; Step 2, which corrects or complements the part found; and Step 3, which outputs the corrected text in a limited format.
[0029] The language model 131 operates according to the system setting prompts to generate a first corrected text by correcting character data (text) corresponding to the input voice data of the user U1 with text that appears in a natural context.
[0030] Fig. 7 is a diagram showing an example of input and output of the language model 131 shown in Fig. 5. Fig. 7 shows a sentence spoken by a patient (aphasic person) (input data of the language model 131) and an output sentence of the language model 131 in association with each other.
[0031] The language model 131 corrects the neologism error "mizumi" to "lake" in the patient's spoken sentence "A boat anchored at the water" (first line in FIG. 7). The language model 131 also corrects the phonological paraphasia, for example, the paraphasia "shinpai" where a consonant is mistaken, to "sanpaku" in the patient's spoken sentence "I was worried about the shrine yesterday" (second line in FIG. 7). The language model 131 also corrects the phonological paraphasia "ojiniri" to "onigiri" in the patient's spoken sentence "I ate ojiniri" (third line in FIG. 7). The language model 131 also corrects the phonological paraphasia, particularly the paraphasia "gyunyuuniku" where a similar word is mistakenly repeated, to "gyunyuuniku" in the patient's spoken sentence "I ate milk" (fourth line in FIG. 7).
[0032] In addition, when there is a failure of word recall due to amnesia-induced aphasia, such as the patient's spoken sentence "I ate a big, round, green, sweet thing," the language model 131 corrects it to "I ate a melon" (line 7 of FIG. 7). In addition, when there is an omitted part of the sentence due to a failure of word recall due to amnesia-induced aphasia, such as the patient's spoken sentence "The rain made me late to the venue today," the language model 131 completes the omitted sentence and corrects it to "It was raining today, so I was late to go to the venue" (line 8 of FIG. 7). Each example in FIG. 7 is described in block B3 as a specific example of a correction of a system setting prompt.
[0033] In addition, in the case of a mixed paraphasia as in the fifth and sixth lines of Figure 7, or a formal paraphasia that makes sense but should be corrected in the same way as a mixed paraphasia, there are cases where the language model 131 cannot correct it. A mixed error is a slip of the tongue that is made using real words and that at first glance makes sense as a sentence, but is strange in context to the listener and is a common symptom of patients.
[0034] As described above, the output UI 30 is the AR glasses 30B or the VR goggles 30C worn by the user U2, the terminal device (e.g., the smartphone 30A) used by the user U2, and / or the speaker used by the user U2. The output UI 30 outputs the first corrected text corrected by the processing device 10 as subtitles, audio, and / or text.
[0035] [Processing method] Next, a description will be given of a processing method according to embodiment 1. Fig. 8 is a diagram showing a processing procedure of the processing method according to embodiment 1.
[0036] The microphone 20 collects voice data uttered by the user U1 (step S1), and transmits the voice data of the user U1 to the processing device 10 (step S2).
[0037] The processing device 10 performs speech recognition on the voice data uttered by the user U1 (step S3) and converts the voice data of the user U1 into text. The processing device 10 performs a correction process to correct the text converted in step S3 using the language model 131 (step S4). The processing device 10 transmits the corrected first corrected text to the output UI 30 used by the user U2, and causes it to be output (steps S5 and S6).
[0038] [Advantages of the first embodiment] In this way, the processing device 10 according to the first embodiment performs speech recognition on the speech data uttered by the user U1 who has aphasia, converts the speech data into text, and corrects the converted text by using the language model 131. The processing device 10 outputs the corrected first corrected text to the output UI 30 used by the user U2 who is the interlocutor of the user U1, and causes the first corrected text to be output.
[0039] As a result, even if the user U2 does not understand the content of the speech of the user U1 who has aphasia, the user U2 can immediately check the content corrected by the language model 131, and can smoothly proceed with the dialogue with the user U1 who has aphasia.
[0040] [Embodiment 2] Next, a description will be given of embodiment 2. In embodiment 2, a language model that is fine-tuned by teacher learning based on teacher data on the correctness and incorrectness of mistakes that are common among people with aphasia is used.
[0041] 9 is a diagram showing an example of the configuration of a processing system according to embodiment 2. A processing system 200 according to embodiment 2 has a processing device 210 instead of the processing device 10 in FIG.
[0042] The processing device 210 has a correction unit 213 having a language model 2131. The language model 2131 is a model that is fine-tuned by teacher learning based on teacher data of correct and incorrect speech mistakes made by aphasic people, based on a language model that has been pre-trained using a large-scale corpus. When text corresponding to the speech data of user U1 is input, the language model 2131 corrects and complements the incorrect parts of the text, and generates a corrected first corrected text.
[0043] The training data is data collected from a large number of cases (actual conversations, etc.) of aphasia patients, and is in which incorrect text is associated with correct text.
[0044] The teacher data may be correct data of mistakes made by aphasics, which are created from various papers (e.g., Reference 1). For example, Reference 1 gives specific examples of mistakes for the correct answer "onigiri" (rice ball), such as "onigiri" (phonological mistake), "akihawaya" (neologism), "sushiyubiwa" (semiotic mistake), "oinari" (mixed mistake), "nokogiri" (formal mistake), "sushi" (semantic mistake), and "yubiwa" (irrelevance mistake). The teacher data may be created using the flow chart of error classification in Reference 1. Reference 1: Yuki Takakura, Yoshinao Nakagawa, Ryusaku Hashimoto, “How should we respond to this aphasia symptom?”, Higher Brain Function Research, Vol. 42, No. 3, 2022. [Retrieved September 22, 2023], Internet <URL:https: / / www.jstage.jst.go.jp / article / hbfr / 42 / 3 / 42_316 / _pdf / -char / ja >
[0045] The language model 2131 looks at the context of the text data corresponding to the input speech of user U1 and outputs as a first corrected text a sentence in which a word that is probabilistically unlikely to occur is replaced with a word that is probabilistically likely to occur, or a sentence in which an omitted sentence is supplemented with a sentence that is probabilistically likely to occur.
[0046] At this time, if the language model 2131 is capable of setting a prompt, for example, the system setting prompt in Fig. 6 may be set, and the language model 2131 may be operated so that text that is likely to appear stochastically is output from the language model 2131. When the system setting prompt is set, the language model 2131 combines words as instructed by the prompt, and outputs sentences with high probability, which is considered to increase the accuracy of correction of the language model 2131.
[0047] [Processing method] Fig. 10 is a diagram showing a processing procedure of the processing method according to the second embodiment. Steps S11 to S13 shown in Fig. 10 are the same processes as steps S1 to S3 in Fig. 8. The processing device 210 performs a correction process for correcting the text converted in step S13 using the language model 2131 (step S14). The processing device 210 transmits the corrected first corrected text to the output UI 30 used by the user U2, and causes it to be output (steps S15 and S16).
[0048] [Effects of the second embodiment] As in the second embodiment, when using a language model 2131 that is fine-tuned by teacher learning based on teacher data on the accuracy of mistakes often made by aphasics, the contents uttered by the user U1 can be corrected and output from the output UI 30. Therefore, according to the second embodiment, the same effects as those of the first embodiment can be achieved.
[0049] [Embodiment 3] Next, a description will be given of embodiment 3. In embodiment 3, a UI is provided that enables a listener, user U2, to correct the first corrected text output from the language model, a correction history corrected by user U2 is accumulated, and the language model is re-trained using the accumulated correction history.
[0050] [Processing system] The configuration of the processing system according to the third embodiment will be described with reference to Fig. 11. Fig. 11 is a diagram showing an example of the configuration of the processing system according to the third embodiment.
[0051] As shown in FIG. 11, a processing system 300 according to the third embodiment includes a processing device 310 instead of the processing device 10 in the processing system 100 in FIG.
[0052] The correction UI 40 can communicate with the processing device 310 and the output UI 30. The correction UI 40 is a UI that allows the user U2 to correct an error in the first corrected text output by the processing device 310 or the output UI 30. The correction UI 40 outputs to the processing device 310 a second corrected text (text corrected by the listener) in which the error in the first corrected text has been corrected by the user U2.
[0053] The correction UI 40 may be provided in the smartphone 30A of the output UI 30, and may be realized by displaying the first corrected text on the screen of the smartphone 30A in a correctable manner. The correction UI 40 may be a device that corrects the first corrected text in response to a correction instruction uttered by the user U2, in addition to correcting text data.
[0054] The processing device 310 includes a history storage unit 314 (storage unit), a correction history DB 315 (database), and an optimization unit 316, in comparison with the processing device 10 of FIG.
[0055] Similarly to the correction units 13 and 213, the correction unit 313 corrects the text converted by the speech recognition unit 12 using a pre-trained language model 3131, and outputs the corrected first corrected text to the output UI 30 used by the user U2. The language model 3131 may be a pre-trained language model, such as a language model that operates when a prompt corresponding to aphasia cases is given (e.g., language model 131) or a language model that has been pre-trained with a large-scale corpus (e.g., language model 2131).
[0056] The history accumulation unit 314 receives an input of a first corrected text output from the language model 3131 (corrected text from the language model) and / or a second corrected text output from the correction UI 40.
[0057] When the user U2 inputs a second corrected text in which an error in the first corrected text has been corrected, the history accumulation unit 314 accumulates the second corrected text as a correction history in the correction history DB 315. When the user U2 does not input a second corrected text, the history accumulation unit 314 accumulates the first corrected text as a correction history in the correction history DB 315.
[0058] The correction history DB 315 stores the text (the spoken sentence of the user U1) converted by the voice recognition unit 12 output from the history storage unit 314, and the first corrected text or the second corrected text as a correction history.
[0059] Fig. 12 is a diagram for explaining the revision history stored in the revision history DB 315 shown in Fig. 11. Fig. 12 also explains the flow of the process of storing the revision history.
[0060] 12, when the system is used, the first corrected text (corrected output from the language model) by the language model 3131 is corrected by the user U2 on the correction UI 40. Then, when the user U2 makes a correction, that is, when the second corrected text is input to the processing device 10, the history accumulation unit 314 records the spoken sentence of the patient (user U1) as input and the output (second corrected text) corrected by the listener user U2 as output (arrows Y31 and Y32).
[0061] On the other hand, if there is no correction by the user U2, that is, if the second corrected text is not input to the processing device 10, the history accumulation unit 314 records the patient's utterance as the input and the output (first corrected text) after correction by the language model 3131 as the output (arrows Y31 and Y33). Note that the data format of the correction history may be a CSV format other than the JSON format in FIG. 12.
[0062] Then, the optimization unit 316 optimizes the parameters of the language model using the correction history stored in the correction history DB 315 as re-learning data. When the optimization unit 316 receives text obtained by converting voice data uttered by the user U1 as input, the optimization unit 316 optimizes the parameters of the language model so that the first correction text or the second correction text associated with the input is output. The optimization unit 316 replaces the language model 3131 with the optimized language model 3131S for a certain period of time.
[0063] Fig. 13 is a diagram showing an example of re-learning data. Fig. 13 shows a correspondence between a spoken sentence (input data of the language model 3131) (column L11) of a user U1 who has aphasia and a sentence (correct data) (column L12) corrected by an input of a listener user U2. Note that Fig. 13 also shows, for reference, an output sentence (not included in the learning data) of a supposed language model (before re-learning) in column L13.
[0064] It is expected that when "a boat anchored in the water" is input from the pre-re-training language model 3131, "a boat anchored in the lake" will be output.
[0065] For example, in the re-learning data, for the spoken sentence (input data) "A ship anchored at the water" by the patient user U1, the correct answer data is "A ship anchored at the port" corrected by the input of the listener user U2 (first line in FIG. 13). Therefore, it is expected that when "A ship anchored at the water" is input, the language model 3131S after re-learning will output "A ship anchored at the port."
[0066] In this way, the language model 3131S can output text in which the sentence errors specific to the user U1 have been corrected by relearning the second corrected text corrected by the listener, user U2.
[0067] Furthermore, in the embodiment, even in the case of a mixed paraphasia or a formal paraphasia that has meaning but should be corrected like a mixed paraphasia, as in the fourth line of Fig. 13, the sentence corrected by the listener, user U2, can be acquired as correct data. Therefore, the language model 3131S after re-learning can correctly correct input data that could not be corrected by the language model 3131 before re-learning.
[0068] [Processing method] Fig. 14 is a diagram showing a processing procedure of the processing method according to the third embodiment. Steps S21 to S23 shown in Fig. 14 are the same processes as steps S1 to S3 in Fig. 8. The processing device 310 performs a correction process for correcting the text converted in step S23 using the language model 3131 (step S24). The processing device 310 transmits the corrected first corrected text to the output UI 30 used by the user U2, and causes it to be output (steps S25, S26).
[0069] The correction UI 40 receives the first corrected text from the output UI 30, for example (step S27). When an error in the first corrected text is corrected by the user U2 through an operation by the user U2 (step S28), the correction UI 40 transmits the corrected second corrected text (output after correction by the listener) to the processing device 310 (step S29).
[0070] The history accumulation unit 314 accepts the input of the first corrected text output from the language model 3131 or the second corrected text output from the correction UI 40, and accumulates it as a correction history in the correction history DB 315 (step S30). As described above, when the second corrected text is input, the history accumulation unit 314 accumulates the second corrected text as a correction history in the correction history DB 315. When the second corrected text is not input, the history accumulation unit 314 accumulates the first corrected text as a correction history in the correction history DB 315.
[0071] Then, the optimization unit 316 optimizes the parameters of the language model using the correction history stored in the correction history DB 315 of the history storage unit 314 as re-learning data (step S31). The optimization unit 316 replaces the language model 3131 with the optimized language model 3131S for a certain period of time (step S32).
[0072] In this way, in the third embodiment, the second corrected text corrected by the user U2, who is the listener of the user U1, for the first corrected text output from the language model 3131 is accumulated as a correction history, and the language model is re-learned using the accumulated correction history. Therefore, according to the third embodiment, it is possible to provide a language model optimized for the symptoms and context of the individual user U1.
[0073] [Embodiment 4] Next, a fourth embodiment will be described. In the fourth embodiment, a second corrected text corrected by a user U2, who is a listener of the user U1, for a first corrected text output from a language model is accumulated as a correction history, and a prompt to be given to the language model is corrected based on the accumulated correction history.
[0074] [Processing system] The configuration of the processing system according to the fourth embodiment will be described with reference to Fig. 15. Fig. 15 is a diagram showing an example of the configuration of the processing system according to the fourth embodiment.
[0075] A processing system 400 according to the fourth embodiment has a processing device 410 instead of the processing device 310 in FIG.
[0076] The processor 410, in comparison with the processor 310 of FIG. 11, includes a history storage unit 414 (storage unit), a correction history DB 415 (database), and a prompt correction unit 416 (correction unit).
[0077] In addition, like the correction unit 313, the correction unit 413 uses a pre-trained language model 4131 to correct the text converted by the speech recognition unit 12, and outputs the corrected first corrected text to the output UI 30 used by the user U2.
[0078] Language model 4131 may be any pre-trained language model, such as a language model that operates in response to prompts corresponding to cases of aphasia (e.g., language model 131) or a language model that has been pre-trained on a large corpus (e.g., language model 2131).
[0079] The language model 4131 is a neural network model with a Transformer architecture, and has, for example, an architecture in which an encoder and a decoder are connected. The language model 4131 has, for example, a natural language processing model such as BERT (Bidirectional Encoder Representations from Transformers) as an encoder part that converts natural language into a vector representation, and a generation model that decodes the vector representation into natural language as a decoder part. The language model 4131 can output vector data obtained by vectorizing the text of an utterance sentence input by the user U1 to the history accumulation unit 414 and the prompt correction unit 416.
[0080] When the second corrected text is input, the history accumulation unit 414 accumulates the second corrected text in the database as a correction history together with an embedding vector obtained by vectorizing the text (the spoken sentence of the user U1) converted by the speech recognition unit 12. When the second corrected text is not input, the history accumulation unit 414 accumulates the first corrected text in the database as a correction history together with an embedding vector obtained by vectorizing the text converted by the speech recognition unit 12.
[0081] The correction history DB 415 stores, as correction history, the text (spoken sentence by user U1) converted by the speech recognition unit 12, its embedding vector, and first corrected text output from the history accumulation unit 314, or the text (spoken sentence by user U1) converted by the speech recognition unit 12, its embedding vector, and second corrected text.
[0082] Fig. 16 is a diagram for explaining the revision history stored in the revision history DB 415 shown in Fig. 15. Fig. 16 also explains the flow of the process of storing the revision history.
[0083] 16, when the system is used, the first corrected text (corrected output from the language model) by the language model 4131 is corrected by the user U2 on the correction UI 40. Then, when the user U2 makes a correction, that is, when the second corrected text is input to the processing device 10, the history accumulation unit 314 records the patient's (user U1) spoken sentence as input, the embedding vector of the patient's spoken sentence as embed, and the output after correction by the listener user U2 (second corrected text) as output in the correction history data D41 (arrows Y41, Y42, Y43).
[0084] On the other hand, if there is no correction by the user U2, that is, if the second corrected text is not input to the processing device 10, the history accumulation unit 414 records the patient's utterance sentence as input, the embedding vector of the patient's utterance sentence as embed, and the output after correction by the language model 4131 (first corrected text) as output in the correction history data D41 (arrows Y41, Y42, Y44). Note that the data format of the correction history may be a csv format other than the JSON format of FIG. 16.
[0085] When new voice data of user U1 is input, the prompt correction unit 416 searches for a correction history including an embedding vector similar to the embedding vector of the newly input voice data of user U1 among the correction histories stored in the correction history DB 415. The prompt correction unit 416 provides a prompt to which the searched correction history is added as reference information to the language model 4131, and causes the input text to be corrected.
[0086] Fig. 17 is a diagram illustrating the flow of a prompt creation process performed by the prompt correction unit 416. Figs. 18 and 19 are diagrams illustrating a specific example of prompt correction.
[0087] When new voice data of user U1 is input, the prompt correction unit 416 calculates the vector distance between the embedding vector of the spoken sentence of user U1 and the embedding vector of user U1 in each correction history stored in the correction history DB 415.
[0088] Then, the prompt correction unit 416 searches for, for example, the top N revision histories that are similar to the spoken sentence of user U1 based on the calculated vector distance ((1) in FIG. 17). For example, the prompt correction unit 416 performs a distance comparison between the embedded vector of the spoken sentence of user U1 input last in the original input prompt P41 and each embedded vector in the revision history DB 415, and cites the top N revision histories that are similar to the spoken sentence of user U1 ((1) in FIG. 18). This citation also includes past cases where there is no problem with the spoken sentence and cases unrelated to the current revision.
[0089] Then, the prompt correction unit 416 provides the language model with prompt P42 in which the searched correction history is added to the prompt as reference information ((2) in FIG. 17). The language model 4131 corrects the prompt P42 by changing "worry" in "I went to the shrine worried about you" to "visit the shrine" based on the reference information in the prompt P42.
[0090] When saving in the correction history DB, the history storage unit 414 saves the final sentence W41 of the input prompt P41 in the location of {input: 'uttered sentence'...} ((1) in FIG. 19).
[0091] [Processing method] Fig. 20 is a diagram showing a processing procedure of the processing method according to embodiment 4. Steps S41 to S43 shown in Fig. 20 are the same processes as steps S1 to S3 in Fig. 8. Steps S44 to S49 shown in Fig. 20 are the same processes as steps S24 to S29 in Fig. 14.
[0092] The history accumulation unit 414 accepts input of the first corrected text output from the language model 3131 or the second corrected text output from the correction UI 40, and accumulates it as a correction history in the correction history DB 315 together with the spoken sentence and embedding vector of the user U1 (step S50).
[0093] When new voice data of user U1 is input, the prompt correction unit 416 searches for a correction history including an embedding vector similar to the embedding vector of the newly input voice data of user U1 among the correction histories stored in the correction history DB 415. The prompt correction unit 416 performs a prompt correction process to create a prompt by adding the searched correction history as reference information (step S51), and then applies the corrected prompt to the language model 4131 (step S52) to execute text correction.
[0094] [Effects of the fourth embodiment] In this way, in the fourth embodiment, the second corrected text corrected by user U2, who is the listener of user U1, for the first corrected text output from the language model 4131 is accumulated as a correction history, and the prompt to be given to the language model 4131 is corrected based on the accumulated correction history. Therefore, according to the fifth embodiment, the language model 4131 can be operated so as to output text optimized for the symptoms and context of the individual user U1.
[0095] [System configuration of the embodiment] Each component of the processing systems 100, 200, 300, and 400 is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the processing systems 100, 200, 300, and 400 is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0096] Further, all or any part of each process performed in the processing systems 100, 200, 300, and 400 may be realized by a CPU, a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU. Further, each process performed in the processing systems 100 and 200 may be realized as hardware using wired logic.
[0097] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the process procedures, control procedures, specific names, and information including various data and parameters described above and shown in the drawings can be changed as appropriate unless otherwise specified.
[0098] [program] 21 is a diagram showing an example of a computer in which each of the components of the processing systems 100, 200, 300, and 400 is realized by executing a program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. Each of these components is connected by a bus 1080.
[0099] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0100] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that defines each process of the processing systems 100, 200 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing the same process as the functional configuration in the processing systems 100, 200 is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced with an SSD (Solid State Drive).
[0101] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary, and executes the program.
[0102] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, but may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or wide area network (WAN)). The program module 1093 and the program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0103] Although the embodiment of the invention made by the present inventor has been described above, the present invention is not limited by the description and drawings that form a part of the disclosure of the present invention according to the present embodiment. In other words, other embodiments, examples, operation techniques, etc. made by those skilled in the art based on the present embodiment are all included in the scope of the present invention. [Explanation of symbols]
[0104] 10,210,310,410 Processing equipment 11 Voice data input section 12 Voice Recognition Unit 13,213,313,413 Correction section 20. Mike 30 Output UI 40 Modified UI 100, 200, 300, 400 Processing Systems 131,2131,3131,4131,3131S Language Model 314,414 History storage unit 315,415 Revision History DB 316 Optimization Department 416 Prompt Correction Unit
Claims
1. an input unit that receives input of speech data uttered by a first user who is an aphasic person; a voice recognition unit that performs voice recognition on the voice data and converts the voice data into text; a correction unit that corrects the text converted by the speech recognition unit using a language model and outputs the corrected first corrected text to a user interface used by a second user who is an interlocutor of the first user; A processing device comprising:
2. the language model generates the first corrected text by correcting the text with text occurring in a natural context in response to a prompt corresponding to an aphasia case; 2. The processing device of claim 1, wherein the prompt instructs the processing device to perform a task of correcting a sentence for the language model by sequentially performing the steps of searching for a portion of input text to be corrected, correcting or completing the portion found, and outputting the corrected text in a limited format, and includes specific examples of corrections corresponding to cases of aphasia.
3. The processing device according to claim 1, characterized in that the language model is a model that has been fine-tuned by teacher learning based on teacher data on correct and incorrect speech errors made by aphasics, the model having been pre-trained using a large-scale corpus.
4. a storage unit that, when a second corrected text in which an error in the first corrected text is corrected is input by the second user, stores the second corrected text in a database as a correction history, and, when the second corrected text is not input, stores the first corrected text in the database as a correction history; an optimization unit that optimizes parameters of the language model using the correction history stored in the database as learning data; 2. The processing apparatus according to claim 1, further comprising:
5. a storage unit which, when a second corrected text in which an error in the first corrected text has been corrected is input by the second user, stores the second corrected text together with an embedded vector obtained by vectorizing the text converted by the speech recognition unit as a correction history in a database, and, when the second corrected text is not input, stores the first corrected text together with an embedded vector obtained by vectorizing the text converted by the speech recognition unit as a correction history in the database; a correction unit that, when the voice data of the first user is newly input, searches for a correction history including an embedding vector similar to an embedding vector of the newly input voice data of the first user among the correction histories stored in the database, and provides a prompt to the language model with the searched correction history added as reference information; 2. The processing apparatus according to claim 1, further comprising:
6. 2. The processing device according to claim 1, characterized in that the user interface is AR (Augmented Reality) glasses or VR (Virtual Reality) goggles worn by the second user, a terminal device used by the second user, and / or a speaker used by the second user.
7. A processing method executed by a processing device, comprising: receiving input of speech data uttered by a first user who is an aphasic; performing speech recognition on the voice data and converting the voice data into text; correcting the text converted in the converting step using a language model, and outputting the corrected first corrected text to a user interface used by a second user who is an interlocutor of the first user; A processing method comprising the steps of:
8. receiving input of speech data uttered by a first user who is an aphasic; performing speech recognition on the voice data and converting the voice data into text; correcting the text converted in the converting step using a language model, and outputting the corrected first corrected text to a user interface used by a second user who is an interlocutor of the first user; A processing program for causing a computer to execute the above.