Processing device, processing method and processing program

The processing device improves subtitle accuracy by using a machine learning model to correct speech and image data, addressing issues with conventional systems that struggle with unclear or imprecise speech, ensuring natural and contextually appropriate subtitle display.

JP2025108249APending Publication Date: 2025-07-23NTT DOCOMO BUSINESS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024002058
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Conventional subtitle display systems struggle with accurately conveying the intended message when speakers stutter, speak in a second language, or have pronunciation issues, as they rely on smooth speech and high-precision speech recognition, leading to incorrect or uncomfortable subtitle displays.

Method used

A processing device that uses a voice input unit, image input unit, and a machine learning model to correct text based on both voice and image data, generating natural and contextually accurate subtitles.

Benefits of technology

Enables smooth communication by correcting speech recognition errors and providing contextually appropriate subtitles, even in situations where speech is unclear or imprecise, enhancing communication for individuals with speech difficulties or language barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108249000001_ABST
    Figure 2025108249000001_ABST
Patent Text Reader

Abstract

To assist communication between speakers.SOLUTION: A processing device 10B includes: a voice recognition unit 12 for converting voice data which a first user and a second user emit into a text; and a correction unit 17B which corrects the text converted by the voice recognition unit 12 by using a machine learning model trained to correct the text on the basis of the text obtained by converting voice data which a speaker emits and data based on an image obtained by imaging the speaker, and outputs a corrected first correction text to a user interface which the first user and the second user use.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processing apparatus, a processing method, and a processing program.

Background Art

[0002] The provision of voice subtitles for TV programs and the like has been spreading mainly for the purpose of information compensation for the hearing impaired. Eventually, when character information such as explanations other than voice subtitles presented in a video began to be utilized in an integrated manner with the voice subtitles, the voice subtitle provision started to play a role of information compensation that complements human voice information not only for the hearing impaired but also for anyone who can read characters.

[0003] In the process of the development of voice subtitles, the words heard and seen by humans have been directly transcribed into sentences and generated and provided in the form of pasting them as characters on the screen. With the development of technology, as the real-time performance of speech recognition technology has improved, a real-time subtitle system has emerged that converts human voice data into text data by speech recognition and renders subtitles simultaneously with the speech in a video being shot or an interface for subtitle display.

[0004] Currently, in this way, voice subtitles are generated simultaneously with the speech by high-speed speech recognition technology, and it has started to play a role of facilitating communication not only for the hearing impaired but also for anyone who can read. For example, as a system that assists language communication by subtitles, there are a system that performs speech recognition and displays subtitles, and a system that displays in real time as subtitles the content explained by a face-to-face narrator (Non-Patent Document 1).

Prior Art Documents

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Conventional subtitle display systems have the following two problems.

[0007] Regarding the first problem. For people who cannot express what they really want to say as intended, such as when they stutter or speak in a second foreign language, it is somewhat difficult to convey their intentions to others. However, systems that display voice subtitles in real time through speech recognition usually assume smooth speech from healthy people and a highly accurate speech recognition system, and display unstable subtitles in situations where such assumptions cannot be met.

[0008] As such, conventional subtitle systems assume healthy people who can speak clearly and highly accurate speech recognition technology. When the speech recognition result is incorrect, or when there are stutters or mispronunciations, words or homonyms that are contextually incorrect are displayed as subtitles as they are.

[0009] A second problem will be described. A real-time speech recognition system is not a multimodal system but a system that converts only speech information into text data. The content of human speech is determined not only by the content of the speech of the conversation partner but also by various pieces of information. However, a normal real-time speech recognition system recognizes the content of a conversation only by converting speech data into text. This causes the display of speech subtitles that are uncomfortable for humans, such as homographs, homophones, and word mistakes due to phonetic errors.

[0010] The present invention has been made in view of the above, and an object thereof is to provide a processing device, a processing method, and a processing program that can assist communication between speakers.

Means for Solving the Problem

[0011] In order to solve the above-described problems and achieve the object, a processing device according to the present invention includes a voice input unit that receives inputs of voice data uttered by a first user and voice data uttered by a second user, an image input unit that receives inputs of an image of the first user and an image of the second user, a voice recognition unit that performs voice recognition on the voice data received by the voice input unit and converts the voice data into text, and a machine learning model trained to correct the text based on the text obtained by converting the voice data uttered by the speaker and data based on the image of the speaker, and corrects the text converted by the voice recognition unit using the machine learning model, and outputs the corrected first corrected text to a user interface used by the first user and the second user.

Effect of the Invention

[0012] According to the present invention, communication between interlocutors can be assisted.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13-1

Figure 13-2

Figure 14-1

Figure 14-2

Figure 15-1

Figure 15-2

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, with reference to the drawings, an embodiment of the present invention will be described in detail. Note that the present invention is not limited by this embodiment. Also, in the description of the drawings, the same parts are denoted by the same reference numerals.

[0015] [Embodiment 1] [Outline of Embodiment 1] With reference to FIG. 1, the outline of the processing in Embodiment 1 will be described. FIG. 1 is a diagram for explaining the outline of the processing in Embodiment 1.

[0016] In Embodiment 1, a processing system 100 that assists communication between speakers will be described. The processing system 100 converts the conversation content between user A (the first user) and user B (the second user) into text, corrects the text of the conversation content, and then displays subtitles of the corrected text on the display UI 40.

[0017] First, in the processing system 100, the voices uttered by user A and user B are collected by the respective microphones 20-1 and 20-2. In the processing system 100, the images of user A and user B during the conversation are captured by the respective imaging devices 30-1 and 30-2.

[0018] The processing device 10 converts the voice data of user A and user B collected by the respective microphones 20-1 and 20-2 into text by voice recognition ((1-1) and (1-2) in FIG. 1). Further, the processing device 10 performs image description to generate scene description text (description text of the image) from the images of user A and user B captured by the respective imaging devices 30-1 and 30-2 ((2-1) and (2-2) in FIG. 1).

[0019] Then, the processing device 10 uses the text of the conversation content between User A and User B and the scene description text of the images of User A and User B as inputs to the speech content correction language model 171 (language model) (see (3-1) and (3-2) in FIG. 1). The correction language model 171 is, for example, a natural language processing model trained using a large amount of text data, that is, a large language model (LLM).

[0020] The processing device 10 uses the correction language model 171 to correct the text of the conversation content between User A and User B, and displays the corrected text as subtitles on the display UI 40. For example, the utterance "This is delicious" by User A is corrected to "This tea is delicious".

[0021] The processing device 10 transmits the corrected text of the conversation content between User A and User B (hereinafter referred to as the first corrected text) to the display UI (user interface) 40 used by User A and User B. The display UI 40 outputs the first corrected text as subtitles. The display UI 40 is, for example, a smartphone 41, AR glasses 42-1 and 42-2, a transparent double-sided display 43, VR goggles 44, MR glasses, etc. The display UI 40 may be any UI through which User A and User B can visually recognize the first corrected text.

[0022] FIGS. 2 to 5 are diagrams for explaining output examples of the first corrected text.

[0023] For example, subtitles W1 of the first corrected text "This is delicious" are rendered on the screens of the AR glasses 42-1 and 42-2 worn on the heads of Users A and B respectively via the smartphones 41 used by Users A and B respectively (see (1) in FIGS. 1 and 2).

[0024] Also, the processing system 100 may display the first corrected text of Users A and B as subtitles on the double-sided display 43 capable of displaying subtitles on both sides (see FIG. 1).

[0025] In the processing system 100, the subtitle W2 of the first corrected text may be displayed on the screens of the VR goggles 44 worn on the heads of users A and B, respectively (Figure 3(1)).

[0026] Also, the smartphones 41 used by users A and B, respectively, may display the camera images of the other user A or B, and render the subtitle W3 of the first corrected text obtained by correcting the conversation content of the other party. (Figure 4(1)). Further, the smartphones 41 used by users A and B, respectively, may display the subtitle W4 of the first corrected text scrolling on the screen. (Figure 4(2)). The processing system 100 displays the subtitle for one of the other users in the conversation on the display UI40.

[0027] Also, the subtitles of the conversations of both users A and B may be displayed on the display UI40. For example, the processing system 100 displays the first corrected texts of users A and B in a thread format on the screen of the smartphone 41 (Figure 5). Of course, the processing system 100 may display the first corrected texts of users A and B in a chat format on both screens of the transparent double-sided display 43, the screens of the AR glasses 42-1 and 42-2, the screens of the VR goggles 44, or the screens of the MR glasses. Further, the voice including the first corrected text may be output from the correction UI50. The display UI40 may be a desktop display. The display UI40 may be any display on which users A and B can view the first corrected text.

[0028] In this way, even when the conversation content of the other user is unclear, users A and B can confirm on the spot the content corrected by the correction language model 171, and the conversation can proceed smoothly.

[0029] [Processing System] The configuration of the processing system according to Embodiment 1 will be described with reference to FIG. 6. FIG. 6 is a diagram showing a configuration example of the processing system according to Embodiment 1.

[0030] As shown in FIG. 1, the processing system 100 according to Embodiment 1 includes microphones 20-1 and 20-2, imaging devices 30-1 and 30-2, a processing device 10, and a display UI 40.

[0031] Microphone 20-1 collects the voice uttered by user A and transmits the voice data of user A to the processing device 10. Microphone 20-1 collects the voice uttered by user B and transmits the voice data of user A to the processing device 10. Microphones 20-1 and 20-2 may be microphones built into the processing device 10.

[0032] Imaging device 30-1 captures an image of user A and transmits the image data of user A to the processing device 10. Imaging device 30-2 captures an image of user B and transmits the image data of user A to the processing device 10. Imaging devices 30-1 and 30-2 capture images of users A and B periodically or continuously.

[0033] The processing device 10 is a device that corrects the voice data of users A and B, and is, for example, a PC (Personal Computer), a notebook PC, a tablet terminal, a smartphone, or the like. The processing device 10 is realized, for example, by a computer including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., with a predetermined program loaded and the CPU executing the predetermined program. Also, the processing device 10 has a communication interface for transmitting and receiving various information to and from other devices (for example, microphones 20-1 and 20-2, imaging devices 30-1 and 30-2, display UI 40) connected via a network or the like.

[0034] The processing device 10 includes a voice input unit 11, a voice recognition unit 12, an image input unit 13, an image description unit 14 (generation unit), and a correction unit 17.

[0035] The voice input unit 11 receives the input of voice data uttered by user A from the microphone 20-1 and outputs it to the voice recognition unit 12. The voice input unit 11 receives the input of voice data uttered by user B from the microphone 20-2 and outputs it to the voice recognition unit 12.

[0036] The voice recognition unit 12 performs voice recognition on the voice data uttered by users A and B, converts the voice data of users A and B into text, and outputs it to the correction unit 17. The voice recognition unit 12 performs text conversion of voice data using existing voice recognition technology.

[0037] The image input unit 13 receives the input of an image of user A from the imaging device 30-1 and outputs it to the image description unit 14. The image input unit 13 receives the input of an image of user B from the imaging device 30-2 and outputs it to the image description unit 14.

[0038] The image description unit 14 generates a description text (caption) of the image received by the image input unit 13. The image description unit 14 is realized using a machine learning model such as a neural network that generates text (scene description text) for explaining the scene of the input image when an image (which may also be a video) is input.

[0039] Figure 7 is a diagram for explaining the processing of the image description unit 14 shown in Figure 6. As shown in Figure 7, for example, when the image G1 is input, the image description unit 14 generates the scene description text C1 of the input image, "A boy holding a PET bottle beverage". The machine learning model of the image description unit 14 uses the image and the scene description text of this image as teacher data, and the parameters are optimized so that the scene description text of the learning image input to the machine learning model approaches the scene description text that is the correct data of this image.

[0040] Note that the image description unit 14 may also include visual descriptions around the video description result in the scene description text. The video description here refers to the entire technology that enables visual information, such as object recognition, object tracking, image captioning, or video captioning by image processing, to be described by sentences or words.

[0041] When basing on object recognition, describe in text what and how many there were in a series. Also, when basing on object tracking, describe in text what, how many, and where it moved from and to in a series. Furthermore, for image description and video description, all captions for each series are utilized. The input data creation unit 16 (described later) synchronizes the environmental recognition result obtained in this way with the speech recognition result and inputs it to the correction language model 171.

[0042] The conversation log integration DB 15 integrates and stores conversation logs. FIG. 8 is a diagram for explaining the conversation logs stored in the conversation log integration DB 15 shown in FIG. 6. As shown in FIG. 8, the conversation log integration DB 15 stores the conversation log L11 in the form of {user's speech, correction by the correction language model 171 (described later)}. The user's speech follows "UserA" or "UserB", and the correction by the correction language model 171 follows "Revised UserA (or B)" including the identification information of the corrected user. Note that FIG. 8 includes content for explaining the processing of the input data creation unit 16 (described later). In this way, the conversation log integration DB 15 integrates the speeches of users A and B supplemented and corrected by the processing device 10 as past speech histories.

[0043] The input data creation unit 16 creates input data for the correction unit 17 based on the text of the speech data output by the speech recognition unit 12, the scene description text output by the image description unit 14, and the conversation logs stored in the conversation log integration DB 15.

[0044] As shown in FIG. 8, the input data creation unit 16 extracts the conversation logs output from the new correction language model 171 from the conversation logs L11 in the conversation log integration DB 15 (step (1) in FIG. 8), and includes them in the input data D11 as past conversation histories (block B1). Subsequently, the input data creation unit 16 includes the text "This is delicious" output by the speech recognition unit 12 in the input data D11 (block B2).

[0045] FIG. 9 is a diagram for explaining the processing of the input data creation unit 16 shown in FIG. 6. As shown in FIG. 9, the input data creation unit 16 includes the scene description text C1 "A boy holding a PET bottle beverage" output by the image description unit 14 in the input data D11 (block B3 in FIG. 9).

[0046] The input data creation unit 16 inputs the input data D11, which uses the new output and new input of the correction language model 171 as one sequence C12 (FIG. 8), to the correction unit 17. The input data repeats sequences such as sequence C12 and includes all the information in this conversation. The input data creation unit 16 inputs the texturized speech recognition result and the scene description text to the correction unit 17 for each sequence.

[0047] The correction unit 17 corrects the text converted by the speech recognition unit 12 based on the text converted by the speech recognition unit 12 and the description text of the image generated by the image description unit 14 using the correction language model 171. The correction unit 17 outputs the corrected first corrected text to the display UI 40 used by user B.

[0048] The language model for correction 171 is a language model (machine learning model) that can be used via the APIs (Application Programming Interfaces) of various operators. The language model for correction 171 is a natural language processing model trained using a large amount of text data, that is, a large language model (LLM: Large-Language-Model), and generates text data (text) in a natural context. The language model for correction 171 is, for example, GPT (Generative Pre-trained Transformer) of OpenAI.

[0049] Then, when the language model for correction 171 is given a prompt adapted to the correction of the conversation content, and the text corresponding to the voice data of users A and B, and the scene description text of the image capturing users A and B are input, it generates the first corrected text that corrects the input text with the text that appears in a natural context.

[0050] The language model for correction 171 is instructed about the conversation support task between users A and B by the system setting prompt and the context prompt. FIG. 10 is a diagram showing an example of the system setting prompt given to the language model for correction 171 shown in FIG. 6.

[0051] The system setting prompt in FIG. 10 instructs the language model for correction 171 to be responsible for the task of complementing and / or correcting the utterances in the conversation between users A and B (block B11), and to perform the correction of the text step by step and the output format (block B12) (frame W11). At the same time, the system setting prompt includes the past conversation history (text) of users A and B (block B13), the scene description text indicating the current utterance and the current situation (block B14), and the output (block B15) (frame W12). Blocks B13 and B14 are based on the input data D11 created by the input data creation unit 16.

[0052] In block B2, the system instructs to correct the text by sequentially performing step 1 (the first step) of complementing the current utterance based on the scene description text indicating the past conversation history and the current situation, step 2 (the second step) of correcting the result of step 1, and step 3 (the third step) of outputting the text (the first text) obtained by correcting the current utterance as the result of step 2. In the system setting prompt, the correction is to replace what seems to be a slip of the tongue or a misspelled word with a contextually natural word.

[0053] By operating according to the system setting prompt, the correction language model 171 generates the first corrected text by correcting the input character data (text) corresponding to the voice data of users A and B with reference to the scene description text of users A and B indicating the current situation of users A and B so as to be text that appears in a natural context.

[0054] In this way, the correction language model 171 receives the input with the speech recognition result as the main result and other information (scene description information, past conversation history) as reference information. Then, based on the reference information, the correction language model 171 corrects mistakes such as word mistakes and pronunciation mistakes like homophones in the subtitles. Also, if there is an important fact in the reference information, the correction language model 171 may complement the content of the subtitles and include information beyond the uttered voice. The correction language model 171 complements and corrects the current utterance into a character string that can be understood just by reading the current utterance based on actions and past utterance history. Note that this example does not include examples of correcting pronunciation mistakes or the incoherence of intention due to word differences. Also, the method of matching the utterance and the action is not specifically specified.

[0055] FIG. 11 is a diagram for explaining an example of the output of the correction language model 171 shown in FIG. 6. In FIG. 11, the system setting prompt (frame W21) shown in FIG. 10 is also shown. For example, to the correction language model 171, together with the past conversation history of users A and B, a scene description text "a boy holding a PET bottle beverage" indicating the situation of user A and the utterance "This is delicious" (frame W22) of user A are input in a state where the timing can be synchronized. The correction language model 171 outputs a first corrected text "This tea is delicious" (frame W23) obtained by correcting the utterance of user A according to the instruction of the system setting prompt.

[0056] As described above, the display UI 40 is a tablet terminal such as a smartphone used by users A and B, AR glasses 42-1 and 42-2 worn by users A and B, a transparent double-sided display 43, VR (Virtual Reality) goggles, MR (Mixed Reality) glasses, or VR goggles 44, a terminal device used by user B (for example, a smartphone 41), and / or a speaker used by user B. The display UI 40 outputs the first corrected text corrected by the processing device 10 in the form of subtitles, text, and / or voice.

[0057] [Processing method] Next, the processing method according to Embodiment 1 will be described. FIG. 12 is a diagram showing the processing procedure of the processing method according to Embodiment 1.

[0058] The microphone 20-1 collects the voice data uttered by user A (step S1) and transmits the voice data of user A to the processing device 10 (step S5). The microphone 20-2 collects the voice data uttered by user A (step S3) and transmits the voice data of user A to the processing device 10 (step S6).

[0059] The processing device 10 performs voice recognition on the voice data uttered by users A and B (step S7) and converts the voice data of user A into text.

[0060] Then, imaging device 30-1 captures an image of user A (step S2) and transmits the image data of user A to processing device 10 (step S8). Imaging device 30-2 captures an image of user B (step S4) and transmits the image data of user A to processing device 10 (step S9).

[0061] Processing device 10 performs image description processing to generate scene description text for the images of users A and B (step S10).

[0062] Based on the text of the voice data output in step S7, the scene description text output in step S11, and the conversation logs stored in conversation log integration DB15, processing device 10 creates input data for correction language model 171 (step S11).

[0063] Processing device 10 performs correction processing to correct the text converted in the speech recognition process by inputting the input data into correction language model 171 (step S12). The input data is data created based on the text converted in the speech recognition process, the description text of the image generated in the image description process, and the conversation logs stored in conversation log integration DB15.

[0064] Processing device 10 transmits the corrected first corrected text to display UI 40 (step S13) and causes it to be displayed (step S14).

[0065] [Effects of Embodiment 1] As described above, processing device 10 according to Embodiment 1 performs speech recognition on the voice data uttered by users A and B respectively, and converts the voice data into text. At the same time, processing device 10 generates scene description text from the images of users A and B at the time of speech. Then, processing device 10 corrects the text of the voice data by inputting the text of the voice data and the scene description text into correction language model 171. Processing device 10 outputs the corrected first corrected text to display UI 40 for users A and B, and causes the first corrected text to be displayed as subtitles.

[0066] By recognizing the subtitles of this first corrected text, even if the pronunciation of Users A and B may not be clear, both parties can check the content of their conversations in the state of the corrected text, so that the communication between Users A and B can proceed smoothly.

[0067] Also, in the processing device 10, an LLM is used as the correction language model 171. For this reason, the processing device 10 displays the first corrected text, which supplements the current speech content based on the context of past conversations, as subtitles. Therefore, according to the processing device 10, even without high-precision speech recognition technology or clear speech, it is possible to generate subtitles that are more natural for humans for the speech content of Users A and B.

[0068] Also, in the processing device 10, the scene description text of the images of Users A and B is input into the correction language model 171 together with the text of the conversation content of Users A and B. In this way, in the processing device 10, the text of the conversation content is supplemented with the scene description text of videos or still images. That is, in the processing device 10, the technology of video description is used to use the description of the speaker state for subtitle supplementation. For this reason, in the processing device 10, the text of the conversation content of Users A and B is displayed as subtitles of the first corrected text supplemented in the situations of Users A and B, so that natural subtitles suitable for the on-site situation can be generated, and the communication between Users A and B can proceed smoothly.

[0069] Specifically, the application scenarios of the processing system 100 will be described. FIGS. 13-1 to 15-2 are diagrams for explaining the application scenarios of the first embodiment.

[0070] [Application Example 1] The case of promoting the understanding of the speech content of a person with speech difficulties and assisting communication will be described. Speech difficulties symptoms, dysarthria, will be described. The inability to pronounce correctly is a common symptom of dysarthria. An example of dysarthria will be described.

[0071] In organic dysarthria, due to the shape of the lips, tongue, palate, etc. after surgery, the sound may be overall muffled, or there may be abnormal forms, resulting in special pronunciation habits. Also, in motor speech disorder, since the tongue etc. cannot move quickly and accurately, the overall sound may sound connected, or the rhythm and speed may be disrupted. Even if one can say a single sound correctly, when it comes to conversation, rhythm and speed are required, so it is likely to become unclear.

[0072] In addition, in auditory dysarthria, distortion is likely to occur in sounds that are difficult to hear, and there are various sounds with depression depending on the type and degree of hearing impairment, such as "only high-pitched sounds are difficult to hear". In functional dysarthria, it may be replaced by other sounds like so-called baby talk (for example, "teacher" becomes "tentei"), or one may acquire pronunciation habits different from normal, presenting a unique sound distortion.

[0073] Figures 13-1 to 15-2 are examples for explaining modified examples of the processing system 100 according to the embodiment.

[0074] In the processing system 100, for example, as shown in Figure 13-1, an example will be described when answering the question of user A and uttering "This is how tentei speaks" to user B. Among the utterance of user B, "This" is corrected to "Today", and subtitles with "tentei" corrected to "sensei" are output. Also, in the processing system 100, for example, as shown in Figure 13-2, "hakimaki that was wrapped around" of user B is corrected to "hakimaki that was wrapped around and studying". Thus, according to Embodiment 1, it is possible to promote the understanding of the utterance content of people who have difficulty speaking due to age, stuttering, ALS / muscular dystrophy, etc., and assist communication.

[0075] [Application Example 2] In addition, in the processing system 100, not limited to dysarthria, general speech mistakes can also be corrected. For example, as shown in Figure 14-1, in the processing system 100, "Did you drink osukuri?" of user A is corrected to "Did you take medicine?". Also, as shown in Figure 14-2, "sanpopo kana" of user B is corrected to "tanpopo kana".

[0076] [Application Example 3] In addition, in the processing system 100, information that cannot be understood from the speech content alone is supplemented by the conversation history and the description of the current speaker's situation. For this reason, high-precision correction is realized. For example, in the processing system 100, as shown in FIG. 15-1, when a scene description text indicating that there is a little snow accumulation in the outdoor background of the speaker is input to the correction language model 171, the utterance "Ah! It's snowing!" of user B is corrected to "Ah! It's snowing!". Also, in the processing system 100, as shown in FIG. 15-2, when a scene description text indicating that user A has the sea and spray as the background is input to the correction language model 171, the utterance "This place is really beautiful" of user A is corrected to "This sea is really beautiful".

[0077] In this way, in the processing system 100, the utterance texts of users A and B are further corrected by using the past history of the conversation, the context, and in addition, the scene description text indicating the current speaker's situation. As a result, in the processing system 100, by supplementing information that cannot be understood only from the subtitling of the speech content with the history of the past conversation and the description of the current speaker's situation, it is possible to reduce the display of subtitles that are uncomfortable for humans. Therefore, according to Embodiment 1, even without high-precision speech recognition technology or clear speech, more natural subtitles can be generated for humans, and the communication between users A and B can proceed smoothly.

[0078] [Application Example 4] Embodiment 1 can also be used as a translator by switching to a prompt that asks to translate the incoming speech into natural English sentences instead of generating subtitles. When a prompt corresponding to the conversation with users A and B is given to the correction language model 171, the input text is corrected to a text that appears in a natural context and then a first corrected text translated into the specified language is generated.

[0079] FIG. 16 is a diagram showing an example of the system setting prompt given to the correction language model 171 when Embodiment 1 is applied to a translator.

[0080] The system setting prompt in FIG. 16 instructs the correction language model 171 to be responsible for the task of complementing and / or correcting while translating the utterances in the conversation between User A and User B (block B21), and to perform the correction of the text step by step and the output format (block B22). At the same time, the system setting prompt includes the past conversation history (text) of User A and B actually input to the correction language model 171 (block B23), the scene description text indicating the current utterance and the current situation actually input to the correction language model 171 (block B24), and the output (block B25).

[0081] In block B22, based on the past conversation history and the scene description text indicating the current situation, step 1 (the first step) of complementing the current utterance, step 2 (the second step) of correcting the result of step 1, step 3 (the third step) of translating the text obtained by correcting the current utterance as the result of step 2 into a specified language different from the last spoken word (for example, English), and step 4 (the fourth step) of outputting the translated text obtained by correcting the current utterance as the result of step 3 are sequentially performed to instruct the correction and translation of the text. In the system setting prompt, the correction is to replace what seems to be a speech error or a word error with a contextually natural word.

[0082] By operating according to the system setting prompt, the correction language model 171 generates a translated text obtained by correcting and then translating the character data (text) corresponding to the input voice data of User A and B with reference to the scene description text of User A and B indicating the current situation of User A and B into text that appears in a natural context.

[0083] At this time, the processing device 10 outputs the translated text translated by the correction language model 171 in text or voice. In this way, the processing device 10 can also be utilized as a translator that translates the incoming voice into another language (for example, English) in natural sentences.

[0084] In this way, in Embodiment 1, even in a situation where there is no high-precision speech recognition technology, or where it is difficult to produce a smooth speech like in the case of stuttering or speaking in a second foreign language, poor speech recognition results affected by acoustic device failures or imperfect pronunciation of the speaker are corrected and translated into natural words and context along the context of the conversation, and are displayed to both parties involved.

[0085] And in Embodiment 1, by presenting the corrected subtitles to both parties having a conversation, it is possible to assist the communication itself between the opposing people beyond the meaning of subtitles for information compensation for hearing-impaired persons.

[0086] Furthermore, Embodiment 1 is a system that not only converts voice data into text data but also performs subtitle display based on multiple human modalities and context. Therefore, according to Embodiment 1, it is possible to complement and display words that may be misheard as other words due to the selection of homophonic words or phonetic mistakes with the most appropriate words that can occur in that situation. As a result, voice subtitles that are not uncomfortable for humans are displayed in automatic voice subtitles in video works, real-time TV calls, etc.

[0087] In this way, in Embodiment 1, the discomfort caused by presenting the speech recognition result as it is is optimized according to the context of the conversation and is displayed as subtitles that are not uncomfortable for humans. Also, in Embodiment 1, speech recognition can be performed without being affected to a certain extent by the quality of the acoustic device or the pronunciation of the speaker, and furthermore, voice subtitles considering the conversation content, expression, and situation are displayed.

[0088] [Modification Example of Embodiment 1] FIG. 17 is a diagram for explaining the outline of the process of the modification example of Embodiment 1. In Modification Example 1 of Embodiment 1, instead of the correction language model 171, a Vision and Language Model (VLM) 171B (machine learning model) is used to correct the utterance texts of Users A and B.

[0089] Generally, a VLM is a machine learning model that takes an image and text as inputs and outputs an image and text. In the modification example of Embodiment 1, machine learning is performed so that the input text is corrected by taking an image and text as inputs.

[0090] In the modification example of Embodiment 1, the voice data of Users A and B are respectively converted into text by voice recognition ((1-1) and (1-2) in FIG. 17). After that, the text of the conversation content and the image of Users A and B are input to the VLM 171B, the text of the conversation content of Users A and B is corrected, and the corrected first text is output to the display UI 40.

[0091] FIG. 18 is a diagram showing a configuration example of the processing system according to the modification example of Embodiment 1. As shown in FIG. 18, the processing system 100B according to the modification example of Embodiment 1 has a processing device 10B as compared with the processing device 10 in FIG. 6.

[0092] The processing device 10B includes a correction unit 17B having a VLM 171B capable of recognizing text and images instead of the correction language model 171 in FIG. 6.

[0093] The VLM 171B is a machine learning model trained to correct text based on the text obtained by converting the voice data spoken by the speaker and the data based on the image of the speaker. The text converted by the voice recognition unit 12 and the images of Users A and B are input to the VLM 171B. The VLM 171B outputs the first text obtained by correcting the text converted by the voice recognition unit 12.

[0094] The correction unit 17B uses VLM171B to correct the text converted by the speech recognition unit 12, and outputs the corrected first corrected text to the display UI 40 used by users A and B.

[0095] Next, a processing method according to a modification example of Embodiment 1 will be described. FIG. 19 is a diagram showing the processing procedure of the processing method according to the modification example of Embodiment 1.

[0096] Steps S21 to S29 shown in FIG. 19 are the same processing as steps S1 to S9 in FIG. 12.

[0097] The processing device 10 performs a correction process of correcting the text converted in the speech recognition process by inputting the text of the speech data output in step S7 and the images of users A and B transmitted in steps S28 and S29 to the correction language model 171 (step S30). Steps S31 and S32 shown in FIG. 19 are the same processing as steps S13 and S14 in FIG. 11.

[0098] [Effect of the modification example of Embodiment 1] As described above, in the modification example of Embodiment 1, by using VLM171B trained to correct text based on the text obtained by converting the speech data spoken by the speaker and the data based on the image of the speaker captured, the image description unit 14 can be omitted.

[0099] [Embodiment 2] Next, Embodiment 2 will be described. In Embodiment 2, a UI is provided in which the first corrected text output from the language model can be corrected by an arbitrary user (for example, users A and B or other users), the correction history corrected by the arbitrary user is accumulated, and the correction language model is re-learned with the accumulated correction history. FIG. 20 is a diagram for explaining the outline of the processing of Embodiment 2.

[0100] In Embodiment 2, a correction UI 50 is provided that allows any user to correct errors in the first text. In this case, the second corrected text corrected by any user is stored in the re-learning data storage DB 219 (Fig. 20(1)).

[0101] And in Embodiment 2, the correction language model 171 (e.g., LLM) is fine-tuned (re-learned) using the second corrected text in the re-learning data storage DB 219 (Fig. 20(2)). The fine-tuned correction language model 171S is replaced with the correction language model 171 at regular intervals (Fig. 20(3)). As a result, a correction language model optimized for the conversations of Users A and B can be applied.

[0102] [Processing System] The configuration of the processing system according to Embodiment 2 will be described with reference to Fig. 21. Fig. 21 is a diagram showing a configuration example of the processing system according to Embodiment 2.

[0103] As shown in Fig. 21, the processing system 200 according to Embodiment 2 has a processing device 210 instead of the processing device 10 compared to the processing system 100 in Fig. 6, and further has a correction UI 50.

[0104] The correction UI 50 can communicate with the processing device 210 and the display UI 40. The correction UI 50 is a UI that can correct errors in the first corrected text output by the processing device 210 or the display UI 40 through the operations of Users A and B. The correction UI 50 outputs the second corrected text, in which the errors in the first corrected text have been corrected by any user, to the processing device 210.

[0105] The correction UI 50 may be provided on the smartphone 41 of the display UI 40 and realized by displaying the first corrected text on the screen of the smartphone 41 so that it can be corrected. Also, the correction UI 50 may be a device that corrects the first corrected text upon receiving a correction instruction uttered by Users A and B in addition to correcting text data.

[0106] The processing device 210 includes a history storage unit 218 (storage unit), a re-learning data storage DB 219 (database), and an optimization unit 220, as compared with the processing device 10 in FIG. 6.

[0107] The history storage unit 218 receives the input of the first corrected text (the corrected text output from the correction language model 171) output from the correction language model 171, and / or the second corrected text output from the correction UI 50.

[0108] When there is an input of a second corrected text in which an error in the first corrected text has been corrected by an arbitrary user, the history storage unit 218 stores the second corrected text as a correction history in the re-learning data storage DB 219. When there is no input of the second corrected text, the history storage unit 218 stores the first corrected text as a correction history in the re-learning data storage DB 219. When there is no input of the second corrected text, the history storage unit 218 may store the text obtained by formatting the first corrected text as a correction history in the re-learning data storage DB 219.

[0109] The re-learning data storage DB 219 stores, as a correction history, the text (utterance texts of users A and B) converted by the speech recognition unit 12, which is output from the history storage unit 218, and the first corrected text or the second corrected text. The re-learning data storage DB 219 stores the first corrected text or the second corrected text as correct answer data for the utterance text of the corresponding user.

[0110] FIG. 22 is a diagram for explaining the correction history stored in the re-learning data storage DB 219 shown in FIG. 21. FIG. 23 is a diagram for explaining the flow of the correction history storage process.

[0111] As shown in the conversation log L21 in FIG. 22, the conversation log integration DB 15 stores the current utterance text and the corrected output (the first corrected text) from the correction language model 171 (LLM) as a conversation log (arrows Y21 and Y22 in FIG. 23).

[0112] Taking as an example the case where there is a correction by an arbitrary user to the first corrected text C21, "Drink this tea," in the conversation log L21 of the conversation log integration DB15 (Fig. 22(1)).

[0113] When there is a correction by an arbitrary user, the second corrected text by the arbitrary user is recorded as a correction history in the re-learning data accumulation DB219 (arrow Y23 in Fig. 23).

[0114] For example, consider the case where the second corrected text C22, "Drink this green tea," corrected by an arbitrary user is input to the processing device 210. The history accumulation unit 218 formats and corrects the corrected history D22 with the utterance sentence of user B, "Drink this," as the input and the second corrected text C22, "Drink this green tea," instead of the first corrected text C21, "Drink this tea," as the output, and records it in the re-learning data accumulation DB219 (Fig. 22(2)).

[0115] In addition, when there is no correction by an arbitrary user, that is, when the second corrected text is not input to the processing device 210, as shown in the correction history D22, the history accumulation unit 218 records the user's utterance sentence as the input and the user's utterance sentence, or the first corrected text or the text obtained by formatting the first corrected text as the output. Also, the description of the image and the image itself are not included in the correction history as being woven into the utterance. Also, the data format of the correction history is the JSON format or the csv format.

[0116] Returning to FIG. 21, the optimization unit 220 will be described. The optimization unit 220 optimizes the parameters of the correction language model 171 using the correction history stored in the re-learning data storage DB 219 as re-learning data. When the optimization unit 220 takes as input the text obtained by converting the voice data uttered by users A and B, it optimizes the parameters of the correction language model so that the first correction text or the second correction text associated with this input is output. The optimization unit 220 replaces the correction language model 171 with the optimized correction language model 171S at regular intervals.

[0117] In this way, by re-learning the second correction text corrected by an arbitrary user, the correction language model 171S can output, for example, text in which sentence errors peculiar to users A and B are correctly corrected. For this reason, the correction language model 171S after re-learning can correctly correct input data that could not be corrected by the correction language model 171 before re-learning.

[0118] [Processing method] FIG. 24 is a diagram showing the processing procedure of the processing method according to Embodiment 2. Steps S41 to S54 shown in FIG. 24 are the same processing as steps S1 to S14 in FIG. 12.

[0119] The correction UI 50 receives the first correction text from the display UI 40, for example (step S55). When the error in the first correction text is corrected by the operation of an arbitrary user (step S56), the correction UI 50 transmits the corrected second correction text to the processing device 210 (step S57).

[0120] The history storage unit 218 receives the input of the first correction text or the second correction text output from the correction language model 171, formats and corrects it, and then stores it in the re-learning data storage DB 219 as a correction history (step S58).

[0121] Then, the optimization unit 220 optimizes the parameters of the correction language model 171 using the correction history stored in the re-learning data accumulation DB 219 as re-learning data (step S59). The optimization unit 220 replaces the correction language model 171 with the optimized correction language model 171S at regular intervals (step S60).

[0122] [Effect of Embodiment 2] As described above, in Embodiment 2, for the first corrected text output from the correction language model 171, the second corrected text corrected by an arbitrary user is accumulated as a correction history, and the language model is re-learned using the accumulated correction history. Therefore, according to Embodiment 2, by constructing the correction language model 171 optimized according to the tendencies and contexts of Users A and B, for example, it is possible to display subtitles for text corrected specifically for expressions unique to the conversations of the two users, which adapts to the conversation habits of Users A and B.

[0123] [Embodiment 3] Next, Embodiment 3 will be described. In Embodiment 3, for the first corrected text output from the correction language model 171, the second corrected text corrected by an arbitrary user is accumulated as a correction history, and based on the accumulated correction history, the prompt given to the language model is corrected. FIG. 25 is a diagram for explaining the outline of the processing of Embodiment 3.

[0124] In Embodiment 3, similar to Embodiment 2, a correction UI 50 is provided that allows an arbitrary user to correct errors in the first text. Then, in Embodiment 3, the second corrected text corrected by an arbitrary user and the first corrected text output by the correction language model 171 (e.g., LLM) are embedded and vectorized, and the history is accumulated in the accumulation DB 319 in the format of {input - input vector - output} ( (1) in FIG. 25). Further, in Embodiment 3, the scene description text (description of the user's situation) is embedded and vectorized, and accumulated in the accumulation DB 319 in the format of {description of the user's situation - vector of the description of the user's situation}.

[0125] Next, in the third embodiment, when new user voice data is input, among the histories stored in the storage DB 319, a history including an embedding vector similar to the embedding vector of the newly input user voice data is searched for ((2) in FIG. 25).

[0126] Then, in the third embodiment, the second corrected text included in the searched history is added to the input to the correction language model 171 ((3) in FIG. 25). Specifically, in the third embodiment, the second text of the searched history and the prompt added as reference information are given to the correction language model 171 to execute the correction of the input text.

[0127] [Processing System] The configuration of the processing system according to the third embodiment will be described with reference to FIG. 26. FIG. 26 is a diagram showing a configuration example of the processing system according to the third embodiment.

[0128] The processing system 300 according to the third embodiment has a processing device 310 instead of the processing device 210 in FIG. 11.

[0129] The processing device 310 has a history storage unit 318 (storage unit), a storage DB 319 (database), and a prompt correction unit 320 (correction unit) as compared with the processing device 210 in FIG. 21.

[0130] The correction language model 171 is a neural network model having a Transformer architecture, and has, for example, an architecture in which an encoder and a decoder are connected. The correction language model 171 has, for example, a natural language processing model that converts natural language such as BERT (Bidirectional Encoder Representations from Transformers) into vector representations as an encoder part, and a generation model that decodes vector representations into natural language as a decoder part. The correction language model 171 can output the text of the input user utterance sentence as vectorized vector data to the history storage unit 318 and the prompt correction unit 320.

[0131] When there is an input of the second corrected text, the history storage unit 318 stores the second corrected text in the database as history together with the embedded vectors obtained by vectorizing the text (utterance sentences of users A and B) converted by the speech recognition unit 12 and the user's situation description (scene description text). When there is no input of the second corrected text, the history storage unit 318 stores the first corrected text in the database as history together with the embedded vectors obtained by vectorizing the text converted by the speech recognition unit 12 and the scene description text (user's situation description).

[0132] The storage DB 319 stores, as history, the text (utterance sentences of users A and B) converted by the speech recognition unit 12 and the user's situation description (scene description text), their embedded vectors, and the first corrected text, output from the history storage unit 218. Or, it stores, as history, the text (utterance sentences of users A and B) converted by the speech recognition unit 12 and the user's situation description (scene description text), their embedded vectors, and the second corrected text.

[0133] FIG. 27 is a diagram for explaining the correction history stored in the storage DB 319 shown in FIG. 26. In FIG. 27, the flow of the correction history storage process is also explained.

[0134] The case where the first corrected text (corrected output from the language model) by the correction language model 171 is corrected on the correction UI 50 by an arbitrary user during system use is explained.

[0135] Thus, when there is a correction by any user, that is, when the second corrected text is input to the processing device 10, the history storage unit 218 sets the utterance of the user (for example, user A) as userA, the embedding vector of this utterance as embed, and the output after correction by any user (the second corrected text) as output, and records them in the history data D31 (arrows Y31, Y32, Y36). The user's utterance is obtained, for example, from the prompt P31 before correction. Also, the embedding vector of the utterance is vector data output from the encoder part of the correction language model 171 at the time of correcting this utterance.

[0136] On the other hand, a case where there is no correction by any user, that is, a case where the second corrected text is not input to the processing device 10 will be described.

[0137] In this case, the history storage unit 318 sets the utterance of the user (for example, user B) as userB, the embedding vector of this user's utterance as embed, and the output after correction by the correction language model 171 (for example, LLM) (the first corrected text) as output, and records them in the history data D31 (arrow Y35).

[0138] Furthermore, the history storage unit 318 sets the situation description (scene description text) of the user (for example, user A) as UserA’s action, the embedding vector of this user's situation description as embed, and records them in the history data D31 (arrows Y33, Y34). Note that the data format of the correction history may be the JSON format in FIG. 27 or the csv format.

[0139] When the voice data of users A and B are newly input, the prompt correction unit 320 searches for histories in the accumulation DB 319 that include embedding vectors similar to the embedding vectors of the newly input voice data of users A and B among the histories accumulated in the accumulation DB 319. The embedding vectors of the newly input voice data of users A and B are output from the encoder part of the correction language model 171. The prompt correction unit 320 gives the prompt with the searched history added as reference information to the correction language model 171 to execute the correction of the input text.

[0140] FIG. 28 is a diagram for explaining the flow of the prompt creation process by the prompt correction unit 320 shown in FIG. 26. FIG. 29 is a diagram for explaining a specific example of the correction of the prompt.

[0141] For example, when the voice data of user A is newly input, the prompt correction unit 320 calculates the vector distance between the embedding vector of the utterance sentence of user A and the embedding vectors of the utterance sentences of users in each history (for example, the history data D31 in FIG. 28) accumulated in the accumulation DB 319.

[0142] Then, the prompt correction unit 320 compares the calculated vector distances and searches for, for example, the top N similar to the utterance sentence of user A or histories with a vector distance higher than a specific threshold ( (1) in FIG. 28, (1) in FIG. 29).

[0143] The prompt correction unit 320 corrects it like the prompt P32 by quoting the current situation (scene description text) and conversation history included in the searched history as reference information in the prompt P31 before correction ( (1) in FIG. 28, (1) in FIG. 29). This quoted part C31 may include, in addition to the current situation (scene description text) and conversation history at this time, things with no problems in the utterance sentence and past cases not related to the current correction.

[0144] Then, the prompt correction unit 320 gives the corrected prompt P32 with the searched history added as reference information to the language model (arrow Y37 in FIG. 28).

[0145] [Processing method] FIG. 30 is a diagram showing the processing procedure of the processing method according to Embodiment 3. Steps S61 to S74 shown in FIG. 30 are the same processing as steps S1 to S14 in FIG. 12. Steps S75 to S77 shown in FIG. 30 are the same processing as steps S55 to S57 in FIG. 24.

[0146] The history accumulation unit 318 receives the input of the first corrected text output from the correction language model 171 or the second corrected text output from the correction UI 50, and accumulates it as a history in the re-learning data accumulation DB 219 together with the user's utterance sentence, the user's situation description (scene description text), and their embedding vectors (step S78).

[0147] When the voice data of users A and B is newly input, the prompt correction unit 320 searches for a history including an embedding vector similar to the embedding vector of the newly input voice data of users A and B among the histories accumulated in the accumulation DB 319. The prompt correction unit 320 performs a prompt correction process of creating a prompt with the searched history added as reference information (step S79). Then, the prompt correction unit 320 applies the corrected prompt to the correction language model 171 (step S80) to cause text correction to be executed.

[0148] [Effects of Embodiment 3] As described above, in Embodiment 3, for the first corrected text output from the correction language model 171, the second corrected text corrected by an arbitrary user is accumulated as a correction history, and based on the accumulated correction history, the prompt given to the correction language model 171 is corrected. Therefore, according to Embodiment 3, the correction language model 171 can be operated so that text optimized for the tendencies and contexts of users A and B is output.

[0149] [Regarding the system configuration of the embodiment] Each component of the processing systems 100, 200, and 300 is conceptually functional and does not necessarily have to be physically configured as shown in the figures. That is, the specific forms of the distribution and integration of the functions of the processing systems 100, 200, and 300 are not limited to those shown in the figures, and all or part of them can be functionally or physically distributed or integrated in any unit according to various loads, usage situations, etc.

[0150] In addition, each process performed in the processing systems 100, 200, and 300 may be realized in whole or in any part by a program analyzed and executed by a CPU, a GPU (Graphics Processing Unit), and the CPU and GPU. Also, each process performed in the processing systems 100 and 200 may be realized as hardware by wired logic.

[0151] Also, among the processes described in the embodiments, all or part of the processes described as being automatically performed can be manually performed. Or, all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the above-described and illustrated processing procedures, control procedures, specific names, and information including various data and parameters can be appropriately changed unless otherwise specified.

[0152] [Program] FIG. 31 is a diagram showing an example of a computer in which each component device of the processing systems 100, 200, and 300 is realized when a program is executed. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0153] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, the mouse 1110 and the keyboard 1120. The video adapter 1060 is connected to, for example, the display 1130.

[0154] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of each component device of the processing systems 100, 200, 300 is implemented as a program module 1093 in which executable code by the computer 1000 is described. The program module 1093 is stored in, for example, the hard disk drive 1090. For example, the program module 1093 for executing the same process as the functional configuration in each component device of the processing systems 100, 200, 300 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0155] Also, the setting data used in the process of the above-described embodiment is stored as program data 1094 in, for example, the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 to the RAM 1012 and executes them as necessary.

[0156] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090. For example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a LAN (Local Area Network) or a WAN (Wide Area Network)). Then, the program module 1093 and the program data 1094 may be read by the CPU 1020 from another computer via the network interface 1070.

[0157] As described above, the embodiments to which the invention made by the present inventor is applied have been described. However, the present invention is not limited by the description and the drawings that form a part of the disclosure of the present invention according to this embodiment. That is, all other embodiments, examples, and operation techniques made by those skilled in the art based on this embodiment are included in the scope of the present invention.

Explanation of Reference Numerals

[0158] 10, 10B, 210, 310 Processing device 11 Voice input unit 12 Voice recognition unit 13 Image input unit 14 Image description unit 15 Conversation log integration DB 16 Input data creation unit 17, 17B Correction unit 20-1, 20-2 Microphone 30-1, 30-2 Imaging device 40 Display UI 50 Correction UI 100, 200, 300 Processing system 171 Correction language model 171B VLM 218, 318 History accumulation unit 219 Relearning data accumulation DB 220 Optimization unit 319 Storage DB 320 Prompt Correction Unit

Claims

1. An audio input unit that receives inputs of audio data uttered by a first user and audio data uttered by a second user; An image input unit that receives inputs of an image of the first user and an image of the second user; An audio recognition unit that performs audio recognition on the audio data received by the audio input unit and converts the audio data into text; A correction unit that corrects the text converted by the audio recognition unit using a machine learning model trained to correct the text based on the text converted from the audio data uttered by the speaker and data based on the image of the speaker, and outputs the corrected first corrected text to a user interface used by the first user and the second user; A processing device, characterized by comprising the above.

2. Further comprising a generation unit that generates a description of the image received by the image input unit; The machine learning model is a language model that outputs a corrected text obtained by correcting the text when the text converted from the audio data uttered by the speaker and the description of the image of the speaker are input; The correction unit corrects the text converted by the audio recognition unit based on the text converted by the audio recognition unit and the description of the image generated by the generation unit using the language model. The processing device according to claim 1, characterized in that.

3. The language model is given a prompt corresponding to the conversation between the first user and the second user, and generates the first corrected text obtained by correcting the input text into text that appears in a natural context; The prompt is responsible for the task of complementing and / or correcting the utterances in the conversation between the first user and the second user by the language model, and based on the description of the image indicating the past conversation history and the current situation, the first step of complementing the current utterance, the second step of correcting the result of the first step, and the third step of outputting the text obtained by correcting the current utterance as the result of the second step are sequentially performed to correct the text, and includes the past conversation history, the current utterance, and the description of the image indicating the current situation. The processing device according to claim 2, characterized in that.

4. The language model is provided with a prompt corresponding to the conversation between the first user and the second user, and generates the first corrected text obtained by correcting the input text into text that appears in a natural context and then translating it into a specified language. The prompt instructs the language model to perform the task of complementing and / or correcting while translating the utterances in the conversation between the first user and the second user, and includes a first step of complementing the current utterance based on the past conversation history and the description of the image indicating the current situation, a second step of correcting the result of the first step, a third step of translating the text obtained by correcting the current utterance as a result of the second step into a specified language, and a fourth step of outputting the translated text obtained by correcting the current utterance as a result of the third step, thereby sequentially performing the correction of the text. The processing device according to claim 2 is characterized by including the past conversation history, the current utterance, and the description of the image indicating the current situation.

5. An accumulation unit that, when there is an input of a second corrected text in which an error in the first corrected text is corrected by an arbitrary user, accumulates the second corrected text in a database as a correction history, and when there is no input of the second corrected text, accumulates the first corrected text in the database as a correction history. An optimization unit that optimizes the parameters of the language model using the correction history accumulated in the database as data for re-learning. The processing device according to claim 2, characterized by comprising the above.

6. An accumulation unit that, when there is an input of a second corrected text in which an error in the first corrected text is corrected by an arbitrary user, accumulates the second corrected text in a database as a history together with the text converted by the speech recognition unit and the embedding vector obtained by vectorizing the description of the image, and when there is no input of the second corrected text, accumulates the first corrected text in the database as a history together with the text converted by the speech recognition unit and the embedding vector obtained by vectorizing the description of the image. When the voice data of the first user and / or the second user is newly input, among each history stored in the database, search for the history including an embedding vector similar to the embedding vector of the newly input voice data, and provide a prompt with the searched history added as reference information to the language model, a correction unit; The processing device according to claim 2, characterized by comprising.

7. The user interface is a terminal device used by the first user and the second user, AR (Augmented Reality) glasses or VR (Virtual Reality) goggles worn by the first user and the second user respectively, and a display on which the first user and the second user can visually recognize the first corrected text. The processing device according to claim 1, characterized in that.

8. A processing method executed by a processing device, A voice input step of receiving an input of voice data uttered by a first user and voice data uttered by a second user; An image input step of receiving an input of an image of the first user and an image of the second user; A voice recognition step of performing voice recognition on the voice data received in the voice input step and converting the voice data into text; Using a machine learning model trained to correct the text based on the text obtained by converting the voice data uttered by the speaker and the data based on the image of the speaker, correct the text converted in the voice recognition step, and output the corrected first corrected text to a user interface used by the first user and the second user. A correction step; A processing method characterized by including.

9. A voice input step of receiving an input of voice data uttered by a first user and voice data uttered by a second user; An image input step of receiving an input of an image of the first user and an image of the second user; A voice recognition step of performing voice recognition on the voice data received in the voice input step and converting the voice data into text; Using a machine learning model trained to correct the text based on the text converted from the voice data spoken by the speaker and the data based on the image capturing the speaker, correct the text converted in the voice recognition step, and output the corrected first corrected text to a user interface used by the first user and the second user. A processing program for causing a computer to execute.