Check system, check program, and check method

The check system addresses the issue of inconsistent speech and text output by a multimodal model by ensuring matching perceptual information is output, enhancing communication accuracy.

JP7778986B1Active Publication Date: 2025-12-02GEN-AX CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025122491
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2025-06-26
Filing Date
2025-07-22
Publication Date
2025-12-02
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

Current speech multimodal models sometimes output speech and text that do not match in content, leading to inconsistencies.

Method used

A check system that includes an acquisition unit, a generation unit, a determination unit, and an output control unit to ensure that only speech and text with matching perceptual information are output by a multimodal AI model.

Benefits of technology

Prevents simultaneous output of speech and text with inconsistent content, ensuring accurate and synchronized voice and text communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778986000001_ABST
    Figure 0007778986000001_ABST
Patent Text Reader

Abstract

Prevents inconsistent audio and text from being output simultaneously by a multimodal AI model. [Solution] The check system (2) includes a generation unit (232) that inputs user voice data into a first multimodal AI model (M1) and obtains the output of the first multimodal AI model as voice data to be judged and text data to be judged, a judgment unit (233) that judges whether or not perceptual information indicated by the voice to be judged based on the voice data to be judged matches perceptual information indicated by the text to be judged based on the text data to be judged, a re-operation instruction unit (234) that operates the generation unit again when the judgment unit judges that the perceptual information does not match, and an output control unit (235) that outputs at least one of the voice data to be judged and the text data to be judged to an output unit (12) when the judgment unit judges that the perceptual information matches.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a check system, a check program, and a check method. [Background technology]

[0002] In recent years, development of a multimodal model capable of simultaneously outputting speech and text (hereinafter referred to as a speech multimodal model) has been progressing (see, for example, Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] "Introducing the Realtime API", [online], October 1, 2024, OpenAI official website, [Retrieved January 15, 2025], Internet <URL:https: / / openai.com / index / introducing-the-realtime-api / ><URL:https: / / platform.openai.com / docs / guides / realtime / overview> Summary of the Invention [Means for solving the problem]

[0004] A check system according to one embodiment of the present disclosure includes an acquisition unit that acquires user voice data including a user's voice input to an input unit; a generation unit that inputs the user voice data to a first multimodal AI model that, when inputted with voice data including the voice, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the output of the first multimodal AI model as voice data to be judged and text data to be judged; a determination unit that determines whether perceptual information indicated by the voice to be judged based on the voice data to be judged matches perceptual information indicated by the text to be judged based on the text data to be judged; a re-operation instruction unit that operates the generation unit again when the determination unit determines that the perceptual information indicated by the voice to be judged does not match the perceptual information indicated by the text to be judged; and an output control unit that outputs at least one of the voice data to be judged and the text data to be judged to an output unit when the determination unit determines that the perceptual information indicated by the voice to be judged matches the perceptual information indicated by the text to be judged.

[0005] A check program according to one embodiment of the present disclosure causes a processor to execute an acquisition process for acquiring user voice data including a user's voice input to an input unit; a generation process for inputting the user voice data into a first multimodal AI model that, when input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquiring the output of the first multimodal AI model as voice data to be judged and text data to be judged; a determination process for determining whether perceptual information indicated by the voice to be judged based on the voice data to be judged matches perceptual information indicated by the text to be judged based on the text data to be judged; a re-operation instruction process for re-executing the generation process when it is determined in the determination process that the perceptual information indicated by the voice to be judged does not match the perceptual information indicated by the text to be judged; and an output process for outputting at least one of the voice data to be judged and the text data to be judged to an output unit when it is determined in the determination process that the perceptual information indicated by the voice to be judged matches the perceptual information indicated by the text to be judged.

[0006] A checking method according to one embodiment of the present disclosure includes an acquisition step in which a processor acquires user voice data including a user's voice input to an input unit; a generation step in which the processor inputs the user voice data to a first multimodal AI model that, when inputted with voice data including the voice, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the output of the first multimodal AI model as voice data to be determined and text data to be determined; a determination step in which the processor determines whether perceptual information indicated by the voice to be determined based on the voice data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction step in which the generation step is performed again if it is determined in the determination step that the perceptual information indicated by the voice to be determined does not match the perceptual information indicated by the text to be determined; and an output step in which, if it is determined in the determination step that the perceptual information indicated by the voice to be determined matches the perceptual information indicated by the text to be determined, an output unit is caused to output at least one of the voice data to be determined and the text data to be determined. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a block diagram illustrating an example of a configuration of a dialogue system according to an embodiment of the present disclosure. [Figure 2] FIG. 10 is a flowchart showing an example of a flow of a process executed by a check system included in the dialogue system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] <Background Technology and Issues> Prior to describing the embodiments, the background art, problems with the conventional art, means for solving the problems, and the principles of the means for solving the problems will be described.

[0009] [Background technology] Because the speech multimodal model can simultaneously input and output speech and text, it can realize natural speech dialogue end-to-end without using conventional pipeline processing such as speech recognition and speech synthesis. A typical application of the speech multimodal model is ChatGPT's Advanced Voice Mode.

[0010] 〔assignment〕 Normally, the speech and text output simultaneously by a speech multimodal model are expected to match in meaning. However, current speech multimodal models sometimes output speech and text that do not match in content. For example, there are cases where the speech output is "12,000 yen," while the text output is "12,000 yen."

[0011] [Overview of the check system] The check system according to the present disclosure includes: an acquisition unit that acquires user voice data including the user's voice input to the input unit; a generation unit that inputs user voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the outputs of the first multimodal AI model as voice data to be determined and text data to be determined; a determination unit that determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction unit that operates the generation unit again when the determination unit determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output control unit that outputs at least one of the speech data to be determined and the text data to be determined to an output unit when the determination unit determines that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Equipped with.

[0012] [Solution principle] In the check system configured as described above, the output control unit outputs only the speech and text to be judged that the judgment unit determines to be a match to the output unit of the user terminal. This makes it possible to prevent speech and text that are output simultaneously by the multimodal AI model and that do not match in content from being output.

[0013] <Embodiment of the dialogue system 100> First, a dialogue system 100 according to an embodiment of the present disclosure will be described. 1 is a block diagram showing an example of the configuration of the dialogue system 100.

[0014] [Configuration of dialogue system 100] 1, the dialogue system 100 includes a user terminal 1 and a check system 2. The check system 2 may be incorporated into the user terminal 1.

[0015] [User terminal 1] The user terminal 1 includes an input unit 11, an output unit 12, and a terminal-side communication unit 13.

[0016] (Input section 11) The user's voice is input to the input unit 11. The input unit 11 then converts the input user's voice into user voice data. The user voice data is voice data containing the user's voice. The input unit 11 can be configured with various microphones, etc.

[0017] (Terminal side communication unit 13) The terminal-side communication unit 13 communicates with the check system 2. The terminal-side communication unit 13 also transmits the user voice data converted in the input unit 11 to the check system 2. The terminal-side communication unit 13 also receives voice data from the check system 2. The terminal-side communication unit 13 according to this embodiment is configured with a wireless communication module, a terminal for wired connection to the check system 2, and the like. The terminal-side communication unit 13 may also be configured to receive text data from the check system 2 together with or instead of the voice data.

[0018] (Output section 12) The output unit 12 outputs a voice based on the voice data received by the terminal-side communication unit 13. The output unit 12 can be configured with various speakers that output voice. Note that, when the terminal-side communication unit 13 is configured to receive text data, the output unit 12 may be configured to output text based on the text data together with or instead of the voice. In this case, the output unit 12 may be configured to include a display that displays the text.

[0019] [Check System 2] The check system 2 includes a system-side communication unit 21, a storage unit 22, and a calculation unit 23.

[0020] (System side communication unit 21) The system-side communication unit 21 communicates with the user terminal 1. The system-side communication unit 21 is configured with, for example, a wireless communication module that communicates with the user terminal 1, a terminal for connecting to the user terminal 1 by wire, and the like.

[0021] (Storage unit 22) The storage unit 22 stores a check program 221. The check program 221 is a program that causes a computer to function as the check system 2. The storage unit 22 according to this embodiment stores a first multimodal AI model M1 and a second multimodal AI model M2. The storage unit 22 according to this embodiment is configured with a semiconductor memory, a hard disk drive, or the like. The storage unit 22 may be configured to store various calculation results of the calculation unit 23. The storage unit 22 may be divided into a storage unit that stores any of the check program 221, the first multimodal AI model M1, and the second multimodal AI model M2, and a storage unit that stores the rest.

[0022] When audio data including a voice is input, the first multimodal AI model M1 outputs audio data corresponding to the content of the voice and text data of the same content. The first multimodal AI model M1 may be stored in an external device (not shown) that communicates with the check system 2. In this case, the storage unit 22 may not store the first multimodal AI model M1. The external device may or may not be included in the check system 2.

[0023] When voice data and text data are input, the second multimodal AI model M2 determines whether the perceptual information indicated by the voice based on the voice data matches the perceptual information indicated by the text based on the text data. The second multimodal AI model M2 may be stored in an external device (not shown) that communicates with the check system 2. In this case, the storage unit 22 may not store the second multimodal AI model M2. The external device may or may not be included in the check system 2.

[0024] (Configuration of Calculation Unit 23) The calculation unit 23 includes an acquisition unit 231, a generation unit 232, a determination unit 233, a re-operation instruction unit 234, and an output control unit 235. The calculation unit 23 according to this embodiment further includes a second determination unit 236. The calculation unit 23 according to this embodiment is configured with a processor and a memory. That is, the check system 2 according to this embodiment is configured with a computer.

[0025] ·Acquisition part 231 The acquisition unit 231 acquires user voice data. The acquisition unit 231 according to this embodiment acquires user voice data received by the system-side communication unit 21. Note that, if the check system 2 and the input unit 11 are integrally configured, the acquisition unit 231 may directly acquire the user voice data generated by the input unit 11.

[0026] ·Generation unit 232 The generation unit 232 inputs user voice data to the first multimodal AI model M1, and acquires the output of the first multimodal AI model M1 as the voice data to be determined and the text data to be determined.

[0027] ·Second judgment section 236 The second determination unit 236 determines whether or not specific words are included in the speech or text to be determined. "Specific words" refer to words that require particular accuracy when conveyed to the user. "Specific words" include, for example, characters that represent numbers (such as "three" and "ten"). Specific words may also include words that indicate monetary amounts, dates, and times. The second determination unit 236 may be configured to determine whether or not content requesting an answer that includes specific words is included in the speech represented by the user speech data.

[0028] ·Judgment section 233 The determination unit 233 determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined. "Perceptual information" refers to information received by the user's eyes or ears. "Perceptual information" includes the meaning (content) of words, the pronunciation of words, etc. "Match" includes not only a perfect match but also a case where the difference is within a predetermined range. As described above, the calculation unit 23 according to this embodiment is equipped with the second determination unit 236. Therefore, when the second determination unit 236 determines that a specific word is included, the determination unit 233 according to this embodiment determines whether perceptual information indicated by the speech to be determined matches perceptual information indicated by the text to be determined. Furthermore, the determination unit 233 according to this embodiment determines whether the meaning of words indicated by the speech to be determined matches the meaning of words indicated by the text to be determined. Note that the determination unit 233 may be configured to determine whether the pronunciation of words indicated by the speech to be determined matches the pronunciation of words indicated by the text to be determined.

[0029] The determination unit 233 according to this embodiment inputs the speech data and text data to be determined to the second multimodal AI model M2 and obtains the output of the second multimodal AI model M2 as the determination result. The determination unit 233 may be configured to determine whether the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined, for example, by extracting text from the speech data to be determined and comparing it with the text to be determined. In this case, the check system 2 does not need to include the second multimodal AI model M2.

[0030] After the first multimodal AI model M1 starts output, the determination unit 233 according to this embodiment performs a determination on the target speech data and target text data output up to the point at which a predetermined determination condition is met, each time the predetermined determination condition is met. The "determination condition" includes, for example, whether a predetermined time has elapsed since the start of output or the previous determination, or whether the first multimodal AI model M1 has output a predetermined number of characters of speech and text. That is, the determination unit 233 determines in real time (in parallel with the output of speech and text by the first multimodal AI model M1) whether the meaning of the words indicated by the target speech matches the meaning of the words indicated by the target text.

[0031] ·Re-operation instruction section 234 The re-operation instructing unit 234 operates the generation unit 232 again when the determination unit 233 determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined. When operating the generation unit 232 again, the re-operation instructing unit 234 may be configured to instruct the generation unit 232 to input modified speech data obtained by modifying the user speech data into the first multimodal AI model M1. When operating the generation unit 232 again, the re-operation instructing unit 234 may be configured to instruct the generation unit 232 to input the user speech data and the modified prompt into the first multimodal AI model M1. When operating the generation unit 232 again, the re-operation instructing unit 234 may be configured to instruct the generation unit 232 to input the unmodified user speech data and the prompt into the first multimodal AI model M1.

[0032] Output control unit 235 When the determination unit 233 determines that the perceptual information indicated by the speech to be determined matches the perceptual information indicated by the text to be determined, the output control unit 235 outputs at least one of the speech data to be determined and the text data to be determined to the output unit 12. As described above, each time a predetermined determination condition is met, the determination unit 233 performs a determination on the speech data to be determined and the text data to be determined that have been output up until the time the determination condition is met. Therefore, the output control unit 235 according to this embodiment outputs at least one of the speech to be determined and the text to be determined, for which the determination has been made, to the output unit 12, each time the determination unit 233 determines that the perceptual information indicated by the speech to be determined matches the perceptual information indicated by the text to be determined.

[0033] Furthermore, when the determination unit 233 determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined, the output control unit 235 according to this embodiment suspends the operation of causing the output unit 12 to output from the point in time when this determination is made. Furthermore, after the determination unit 233 determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined, and before the re-operation instruction unit 234 operates the generation unit 232 again, the output control unit 235 according to this embodiment causes the output unit 12 to output a predetermined guidance voice or guidance text. The "guidance voice (guidance text)" is, for example, a voice (text) indicating that there is a delay in processing, that it will take x seconds until output, that the user is requested to wait, etc.

[0034] (Processing performed by the calculation unit 23) Next, the flow of the check process S1 (an example of a check method) executed by the check system 2 will be described. FIG. 2 is a flow diagram showing the flow of the check process S1. The functions of the above-mentioned control blocks 231 to 236 are realized by the calculation unit 23 executing the check process S1 as shown in FIG. 2 in accordance with the check program 221 stored in the storage unit 22. The check process S1 includes an acquisition process S11, a generation process S12, a determination process S13, a re-operation instruction process S14, and an output process S15. The check process S1 according to this embodiment further includes a second determination process S16, a reception determination process S17, and an end determination process S18.

[0035] Reception determination process S17 In the initial reception determination process S17, the system side communication unit 21 repeatedly determines whether or not it has received user voice data until it determines that it has received user voice data (S17: YES).

[0036] Acquisition process S11 If it is determined in the reception determination process S17 that the system-side communication unit 21 has received user voice data (S17: YES), the process proceeds to the acquisition process S11. In the acquisition process S11, the acquisition unit 231 acquires user voice data including the user's voice input to the input unit 11. The execution of the acquisition process S11 by the acquisition unit 231 corresponds to the acquisition step in the check method.

[0037] Generation process S12 After the acquisition process S11, the process proceeds to the generation process S12. In the generation process S12, the generation unit 232 inputs user voice data to the first multimodal AI model M1 and acquires the output of the first multimodal AI model M1 as the voice data and text data to be judged. Execution of the generation process S12 by the generation unit 232 corresponds to the generation step in the check method.

[0038] Second determination process S16 In the check process S1 according to this embodiment, the generation process S12 is followed by a second determination process S16. In the second determination process S16, the second determination unit 236 determines whether a specific word is included in the speech or text to be determined. Execution of the second determination process S16 by the second determination unit 236 corresponds to the second determination step in the check method.

[0039] Judgment process S13 If the second determination process S16 determines that the target speech or target text contains a specific word (S16: YES), or after the generation process S12, the process proceeds to determination process S13. In determination process S13, the determination unit 233 determines whether the perceptual information indicated by the target speech based on the target speech data matches the perceptual information indicated by the target text based on the target text data. Execution of determination process S13 by the determination unit 233 corresponds to the determination step in the check method. This determination process S13 is executed every time a predetermined determination condition is met.

[0040] Re-operation instruction processing S14 If it is determined in the determination process S13 that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined (S13: NO), the process proceeds to the re-operation instruction process S14. In the re-operation instruction process S14, the re-operation instruction unit 234 causes the generation unit 232 to execute the generation process S12 again. The execution of the re-operation instruction process S14 by the re-operation instruction unit 234 corresponds to the re-operation instruction step in the check method.

[0041] Output process S15 If the second determination process S16 determines that the target speech or target text does not contain a specific word (S16: NO), or if the determination process S13 determines that the perceptual information indicated by the target speech matches the perceptual information indicated by the target text (S13: YES), the process proceeds to output process S15. In output process S15, the output control unit 235 causes the output unit 12 to output at least one of the target speech data and the target text data. Execution of output process S15 by the output control unit 235 corresponds to the output step in the checking method. This output process S15 is executed every time the determination process S13 determines that the perceptual information indicated by the target speech matches the perceptual information indicated by the target text.

[0042] Termination decision process S18 After the output process S15, the process proceeds to the end determination process S18. In the end determination process S18, it is determined whether or not the user's speech has ended. If it is determined that the user's speech has not ended (S18: NO), the process returns to the reception determination process S17. On the other hand, if it is determined that the user's speech has ended (S18: YES), the check process S1 ends.

[0043] [Effects of Check System 2 (Dialogue System 100)] In the check system 2 described above, the output control unit 235 outputs only the speech to be judged based on the speech data to be judged generated by the first multimodal AI model M1 and the text to be judged based on the text data to be judged that the judgment unit 233 judges to match, to the output unit 12 of the user terminal 1. Therefore, the check system 2 and the dialogue system 100 including the check system 2 can prevent speech and text with inconsistent content from being output simultaneously by the multimodal AI model.

[0044] [Modification of Check System 2] For example, the calculation unit 23 of the check system 2 according to the above embodiment includes the second determination unit 236. However, the calculation unit 23 does not have to include the second determination unit 236. That is, the check system 2 may be configured such that the determination unit 233 determines whether or not the perceptual information matches for all of the speech data and text data to be determined that are output by the first multimodal AI model M1.

[0045] Furthermore, the check program 221 may be stored non-transitory on one or more computer-readable storage media. The storage media may or may not be included in the device. In the latter case, the check program 221 may be supplied to the device via any wired or wireless transmission medium.

[0046] In addition, some or all of the functions of each of the control blocks can be realized by logic circuits. For example, integrated circuits in which logic circuits that function as each of the control blocks are formed are also included in the scope of the present disclosure. In addition, the functions of each of the control blocks can also be realized by, for example, a quantum computer.

[0047] <Summary> The present disclosure describes at least the following configurations.

[0048] (Configuration 1) an acquisition unit that acquires user voice data including the user's voice input to the input unit; a generation unit that inputs user voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the outputs of the first multimodal AI model as voice data to be determined and text data to be determined; a determination unit that determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction unit that operates the generation unit again when the determination unit determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output control unit that outputs at least one of the speech data to be determined and the text data to be determined to an output unit when the determination unit determines that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Equipped with Check system.

[0049] According to the above configuration, it is possible to prevent voice and text with inconsistent content from being output simultaneously by the first multimodal AI model.

[0050] (Configuration 2) the determination unit determines whether or not the meaning of a word indicated by the speech to be determined matches the meaning of a word indicated by the text to be determined. The checking system according to claim 1.

[0051] According to the above configuration, words and phrases that have the same meaning but different expressions are not judged as being inconsistent.

[0052] (Configuration 3) the determination unit determines whether or not the pronunciation of words indicated by the target speech matches the pronunciation of words indicated by the target text; 3. The check system according to configuration 1 or 2.

[0053] According to the above configuration, it is possible to prevent confusion between the on-yomi and kun-yomi readings of the same character (for example, reading "tai ne" as "gaine").

[0054] (Configuration 4) the determination unit inputs the target voice data and the target text data to a second multimodal AI model that, when inputting voice data and text data, determines whether perceptual information indicated by voice based on the voice data matches perceptual information indicated by text based on the text data, and obtains an output of the second multimodal AI model as a determination result; 4. The check system according to claim 1.

[0055] According to the above configuration, it is possible to quickly and accurately perform a judgment by simply inputting the speech data to be judged and the text data to be judged.

[0056] (Configuration 5) the determination unit, after the first multimodal AI model starts outputting, each time a predetermined determination condition is satisfied, performs a determination on the determination-target voice data and the determination-target text data that have been output up until the time the determination condition is satisfied; The output control unit each time the determination unit determines that the perceptual information indicated by the speech to be determined matches the perceptual information indicated by the text to be determined, the output unit is caused to output at least one of the speech to be determined and the text to be determined, When the determination unit determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined, the output unit stops outputting the speech from the point in time when the determination unit made the perceptual information indicated by the speech to be determined. The check system according to any one of configurations 1 to 4.

[0057] According to the above configuration, a real-time judgment can be made in parallel with the speech and text output by the first multimodal AI model, and the output can be stopped immediately after a mismatch in perceptual information is found.

[0058] (Configuration 6) a second determination unit that determines whether a specific word is included in the speech to be determined or the text to be determined; When the second determination unit determines that the specific word is included, the determination unit determines whether or not perceptual information indicated by the target speech coincides with perceptual information indicated by the target text. 6. The check system according to any one of configurations 1 to 5.

[0059] According to the above configuration, judgment can be omitted for voice and text that are unlikely to have content inconsistencies, thereby reducing the time lag that occurs in the output of voice and text and reducing the running costs of the system.

[0060] (Configuration 7) The re-operation instruction unit instructs the generation unit to input modified voice data obtained by modifying the user voice data into the first multimodal AI model. The check system according to any one of configurations 1 to 6.

[0061] The above arrangement increases the likelihood that speech and text that match in content will be reproduced.

[0062] (Configuration 8) the re-operation instruction unit instructs the generation unit to input the user voice data and the modified prompt into the first multimodal AI model; The check system according to any one of configurations 1 to 7.

[0063] The above arrangement increases the likelihood that speech and text that match in content will be reproduced.

[0064] (Configuration 9) the output control unit causes the output unit to output a predetermined guidance voice or guidance text before the re-operation instruction unit causes the generation unit to operate again. The check system according to any one of configurations 1 to 8.

[0065] According to the above configuration, even if the regeneration causes the system operation to slow down or pause, the user can be reassured by the guidance.

[0066] (Configuration 10) The processor an acquisition process for acquiring user voice data including the user's voice input to the input unit; a generation process of inputting the user's voice data into a first multimodal AI model that, when inputted with voice data including a voice, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquiring the output of the first multimodal AI model as the voice data to be determined and the text data to be determined; a determination process for determining whether or not perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction process for re-executing the generation process when it is determined in the determination process that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output process for outputting at least one of the speech data to be determined and the text data to be determined to an output unit when the determination process determines that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Execute Check program.

[0067] According to the above configuration, the same effects as those of the first configuration are achieved.

[0068] (Configuration 11) an acquiring step in which a processor acquires user voice data including the user's voice input to an input unit; a generation step in which a processor inputs user voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the outputs of the first multimodal AI model as voice data to be determined and text data to be determined; a determining step in which a processor determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction step of causing the generation step to be performed again when it is determined in the determination step that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output step of outputting at least one of the speech data to be determined and the text data to be determined to an output unit when it is determined in the determination step that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Including, How to check.

[0069] According to the above configuration, the same effects as those of the first configuration are achieved.

[0070] (Additional notes) The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present disclosure.

[0071] The configuration disclosed herein can prevent inconsistent voice and text output from a multimodal AI model. This effect contributes to achieving, for example, Goal 8 "Decent Work and Economic Growth" and Goal 9 "Industry, Innovation and Infrastructure" of the United Nations' Sustainable Development Goals (SDGs). [Explanation of symbols]

[0072] 100 Dialogue Systems 1. User terminal 11 Input section 12 Output section 13 Terminal communication unit 2 Check System 21 System side communication section 22 Memory section 221 Check Program M1 First multimodal AI model M2 Second Multimodal AI Model 23 Arithmetic section 231 Acquisition Department 232 Generation part 233 Judgment section 234 Re-operation instruction section 235 Output control section 236 Second Judgment Department S1 Check Processing S11 Acquisition process S12 Generation process S13 Judgment process S14 Re-operation instruction processing S15 Output Processing S16 Second determination process S17 Reception determination process S18 End judgment process

Claims

1. an acquisition unit that acquires user voice data including the user's voice input to the input unit; a generation unit that inputs user voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the output of the first multimodal AI model as voice data to be determined and text data to be determined; a determination unit that determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction unit that operates the generation unit again when the determination unit determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output control unit that outputs at least one of the speech data to be determined and the text data to be determined to an output unit when the determination unit determines that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Equipped with Check system.

2. the determination unit determines whether or not the meaning of a word indicated by the speech to be determined matches the meaning of a word indicated by the text to be determined. The checking system according to claim 1 .

3. the determination unit determines whether or not the pronunciation of words indicated by the target speech matches the pronunciation of words indicated by the target text; The checking system according to claim 1 .

4. the determination unit inputs the voice data and the text data to be determined into a second multimodal AI model that, when inputting voice data and text data, determines whether perceptual information indicated by voice based on the voice data matches perceptual information indicated by text based on the text data, and obtains an output of the second multimodal AI model as a determination result; The check system according to any one of claims 1 to 3.

5. the determination unit, after the first multimodal AI model starts outputting, each time a predetermined determination condition is satisfied, performs a determination on the determination target voice data and the determination target text data that have been output up until the time the determination condition is satisfied; The output control unit each time the determination unit determines that the perceptual information indicated by the speech to be determined matches the perceptual information indicated by the text to be determined, the output unit is caused to output at least one of the speech to be determined and the text to be determined, When the determination unit determines that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined, the output unit stops outputting the speech from the point in time when the determination unit made the perceptual information indicated by the speech to be determined. The checking system according to claim 1 .

6. a second determination unit that determines whether a specific word is included in the speech to be determined or the text to be determined; When the second determination unit determines that the specific word is included, the determination unit determines whether or not perceptual information indicated by the target speech coincides with perceptual information indicated by the target text. The checking system according to claim 1 .

7. the output control unit causes the output unit to output a predetermined guidance voice or guidance text before the re-operation instruction unit causes the generation unit to operate again. The checking system according to claim 1 .

8. The processor an acquisition process for acquiring user voice data including the user's voice input to the input unit; a generation process of inputting the user's voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquiring the output of the first multimodal AI model as the voice data to be determined and the text data to be determined; a determination process for determining whether or not perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction process for re-executing the generation process when it is determined in the determination process that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output process for outputting at least one of the speech data to be determined and the text data to be determined to an output unit when the determination process determines that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Execute Check program.

9. an acquiring step in which a processor acquires user voice data including the user's voice input to an input unit; a generation step in which a processor inputs the user's voice data into a first multimodal AI model that, when voice data including a voice is input, outputs voice data having content corresponding to the content of the voice and text data having the same content, and acquires the output of the first multimodal AI model as voice data to be determined and text data to be determined; a determining step in which a processor determines whether perceptual information indicated by the speech to be determined based on the speech data to be determined matches perceptual information indicated by the text to be determined based on the text data to be determined; a re-operation instruction step of causing the generation step to be performed again when it is determined in the determination step that the perceptual information indicated by the speech to be determined does not match the perceptual information indicated by the text to be determined; an output step of outputting at least one of the speech data to be determined and the text data to be determined to an output unit when it is determined in the determination step that the perceptual information indicated by the speech data to be determined matches the perceptual information indicated by the text to be determined; Including, How to check.

Citation Information

Patent Citations

  • Information processing system, image processing apparatus, method for controlling information processing system, and program

    JP2022096305A

  • Foreign language song pronunciation learning support device, method and computer program of the same

    JP2024117979A

  • Authenticity determination support device, authenticity determination support method, program, and synthesis degree determination support device

    JP7591854B1

  • JPP7591854B