Language processing system, language processing method, and language processing program

The language processing system addresses slow response times in LLM dialogues by using parallel inferences to generate timely and accurate responses based on user inputs, improving human-machine communication efficiency.

JP2026006947APending Publication Date: 2026-01-16DONUT ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024106332
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing language processing systems face challenges in improving response speed, particularly in dialogues involving large-scale language models (LLMs), where users have to wait significant time for machine responses.

Method used

A language processing system that processes user utterances and machine responses by outputting repeated inference instructions at predetermined intervals, generating predicted sentences with high accuracy, and dynamically executing judgment, predictive, and response inferences in parallel to enhance response speed and quality.

Benefits of technology

The system significantly reduces waiting times for machine responses by processing inferences in parallel, allowing for timely and accurate responses based on user inputs, including facial expressions and past interactions, thereby enhancing human-machine communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026006947000001_ABST
    Figure 2026006947000001_ABST
Patent Text Reader

Abstract

To provide a technology for improving a response speed in language processing related to interaction between a user and a machine.SOLUTION: An inference instruction is repeatedly outputted on the basis of the uttered sentence of a user updated at every prescribed time, a predicted sentence following the uttered sentence and its prediction accuracy are generated according to the inference instruction, and an answer sentence to the predicted sentence whose prediction accuracy is judged to be high is outputted.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a language processing system, a language processing method, and a language processing program that perform language processing of utterances by a user and responses by a machine. [Background technology]

[0002] With the development of natural language processing technology, the accuracy of dialogues in human-machine interfaces is improving. Also, to realize more natural dialogues, it is desirable to improve the response speed of machines.

[0003] Patent Document 1 discloses a technology that shortens the waiting time between a user's utterance and a response from a device and realizes a smooth dialogue between the user and the device. Patent Document 1 discloses a technology that includes the steps of generating a first response sentence that constitutes the opening part of a response to the utterance, outputting the first response sentence by voice, acquiring information related to text data in parallel with the voice output of the first response sentence and generating a second response sentence that constitutes the answer part of the response to the utterance based on the acquired information, and outputting the second response sentence by voice after the voice output of the first response sentence is completed, thereby realizing a smooth dialogue between the user and the device.

[0004] Patent Document 2 discloses a technology that can improve user convenience by estimating the content of an utterance made by a user U following a received utterance and providing the utterance content including the estimated utterance content to the user U. Patent Document 2 discloses generating the trained model using training data that includes divided audio information obtained by dividing the audio information of the entire user's utterance from the beginning to different positions or divided text information corresponding to the divided audio information and text information of the entire user's utterance, and estimating the content of the user's utterance following the utterance using the trained model.

[0005] Non-Patent Document 1 discloses an AI chatbot built based on a large-scale language model (LLM). The AI ​​chatbot described in Non-Patent Document 1 is characterized by generating answers that humans can recognize naturally in response to questions, requests, calls, etc. input by humans. As Non-Patent Document 1 is able to provide highly accurate answers to all inputs, it is increasingly being used in information gathering, text summarization, writing articles and scripts, and automatic response systems via websites. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-107078 [Patent Document 2] Japanese Patent Application Publication No. 2023-129020 [Non-patent literature]

[0007] [Non-Patent Document 1] "Introducing ChatGPT," [Retrieved April 10, 2024], Internet<URL:https: / / openai.com / blog / chatgpt> Summary of the Invention [Problem to be solved by the invention]

[0008] In services that utilize natural language processing technology such as those mentioned above, improvements in machine response speed are desirable. Although machine response speeds are on the rise, if it takes a long time to generate an answer, users have to wait, which hinders smooth communication. Furthermore, in dialogues that utilize LLMs, for example, the LLM begins generating output after the user enters input, which means that users have to wait a considerable amount of time before receiving the output from the LLM.

[0009] An object of the present invention is to provide a technology for improving response speed in language processing of user utterances and machine responses. [Means for solving the problem]

[0010] [1] A language processing system that processes user utterances and machine responses, outputting repeated inference instructions based on user utterances updated at predetermined intervals; A predicted sentence following the uttered sentence and its prediction accuracy are generated according to the inference instruction; A language processing system that outputs a response sentence to the predicted sentence that is determined to have a high prediction accuracy. [2] The language processing system according to [1], wherein the inference instruction is output to an inference model. [3] If the prediction accuracy is high, output a response sentence to the predicted sentence; A language processing system as described in [1] or [2], which, if the prediction accuracy is low, outputs an inference instruction including an updated spoken sentence at a subsequent predetermined time, and outputs an agreement sentence with the predicted sentence. [4] determining whether the predicted sentence is a response request from the user; If the response request is received, a response sentence is output in response to the predicted sentence; The language processing system according to any one of [1] to [3], wherein if the predicted sentence is not a response request, a corresponding sentence is output in response to the predicted sentence. [5] When the end of the user's utterance is detected, the series of utterances is output. determining whether the predicted sentence and the series of spoken sentences are different; If there is a difference, output a response sentence to the series of utterance sentences; The language processing system according to any one of [1] to [4], wherein if there is no discrepancy, a response sentence or a concurring sentence to the predicted sentence is output. [6] Output an inference instruction as to whether the facial information obtained by analyzing the image of the user corresponds to the response request. The language processing system according to any one of [1] to [5], wherein if the response request is received, a spontaneous opinion sentence is output. [7] The language processing system according to [6], wherein if the response is not a response request, an inference instruction relating to the response request is output after a predetermined waiting time has elapsed. [8] outputting an inference instruction including information about the user's state, which is the result of analyzing the image; A language processing system according to [6] that outputs spontaneous opinion sentences. [9] When the end of the user's speech is detected, the elapsed time is measured. The language processing system according to any one of [1] to [8], which outputs a spontaneous opinion sentence according to the elapsed time.

[10] The language processing system according to any one of [1] to [9], which outputs a translation of the predicted sentence determined to have a high prediction accuracy.

[11] Output a translation instruction including the predicted sentence determined to have a high prediction accuracy and a specified language; The language processing system according to

[10] , wherein a translation of the predicted sentence is generated in accordance with the translation instruction, and the translation is output.

[12] Contextual information corresponding to past interactions with the user is associated with an index personalized for the user and used to reference the interactions, and the associated information is stored in a database; When the context information included in the predicted sentence is extracted, the corresponding index stored in the database is referenced; The language processing system according to any one of [1] to

[11] , which outputs an answer sentence in accordance with an inference instruction including the index and the predicted sentence.

[13] The database stores, as a long-term memory index relating to a memory, context information, a vector representation representing the memory, an importance index of the memory, and an emotion tag of the user in the interaction, in association with each other; When the context information included in the predicted sentence is extracted, the long-term memory index is referenced; The language processing system according to

[12] , which outputs an answer sentence according to an inference instruction including the long-term memory index and the predicted sentence.

[14] The database stores, as a skill index relating to a skill, context information, the content of the skill, a relevance index, and an execution code relating to a predetermined interaction, in association with each other; When the context information included in the predicted sentence is extracted, the skill index is referenced, The language processing system according to

[12] , which outputs an answer sentence according to an inference instruction including the skill index and the predicted sentence.

[15] The inference instruction outputs a false statement and a synchronizing statement when the prediction accuracy is low; If the prediction accuracy is high, output true and the predicted sentence. The language processing system according to any one of [1] to

[14] , wherein the output is repeated until the true is output.

[16] A language processing system that processes user utterances and machine responses, outputting an inference instruction as to whether a face image or state information obtained as a result of analyzing an image of the user corresponds to a response request; If the response request is received, the language processing system outputs a spontaneous opinion sentence.

[17] A language processing method for processing utterances by a user and responses by a machine, outputting repeated inference instructions based on user utterances updated at predetermined intervals; A predicted sentence following the uttered sentence and its prediction accuracy are generated according to the inference instruction; A language processing method in which a computer executes each process to output a response sentence to the predicted sentence that is determined to have a high prediction accuracy.

[18] A language processing program that causes a computer to function so as to execute the language processing method described in

[17] .

[0011] According to the invention described in [1], it is possible to improve the response speed in language processing in human-machine communication.

[0012] According to the invention described in [2], multiple inferences can be processed.

[0013] According to the invention described in [3], an appropriate answer sentence can be output according to the prediction accuracy.

[0014] According to the invention described in [4], it is possible to determine whether or not the user is requesting a response, and then generate an appropriate reply sentence.

[0015] According to the invention of [5], if the predicted sentence is inappropriate, an appropriate answer can be given.

[0016] According to the inventions in [6], [7], and [8], it is possible to determine whether a response is desired from the user's facial expression and behavior, and have the machine make spontaneous statements accordingly.

[0017] According to the invention described in [9], it is possible to make a machine spontaneously speak in situations where silence continues.

[0018] According to the inventions

[10] and

[11] , the response speed in translation can be improved.

[0019] According to the invention of

[12] , it is possible to refer to the user's past interactions and have the machine make more appropriate statements with a high response speed.

[0020] According to the invention of

[13] , it is possible to refer to the emotionally important memories of the user and have the machine make appropriate statements.

[0021] According to the invention of

[14] , it is possible to refer to the user's preferences and have the machine take appropriate action.

[0022] According to the invention described in

[15] , it is possible to generate a reply sentence more quickly.

[0023] According to the invention of

[16] , it is possible to infer from the image analysis results the situations in which the user wants a response, and to have the machine make spontaneous statements as necessary. [Effects of the Invention]

[0024] According to the present invention, it is possible to provide a technology for improving response speed in language processing of user utterances and machine responses. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 is a block diagram of a system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a hardware configuration of the present embodiment. [Figure 3] 10 is a Gantt chart according to the present embodiment. [Figure 4] FIG. 2 is a schematic diagram of the flow of parallel inference according to the present embodiment. [Figure 5] FIG. 2 is a schematic diagram of the flow of opinion inference according to the present embodiment. [Figure 6] 10 shows an example of index data according to the present embodiment. [Figure 7] 10 is a flowchart of an index generation process according to the present embodiment. [Figure 8] 10 is a flowchart showing the use of an index according to the present embodiment. [Figure 9] 10 is a flowchart of a translation process according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] A language processing system according to an embodiment of the present invention will be described below with reference to the accompanying drawings. Note that the embodiment shown below is an example of the present invention, and the present invention is not limited to the embodiment below, and various configurations can be adopted.

[0027] In this embodiment, the configurations, operations, etc. of a language processing system, a language processing device, and a language processing program are described, but similarly configured methods, computer programs, and program recording media on which the programs are recorded also achieve similar effects. For example, by using a program recording media, the programs can be installed on a computer. The series of processes according to this embodiment described below are provided as a computer-executable program, and can be provided via a non-transitory computer-readable recording medium such as a CD-ROM or a flexible disk, or even via a communication line.

[0028] The language processing system is composed of a computer device. The computer device has an arithmetic unit such as a CPU (Central Processing Unit) and a storage device. The computer device can function as a language processing device by executing a language processing program stored in the storage device using the arithmetic unit. The language processing method is realized by processing the computer device including the language processing device.

[0029] In this embodiment, the language processing system realizes a dialogue between a user and a language processing device (machine). The dialogue includes voice communication and written communication between the user and the machine. It also includes a situation where one party communicates by voice and the other by written communication. In this embodiment, an example is described in which both the user and the machine communicate by voice, but this is not limited to this. Furthermore, the communication is not limited to one-to-one communication between the user and the machine, and may be one-to-many communication between multiple users and machines.

[0030] In this embodiment, a dialogue action from a user to a language processing device is defined as an "utterance." Furthermore, a dialogue action from a language processing device to a user is defined as a "reply." These definitions of dialogue actions are provided with the purpose of clarifying the subject of dialogue actions. Therefore, dialogue actions performed spontaneously by a machine may also be referred to as a reply in this specification.

[0031] In this embodiment, a user interacts with a machine through a user terminal. The user terminal is configured as a dialogue interface between the user and the machine. The user terminal functions as a dialogue interface in the dialogue with the machine. The user terminal has an input interface for input information including voice, text, and image. The user terminal has an output interface for output information including voice and text. In this embodiment, the user terminal realizes the dialogue with the machine through voice input and output. In one embodiment, the user terminal also inputs images to realize the dialogue with the machine.

[0032] FIG. 1 shows a block diagram of a language processing system 1. The language processing system 1 comprises a language processing device 2 and a user terminal 3, which are communicatively connected via a communication network (not shown). The language processing device 2 is also connected to an inference device 4. The inference device 4 can be installed inside or outside the language processing device 2. Here, for simplicity of explanation, an example is shown in which there is one user terminal 3, but there can be multiple user terminals 3.

[0033] The language processing device 2 includes, as functional components, an acquisition unit 21, an instruction generation unit 22, an instruction unit 23, a response processing unit 24, and an index generation unit 25. The language processing device 2 is connected to a database DB for storing indexes used to refer to interactions so as to be able to communicate data. The database DB includes a long-term memory index DB1 related to memory, and a skill index DB2 related to skills. In this embodiment, these databases DB may be further connected to the inference device 4.

[0034] The language processing device 2 receives data from the user terminal 3. The language processing device 2 generates response data based on the received data. The language processing device 2 transmits the generated response data to the user terminal 3. Here, the language processing device 2 transmits the generated response data to the user terminal 3 at any time.

[0035] The user terminal 3 has, as functional components, a voice recognition unit 31, a voice processing unit 32, an image acquisition unit 33, an image processing unit 34, a text processing unit 35, and an output unit 36. The user terminal 3 serves as an input / output interface for data input by the user and data output related to answer data that is the processing result of the language processing device 2.

[0036] In this embodiment, the user terminal 3 acquires the user's voice or image as input data. The input data may include text. The user terminal 3 acquires text as output data from the language processing device 2 as a processing result. The user terminal 3 outputs voice or text based on the acquired text. The user terminal 3 may also control an actuator (not shown) based on the acquired text.

[0037] The inference device 4 executes a process of generating a response sentence to the data included in the inference instruction. The inference device 4 receives the inference instruction from the language processing device 2, generates a response sentence, and transmits the response sentence to the language processing device 2, which is the source of the response sentence. The inference device 4 has an inference model 40, and by inputting the data included in the inference instruction into the inference model 40, the inference device 4 obtains the response sentence as an output.

[0038] In this embodiment, the inference model 40 is a natural language processing model (NLP). The inference model 40 is preferably a large-scale language model (LLM). The natural language processing model realizes computer processing of data input as natural language. In this embodiment, the type of natural language processing model employed is not limited.

[0039] The natural language processing model can output text classification, sentiment analysis, text summarization, question answering, and the like as a result of natural language processing of data input in natural language. In this embodiment, the inference model 40 functions like a chatbot and outputs processing results for data input in natural language. Note that data input in natural language may be in natural language at the input stage but may no longer be in natural language before or during processing in the natural language processing model. For example, data input in natural language may be converted into data in a format that can be processed by the natural language processing model and then processed.

[0040] The inference model 40 may be implemented in a form other than a natural language processing model as long as it can output a response as a processing result for a natural language input. There may also be multiple inference models 40. The inference model 40 may be a combination of machine-learned models that correspond to the format of the data to be inferred or the subject matter to be inferred.

[0041] When the inference model 40 is an LLM, it outputs an inference result according to an inference instruction that includes text and an image. The inference instruction is a prompt, and the instruction content to be inferred is written in natural language. For example, in an inference instruction that includes text, the instruction "Please create a response to [utterance]" outputs a natural response to the utterance as the inference result. Note that [utterance] is any text. Also, in an inference instruction that includes an image, the instruction "Please explain the content of [image] in text" outputs text that explains the content of the image as the inference result. Note that [image] is any image.

[0042] The user terminal 3 functions as an interface for inputting and outputting information. The voice recognition unit 31 acquires the user's spoken voice obtained through a microphone. The voice processing unit 32 converts the acquired spoken voice into a spoken sentence, which is text data. The voice processing unit 32 transmits the spoken sentence to the language processing device 2 at any time. The image acquisition unit 33 acquires an image including the user and their surrounding environment captured by a camera. The image processing unit 34 transmits the image to the language processing device 2 at any time. In this embodiment, the image processing unit 34 has a face analysis function and generates facial information of the user in the acquired image as an analysis result. The image processing unit 34 transmits the facial information to the language processing device 2 at any time.

[0043] The language processing device 2 processes various input information sent from the user terminal 3 and transmits the processing results to the user terminal. The acquisition unit 21 acquires input information including utterances and images from the user terminal 3. The instruction generation unit 22 generates an inference instruction using the acquired input information. The instruction unit 23 transmits the inference instruction to the inference device 4. The inference model 40 of the inference device 4 executes inference processing in accordance with the inference instruction and transmits the inference result to the language processing device 2. The answer processing unit 24 acquires the inference result from the inference device 4 and transmits an answer sentence based on the inference result to the user terminal 3 at any time.

[0044] The index generation unit 25 acquires various input information from the acquisition unit 21, generates a long-term memory index and a skill index, and stores them in the database DB. The inference device 4 determines whether or not it is necessary to refer to the index by inference, and can refer to the long-term memory index DB1 or the skill index DB2.

[0045] The text processing unit 35 of the user terminal 3 acquires the answer sentence, which is text data, from the answer processing unit 24. The text processing unit 35 executes a process of converting the answer sentence into voice, thereby generating a voiced answer. The output unit 36 ​​outputs the voiced answer from a speaker, thereby realizing a voice dialogue between the user and the machine. The output unit 36 ​​may output the answer sentence, which is text data, by displaying it on a display or the like. The output unit 36 ​​may also display an image of the answer sentence.

[0046] FIG. 2(a) shows a hardware configuration diagram of the language processing device 2. The language processing device 2 includes a control device 201, a storage device 202, and a communication device 203 as its hardware configuration. In this embodiment, the language processing device 2 can be a computer device such as a server device or a personal computer. Note that the language processing device 2 may be configured by multiple computer devices, and is not limited to the configuration shown in FIG. 2(a) as long as it can realize the above-mentioned functional components (21-25) as a whole.

[0047] The control device 201 is configured with one or more processors such as a CPU, and controls the overall processing in the language processing device 2 by executing a language processing program, an OS (Operating System), and other applications. The storage device 202 is a hard disk drive (HDD), solid state drive (SSD), flash memory, RAM (Random Access Memory), etc., and stores the language processing program and various data. The communication device 203 controls communication with a communication network, and realizes data communication with the user terminal 3 and external devices.

[0048] FIG. 2(b) shows a hardware configuration diagram of the user terminal 3. The user terminal 3 includes, as its hardware configuration, a control device 301, a storage device 302, a communication device 303, a microphone 304, a speaker 305, a camera 306, and a display 307. The user terminal 3 may further include an input interface such as a keyboard, a mouse, or a touch panel. In this embodiment, the user terminal 3 may be a smartphone, a personal computer, a tablet terminal, or the like. The user terminal 3 may also be a robot device.

[0049] The control device 301 is composed of one or more processors such as a CPU, and controls the overall processing of the user terminal 3 by executing a terminal language processing program, an OS, other applications, etc. The storage device 302 is an HDD, SSD, flash memory, RAM, etc., and stores a browser application and various data. The communication device 303 controls communication with the communication network and realizes data communication with at least the language processing device 2. The microphone 304 accepts voice input. The speaker 305 outputs voice. The camera 306 captures images of the user and their surrounding environment. The display 307 displays and outputs text data and images. The user terminal 3 may further be equipped with a GPS communication unit that acquires location information.

[0050] In this embodiment, the user terminal 3 stores a terminal language processing program in the storage device 302 and executes the program using the control device 301, thereby realizing the functional components (31-37). The user terminal 3 may also function as the functional components (31-37) by accessing the language processing system 1 via a web browser. The user terminal 3 may also be configured to realize part of the functional components (21-24) of the language processing device 2.

[0051] In this embodiment, inferences are classified into judgment inferences, predictive inferences, and answer inferences. These inferences are executed by inputting information into the inference model 40 of the inference device 4. In a large-scale language model, inferences with different properties can be executed, and each inference can also be processed in parallel.

[0052] Judgment inference is inference that determines whether subsequent processing should be performed in response to information input. In this embodiment, judgment inference includes prediction accuracy judgment, question mark judgment, response request judgment, and discrepancy judgment. Judgment inference outputs a judgment result using a true / false value or a graded value. In this embodiment, judgment inference uses a true / false value, and if a true judgment is made, processing such as subsequent inference is performed, and if a false judgment is made, processing different from true is performed. Note that if graded values ​​are used, processing can be specified for each grade.

[0053] Predictive inference is an inference that predicts what kind of sentence will follow a certain sentence that has been input. Here, the certain sentence that has been input is, for example, a sentence that is being spoken by a user, and is not a complete sentence but an interrupted sentence. Predictive inference generates a predicted sentence that will follow the interrupted sentence. It is preferable that the predicted sentence be a complete sentence, but as long as the meaning of the sentence can be read and a response can be made to it, completeness as a sentence is not an issue.

[0054] Answer inference is an inference that indicates what kind of answer sentence should be produced in response to information input. In this embodiment, the answer sentence includes a response sentence, an agreement sentence, an opinion sentence, and a translation sentence. The answer sentence obtained by answer inference is text data, and is output as voice or text on the user terminal 3. The answer sentence can also include a control command for an actuator of the user terminal 3, and the user terminal 3 can output a predetermined operation by driving the actuator in accordance with the control command. The actuator, for example, is involved in driving the arm or wheel of the user terminal 3 configured as a robot device, and can control the grasping and movement of an object and the movement of the user terminal 3.

[0055] The reply may include machine-generated emotional data. The emotional data includes joy, anger, sadness, happiness, surprise, admiration, etc. The emotional data is linked to a facial expression corresponding to the emotion, and the user terminal 3 can convey the emotion to the user by displaying an image corresponding to the facial expression on the display.

[0056] The reply sentence is text data output to the user terminal 3. In this embodiment, the reply sentence includes a response sentence indicating a response to an utterance from the user, an agreement sentence indicating agreement with an utterance from the user, an opinion sentence spontaneously uttered by the machine, a translation of the content of the utterance from the user, and the like.

[0057] <Embodiment 1> In the first embodiment, parallel inference, which is a mechanism for improving the response speed of a machine, will be described. Parallel inference is realized by a combination of at least judgment inference and response inference. Parallel inference is preferably realized by a combination of judgment inference, predictive inference, and response inference.

[0058] In a conversation between humans, the speaker speaks, and the listener listens to the utterance and infers what response will be given. This allows the listener to give an appropriate response as soon as the speaker finishes speaking. In contrast, in conventional human-machine communication, the speaker (user) inputs a series of utterances into the machine, and the machine then performs an inference process to determine what response to give. As a result, compared to a conversation between humans, it takes a certain amount of time for the machine to begin responding, posing a challenge to achieving smooth communication. Parallel inference aims to shorten the waiting time between the end of the user's utterance and the start of the machine's response by dynamically executing inference processes in parallel to determine what response the machine will give while the user is speaking.

[0059] FIG. 3 is an example of a Gantt chart showing the flow of processing related to parallel inference in chronological order. In FIG. 3, time t0 to time t9 elapses from left to right. Each time t indicates a predetermined time. The predetermined time may be any time, such as one second, and is not particularly limited.

[0060] At time t0, voice input of an utterance begins via the microphone 304 of the user terminal 3. FIG. 3 illustrates an example in which a user attempts to obtain an answer from a machine for the utterance, "Please tell me some recommended ramen restaurants in Tokyo." This utterance is dynamically input into the microphone over the following times t1 to t5. The utterance input into the microphone is dynamically transmitted to the language processing device 2 for processing. The language processing device 2 dynamically transmits the received utterance to the inference device 4 and acquires the inference result. The language processing device 2 transmits an answer sentence based on the inference result to the user terminal 3, allowing the user to obtain the answer.

[0061] The utterance sentence input via the microphone is subjected to a judgment inference relating to prediction accuracy judgment. The prediction accuracy judgment determines the prediction accuracy with which a predicted sentence can be generated from the input utterance sentence. In the prediction accuracy judgment, an inference instruction including the utterance sentence is sent to the inference model 40, and the prediction accuracy can be generated as the inference result. In this embodiment, if the prediction accuracy is high, true is generated, and if the prediction accuracy is low, false is generated.

[0062] At time t1, the utterance sentence "In Tokyo" is input, and at time t2, the utterance sentence "Recommended in Tokyo" is input, and false is generated as the prediction accuracy is low. In this embodiment, when the prediction accuracy is low, an agreement sentence is generated. The generated agreement sentence is temporarily stored in the language processing device 2 at least until the end of the utterance.

[0063] An accompaniment sentence indicates content that is in agreement with the utterance sentence. An accompaniment sentence is used to return some kind of reaction, such as an interjection, when the content of the utterance sentence is unclear or when no response is required. An accompaniment sentence is generated that expresses joy, anger, sadness, happiness, admiration, surprise, etc., depending on the content of the utterance sentence. In addition, one or more accompaniment sentences may be set in advance depending on the content of the utterance sentence, and a sentence may be selectively generated from them.

[0064] At time t3, the utterance sentence "Recommended ramen in Tokyo" is input, and true is generated as the prediction accuracy is high. In this embodiment, when the prediction accuracy is high, a predicted sentence is generated. The generated predicted sentence is temporarily stored in the language processing device 2 at least until the end of the utterance. The predicted sentence is generated as the inference result by sending an inference instruction related to predictive inference to the inference model 40.

[0065] At time t4, a predicted sentence, "Please tell me some recommended ramen restaurants in Tokyo," is generated. At time t4, a response request judgment (judgment inference) is further performed on the predicted sentence. The response request judgment determines the possibility that the predicted sentence is a response request. In the response request judgment, an inference instruction related to the response request judgment including the predicted sentence is sent to the inference model 40, and the possibility that the predicted sentence is a response request can be generated as the inference result. In this embodiment, if it is a response request, true is generated, and if it is not a response request, false is generated.

[0066] A response request means that the user is requesting a response from the machine. It is important for the quality of human-machine communication that the machine speak when the user needs it. It is believed that the quality of communication can be improved by creating a system in which the machine speaks when a response request is received from the user. Specific examples of response requests include speaking when a user wants a response, speaking when called upon, speaking when the user is worried or deep in thought, speaking when needed to overcome a negative situation, not just during a conversation, prompting the user to speak or take action during long periods of silence, and suggesting topics to enhance the user's sense of happiness during long periods of silence. The strictness of the response request determination can be set by including these specific examples in the inference instructions for the inference model 40 as needed.

[0067] If it is a response request, a response inference is made regarding a response sentence to the predicted sentence. In the response inference regarding the response sentence, an inference instruction indicating what kind of response sentence should be made to the predicted sentence is sent to the inference model 40, and a response sentence can be generated as the inference result.

[0068] If the response is not a response request, a response inference is made regarding a synchronizing sentence for the predicted sentence. In the response inference regarding the synchronizing sentence, an inference instruction indicating what kind of synchronizing sentence should be made for the predicted sentence is sent to the inference model 40, and a synchronizing sentence can be generated as the inference result.

[0069] The text of the response sentence or the agreement sentence generated by answer inference is sent to the user terminal 3 at any time. The user terminal 3 converts the received text into voice using a TTS (Text-to-Speech) function and outputs it. In the example of Figure 3, the user's utterance ends at time t5, and the voice is output from the speaker at time t6.

[0070] In Figure 3, inference instructions related to prediction accuracy judgment are repeatedly output to the inference model 40 until time t3 when it is determined that the prediction accuracy is high. This allows the response request judgment and response inference that continue in the early stage when the prediction accuracy is determined to be high to proceed in parallel, thereby improving the machine's response speed.

[0071] Furthermore, the inference instruction for the prediction accuracy judgment may be repeatedly output to the inference model 40 until the end of the utterance. For example, at time t4, the prediction accuracy of the predicted sentence is further improved, and the quality of the predicted sentence is improved. Furthermore, the quality of the response request judgment of the predicted sentence and the subsequent response sentence or agreement sentence is also improved. Therefore, it is preferable that the language processing device 2 temporarily stores multiple predicted sentences and multiple response sentences or agreement sentences in response thereto. The language processing device 2 can provide an answer of improved quality by transmitting the optimal answer sentence selected from the multiple response sentences or agreement sentences to the user terminal 3.

[0072] At time t6, the end of the speech by the user is detected. The end of the speech can be detected by any of the user terminal 3, the language processing device 2, or the inference model 40. The language processing device 2 acquires the end-of-speech detection event and confirms the utterance sentence "Please tell me some recommended ramen shops in Tokyo" acquired up to time t5. For the sake of distinction, the confirmed utterance sentence is called a confirmed utterance sentence.

[0073] At time t7, the language processing device 2 determines whether the predicted sentence generated at time t4 and the definite uttered sentence are different. The difference determination may be performed by sending an inference request including the predicted sentence and the definite uttered sentence to the inference model 40 and obtaining the inference result. Note that the difference between the predicted sentence and the definite uttered sentence here does not require a perfect match, but is determined to be no difference if the meanings of the sentences are considered to be equivalent. If there is a difference, the difference determination result is true, and if there is no difference, false is generated.

[0074] If the predicted sentence and the definite utterance sentence do not differ, the definite utterance sentence is discarded, and a response sentence or synchronizing sentence based on the predicted sentence is adopted as the answer sentence. If the predicted sentence and the definite utterance sentence differ, an inference instruction for a response request determination including the definite utterance sentence is sent to the inference model 40. A response sentence or synchronizing sentence is generated according to the result of the response request, and an answer sentence based on it is finally output to the user terminal 3. In this case, a speaker output is obtained at time t9 in the user terminal 3.

[0075] In parallel inference, an answer sentence based on the predicted sentence can be output to the speaker at time t6. If the discrepancy determination is made after time t6, it is determined that the predicted sentence differs from the definitive utterance sentence while the answer sentence is being output to the speaker. In this case, the inference model 40 can correct the content output to the speaker, generate an answer sentence that answers with the correct content, and output the answer sentence to the speaker. At this time, the answer sentence that has been partially output to the speaker may be included in the inference instruction to generate the answer sentence.

[0076] As explained above, the speaker output based on parallel inference is obtained at time t6, and the speaker output based on inference after detecting the end of the conversation is obtained at time t9. It is understood that parallel inference has the effect of improving response time. Furthermore, even when discrepancy discrimination is performed, if the predicted sentence and the definite utterance sentence do not differ, it is understood that the effect of improving response time is also obtained because a response sentence to the predicted sentence has been generated in advance. Furthermore, discrepancy discrimination can be omitted if the prediction accuracy is higher.

[0077] The acquisition unit 21 acquires utterances by a user that are updated at predetermined intervals. Here, the predetermined intervals refer to the times t shown in Fig. 3. The acquisition unit 21 updates the utterances as a series of sentences every time the predetermined interval t has elapsed since the start of the utterance.

[0078] The instruction generation unit 22 generates an inference instruction based on the acquired utterance sentence. In LLM, the inference instruction is a prompt, and includes the acquired information and instruction information. The instruction unit 23 transmits the inference instruction to the inference device 4. The answer processing unit 24 acquires the inference result from the inference device 4. The answer processing unit 24 transmits an answer sentence to the user terminal 3 based on the inference result.

[0079] 4 shows an overview of the flow of inference. In this embodiment, the inference process is executed by an inference model 40. The instruction generation unit 22 generates an inference instruction for causing the inference model 40 to appropriately execute the inference process shown in FIG.

[0080] The inference model 40 executes an inference process related to determining the prediction accuracy (S201). In S201, if the prediction accuracy is high, the inference model 40 generates a predicted sentence. The inference model 40 executes an inference process to determine whether the predicted sentence contains a question mark (S202). A question mark is inferred to be present if the predicted sentence contains content that indicates a doubt or question. If the inference model 40 determines that a question mark is not present, it further executes a response request determination for the predicted sentence (S203).

[0081] If the inference model 40 determines that there is a question mark (YES in S202) or that there is a response request (YES in S203), it generates a response sentence for the predicted sentence (S204). If the prediction accuracy is low (NO in S201) or if it determines that there is no response request (NO in S203), the inference model 40 generates a default synchronizing sentence or a synchronizing sentence for the predicted sentence (S205). The generated synchronizing sentence is temporarily stored.

[0082] When the language processing device 2 acquires an end-of-speech detection event (YES in S206), it outputs a response sentence or an agreement sentence as a reply sentence to the user terminal 3 (S207). In S207, the language processing device 2 can select a response sentence to be adopted as a reply sentence depending on the result of determining the difference between the predicted sentence and the confirmed utterance sentence. Note that the reply sentence may be output to the user terminal 3 before acquiring an end-of-speech detection event. When the language processing device 2 does not detect an end-of-speech detection event (NO in S206), it generates an inference instruction using the updated utterance sentence and repeatedly executes inference based on the inference instruction from S201.

[0083] The inference process relating to S201 to S203 is not limited to this flow. Furthermore, the inference model 40 can simultaneously execute the inference process relating to the discriminant inference of S201 to S203. The inference instruction can specify the order of the discriminant inference of S201 to S203.

[0084] An example of an inference instruction according to this embodiment is shown in the following paragraph. In the example inference instruction, the instruction information is, "If the user's utterance is a question or an instruction, the content of the utterance can be reliably determined, and there is a high possibility that it is a response request, please return true." This instruction instructs inferences related to a question mark judgment, a prediction accuracy judgment, and a response request judgment. Furthermore, if the result of these judgment inferences is true, a predicted sentence is generated in [the content predicted for the continuation of the user's utterance]. Note that this inference instruction only needs to be input to the inference model 40 at the beginning, and subsequent inference instructions may include only the user's utterance (spoken sentence). From the perspective of improving response speed, it is preferable that the inference instruction be simple. Furthermore, the inference instruction may include an instruction to generate a response sentence to the predicted sentence.

[0085] If the user's statement is a question or instruction, the content of the statement can be determined with certainty, and there is a high possibility that it requires a response, return true. Returns the following format: [true][Prediction of what the user will say next] If it is false, do the following: [false][Synchronized response to user's comment] The synchronized response outputs: "The same opinion as me" ·yes "Impressive content" ·I see "Amazing content" ·amazing! Really? "Interesting content" ·funny ·n / a

[0086] <Embodiment 2> In the second embodiment, we will explain opinion reasoning, which is a mechanism for improving the quality of machine-based communication. Opinion reasoning is realized by combining judgment reasoning and response reasoning.

[0087] As explained in the first embodiment, the quality of communication can be improved by providing a mechanism in which the machine makes a statement when a response request is received from the user. In the first embodiment, the response request was determined by determining whether to respond to the sentence uttered by the user or a predicted sentence. On the other hand, in order to improve the quality of communication, when the conversation partner (user) is troubled or deep in thought, it is necessary to actively speak to them in some way to try to resolve the situation. It is also necessary to provide speech guidance to encourage the user to speak or take action during long periods of silence. Opinion inference aims to improve the quality of communication by having the machine spontaneously infer the appropriate timing, without being limited to the timing of the user's utterance.

[0088] In opinion inference, the instruction unit 23 sends an inference instruction including input information to the inference model 40. The answer processing unit 24 obtains an opinion sentence as an inference result and generates an answer sentence. Here, the input information includes image analysis results or situation information.

[0089] The image is captured by the camera of the user terminal 3, capturing the user's face, surrounding environment, etc. In this embodiment, the image is analyzed by the face analysis function of the user terminal 3 to obtain face information. The face information includes gaze, facial direction, facial expression, etc. The acquisition unit 21 can acquire the face information that is the result of analyzing the image and use it as input information. Note that the face analysis function may be provided in the language processing device 2, and the acquisition unit 21 may acquire the image from the user terminal 3 and analyze the acquired image using the face analysis function to obtain the face information.

[0090] The facial information may be used in the first embodiment. The instruction generation unit 22 can generate an inference instruction including a spoken sentence and facial information. This makes it possible to infer the user's emotions from their facial expressions and generate a response sentence or an agreement sentence, for example.

[0091] The acquisition unit 21 also acquires an image from the user terminal 3. The instruction generation unit 22 generates an inference instruction including the acquired image and instruction information for converting the appearance of the user in the image into text. The instruction unit 23 transmits the inference instruction to the inference device 4, and the response processing unit 24 can acquire the appearance information of the user as text data. The appearance information includes the user's gestures, surrounding environment, etc. that are not included in the face information.

[0092] The acquisition unit 21 acquires situation information. In this embodiment, the situation information includes the time elapsed since the call end detection event, the current date and time, location information, news information obtainable via a communication network, SNS (Social Networking Service) information, etc. The acquisition unit 21 can acquire the location information and SNS information via the user terminal 3.

[0093] 5 shows an overview of the flow of opinion inference. In this embodiment, the inference process is executed by an inference model 40. The instruction generation unit 22 generates inference instructions for causing the inference model 40 to appropriately execute the inference process.

[0094] The instruction generation unit 22 generates an inference instruction every predetermined waiting time. The waiting time is set to, for example, about 10 to 60 seconds, but is not limited to this. The waiting time may vary depending on the frequency of interactions. When it is determined that the predetermined waiting time has elapsed (YES in S301), the instruction generation unit 22 generates an inference instruction.

[0095] The instruction generation unit 22 generates an inference instruction related to a response request determination that includes at least facial information, and the instruction unit 23 transmits the inference instruction to the inference model 40. If there is no response request (NO in S302), the inference model 40 generates an inference result of false. When the response processing unit 24 obtains a false inference result, it does not generate a response sentence and transitions to a standby state again. If there is a response request (YES in S303), the inference device 4 performs a response inference related to the opinion sentence and generates the opinion sentence (S306).

[0096] The acquisition unit 21 acquires images from the user terminal 3 at any time (S303). The instruction generation unit 22 transmits an inference instruction for the image, including the acquired image and instruction information for converting the image into text, to the inference device 4. The inference device 4 generates state information for the image in accordance with the inference instruction (S304). The response processing unit 24 acquires the state information (S305).

[0097] The instruction generation unit 22 generates an inference instruction related to the opinion sentence including the state information acquired in S305 (S306). The instruction generation unit 22 may generate an inference instruction including, in addition to the state information, the time elapsed since the call end detection event and situation information. The acquisition unit 21 measures and acquires the time elapsed since the call end detection event.

[0098] In S306, the instruction generation unit 22 generates an inference instruction for generating an opinion sentence, including at least one piece of information selected from face information, appearance information, and situation information. The inference instruction for generating the opinion sentence may be set on the condition that a response request is present. The response request may be determined by referring to a relationship model with the user, which will be described later.

[0099] The inference instruction related to the opinion sentence can refer to the dialogue log up to that point and generate an opinion sentence that induces the user to speak about a topic related to the dialogue log. For example, when a predetermined amount of time has elapsed since the end-of-conversation detection event, the instruction generation unit 22 generates an inference instruction related to the opinion sentence that induces the user to speak. Furthermore, the instruction generation unit 22 can refer to news information and SNS information to generate an inference instruction related to the opinion sentence that induces the user to speak about a topic that interests the user.

[0100] <Embodiment 3> In the third embodiment, an index function will be described, which is a mechanism for providing high-quality machine-based communication with a high response speed. The index function is a function for generating an answer sentence by referring to an index in answer inference, and can be applied to the answer inference in the first and second embodiments.

[0101] In communication between humans, the relationship between the two people changes depending on the number, frequency, and content of communication. As the relationship changes, so does the communication between the two people. For example, as the relationship deepens, the content of communication desired by the interlocutor changes, such as bringing up topics that suit the other person's preferences and interests, or developing a conversation from a past topic into a deeper one. However, in human-machine communication, it has not been easy for a machine to recognize its relationship with a user or to recognize what changes in communication are required depending on that relationship. The index function is a mechanism that stores interactions, including past communications between the user and the machine, in a format that can be referenced as an index, and can be referenced whenever necessary, thereby enabling the machine to realize the changes in communication desired by the user with high response speed.

[0102] In this embodiment, the indexes are classified into long-term memory indexes and skill indexes. The long-term memory index is data that stores past interactions between a user and a machine as long-term memory for each user, and enables rapid access when it is necessary to refer to the corresponding long-term memory in a real-time dialogue between a user and a machine. The skill index is data that stores the relationship between a user and a machine, which changes depending on the past interactions between the user and the machine, as a skill for each user, and enables rapid and appropriate access by referring to an appropriate skill depending on the relationship in a real-time dialogue between the user and the machine.

[0103] The database DB includes a long-term memory index DB1 that stores long-term memory indexes, and a skill index DB2 that stores skill indexes. In this embodiment, the long-term memory index specifically indicates a dialogue memory function that the inference model 40 uses for answer inference. In this embodiment, the skill index specifically indicates a rule memory function that the inference model 40 uses for answer inference.

[0104] The index is linked to context information, which refers to the situation of the interaction between the user and the machine and the main context of the communication. The analysis results of the spoken sentences and images can be used to extract context information through inference using an inference model 40.

[0105] The inference model 40 can perform call inference according to the context information. The call inference can utilize the API (Application Programming Interface) of the inference device 4. The call inference is an inference regarding whether or not a function corresponding to the extracted context information needs to be called (which may be a true / false value). If the result of the call inference is true, the inference model 40 refers to an index in the database DB. By referring to the index corresponding to the context information, the inference model 40 can perform answer inference using long-term memory or skills. Note that the index reference may be performed by the language processing device 2.

[0106] The index generation unit 25 of the language processing device 2 generates an index based on various input information acquired by the acquisition unit 21 and stores it in the database DB. The inference device 4 executes call inference according to the context information included in the inference instruction, and can refer to the index stored in the database DB according to the inference result. The index generation unit 25 executes generation, update, and deletion of long-term memory indexes and skill indexes.

[0107] The index generation unit 25 associates context information corresponding to past interactions with a user with an index generated for each user and used to refer to the interaction, and stores the associated information in the database DB. Interactions include, but are not limited to, dialogues between a user and a machine. In this embodiment, context information indicates the context of a dialogue-related interaction, and can also be described as the central subject of a conversation or a characteristic subject of a conversation. The context information may include facial information, appearance information, and situation information of the user extracted from the image analysis results.

[0108] FIG. 6(a) shows an example of the data structure of a long-term memory index. The long-term memory index includes a memory ID, a memory name, a vector representation of the memory, an importance index for the memory, and metadata. The long-term memory index is linked to the user's identification information. The metadata includes context information, an emotion tag of the user in the interaction, and the creation date and time. The vector representation represents data obtained by converting the content of the conversation related to the memory into a high-dimensional vector. Here, the context information can be considered a summary of the content of the conversation.

[0109] In this embodiment, the importance index is a numerical index that particularly represents the importance of a user's emotion. A user's emotion is indicated by an emotion tag and is classified into joy, anger, sadness, happiness, surprise, admiration, etc. The importance index is determined by the magnitude of the user's emotion indicated by the emotion tag. The importance index determines the priority to be referenced as an index, and is referenced in retrieval inference in order from the long-term memory index that is considered "high."

[0110] In the long-term memory index, either vector representation or context information is referenced as context information. In the call inference, the long-term memory index DB1 is referenced using the context information extracted from the utterance sentence or the vector representation obtained by vectorizing the utterance sentence, and the long-term memory index of the corresponding memory ID is obtained as the inference result. For example, by referencing the long-term memory index associated with "memory ID: M001" in Figure 6(a) from the context information related to movies, the inference model 40 can infer that the user has a high interest in movies, and generate an answer sentence based on that information.

[0111] FIG. 6(b) shows an example of the data structure of a skill index. The skill index includes a skill ID, a skill name, a skill description, a relevance index, an execution code, and a usage history. The skill index is linked to a user's identification information. The execution code corresponds to a specific interaction and executes the corresponding skill. The execution code may include supplemental instructions to execute the skill according to the user's preferences, situation, and mood. In the skill index, context information is extracted by analyzing the skill name or skill description. Here, the analysis may be performed by inference using an inference model 40.

[0112] In this embodiment, the relevance index is an index that quantifies the relevance between a user's interaction and a skill in the range of 0 to 1. Skill SK001 in FIG. 6(b) is a skill related to how to make coffee, and the relevance index associates a numerical value of the relevance with the skill to the interaction that the user likes coffee. Skill SK002 in FIG. 6(b) is a skill related to recommending relaxation music, and the relevance index associates a high numerical value of the relevance with the interaction that the user desires to relax. When a certain user interaction is recognized, the skill index can refer to the execution code of a skill that is highly relevant to the interaction, thereby providing the skill desired by the user. The execution code may search for music that matches the user's mood via the Internet as needed.

[0113] The execution code is used to infer a response sentence on a specific topic. The execution code may also include control data for the user terminal 3. When the user terminal 3 is configured as a robotic device, the control data may include data for controlling the actuators of its arms and movement mechanism. According to the control data, the robotic device can not only provide instruction on how to make coffee, for example, but also make coffee on behalf of the user.

[0114] In this embodiment, the database DB further includes a user profile DB 3. The user profile DB 3 stores user profiles such as user preferences, interests, and behavioral patterns.

[0115] The user profile also includes a relationship model that models the relationship between the user and the machine. The relationship model is constructed in response to changes in the user's behavior, choices, emotional expressions, etc. during the interaction between the user and the machine. By referencing the relationship model in response inference, the inference model 40 can adjust the form, frequency, and wording of communication and generate a response sentence.

[0116] The relationship model is updated over time and adjusted to deepen the relationship through long-term interactions between users and machines. The relationship model is not limited to relationships between machines and people, but can also build relationships with animals, plants, objects, spaces, and organizations to which users belong.

[0117] The relationship model is used to evaluate the relationship with a user by integrating features extracted from interactions with the user and a relationship score. Features include the user's emotions, speech content, and tone of voice contained in the interaction. The relationship score is calculated from the intimacy, trust, and frequency of interaction between the user and the machine. For example, answer inference can adjust greeting expressions depending on intimacy, or provide more specialized advice depending on trust. The relationship model defines how each element of the model is reflected in the inference in inference instructions and call functions.

[0118] The inference instruction may include an instruction to refer to a user profile together with the skill index. For example, the relationship index of the skill index may be preferentially referenced to an index corresponding to the user's preferences or interests in the user profile.

[0119] 7 shows a process flowchart for generating an index. The acquisition unit 21 acquires various input information from the user terminal 3 (S401). The index generation unit 25 extracts context information included in the acquired input information (S402). The context information may be extracted by inference using the inference model 40. The index generation unit 25 generates an index based on the input information including the extracted context information, and stores the index in the database DB (S403).

[0120] In S403, if the index generation unit 25 can extract an emotion tag in addition to the context information, it generates a long-term memory index based on the input information and stores it in the long-term memory index DB 1. The index generation unit 25 inputs the input information to the inference model 40 and can generate a vector representation and an emotion tag through the inference. The index generation unit 25 may also cause the inference model 40 to infer an importance index for the generated emotion tag.

[0121] The index generating unit 25 classifies the generated vector representations using a k-nearest neighbor method or a nearest neighbor search algorithm, thereby constructing an index with improved search efficiency.

[0122] In S403, if the index generation unit 25 can extract information about the skill content in addition to the context information, it generates a skill index based on the input information and stores it in the skill index DB 2. The index generation unit 25 inputs the input information to the inference model 40 and can generate a skill name, skill description, and corresponding execution code from the skill content through inference. At this time, the index generation unit 25 can further refer to the user profile in the inference model 40 to make inferences including relevance indices according to the user's preferences, hobbies, etc.

[0123] When generating a new skill index, the index generation unit 25 refers to existing skill indexes. The index generation unit 25 can generate a new skill index by combining existing skill indexes. At this time, the skill content can be improved by referring to the success rate from the usage history of the skill index to be combined.

[0124] FIG. 8 shows a processing flowchart related to the use of indexes. The acquisition unit 21 acquires various input information from the user terminal 3 (S501). The instruction generation unit 22 generates an inference instruction related to answer inference including the input information, and the instruction unit 23 sends the inference instruction to the inference device 4 (S502). The inference device 4 performs call inference to determine whether the input information includes context information corresponding to the index. The inference device 4 can generate an answer sentence by referencing the corresponding index through call inference. The call inference can also refer to the user profile.

[0125] <Embodiment 4> In the fourth embodiment, an example of applying parallel inference to a translation function will be described. The translation function refers to a function for translating a first language into a second language. In conventional translation functions, translation processing is performed after a series of spoken sentences are input, resulting in a certain waiting time before the user receives the translated sentences. The translation function using the parallel inference mechanism according to the fourth embodiment aims to reduce the waiting time for the user by performing processing in parallel with the user's speech input.

[0126] FIG. 9 shows a processing flowchart for the translation function. First, the acquisition unit 21 acquires a designated language (S601). The designated language includes a first language, which is a language from which the text is to be translated, and a second language, which is a language to which the text is to be translated. It is also possible to set multiple second languages. The designated language only needs to be set at least once at the beginning, and may also be set in advance.

[0127] The acquisition unit 21 repeatedly acquires utterance sentences at predetermined time intervals (S602). The utterance sentences acquired until the end-of-speech detection event is acquired are updated as needed. Here, the utterance sentences at each predetermined time interval are divided sentences, and when a subsequent utterance sentence is acquired at the next predetermined time interval, it is combined with the previous utterance sentence to update the utterance sentences.

[0128] The instruction generating unit 22 generates an inference instruction including an utterance sentence for each predetermined time period and instruction information for judgment inference related to the prediction accuracy of the utterance sentence, and the instruction unit 23 transmits the inference instruction to the inference device 4 (S603). The inference instruction is repeatedly transmitted until a speech end detection event is acquired or until it is determined that the prediction accuracy is high.

[0129] The inference device 4 determines the prediction accuracy of the utterance sentence included in the inference instruction (S604). If the prediction accuracy is low (NO in S604), the process returns to S602 and executes the process of acquiring the next utterance sentence at the next predetermined time.

[0130] If the prediction accuracy is high (YES in S604), the inference device 4 generates a predicted sentence (S605). The response processing unit 24 acquires the predicted sentence generated by the inference device 4 and generates a translation by translating the predicted sentence into a second language (S606). The generated translation is sent to the user terminal 3, thereby providing the user with a translation of the spoken sentence. The generation of the translation may be executed by an external device (not shown) and sent to the user terminal 3 via the external device. Alternatively, the generation of the translation may be executed by the inference device 4.

[0131] After acquiring the end-of-speech detection event, the language processing device 2 may execute a difference determination between the finalized utterance sentence and the predicted sentence. The difference determination is the same as in the first embodiment, and if the predicted sentence differs from the finalized utterance sentence, a translation obtained by translating the finalized utterance sentence is output to the user terminal 3. If the predicted sentence does not differ from the finalized utterance sentence, a translation obtained by translating the predicted sentence is output to the user terminal 3.

[0132] As shown in embodiment 4, by performing translation processing based on predicted sentences, it is possible to improve response speed compared to when translation processing is performed after a series of utterance sentences are finalized. Furthermore, the translation function can be implemented simultaneously with parallel inference and other embodiments.

[0133] The above-described first to fourth embodiments can be implemented independently of each other, or can be implemented in combination.

[0134] In this embodiment, the inference elements executed by the inference model 40 are defined as judgment inferences and answer inferences. A judgment inference is an element that determines whether a subsequent inference should be executed for the information input and whether it is true or false. An answer inference is an element that generates text when the judgment inference is judged to be at least true. These inference elements are incorporated as instruction information in natural language in the inference instructions (prompts) of the inference model 40. The instruction generation unit 22 generates an inference instruction that includes an inference element based on the input information and instruction information corresponding to the inference element.

[0135] The instructing unit 23 outputs the generated inference instruction to the inference model 40. In this embodiment, the instructing unit 23 outputs the inference instruction by either the following procedure (a) or (b). (a) Outputting inference instructions including judgment inferences and answer inferences. (b) After outputting an inference instruction including a judgment inference, output an inference instruction including an answer inference.

[0136] The inference elements further include predictive inference that predicts predicted content, which is input information that follows the input information in chronological order. In this embodiment, the predicted content indicates a predicted sentence, but is not limited to this and may include, for example, predicted content of changes over time in facial information, appearance information, and situation information. The instruction unit 23 further outputs an inference instruction to the large-scale language model using any of the procedures (c) to (e). (c) Outputting inference instructions including judgment inference, predictive inference, and answer inference. (d) After outputting an inference instruction including a judgment inference and a predictive inference, output an inference instruction including a response inference. (e) After outputting an inference instruction including a judgment inference, output an inference instruction including a predictive inference and a response inference.

[0137] The inference model 40 according to this embodiment employs a large-scale language model, allowing multiple inferences to be performed simultaneously. Note that in this embodiment, judgment inference, predictive inference, and response inference may each be performed using a different inference model. The inference model 40 may be a single model or a combination of multiple models, and is not limited by the model configuration. [Explanation of symbols]

[0138] 1 Language Processing System 2 Language Processing Unit 21 Acquisition Department 22 Instruction generation section 23 Instruction section 24 Response processing section 25 Index Generation Unit 3. User terminal 31 Voice Recognition Unit 32 Audio processing unit 33 Image acquisition unit 34 Image processing section 35 Text Processing Unit 36 Output section 4 Reasoning device 40 Inference Model

Claims

1. A language processing system that processes utterances by a user and responses by a machine, outputting repeated inference instructions based on user utterances updated at predetermined intervals; A predicted sentence following the uttered sentence and its prediction accuracy are generated according to the inference instruction; A language processing system that outputs a response sentence to the predicted sentence that is determined to have a high prediction accuracy.

2. The language processing system according to claim 1 , wherein the inference instructions are output to an inference model.

3. If the prediction accuracy is high, output a response sentence to the predicted sentence; 3. The language processing system according to claim 1, wherein when the prediction accuracy is low, an inference instruction including an utterance sentence updated at a subsequent predetermined time is output, and an agreement sentence with the predicted sentence is output.

4. determining whether the predicted sentence is a response request from a user; If the response request is received, a response sentence is output in response to the predicted sentence; 3. The language processing system according to claim 1, wherein if the predicted sentence is not a response request, a corresponding sentence is output.

5. When the end of the user's utterance is detected, the series of utterances is output. determining whether the predicted sentence and the series of spoken sentences are different; If there is a difference, output a response sentence to the series of utterance sentences; 3. The language processing system according to claim 1, wherein if there is no difference, a response sentence or a concurring sentence to the predicted sentence is output.

6. outputting an inference instruction as to whether the facial information obtained as a result of analyzing the image of the user corresponds to the response request; 3. The language processing system according to claim 1, wherein if the response is a request, a spontaneous opinion sentence is output.

7. 7. The language processing system according to claim 6, wherein, if the response is not a response request, an inference instruction relating to the response request is output after a predetermined waiting time has elapsed.

8. outputting an inference instruction including information about the user's state, which is the result of analyzing the image; The language processing system according to claim 6, wherein the system outputs spontaneous opinion sentences.

9. When the end of the user's speech is detected, the elapsed time is measured, 3. The language processing system according to claim 1, wherein a spontaneous opinion sentence is output in accordance with the elapsed time.

10. 3. The language processing system according to claim 1, wherein a translation of the predicted sentence determined to have a high prediction accuracy is output.

11. outputting a translation instruction including the predicted sentence determined to have a high prediction accuracy and a specified language; The language processing system according to claim 10 , wherein a translation of the predicted sentence is generated in accordance with the translation instruction, and the translation is output.

12. storing in a database context information corresponding to past interactions with the user and an index personalized for the user and used to reference the interactions; When the context information included in the predicted sentence is extracted, the corresponding index stored in the database is referenced; 3. The language processing system according to claim 1, wherein an answer sentence is output in accordance with an inference instruction including the index and the predicted sentence.

13. the database stores, as a long-term memory index relating to a memory, context information, a vector representation representing the memory, an importance index of the memory, and an emotion tag of the user in the interaction, in association with each other; When the context information included in the predicted sentence is extracted, the long-term memory index is referenced; The language processing system according to claim 12 , wherein an answer sentence is output in accordance with an inference instruction including the long-term memory index and the predicted sentence.

14. the database stores, as a skill index relating to a skill, context information, a content of the skill, a relevance index, and an execution code relating to a predetermined interaction in association with each other; When the context information included in the predicted sentence is extracted, the skill index is referenced, The language processing system according to claim 12 , wherein an answer sentence is output in accordance with an inference instruction including the skill index and the predicted sentence.

15. The inference instruction outputs false and an agreement statement when the prediction accuracy is low; If the prediction accuracy is high, output true and the predicted sentence.

3. The language processing system according to claim 1, wherein the output is repeated until the true value is output.

16. A language processing system that processes utterances by a user and responses by a machine, outputting an inference instruction as to whether a face image or state information obtained as a result of analyzing an image of the user corresponds to a response request; If the response request is received, the language processing system outputs a spontaneous opinion sentence.

17. A language processing method for processing utterances by a user and responses by a machine, outputting repeated inference instructions based on user utterances updated at predetermined intervals; A predicted sentence following the uttered sentence and its prediction accuracy are generated according to the inference instruction; A language processing method in which a computer executes each process to output a response sentence to the predicted sentence that is determined to have a high prediction accuracy.

18. A language processing program that causes a computer to function so as to execute the language processing method according to claim 17.

Citation Information

Patent Citations

  • Voice interactive method, voice interactive device, and voice interactive program

    JP2017107078A

  • Terminal device, information processing method, and information processing program

    JP2023129020A