Information Processing Apparatus, Information Processing Method, and Program

The information processing apparatus enhances AI agent call efficiency by determining and outputting response information based on utterance information from the target person, addressing the issue of response delay and improving call efficiency.

JP7689787B1Active Publication Date: 2025-06-09RECHO CO LTD

Patent Information

Application Number
JP2025046733
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-09
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Existing AI agent call systems do not adequately improve call efficiency, particularly in terms of reducing response delay to the target person.

Method used

An information processing apparatus that acquires utterance information from a target person, determines first and second response information using a large language model, and outputs the responses accordingly to enhance call efficiency.

Benefits of technology

The solution effectively reduces the perceived latency in AI agent responses, thereby improving the overall efficiency of calls with AI agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689787000001_ABST
    Figure 0007689787000001_ABST
Patent Text Reader

Abstract

To improve the efficiency of calls by an AI agent. 【Solution means】An acquisition unit 100 that acquires utterance information regarding the utterance of the target person, where the utterance information includes first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; a first response determination unit 102a that determines first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; a second response determination unit 102b that determines second response information regarding another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person when a predetermined condition regarding the utterance is satisfied; and an output unit 106 that outputs the first response information when the predetermined condition is not satisfied and outputs at least one of the first response information and the second response information when the predetermined condition is satisfied. An information processing apparatus 2 is provided with these components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, a technique using speech recognition for calls has been known. For example, Patent Document 1 describes a system in which a generative AI (Artificial Intelligence) responds when receiving a call from an unregistered or unannounced phone number, and after the call ends, the matter is documented and transmitted, and provides a technique for providing means for confirming family members and acquaintances and means for police cooperation with numbers having a history of abuse.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, with the technique described in Patent Document 1, it is not possible to sufficiently improve the efficiency of calls by an AI agent. For example, there is room for consideration regarding suppressing the delay in the response to the target person.

[0005] The present disclosure provides an information processing apparatus, an information processing method, and a program capable of improving the efficiency of calls by an AI agent.

Means for Solving the Problems

[0006] An information processing apparatus according to one aspect of the present disclosure includes an acquisition unit that acquires utterance information regarding an utterance of a target person, where the utterance information includes first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; a first response determination unit that determines first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; a second response determination unit that, when a predetermined condition regarding the utterance is satisfied, determines second response information regarding another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person; and an output unit that outputs the first response information when the predetermined condition is not satisfied and outputs at least one of the first response information and the second response information when the predetermined condition is satisfied.

[0007] An information processing method according to another aspect of the present disclosure includes: an information processing apparatus acquiring utterance information regarding an utterance of a target person, where the utterance information includes first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; determining first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; when a predetermined condition regarding the utterance is satisfied, determining second response information regarding another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person; and outputting the first response information when the predetermined condition is not satisfied and outputting at least one of the first response information and the second response information when the predetermined condition is satisfied.

[0008] A program according to another aspect of the present disclosure causes an information processing apparatus to acquire utterance information regarding an utterance of a target person, where the utterance information includes first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; determine first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; when a predetermined condition regarding the utterance is satisfied, input a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person, to determine second response information regarding another response to at least a part of the utterance; output the first response information when the predetermined condition is not satisfied, and output at least one of the first response information and the second response information when the predetermined condition is satisfied.

Effect of the Invention

[0009] According to the present disclosure, it is possible to improve the efficiency of a call by an AI agent.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Mode for Carrying Out the Invention

[0011] 1 Summary In the present embodiment, it is assumed that the subject makes a call with an AI agent using his / her own terminal device. That is, the terminal device of the subject generates voice information by receiving the speech of the subject via a voice input device (e.g., a microphone, etc.), transmits the voice information to an information processing device (e.g., a server device, etc.), and acquires response information regarding the response by the AI agent. This response information is output as voice at the terminal device. As a result, the subject can feel as if he / she is having a conversation with the AI agent via the terminal device. One of the problems that the system 1 according to the present embodiment (hereinafter simply referred to as "system 1") attempts to solve is to improve the efficiency of the call with the AI agent in such a situation.

[0012] With reference to FIG. 1, the outline of the system 1 will be described. The system 1 includes a terminal device 3, an information processing device 2, and an LLM server device 4. The terminal device 3 is a device used by the subject. The information processing device 2 is a device that executes at least a part of the processing related to improving the efficiency of the call by the AI agent. The LLM server device 4 is a device that provides services based on a large language model (LLM: Large Language Model, hereinafter referred to as "LLM").

[0013] The terminal device 3 generates first audio information by receiving the speech of the target person via the audio input device, and transmits it to the information processing device 2 (S1). The information processing device 2 acquires first speech information based on the first audio information. The first speech information may include text data obtained by speech recognition of the first audio information, or may include the first audio information itself (for example, audio data, etc.).

[0014] Next, the information processing device 2 transmits a first response determination instruction including the first speech information to the LLM server device 4 (S2). The first response determination instruction may include an instruction for determining first response information regarding a response to at least a part of the speech of the target person. The LLM server device 4 generates first response information based on the first response determination instruction, and transmits the first response information to the information processing device 2 (S3).

[0015] Note that the information processing device 2 does not output the first response information to the target person immediately after step S3. The information processing device 2 waits for the input of additional audio information from the terminal device 3 while holding the first response information as a provisional response. Then, when a predetermined condition regarding the speech of the target person is not satisfied (in one example, when it is determined that there is no input of additional audio information and the speech is not continued), the information processing device 2 outputs the first response information (not shown). On the other hand, regarding the case where a predetermined condition regarding the speech of the target person is satisfied (in one example, when it is determined that there is input of additional audio information and the speech is continued), it will be described below.

[0016] The terminal device 3 further generates second audio information by receiving the speech of the target person via the audio input device, and transmits it to the information processing device 2 (S4). The information processing device 2 acquires second speech information based on the second audio information. Similar to the first speech information, the second speech information may include text data obtained by speech recognition of the second audio information, or may include the second audio information itself (for example, audio data, etc.).

[0017] Next, the information processing device 2 transmits a second response determination instruction including the second utterance information, the already acquired first utterance information, and / or the first response information held as a provisional response to the LLM server device 4 (S5). The second response determination instruction may include an instruction for determining second response information regarding a response to at least a part of the utterance of the target person. The LLM server device 4 generates second response information based on the second response determination instruction and transmits the second response information to the information processing device 2 (S6). The second response information can also be information obtained by updating the first response information, which is a candidate for a provisional response, based on the second utterance information corresponding to additional input by the target person.

[0018] Next, the information processing device 2 outputs the second response information (S7). The information processing device 2 may output the second response information after determining that there is no additional utterance by the target person.

[0019] According to the system 1, it is possible to suppress the time interval that is felt as the latency of the response from the perspective of the target person. The system 1 acquires first response information, which is a provisional response, based on the first utterance information regarding the utterance of the target person (see S3), and then outputs the first response information when a predetermined condition regarding the utterance is not satisfied, and acquires second response information based on the first utterance information and / or the first response information when the predetermined condition is satisfied (see S6). According to such a configuration, in one example, the response information is continuously updated while the user's utterance continues, and when the utterance ends, the most recently updated response information is promptly output. That is, for example, compared with the prior art in which response generation starts after the utterance by the target person ends, it is possible to suppress the latency felt by the target person. As a result, the call by the AI agent is made more efficient.

[0020] In this embodiment, for the sake of convenience of explanation, terms such as "update" and "provisional" may be used, but these terms do not limit the actual processing by the computer.

[0021] Also, hereinafter, when the first audio information and the second audio information are not particularly distinguished, or when they are collectively referred to, they are referred to as "audio information". Similarly, when the first utterance information and the second utterance information are not particularly distinguished, or when they are collectively referred to, they are referred to as "utterance information". Similarly, when the first response information and the second response information are not particularly distinguished, or when they are collectively referred to, they are referred to as "response information". Similarly, when the first response decision instruction and the second response decision instruction are not particularly distinguished, or when they are collectively referred to, they are referred to as "response decision instruction".

[0022] Hereinafter, with reference to FIGS. 2 to 9, the detailed aspects of the system 1 will be exemplarily described.

[0023] 2 Functional Configuration With reference to FIG. 2, the functional configuration of the system 1 of the present embodiment will be described. The system 1 includes an information processing device 2, a terminal device 3, an LLM server device 4, and a communication network 5. The information processing device 2, the terminal device 3, and the LLM server device 4 are configured to be communicable via the communication network 5.

[0024] 2.1 Information Processing Device 2 The information processing device 2 executes at least a part of the processing related to improving the efficiency of calls by the AI agent. In one embodiment, the information processing device 2 is a server device when the terminal device 3 is a client device. In one embodiment, the information processing device 2 is a cloud server device. Note that the information processing device 2 may be a device including, for example, one or more virtual or physical web server devices and one or more virtual or physical database server devices.

[0025] The information processing device 2 includes a control unit 10, a storage unit 12, a network interface unit 14, and a bus 16. The control unit 10, the storage unit 12, and the network interface unit 14 are electrically connected via the bus 16.

[0026] 2.1.1 Control Unit 10 By executing various programs stored in the storage unit 12 described later, the control unit 10 can function as an acquisition unit 100, a determination unit 102, a determination unit 104, and an output unit 106.

[0027] 2.1.1.1 Acquisition Unit 100 The acquisition unit 100 acquires speech information regarding the speech of the target person. The speech information includes first speech information regarding the first speech and second speech information regarding the second speech after the first speech.

[0028] In one embodiment, the speech information may include text data obtained by transcribing the speech of the target person based on the voice information received by the information processing device 2 from the terminal device 3. That is, the first speech information may include text data obtained by transcribing the first speech part of the speech of the target person, and the second speech information may include text data obtained by transcribing the second speech part after the first speech of the speech of the target person. The acquisition unit 100 can transcribe the voice information by a voice recognition program selected based on the knowledge of those skilled in the art.

[0029] In one example, when the speech of the target person is "Umm, I want to make a reservation for two at 7 o'clock tonight... Ah, my name is ○○.", the voice information can include voice data corresponding to this speech, the first speech information can include text data "Umm, I want to make a reservation for two at 7 o'clock tonight...", and the second speech information can include text data "Ah, my name is ○○.".

[0030] Note that, among the utterances of the target person, where to set the first utterance and where to start the second utterance can be determined based on the determination result of the determination unit 104 described later. The determination unit 104 determines whether a break has occurred in the utterance of the target person. The acquisition unit 100 can acquire the first utterance information and the second utterance information by setting at least a part of the utterance of the target person up to the point where a break is determined to have occurred as the first utterance, and at least a part of the utterance after the point where a break is determined to have occurred as the second utterance. In the above, the case where a break is determined to have occurred between "Well, I would like to make a reservation for two at 7:00 tonight" (corresponding to the first utterance information) and "Ah, my name is ○○." (corresponding to the second utterance information) is illustrated.

[0031] In one embodiment, the utterance information may include the voice information itself. That is, the first utterance information may include the voice information itself of the part of the first utterance among the utterances of the target person, and the second utterance information may include the voice information itself of the part of the second utterance after the first utterance among the utterances of the target person.

[0032] 2.1.1.2 Decision Unit 102 The decision unit 102 includes a first response decision unit 102a, a second response decision unit 102b, and an echo decision unit 102c.

[0033] 2.1.1.2.1 First Response Decision Unit 102a The first response decision unit 102a determines first response information regarding a response to at least a part of the utterance of the target person by inputting a first response decision instruction including the first utterance information to the LLM.

[0034] In one embodiment, inputting the first response decision instruction to the LLM may include transmitting an HTTP request including the first response decision instruction to the LLM server device 4. Determining the first response information may include obtaining an HTTP response to the HTTP request.

[0035] In one embodiment, the first response determination instruction may include, in addition to the first utterance information, for example, a system prompt and the conversation history between the target person and the AI agent (that is, information regarding the utterance of the target person before determining the first response information and information regarding the response by the information processing apparatus 2 thereto), and the like.

[0036] In one embodiment, when the determination unit 104 (described later) determines that a first delimiter has occurred in the utterance of the target person, the first response determination unit 102a inputs the first response determination instruction to the LLM. In one example, when the utterance of the target person is "Well, I would like to make a reservation for two at 7:00 p.m. today... Ah, the name is ○○.", the determination unit 104 may determine that a first delimiter has occurred between "I would like to make a reservation for two at 7:00 p.m. today" and "Ah, the name is ○○." The first response determination unit 102a may transmit the first response determination instruction to the LLM server device 4 at the time when it is determined that the first delimiter has occurred, and may determine the first response information.

[0037] In one embodiment, when determining the first response information, the first response determination unit 102a may further determine information for outputting the first response information by voice. The information for outputting the response information by voice may be generated based on a voice generation program that can be arbitrarily selected by those skilled in the art. Note that "when determining the response information" may be immediately before determining the response information, may be parallel to the process of determining the response information, may be performed integrally and / or continuously with the process of determining the response information, or may be immediately after the process of determining the response information.

[0038] In one embodiment, the storage unit 12 stores information for outputting each of one or more texts by voice. When the first response information matches at least one of the one or more texts, determining information for outputting the first response information by voice includes obtaining, from the storage unit 12, information for outputting the matching text by voice. When the first response information does not match any of the one or more texts, determining information for outputting the first response information by voice includes generating, based on a predetermined voice generation program, information for outputting the first response information by voice. That is, when the response information typically includes "common" words and / or phrases, the first response determination unit 102a can determine the voices of the words and / or phrases prepared in advance. In one example, when the storage unit 12 stores information for outputting by voice the text "Your reservation has been received", if the text "Your reservation has been received" is included in the first response information, the information for outputting that part of the first response information by voice is obtained from the storage unit 12. In another example, when the storage unit 12 stores information for outputting by voice the text "Your reservation has been received" as described above, if the text "Your reservation has been received" is not included in the first response information (that is, when information for outputting the corresponding voice from the storage unit 12 cannot be obtained), the first response determination unit 102a generates, by a predetermined voice generation program, information for outputting "Your reservation has been received" by voice. According to this configuration, since a part of the text included in the first response information can be output by a prepared voice, the latency felt by the subject can be suppressed.

[0039] 2.1.1.2.2 Second response determination unit 102b When a predetermined condition regarding the speech of the subject is satisfied (in one example, when it is determined that the speech of the subject continues), the second response determination unit 102b inputs a second response determination instruction including the first speech information and / or the first response information and the second speech information to the LLM before outputting the first response information to the subject, thereby determining second response information regarding other responses to at least a part of the speech.

[0040] In one embodiment, inputting the second response determination instruction into the LLM may include transmitting an HTTP request including the second response determination instruction to the LLM server device 4, similar to inputting the first response determination instruction into the LLM. Determining the second response information may include obtaining an HTTP response to the HTTP request.

[0041] In one embodiment, similar to the first response determination instruction, the second response determination instruction may include, in addition to the first response information and the second utterance information, for example, a system prompt, and a conversation history between the subject and the AI agent (that is, information regarding the subject's utterance before determining the second response information and information regarding the response by the information processing device 2 thereto).

[0042] As illustrated with reference to FIG. 1, the first response information may be a tentative response. When a predetermined condition regarding the subject's utterance is satisfied, the second response determination unit 102b may determine the second response information before outputting the first response information to the subject (in one example, without outputting such a tentative response). In one embodiment, the period before outputting the first response information to the subject may be the period before the first response information is output by voice or the like in the terminal device 3.

[0043] In one embodiment, when determining the second response information, the second response determination unit 102b may further determine information for outputting the second response information by voice.

[0044] In one embodiment, the storage unit 12 stores information for outputting each of one or more texts by voice, and when the second response information matches at least one of the one or more texts, determining information for outputting the second response information by voice includes obtaining, from the storage unit 12, information for outputting the matching text by voice; and when the second response information does not match any of the one or more texts, determining information for outputting the second response information by voice includes generating, based on a predetermined voice generation program, information for outputting the second response information by voice.

[0045] In one embodiment, when the determination unit 104 determines that a second delimiter has occurred in the utterance, the second response determination unit 102b inputs a second response determination instruction to the LLM. In one example, when the utterance of the subject is "Umm, I would like to make a reservation for two at 7:00 pm today... Ah, the name is ○○.", the determination unit 104 determines that a first delimiter has occurred between "I would like to make a reservation for two at 7:00 pm today" and "Ah, the name is ○○" as described above, and may further determine that a second delimiter has occurred after "Ah, the name is ○○". The second response determination unit 102b may transmit a second response determination instruction to the LLM server device 4 when it is determined that a second delimiter has occurred, and determine the second response information. Note that after obtaining the first utterance information, if a predetermined condition regarding the utterance of the subject is not satisfied (in one example, when it is determined that the utterance of the subject is not continued), the first response information may be output, so the second response determination unit 102b does not have to determine the second response information.

[0046] 2.1.1.2.3 Backchannel Determination Unit 102c When determining the first response information, the backchannel determination unit 102c determines backchannel information regarding the backchannel corresponding to the first utterance information. Note that "when determining the first response information" may be before determining the first response information, during the execution of the process of determining the first response information, or after determining the first response information.

[0047] In one embodiment, the response determination unit 102c may determine response information based on whether the first utterance information includes information regarding a positive utterance of the target person or information regarding a negative utterance. In one example, when the first utterance information includes information regarding a positive utterance such as "Yes, it is possible", the response determination unit 102c may determine response information including text data such as "Thank you". In another example, when the first utterance information includes information regarding a negative utterance such as "I'm sorry", the response determination unit 102c may determine response information including text data such as "Understood". In this case, the response determination unit 102c may refer to, for example, a list of words that may be included in a positive utterance and a list of words that may be included in a negative utterance, and determine response information corresponding to any of the lists.

[0048] In one embodiment, the response determination unit 102c may determine response information based on the form of the question included in the first utterance information. In one example, when the first utterance information includes a question that can be answered with YES / NO such as "May I ~?" or a question that can be answered in a closed form such as "Which is better, A or B?", the response determination unit 102c may determine response information including text data such as "Thank you for your confirmation". In another example, when the first utterance information includes a question that can be answered in an open form such as "Please say it again", the response determination unit 102c may determine response information including text data such as "Understood". In this case, the response determination unit 102c may determine the form of the question included in the first utterance information based on a natural language processing algorithm selected based on the knowledge of those skilled in the art, and determine response information based on the determination result.

[0049] In one embodiment, when determining the response information, the response determination unit 102c may further determine information for outputting the response information by voice.

[0050] In one embodiment, the memory unit 12 stores information for outputting each of one or more texts by voice. When the response information matches at least one of the one or more texts, determining information for outputting the response information by voice includes obtaining, from the memory unit 12, information for outputting the matching text by voice. When the response information does not match any of the one or more texts, determining information for outputting the response information by voice includes generating information for outputting the response information by voice based on a predetermined voice generation program.

[0051] In one embodiment, the response determination unit 102c estimates the emotion in the speech of the target person based on the first speech information and / or the voice of the speech corresponding to the first speech information, and determines the response information based on the estimated emotion. As an example, assume that the speech corresponding to the first speech information is the voice "Why can't this be done?" At this time, when it is estimated based on the voice that the target person is angry, the response determination unit 102c may determine a response of "I'm very sorry." On the other hand, when it is estimated based on the voice that the target person is not angry, the response determination unit 102c may determine a response of "Thank you for your question." According to this configuration, a more natural response can be output for the target person. Note that the estimation of the emotion in the speech of the target person based on the first speech information and / or the voice of the speech corresponding to the first speech information can be executed based on an emotion estimation algorithm that can be arbitrarily selected by those skilled in the art.

[0052] In one embodiment, the response determination unit 102c determines the response information by inputting a response determination instruction including the first speech information to the LLM.

[0053] In one embodiment, inputting the turn-taking decision instruction into the LLM may include transmitting an HTTP request including the turn-taking decision instruction to the LLM server device 4, similar to inputting the response decision instruction into the LLM. Determining the turn-taking information may include obtaining an HTTP response to the HTTP request.

[0054] In one embodiment, the turn-taking decision instruction may include, in addition to the first utterance information, for example, a system prompt or the like, similar to the response decision instruction.

[0055] In one embodiment, when determining the second response information, the turn-taking decision unit 102c determines other turn-taking information regarding other turn-takings corresponding to the second utterance information. At this time, the turn-taking decision unit 102c may determine the other turn-taking information by inputting a turn-taking decision instruction including the second utterance information into the LLM.

[0056] 2.1.1.3 Determination Unit 104 The determination unit 104 determines whether a first delimiter has occurred in the utterance. The first utterance information includes information regarding the portion of the target person's utterance up to the first delimiter.

[0057] In one embodiment, the determination unit 104 further determines whether a second delimiter has occurred after the first delimiter of the utterance. The second utterance information includes information regarding the portion of the target person's utterance up to the second delimiter.

[0058] In one embodiment, the determination unit 104 determines whether a delimiter has occurred in the utterance based on the continuity of the voice of the target person's utterance. In one example, the determination unit 104 sequentially determines whether the target person is speaking based on each of a plurality of voice chunks included in the voice information (for example, data obtained by dividing the voice information every 0.1 seconds), and may determine that a delimiter has occurred when it is determined that the target person is not speaking in a continuous predetermined number of voice chunks.

[0059] In one embodiment, the determination unit 104 determines whether a break has occurred in the speech based on the semantic continuity of the speech of the target person. In one example, when the speech information includes text data obtained by speech-to-text conversion of voice information, the determination unit 104 may divide the text data into a part related to a first theme and a part related to a second theme, and determine that a break has occurred between them. For example, when the speech information includes text data such as "I would like to make a reservation for four people at 7 pm tomorrow. Can the all-you-can-drink option be added to the course?", the determination unit 104 divides it into a part related to the first theme "I would like to make a reservation for four people at 7 pm tomorrow. (In this example, a question about the availability of the reservation)" and a part related to the second theme "Can the all-you-can-drink option be added to the course? (In this example, a question about the content of the course)", and may determine that a break has occurred between them.

[0060] In one embodiment, the determination unit 104 further determines whether a predetermined time related to the output of the response has elapsed after at least one of the first break and the second break of the speech. In one embodiment, the predetermined time is determined based on the feedback information. In one example, the predetermined time is determined based on the playback time when the feedback information is output by voice. For example, if the feedback information includes text data such as "Understood" and it takes 0.8 seconds to output this by voice, the predetermined time may be 0.8 seconds (or the number of seconds obtained by adding about 0.1 seconds as a buffer to this).

[0061] 2.1.1.4 Output Unit 106 When a predetermined condition regarding the speech of the target person is not satisfied (for example, when the speech continues), the output unit 106 outputs first response information, and when the predetermined condition is satisfied (for example, when the speech ends), the output unit 106 outputs at least one of the first response information and the second response information. Outputting the response information includes transmitting information for outputting the response information to the terminal device 3.

[0062] In one embodiment, the output unit 106 controls the terminal device 3 of the target person to output the response information by voice. At this time, the output unit 106 may transmit information for outputting the voice corresponding to the response information to the terminal device 3. The information for outputting the voice corresponding to the response information may be generated based on a voice generation program selected based on the knowledge of those skilled in the art.

[0063] In one embodiment, the output unit 106 further outputs the backchannel information. In one embodiment, the output unit 106 controls the terminal device 3 of the target person to output the backchannel information by voice. At this time, the output unit 106 may transmit information for outputting the voice corresponding to the backchannel information to the terminal device 3. The information for outputting the voice corresponding to the backchannel information may be generated based on a voice generation program selected based on the knowledge of those skilled in the art.

[0064] In one embodiment, the output unit 106 outputs at least one of the first response information and the second response information after outputting the backchannel information. Outputting at least one of the first response information and the second response information after outputting the backchannel information includes controlling the terminal device 3 so that the response information is output by voice at the terminal device 3 after the backchannel information is output by voice at the terminal device 3.

[0065] In one embodiment, the output unit 106 outputs the second response information when the determination unit 104 determines that a predetermined time has elapsed after the second delimiter.

[0066] 2.1.2 Storage Unit 12 The storage unit 12 stores various information for the information processing device 2 to operate. In one embodiment, the storage unit 12 stores the program executed by the control unit 10.

[0067] 2.1.3 Network Interface Unit 14 The network interface unit 14 realizes communication with other devices via the communication network 5.

[0068] 2.2 Terminal Device The terminal device 3 is a communication device used by the subject. The terminal device 3 may be, for example, a smartphone, a personal computer, a tablet terminal, a wearable terminal, or the like. The terminal device 3 includes an input interface, an output interface, and a communication interface.

[0069] The input interface is an interface for the terminal device 3 to receive inputs from the subject. The input interface may be, for example, a touch panel, a microphone, a camera, a keyboard, and a mouse.

[0070] The output interface is an interface for transmitting information to the subject by means of images, sounds, etc. The output interface is, for example, a display (which may also serve as a touch panel) and a speaker.

[0071] The communication interface is an interface for realizing communication with other devices via the communication network 5. The communication interface may be a wireless communication interface or a wired communication interface.

[0072] The terminal device 3 may be able to access the services provided by the information processing device 2 via, for example, a web browser, or may be able to access the services by installing dedicated software.

[0073] 2.3 LLM Server Device 4 The LLM server device 4 is a device that provides services by means of an LLM. The LLM may be a deep learning model that has hundreds of millions or more parameters and has learned data on natural language of hundreds of GB or more. The LLM may be, for example, gpt-4о or the like. In one example, the LLM server device 4 provides services using the LLM via an API (Application Programming Interface).

[0074] In one embodiment, the LLM server device 4 receives an input of an instruction (which can also be referred to as a prompt) from another device and returns a response along with the instruction to the other device. In one example, both the instruction and the response are texts.

[0075] 2.4 Communication network 5 The communication network 5 realizes communication between each device included in the system 1. The communication network 5 realizes communication between each device based on, for example, the TCP / IP protocol.

[0076] 3 Operation With reference to FIGS. 3 to 8, an example of the operation of the system 1 will be described.

[0077] 3.1 First embodiment With reference to FIGS. 3 to 5, the operation of the system 1 according to the first embodiment will be described. In the first embodiment, the basic mode of the system 1 will be exemplarily described.

[0078] 3.1.1 Flowchart FIG. 3 is a flowchart for explaining the operation of the information processing device 2 according to the first embodiment. Steps S100 to S112 are described in the flowchart of FIG. 3. However, the information processing device 2 continuously and sequentially executes acquisition of an audio chunk, determination of whether or not the device is in the middle of voice synthesis based on the audio chunk, and speech recognition of voice information in parallel with these processes. This process is an example of the acquisition unit 100 acquiring utterance information regarding the utterance of the target person.

[0079] The information processing device 2 determines whether or not there is a break in the utterance by the target person at that time (S100). In one example, the information processing device 2 executes this determination based on whether or not a predetermined number of continuously voice chunks determined as having no voice are continuous at that time. This process is an example of the determination unit 104 determining whether or not a first break has occurred in the utterance of the target person.

[0080] When there is no break in the speech (S100 NO), the information processing apparatus 2 waits until a break occurs in the speech (S102). As described above, even during the waiting period, the information processing apparatus 2 continuously and sequentially executes acquisition of an audio chunk, determination as to whether or not voice is being uttered based on the audio chunk, and speech recognition of audio information.

[0081] On the other hand, when there is a break in the speech (S100 YES), the information processing apparatus 2 executes response determination processing. The response determination processing at this point may include the following processes (1) to (3). (1) Sending a first response determination instruction including first speech information (in one example, the text data of the speech recognition obtained up to that point) to the LLM server apparatus 4. (2) Obtaining first response information from the LLM server apparatus 4. (3) Generating information for outputting a voice corresponding to the first response information.

[0082] The text data of the speech recognition obtained up to this point is an example of information regarding the portion up to the first break in the speech of the target person.

[0083] Also, the response determination processing at this point is an example in which the first response determination unit 102a determines first response information regarding a response to at least a part of the speech of the target person by inputting a first response determination instruction including the first speech information to the LLM.

[0084] Also, the response determination processing at this point is an example in which the first response determination unit 102a inputs a first response determination instruction to the LLM when the determination unit 104 determines that a first break has occurred in the speech of the target person.

[0085] Next, the information processing apparatus 2 determines whether or not the speech has ended (S110). In one example, the information processing apparatus 2 executes this determination based on whether or not a predetermined number or more of audio chunks determined to have no voice up to that point are continuous from the time of step S100.

[0086] When it is determined that the speech has ended (S110 YES), as will be described later, the information processing apparatus 2 can output first response information. On the other hand, when it is not determined that the speech has ended, that is, when the speech by the target person continues (S110 NO), the information processing apparatus 2 waits until a break occurs in the speech (S102). Then, when a break occurs in the speech in a subsequent loop (S100 YES), the information processing apparatus 2 executes the response determination process again. The response determination process at this point may include the following processes (1) to (3). (1) Sending a second response determination instruction including the first speech information and / or the first response information and the second speech information (in one example, the text data of the speech-to-text obtained up to that point) to the LLM server device 4. (2) Obtaining second response information from the LLM server device 4. (3) Generating information for outputting the voice corresponding to the second response information.

[0087] The second step S100 is an example in which the determination unit 104 further determines whether a second break has occurred after the first break in the speech of the target person.

[0088] The text data of the speech-to-text obtained up to this point is an example of information regarding the portion of the speech of the target person up to the second break.

[0089] In addition, in this response determination process, when a predetermined condition regarding the speech of the target person is satisfied, the second response determination unit 102b inputs a second response determination instruction including the first speech information and / or the first response information and the second speech information to the LLM before outputting the first response information to the target person, thereby determining second response information regarding other responses to at least a part of the speech of the target person. This is an example.

[0090] In addition, this response determination process is an example in which the second response determination unit 102b inputs a second response determination instruction to the LLM when the determination unit 104 determines that a second break has occurred in the speech of the target person.

[0091] The information processing apparatus 2 repeatedly executes the loop of steps S100 to S110 while the speech of the target person continues. Then, in step S104, the response information is updated each time by the response determination process.

[0092] Step S110 immediately after the second execution of step S100 is an example in which the determination unit 104 further determines whether or not a predetermined time regarding the output of the response has elapsed after the second delimiter of the speech of the target person.

[0093] Thereafter, when it is determined that the speech has ended (S110 YES), the information processing apparatus 2 transmits information for outputting a voice corresponding to the response information determined in the response determination process of the most recent step S104 to the terminal device 3 (S112). The "response information determined in the response determination process of the most recent step S104" may be, for example, the response information determined in the first response determination process if it has never been determined as NO in step S110, or the response information determined in the second response determination process if it has been determined as NO only once in step S110.

[0094] Transmitting information for outputting a voice corresponding to the response information to the terminal device 3 is an example of the output unit 106 outputting at least one of the first response information and the second response information.

[0095] Also, outputting the response information when it is determined that the speech has ended is an example of the output unit 106 outputting the second response information when it is determined by the determination unit 104 that a predetermined time has elapsed after the second delimiter.

[0096] 3.1.2 Table (When Outputting the First Response Information) FIG. 4 is a diagram for explaining the operation of the information processing apparatus 2 according to the first embodiment from different viewpoints. In the following description, "Th1" is a value corresponding to the length of time for determining that a break has occurred in speech, and "Th2" is a value corresponding to the length of time for determining that speech has ended.

[0097] The information processing apparatus 2 sequentially acquires voice chunks (voice chunks v 1 ~t N+Th1+Th2+1 ) at each of the time points t 1 ~v N+Th1+Th2+1 ) and executes a voice determination for each voice chunk.

[0098] Based on each of the N voice chunks v 1 ~v N , it is determined that the subject is speaking at the time points t 1 ~t N , and speech recognition is executed. The text data obtained by performing speech recognition on the voice chunks v 1 ~v N is an example of first speech information regarding the first speech. This process corresponds to the process described in steps S100 NO to S102 of FIG. 3.

[0099] Next, based on Th1 voice chunks v N+1 ~v N+Th1 , it is determined that the subject is not speaking at the time points t N+1 ~t N+Th1 . Then, based on the fact that there is no voice of the subject during this time period, the information processing apparatus 2 determines that a break has occurred in the speech and executes a response determination process. As a result, the information processing apparatus 2 determines the first response information. These processes correspond to the processes described in steps S100 YES to S104 of FIG. 3. The response determination process continues from the time point t N+Th1+1 to the time point t N+Th1+Th2 for Th2 time intervals.

[0100] Next, based on the fact that there has been no utterance from the subject from time point t in Th2 time intervals, the information processing apparatus 2 determines that the speech has ended and outputs first response information. These processes correspond to the processes described in steps S110 YES to S112 in FIG. 3. N+Th1+1 to time point t N+Th1+Th2 Based on the fact that there has been no utterance from the subject until, it is determined that the speech has ended and first response information is output. These processes correspond to the processes described in steps S110 YES to S112 in FIG. 3.

[0101] 3.1.3 Table (When outputting second response information) FIG. 5 is a diagram for further explaining the operation of the information processing apparatus 2 according to the first embodiment from a different perspective. The information processing apparatus 2 sequentially acquires voice chunks (voice chunks v 1 to v M+Th1+Th2+1 ) at each of time points t 1 to v M+Th1+Th2+1 ) and executes voice detection for each voice chunk. Since the operations from time point t 1 to t N+Th1 are common to the example in FIG. 4, the description thereof will be omitted below.

[0102] In FIG. 4, an example in which it is determined that the subject is not speaking at time points t N+Th1+1 to t N+Th1+Th2 was described. In FIG. 5, an example in which the voice of the subject is detected at a time point t N+Th1+1 to t N+Th1+Th2 between (that is, after the information processing apparatus 2 determines that there is a break in the speech at time point t N+Th1+k+1 , and before the first response information is output at time point t N+Th1+1 , the subject resumes speaking at time point t N+Th1+Th2 to t N+Th1+k+1 ) will be described.

[0103] Based on the voice chunks v N+Th1+k+1 to v M , the information processing apparatus 2 determines that the subject is speaking at time points t N+Th1+k+1 to t M and executes speech-to-text conversion. Voice chunks v N+Th1+k+1 to v MThe text data obtained by speech recognition is an example of second utterance information regarding a second utterance after the first utterance. This process corresponds to the process described in steps S100 NO to S102 of FIG. 3.

[0104] Next, the information processing apparatus 2 determines that the subject is not uttering based on Th1 voice chunks v M+1 ~v M+Th1 at time points t M+1 ~t M+Th1 And based on the fact that the subject is not uttering during this time period, the information processing apparatus 2 determines that there is a break in the utterance and executes response determination processing. Thereby, the information processing apparatus 2 determines the second response information. These processes correspond to the processes described in steps S100 YES to S104 of FIG. 3. The response determination processing continues from time point t M+Th1+1 to time point t M+Th1+Th2 for Th2 time intervals.

[0105] Next, based on the fact that the subject is not uttering from time point t M+Th1+1 to time point t M+Th1+Th2 for Th2 time intervals, the information processing apparatus 2 determines that the utterance has ended and outputs the second response information. These processes correspond to the processes described in steps S110 YES to S112 of FIG. 3.

[0106] 3.2 Second Embodiment With reference to FIGS. 6 to 8, the operation of the system 1 according to the second embodiment will be described. In the second embodiment, an example of the mode of the system 1 when the information processing apparatus 2 outputs backchannel information will be exemplarily described.

[0107] 3.2.1 Flowchart FIG. 6 is a flowchart for explaining the operation of the information processing apparatus 2 according to the sixth embodiment. Steps S200 to S212 are described in the flowchart of FIG. 6. The information processing apparatus 2 continuously and sequentially executes acquisition of audio chunks, determination of whether or not voice is being uttered based on the audio chunks, and speech recognition of audio information in parallel with these processes.

[0108] Since steps S200 to S204 and steps S210 to S212 in FIG. 6 may respectively correspond to steps S100 to S104 and steps S110 to S112 in FIG. 3, the following will describe steps S206 to S208 in FIG. 6.

[0109] While executing response determination processing (S204), the information processing apparatus 2 may further execute backchannel determination processing (S206). The backchannel determination processing may include the following processes (1) to (3). (1) Sending a backchannel determination instruction including first utterance information (in one example, the text data of the speech recognition obtained up to that point) to the LLM server device 4. (2) Obtaining backchannel information from the LLM server device 4. (3) Generating information for outputting a voice corresponding to the backchannel information. This backchannel determination processing is an example in which the backchannel determination unit 102c determines backchannel information regarding the backchannel corresponding to the first utterance information when determining the first response information. Note that the backchannel determination processing is not limited to being executed only when determining the first response information, and may be executed for each loop from step S200 to step S210.

[0110] Next, the information processing apparatus 2 outputs backchannel information (S208). That is, while executing the loop of steps S200 to S210 (the process repeated while the target person continues speaking), the information processing apparatus 2 can output backchannel information before outputting the response information in step S212. Thereby, the target person can experience a more natural conversation with the AI agent. Note that the fact that the response information is output in step S212 after the backchannel information is output in step S208 is an example of the output unit 106 outputting the second response information after outputting the backchannel information.

[0111] Note that in FIG. 6, step S204 and steps S206 to S208 are described in series for convenience, but these processes may be executed in parallel. That is, the information processing apparatus 2 may execute the backchannel determination process and the output of the backchannel information while executing the response determination process. According to this configuration, since the backchannel information is first output to the target person while the response determination process is being executed, the latency experienced by the target person can be further suppressed.

[0112] 3.2.2 Table (when outputting the first response information) FIG. 7 is a diagram for explaining the operation of the information processing apparatus 2 according to the second embodiment from a different perspective. In FIG. 4, an example was described in which the information processing apparatus 2 executes the response determination process at time t N+Th1+1 ~ time t N+Th1+Th2 and does not output information during that time. In contrast, in the example of FIG. 7, based on the fact that there was no utterance from the target person at time t N+1 ~ t N+Th1 the information processing apparatus 2 determines that there is a break in the conversation and executes the response determination process and the backchannel determination process. Thereby, the information processing apparatus 2 determines the first response information and the backchannel information and outputs the backchannel information. These processes correspond to the processes described in steps S200 YES to S208 of FIG. 6.

[0113] 3.2.3 Table (when outputting the second response information) FIG. 8 is a diagram for further explaining the operation of the information processing apparatus 2 according to the second embodiment from a different perspective. FIG. 5 shows time point t N+Th1+1 to time point t N+Th1+k and time point t M+Th1+1 to time point t M+Th1+Th2 wherein the information processing apparatus 2 executes response determination processing and does not output information during that period. An example was described above.

[0114] In contrast, in the example of FIG. 8, based on the fact that there is no utterance from the target person at time point t N+1 to t N+Th1 , the information processing apparatus 2 determines that there is a break in the conversation and executes response determination processing and backchannel determination processing. The information processing apparatus 2 further outputs backchannel information at time point t N+1 to t N+Th1+k .

[0115] Also, based on the fact that there is no utterance from the target person at time point t M+1 to t M+Th1 , the information processing apparatus 2 determines that there is a break in the conversation and executes response determination processing and backchannel determination processing. The information processing apparatus 2 further outputs backchannel information at time point t M+1 to t M+Th1+Th2 .

[0116] 3.3 Specific Example Hereinafter, specific examples of the first response determination instruction, the first response information, the second response determination instruction, and the second response information will be described when it is assumed that the utterance by the target person is "Umm, I would like to make a reservation for two at 7:00 tonight... Oh, my name is ○○". Note that the first utterance information includes text data "Umm, I would like to make a reservation for two at 7:00 tonight...", and the second utterance information includes text data "Oh, my name is ○○".

[0117] In one example, the first response determination instruction may include the following text data (1) to (3). (1) System prompt: "You are an excellent AI agent for accepting reservations at a restaurant." (2) Conversation history: "Target person 'Hello?' → AI agent 'Yes, this is [store name].'" (3) Instruction including the first utterance information: "Recently, the target person said, 'Umm, I would like to make a reservation for two at 7 pm tonight.' Please create a response to this."

[0118] The LLM server device 4 can generate first response information including text data such as "A reservation for two at 7 pm tonight, right?" for this example's first response determination instruction, and transmit it to the information processing device 2.

[0119] In one example, the second response determination instruction may include the following text data (1) to (4). (1) System prompt: "You are an excellent AI agent for accepting reservations at a restaurant." (2) Conversation history: "Target person 'Hello?' → AI agent 'Yes, this is [store name].' → Target person 'Umm, I would like to make a reservation for two at 7 pm tonight.'" (3) First response information: "A reservation for two at 7 pm tonight, right?" (4) Instruction including the second utterance information: "The target person additionally said, 'Oh, my name is [name].' Based on the first response information, please create a response."

[0120] The LLM server device 4 can generate second response information including text data such as "A reservation for two at 7 pm tonight, right? Is your name [name], right?" for this example's second response determination instruction, and transmit it to the information processing device 2. Note that "Based on the first response information, create a response" may mean that the text data included in the first response information is also included in the second response information, or that text data corresponding to the continuation of the text data included in the first response information is included in the second response information, or that the second response information is determined while referring to the first response information without being restricted by it.

[0121] 4 Hardware Configuration Referring to FIG. 9, an example of the hardware configuration when the devices included in the system 1 described above are realized by a computer 70 will be described. Note that the functions of each device can also be realized by dividing them among a plurality of devices.

[0122] As shown in FIG. 9, the computer 70 includes a processor 700, a storage device 702, an input I / F 704, a data I / F 706, a communication I / F 708, and a display device 710.

[0123] The processor 700 controls various processes in the computer 70 by executing programs stored in the storage device 702. For example, each functional unit included in the control unit 10 of the information processing device 2 can be realized by the processor 700 executing a program stored in the storage device 702.

[0124] The storage device 702 is a storage medium such as a RAM (Random Access Memory), for example. The RAM temporarily stores the program code of the program executed by the processor 700 and the data required during the execution of the program.

[0125] The storage device 702 is also a non-volatile storage medium such as a hard disk drive (HDD) or a flash memory, for example. The storage device 702 stores an operating system and various programs for realizing the above-described configurations. The storage medium storing the various programs may be a non-transitory computer readable medium readable by a computer. In addition, the storage device 702 can also store a table for registering various information and a DB for managing the table. Such programs and data are referred to by the processor 700 by being loaded into the storage device 702 as needed.

[0126] The input I / F 704 is a device for receiving an input from a subject. Specific examples of the input I / F 704 include a camera, buttons, a microphone, a keyboard, a mouse, a touch panel, various sensors, wearable devices, and the like. The input I / F 704 may be connected to the computer 70 via an interface such as USB (Universal Serial Bus).

[0127] The data I / F 706 is a device for inputting data from outside the computer 70. Specific examples of the data I / F 706 include a drive device for reading data stored in various storage media. The data I / F 706 may be provided outside the computer 70. In that case, the data I / F 706 is connected to the computer 70 via an interface such as USB.

[0128] The communication I / F 708 is a device for performing data communication via the communication network 5, either wired or wirelessly, with a device outside the computer 70. The communication I / F 708 may be provided outside the computer 70. In that case, the communication I / F 708 is connected to the computer 70 via an interface such as USB.

[0129] The display device 710 is a device for displaying various information. Specific examples of the display device 710 include, for example, a liquid crystal display, an organic EL (Electro-Luminescence) display, a display of a wearable device, and the like. The display device 710 may be provided outside the computer 70. In that case, the display device 710 is connected to the computer 70 via, for example, a display cable. Also, when a touch panel is adopted as the input I / F 704, the display device 710 can be configured integrally with the input I / F 704.

[0130] In addition, the components included in the device included in the system 1 described in the above embodiment are assumed to be such that the programs stored in the storage device 702 are executed by the processor 700, and the defined processing is realized in cooperation with other hardware. In other words, these components are assumed to be either software or firmware, or the corresponding hardware, and in both concepts, they are also described as "function", "means", "section", "processing circuit", "unit", or "module", etc., and can be read and replaced with each other.

[0131] 5 Variations The embodiments described above are for facilitating the understanding of the present disclosure and are not for limiting and interpreting the present disclosure. The configurations that the embodiments may include are not limited to those exemplified and can be changed as appropriate. Also, it is possible to partially replace or combine the configurations shown in different embodiments.

[0132] The matters described with the prefixes "first" and "second" in the above embodiment can be extended and understood in the relationship of "first" to "Nth" (where N is a natural number) based on the knowledge of those skilled in the art.

[0133] In one example, the acquisition unit 100 may acquire speech information regarding the speech of the target person, and the speech information may include first speech information regarding the first speech to Nth speech information regarding the Nth speech.

[0134] In one example, for natural numbers n = 2 to N, when a predetermined condition regarding the speech is satisfied, the nth response determination unit inputs a nth response determination instruction including at least one of the first speech information to the (n - 1)th speech information and the first response information to the (n - 1)th response information and the nth speech information to the large language model before outputting the (n - 1)th response information to the target person, and may determine the nth response information regarding other responses to at least a part of the speech.

[0135] In one example, when it is determined that the utterance by the subject has ended between when the (n - 1)-th response information is determined and when the n-th response information is determined for natural numbers n = 2 to N, the output unit 104 may output at least one of the first response information to the (n - 1)-th response information, and when it is determined that the utterance by the subject has ended after the n-th response information is determined, may output at least one of the first response information to the n-th response information.

[0136] In the above embodiment, the LLM has been described as being hosted by the LLM server device 4, but it is not limited thereto. The LLM may be hosted by the information processing device 2 or may be hosted by the terminal device 3.

[0137] In the above embodiment, the information processing device 2 has been described as acquiring voice information from the terminal device 3 and executing a process of streamlining the call by the AI agent based on the voice information, but it is not limited thereto. At least a part of the functions described as being provided in the information processing device 2 in the above embodiment may be provided in the terminal device 3.

[0138] In the above embodiment, the information processing device 2 has been described as communicating with the terminal device 3, but it is not limited thereto. There may be a predetermined intermediate server between the information processing device 2 and the terminal device 3.

[0139] In the above-described embodiment, when the response information and the second response information are determined, an example in which the second response information is output and the response information is also further output has been described, but the present invention is not limited thereto. For example, in the case where the response determination process related to the response determination process and the determination of the second response information are executed in parallel, whether to output the response information can be determined based on the timing at which the response determination process is completed. In one example, when the response determination process is completed before the response determination process is completed, the information processing apparatus 2 may output the second response information without outputting the response information. On the other hand, when the response determination process is completed after the response determination process is completed, the information processing apparatus 2 may output the response information during the execution of the response determination process and then output the second response information. According to such a configuration, it is possible to avoid a situation in which the response information is output earlier than the response information even though the response information has already been determined, and as a result, it may be possible to output the response information to the target person earlier.

[0140] 6 Supplementary The language in the present embodiment can be understood as follows within a range where no contradiction occurs.

[0141] In the present embodiment, "executing a predetermined process based on predetermined information" may be any one of executing the predetermined process based on at least a part of the predetermined information, executing the predetermined process based on at least the predetermined information, and executing the predetermined process probabilistically based on the predetermined information. That is, "executing a predetermined process based on predetermined information" is not limited to executing the predetermined process based only on the predetermined information.

[0142] In this embodiment, "executing another process based on a predetermined process" may be any of the following: executing the other process after the predetermined process is executed; continuously executing the predetermined process and the other process; executing the other process based on the information determined by the predetermined process; executing the other process on the condition that the predetermined process has been executed; or executing the other process by means of the predetermined process. Note that "executing another process by a predetermined process" may be understood in the same way as "executing another process based on a predetermined process".

[0143] In this embodiment, "predetermined information includes other information" may be either that at least a part of the predetermined information is the other information or that the other information can be obtained based on the predetermined information.

[0144] In this embodiment, "a predetermined process includes another process" may be either that at least a part of the predetermined process is the other process (i.e., the other process is performed in the process of obtaining the result of the predetermined process) or that one aspect of the predetermined process is the other process.

[0145] In this embodiment, "a predetermined object corresponds to another object" may be any of the following: the predetermined object and the other object are in a one-to-one relationship; the other object is included in a predetermined set specified based on the predetermined object; or the other object can be specified based on the predetermined object. Note that "a predetermined object corresponds to another object" is not limited to being managed, for example, on a database. Also, "a predetermined object is associated with another object" may be understood in the same way as "a predetermined object corresponds to another object".

[0146] In this embodiment, "acquiring information" includes making the information processable by the control unit 10. "Acquiring information" may be, for example, receiving the information from another device, obtaining the information by a predetermined process, and reading the information from the storage unit 12, and the like.

[0147] In this embodiment, "generating information" may be either making the information obtained by a predetermined process processable by the control unit 10 or storing the information obtained by a predetermined process in the storage unit 12.

[0148] In this embodiment, "determining information" may be either selecting at least one from among one or more pieces of information or newly generating the information.

[0149] In this embodiment, "outputting information" may be either transmitting the information to another device or outputting the information by voice or video.

[0150] 7 Configuration Example This disclosure includes the following technologies.

[0151] [Appendix 1] An acquisition unit 100 that acquires speech information regarding the speech of a target person, where the speech information includes first speech information regarding a first speech and second speech information regarding a second speech that is after the first speech; a first response determination unit 102a that determines first response information regarding a response to at least a part of the speech by inputting a first response determination instruction including the first speech information into a large language model; a second response determination unit 102b that, when a predetermined condition regarding the speech is satisfied, determines second response information regarding another response to at least a part of the speech by inputting a second response determination instruction including the first speech information and / or the first response information and the second speech information into the large language model before outputting the first response information to the target person; and an output unit 106 that outputs the first response information when the predetermined condition is not satisfied and outputs at least one of the first response information and the second response information when the predetermined condition is satisfied. The information processing apparatus 2 is provided with these components.

[0152] [Appendix 2] The information processing apparatus 2 further includes a backchannel determination unit 102c that determines backchannel information regarding a backchannel corresponding to the first speech information when determining the first response information, and the output unit 106 further outputs the backchannel information. The information processing apparatus 2 according to Appendix 1.

[0153] [Appendix 3] The information processing apparatus 2 according to Appendix 2, wherein when the predetermined condition is not satisfied, the output unit 106 outputs the backchannel information and then outputs the first response information, and when the predetermined condition is satisfied, the output unit 106 outputs the backchannel information and then outputs at least one of the first response information and the second response information.

[0154] [Appendix 4] The information processing apparatus 2 according to Appendix 2 or 3, wherein the backchannel determination unit 102c determines the backchannel information by inputting a backchannel determination instruction including the first speech information into a large language model.

[0155] [Appendix 5] The information processing apparatus 2 according to Supplementary Note 1, further comprising a determination unit 104 that determines whether a first delimiter has occurred in the speech, wherein the first speech information includes information regarding the portion of the speech up to the first delimiter, and the first response determination unit 102a inputs a first response determination instruction to the large language model when the determination unit 104 determines that a first delimiter has occurred in the speech.

[0156] [Supplementary Note 6] The determination unit 104 further determines whether a second delimiter has occurred after the first delimiter of the speech. The second speech information includes information regarding the portion of the speech up to the second delimiter. The second response determination unit 102b inputs a second response determination instruction to the large language model when the determination unit 104 determines that a second delimiter has occurred in the speech. The information processing apparatus 2 according to Supplementary Note 5.

[0157] [Supplementary Note 7] The determination unit 104 further determines whether a predetermined time regarding the output of the response has elapsed after the second delimiter of the speech. The output unit 106 outputs at least one of the first response information and the second response information when it is determined that a predetermined time has elapsed after the second delimiter. The information processing apparatus 2 according to Supplementary Note 6.

[0158] [Supplementary Note 8] The information processing apparatus 2 according to Supplementary Note 7, further comprising a response determination unit 102c that determines response information regarding the response corresponding to the first speech information when determining the first response information. The output unit 106 further outputs the response information, and the predetermined time is determined based on the response information.

[0159] [Supplementary Note 9] The output unit 106 controls so that the terminal device of the target person outputs the first response information by voice when a predetermined condition is not satisfied, and controls so that at least one of the first response information and the second response information is output by voice when the predetermined condition is satisfied. The information processing apparatus 2 according to any one of Supplementary Notes 1 to 8.

[0160] [Supplementary Note 10] When determining the first response information, the first response determination unit 102a further determines information for outputting the first response information by voice. When determining the second response information, the second response determination unit 102b further determines information for outputting the second response information by voice. The information processing apparatus 2 according to Supplementary Note 9.

[0161] [Supplementary Note 11] The information processing apparatus 2 according to Supplementary Note 10 further includes a storage unit 12 that stores information for outputting each of one or more texts by voice. When the first response information matches at least one of the one or more texts, determining the information for outputting the first response information by voice includes obtaining, from the storage unit, the information for outputting the matching text by voice. When the first response information does not match any of the one or more texts, determining the information for outputting the first response information by voice includes generating, based on a predetermined voice generation program, the information for outputting the first response information by voice. When the second response information matches at least one of the one or more texts, determining the information for outputting the second response information by voice includes obtaining, from the storage unit, the information for outputting the matching text by voice. When the second response information does not match any of the one or more texts, determining the information for outputting the second response information by voice includes generating, based on a predetermined voice generation program, the information for outputting the second response information by voice.

[0162] [Supplementary Note 12] An information processing method, wherein an information processing apparatus 2 acquires utterance information regarding an utterance of a target person, the utterance information including first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; determines first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; when a predetermined condition regarding the utterance is satisfied, determines second response information regarding another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person; outputs the first response information when the predetermined condition is not satisfied, and outputs at least one of the first response information and the second response information when the predetermined condition is satisfied.

[0163] [Appendix 13] A program for causing an information processing apparatus 2 to acquire utterance information regarding an utterance of a target person, the utterance information including first utterance information regarding a first utterance and second utterance information regarding a second utterance after the first utterance; determine first response information regarding a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large language model; when a predetermined condition regarding the utterance is satisfied, determine second response information regarding another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into the large language model before outputting the first response information to the target person; output the first response information when the predetermined condition is not satisfied, and output at least one of the first response information and the second response information when the predetermined condition is satisfied.

Explanation of Signs

[0164] 1... System, 2... Information processing device, 3... Terminal device, 4... LLM server device, 10... Control unit, 12... Storage unit, 70... Computer, 100... Acquisition unit, 102... Decision unit, 102a... First response decision unit, 102b... Second response decision unit, 102c... Interjection decision unit, 104... Judgment unit, 106... Output unit, 700... Processor

Claims

1. an acquisition unit that acquires speech information related to an utterance of a target person, the speech information including first utterance information related to a first utterance and second utterance information related to a second utterance subsequent to the first utterance; a first response determination unit that determines first response information related to a response to at least a part of the utterance by inputting a first response determination instruction including the first utterance information into a large-scale language model; a second response determination unit that determines second response information related to another response to at least a part of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into a large-scale language model before outputting the first response information to the target person when a predetermined condition related to the utterance is satisfied; an output unit that outputs the first response information when the predetermined condition is not satisfied, and outputs at least one of the first response information and the second response information when the predetermined condition is satisfied; An information processing device comprising:

2. A backchannel determination unit that determines backchannel information related to a backchannel corresponding to the first utterance information when determining the first response information, The output unit further outputs the backchannel information. The information processing device according to claim 1 .

3. When the predetermined condition is not satisfied, the output unit outputs the backchannel information and then outputs the first response information, and when the predetermined condition is satisfied, the output unit outputs the backchannel information and then outputs at least one of the first response information and the second response information. The information processing device according to claim 2 .

4. The backchannel determination unit determines the backchannel information by inputting a backchannel determination instruction including the first utterance information into a large-scale language model. The information processing device according to claim 2 .

5. a determination unit that determines whether a first division has occurred in the utterance, the first utterance information includes information on a portion of the utterance up to the first break, the first response determination unit inputs the first response determination instruction to a large-scale language model when the determination unit determines that the first break has occurred in the utterance. The information processing device according to claim 1 .

6. The determination unit further determines whether a second segment of the utterance occurs after the first segment of the utterance; the second utterance information includes information on a portion of the utterance up to the second break, the second response determination unit inputs an instruction to determine the second response to a large-scale language model when the determination unit determines that the second break has occurred in the utterance. The information processing device according to claim 5 .

7. The determination unit further determines whether or not a predetermined time for outputting a response has elapsed after the second division of the utterance; the output unit outputs at least one of the first response information and the second response information when the determination unit determines that the predetermined time has elapsed after the second division. The information processing device according to claim 6.

8. A backchannel determination unit that determines backchannel information related to a backchannel corresponding to the first utterance information when determining the first response information, The output unit further outputs the backchannel information, The predetermined time is determined based on the backchannel information. The information processing device according to claim 7.

9. The output unit controls the terminal device of the subject to output the first response information by voice when the predetermined condition is not satisfied, and controls the terminal device of the subject to output at least one of the first response information and the second response information by voice when the predetermined condition is satisfied. The information processing device according to claim 1 .

10. The first response determination unit further determines information for outputting the first response information by voice when determining the first response information, The second response determination unit further determines information for outputting the second response information by voice when determining the second response information. The information processing device according to claim 9.

11. a storage unit configured to store information for outputting each of the one or more texts by voice, When the first response information is consistent with at least one of the one or more texts, determining information for outputting the first response information by voice includes obtaining information for outputting the consistent text by voice from the storage unit; When the first response information does not match any of the one or more texts, determining information for outputting the first response information by voice includes generating information for outputting the first response information by voice based on a predetermined voice generation program; When the second response information is consistent with at least one of the one or more texts, determining information for outputting the second response information by voice includes obtaining information for outputting the consistent text by voice from the storage unit; When the second response information does not match any of the one or more texts, determining information for outputting the second response information by voice includes generating information for outputting the second response information by voice based on the predetermined voice generation program. The information processing device according to claim 10.

12. An information processing device, acquiring speech information relating to an utterance of a target person, the speech information including first utterance information relating to a first utterance and second utterance information relating to a second utterance subsequent to the first utterance; determining first response information related to a response to at least a portion of the utterance by inputting a first response determination instruction including the first utterance information into a large-scale language model; determining second response information related to another response to at least a portion of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into a large-scale language model before outputting the first response information to the target person when a predetermined condition related to the utterance is satisfied; outputting the first response information when the predetermined condition is not satisfied, and outputting at least one of the first response information and the second response information when the predetermined condition is satisfied; An information processing method.

13. In the information processing device, acquiring speech information relating to an utterance of a target person, the speech information including first utterance information relating to a first utterance and second utterance information relating to a second utterance subsequent to the first utterance; determining first response information related to a response to at least a portion of the utterance by inputting a first response determination instruction including the first utterance information into a large-scale language model; determining second response information related to another response to at least a portion of the utterance by inputting a second response determination instruction including the first utterance information and / or the first response information and the second utterance information into a large-scale language model before outputting the first response information to the target person when a predetermined condition related to the utterance is satisfied; outputting the first response information when the predetermined condition is not satisfied, and outputting at least one of the first response information and the second response information when the predetermined condition is satisfied; A program to execute.

Citation Information

Patent Citations

  • Voice interactive system

    JP1994259090A

  • Response voice generating method and voice interactive system

    JP1996263092A

  • Using large language models in generating automated assistant responses

    JP2024521053A

  • system

    JP7550335B1

Cited By

  • Information processing systems, information processing methods, and programs

    JP7921687B1