Information processing device, method, program, and dialogue system
The information processing apparatus addresses the challenge of conventional dialogue systems by using a combination of models and dialogue history to create appropriate responses, enhancing interaction quality between humans and computers.
Patent Information
- Application Number
- JP2024147259
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-09
AI Technical Summary
Conventional dialogue systems struggle to respond appropriately based on the content of the dialogue, leading to ineffective interactions between humans and computers.
An information processing apparatus that includes a model storage unit for multiple models, such as a large language model and a task specialization model, a dialogue history storage unit, and a dialogue control unit. The dialogue control unit selects models based on dialogue history and creates response messages using the output data from the selected models.
Enables the system to respond appropriately and effectively to the content of the dialogue, improving the interaction quality between humans and computers.
Smart Images

Figure 2025086860000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, method, program, and dialogue system.
Background Art
[0002] Conventionally, a dialogue system in which a human can interact with a computer is known. For example, the messages that the computer responds with are created using a machine learning model.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, in conventional dialogue systems, it has been difficult for the computer to respond appropriately according to the content of the dialogue. Therefore, an object of the present invention is to respond appropriately according to the content of the dialogue.
Means for Solving the Problems
[0004] An information processing apparatus according to an embodiment of the present invention includes a model storage unit that stores a plurality of models including a large language model and a task specialization model that is machine-learned to specialize in a specific task different from the large language model, a dialogue history storage unit that stores the history of a dialogue in which a dialogue agent participates, and a dialogue control unit that selects one or more models from among the plurality of models based on the history of the dialogue and creates a response message for the dialogue agent using the output data of the selected model.
Advantages of the Invention
[0005] According to the present invention, it is possible to respond appropriately according to the content of the dialogue.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
MODE FOR CARRYING OUT THE INVENTION
[0007] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the present invention, different from the prior art, the problem of selecting different machine learning models from a large language model and appropriately responding according to the conversation history can be solved.
[0008] FIG. 1 is a diagram for explaining a usage scenario according to an embodiment of the present invention. The present invention is assumed to be used during business negotiations (for example, when a salesperson 2 recommends a product to a customer 3). Hereinafter, the overall flow will be described.
[0009] In step 1 (S1), the user (salesperson) 2 asks the conversation partner (customer) 3 about the needs (also referred to as issues) of the conversation partner (customer) 3.
[0010] In step 2 (S2), the conversation partner (customer) 3 talks about the needs of the conversation partner (customer) 3 in response to what was asked by the user (salesperson) 2 in S1.
[0011] In step 3 (S3), the user (salesperson) 2 asks the conversation system 1 to find a solution to the needs of the conversation partner (customer) 3.
[0012] In step 4 (S4), the conversation system 1 outputs a response message (specifically, a message for a solution to the needs of the conversation partner (customer) 3, or an understanding of the needs (for example, potential needs), or a question regarding the needs (for example, a question to obtain additional information)) in response to the query in S3.
[0013] FIG. 2 is a diagram for explaining an example of a conversation according to an embodiment of the present invention.
[0014] Step 11 (S11) shows an example of the content spoken by the user (salesperson) 2 in S1 of FIG. 1 (“Do you have any problems?”).
[0015] Step 12 (S12) shows an example of the content spoken by the interlocutor (customer) 3 in S2 of FIG. 1 ("I record equipment with a photo when entering and leaving, and check it when leaving, but it takes time to search for the photo.").
[0016] Step 13 (S13) shows an example of the content spoken by the user (salesperson) 2 in S3 of FIG. 1 ("Alfred, do you have any solutions?"). Note that the dialogue system 1 shall start creating a response message according to a predetermined inquiry message (for example, "Alfred" in the example of FIG. 2).
[0017] Step 14 (S14) shows an example of the content spoken by the dialogue system 1 in S4 of FIG. 1 ("You want to reduce time-consuming work and labor, and improve work efficiency. How about the ABC system that can centrally manage admission control?"). Note that the dialogue system 1 may output a response message by playing sound, or may output a response message by displaying text, or may output both sound and text as shown in FIG. 2.
[0018] <Overall Configuration> FIG. 3 is an overall configuration diagram (Example 1) according to an embodiment of the present invention. The dialogue system 1 includes an information processing device (server) 11 and a user terminal 12. In this specification, the information processing device 11 and the user terminal 12 are described as separate devices, but the information processing device 11 and the user terminal 12 may be implemented by one device (that is, the information processing device 11 may have the functions of the user terminal 12).
[0019] <<Information Processing Device>> The information processing device 11 selects one or more machine learning models from two or more machine learning models based on the dialogue history, and creates a response message using the output data of the selected machine learning model. The information processing device 11 is composed of one or more computers. For example, the information processing device 11 is a server.
[0020] The device group described in the embodiments merely represents one of the multiple computing environments for implementing the embodiments disclosed in this specification. In one embodiment, the information processing device 11 includes multiple computing devices such as a server cluster. The multiple computing devices are configured to communicate with each other via any type of communication link including a network or a shared memory, and implement the processing disclosed in this specification.
[0021] <<User Terminal>> The user terminal 12 transmits the data of the voice spoken by a human to the information processing device 11, and receives and outputs (such as voice playback, text display, etc.) the response message created by the information processing device 11. The user terminal 12 is equipped with a microphone function, a speaker function, and a display function. For example, the user terminal 12 is a tablet, a smartphone, a personal computer, etc.
[0022] <System Configuration> FIG. 4 is an overall configuration diagram (Example 2) according to an embodiment of the present invention. In the example of FIG. 4, the communication support system 1 includes a communication support device 10 connected to a communication network N such as the Internet and a LAN (Local Area Network), and a terminal device 100.
[0023] As an example of the usage scenario, the salesperson 2 brings a terminal device 100 that displays the virtual sales 110 to the business negotiation with the customer 3, and conducts the negotiation with the virtual sales 110 involved. Here, the business negotiation is an example of communication. The salesperson 2 is an example of a host participating in the communication, and the customer 3 is an example of a guest participating in the communication. The virtual sales 110 is an example of an interactive agent. The interactive agent is a virtual sales that supports the business negotiation.
[0024] The terminal device 100 is an information terminal such as a PC (Personal Computer), a tablet terminal, or a smartphone that is used by the salesperson 2. The terminal device 100 acquires the speech voices of the salesperson 2, who is a participant in the business negotiation, and the customer 3, and transmits the acquired speech voices (voice data) to the communication support device 10. Note that the speech voice is an example of a statement in communication. Statements in communication include, for example, statements made using text data such as chat. Here, the following explanation will be given assuming that the statements in communication are speech.
[0025] The communication support device (server device) 10 is, for example, an information processing device having a computer configuration or a system including a plurality of computers. The communication support device 10 acquires the speech voices transmitted by the terminal device 100, analyzes the acquired speech voices, and generates a response according to the needs of the customer 3. This response includes, for example, proposing recommendation information, proposing a specific commercial product, responding to small talk, etc.
[0026] The communication support device 10 controls the virtual sales 110 displayed on the terminal device 100 according to the content of this response. The communication support device 10 controls body language, gestures, behavior, and speech.
[0027] The terminal device 100 displays the virtual sales 110 controlled by the communication support device 10 and outputs the speech of the virtual sales 110. As a result, the virtual sales 110 can, for example, present information on specific commercial products according to the needs of the customer 3, present recommendation information, or present topics such as small talk to the customer 3 and the salesperson 2 according to the flow of the business negotiation.
[0028] As a specific example, when it is determined that the need of the customer 3 is "electronic invoicing", the virtual sales 110 proposes a commercial product related to the electronic invoice issuance system, such as "What do you think of product A if you want to digitize your invoices?".
[0029] As another specific example, if it is determined that the need of Customer 3 is a potential need to solve a vague problem such as "AI and digital transformation are popular these days, but I don't know what to do about it in our company", Virtual Sales 110 presents potential problems that the customer is likely to have, such as "For example, are there any issues like time-consuming invoice processing? In such cases, Commercial Product A is recommended.", and then proposes commercial products related to the problem.
[0030] As another specific example, when Salesperson 2 asks Virtual Sales 110 for small talk such as "It's been hot lately, right?", Virtual Sales 110 responds to the small talk, such as "Yes, the temperature has risen to 〇 degrees today, and it's very hot."
[0031] <Hardware Configuration> FIG. 5 is a hardware configuration diagram of an information processing apparatus (server) 11 according to an embodiment of the present invention.
[0032] As shown in FIG. 5, the information processing apparatus (server) 11 is constructed by a computer and includes a CPU 1001, a ROM 1002, a RAM 1003, an HD 1004, an HDD (Hard Disk Drive) controller 1005, a display 1006, an external device connection I / F (Interface) 1007, a network I / F 1008, a data bus 1009, a keyboard 1010, a pointing device 1011, a DVD-RW (Digital Versatile Disk Rewritable) drive 1013, and a media I / F 1015, as shown in FIG. 5.
[0033] Among these, the CPU 1001 controls the operation of the entire information processing apparatus (server) 11. The ROM 1002 stores programs used for driving the CPU 1001 such as the IPL. The RAM 1003 is used as a work area for the CPU 1001. The HD 1004 stores various data such as programs. The HDD controller 1005 controls the reading or writing of various data to and from the HD 1004 according to the control of the CPU 1001. The display 1006 displays various information such as a cursor, menu, window, characters, or images. The external device connection I / F 1007 is an interface for connecting various external devices. The external devices in this case are, for example, a USB (Universal Serial Bus) memory, a printer, and the like. The network I / F 1008 is an interface for performing data communication using a communication network. The bus line 1009 is an address bus, a data bus, etc. for electrically connecting each component such as the CPU 1001 shown in FIG. 5.
[0034] Also, the keyboard 1010 is a type of input means having a plurality of keys for inputting characters, numerical values, various instructions, and the like. The pointing device 1011 is a type of input means for selecting and executing various instructions, selecting a processing target, moving a cursor, and the like. The DVD-RW drive 1013 controls the reading or writing of various data to and from the DVD-RW 1012 as an example of a removable recording medium. Note that it is not limited to the DVD-RW, and a DVD-R or the like may also be used. The media I / F 1015 controls the reading or writing (storage) of data to and from the recording medium 1014 such as a flash memory.
[0035] FIG. 6 is a hardware configuration diagram of the user terminal 12 according to an embodiment of the present invention.
[0036] As shown in FIG. 6, the user terminal 12 includes a CPU 2001, a ROM 2002, a RAM 2003, an EEPROM 2004, a CMOS sensor 2005, an imaging device I / F 2006, an acceleration / azimuth sensor 2007, a media I / F 2009, and a GPS receiver 2011.
[0037] Among these, the CPU 2001 controls the operation of the entire user terminal 12. The ROM 2002 stores programs used for driving the CPU 2001 such as the CPU 2001 and IPL. The RAM 2003 is used as a work area for the CPU 2001. The EEPROM 2004 reads or writes (stores) various data such as smartphone programs according to the control of the CPU 2001. The CMOS (Complementary Metal Oxide Semiconductor) sensor 2005 is a type of built-in imaging means that captures a subject (mainly a self-portrait) according to the control of the CPU 2001 to obtain image data. Note that an imaging means such as a CCD (Charge Coupled Device) sensor may be used instead of the CMOS sensor. The imaging device I / F 2006 is a circuit that controls the driving of the CMOS sensor 2005. The acceleration / azimuth sensor 2007 is various sensors such as an electronic magnetic compass that detects geomagnetism, a gyrocompass, and an acceleration sensor. The media I / F 2009 controls the reading or writing (storage) of data to / from a recording medium 2008 such as a flash memory. The GPS receiver 2011 receives GPS signals from GPS satellites.
[0038] In addition, the user terminal 12 includes a long-distance communication circuit 2012, a CMOS sensor 2013, an imaging device I / F 2014, a microphone 2015, a speaker 2016, an audio input / output I / F 2017, a display 2018, an external device connection I / F (Interface) 2019, a short-distance communication circuit 2020, an antenna 2020a of the short-distance communication circuit 2020, and a touch panel 2021.
[0039] Among these, the long-distance communication circuit 2012 is a circuit that communicates with other devices via a communication network. The CMOS sensor 2013 is a type of built-in imaging means that captures a subject according to the control of the CPU 2001 to obtain image data. The imaging device I / F 2014 is a circuit that controls the driving of the CMOS sensor 2013. The microphone 2015 is a built-in circuit that converts sound into an electrical signal. The speaker 2016 is a built-in circuit that converts an electrical signal into physical vibrations to produce sounds such as music and voices. The audio input / output I / F 2017 is a circuit that processes the input and output of audio signals between the microphone 2015 and the speaker 2016 according to the control of the CPU 2001. The display 2018 is a type of display means such as a liquid crystal or an organic EL (Electro Luminescence) that displays an image of a subject, various icons, etc. The external device connection I / F 2019 is an interface for connecting various external devices. The short-distance communication circuit 2020 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 2021 is a type of input means for operating the user terminal 12 when the user presses the display 2018.
[0040] Also, the user terminal 12 includes a bus line 2010. The bus line 2010 is an address bus, a data bus, etc. for electrically connecting each component such as the CPU 2001 shown in FIG. 6.
[0041] <Functional Configuration> FIG. 7 is a functional block diagram of the information processing apparatus 11 according to an embodiment of the present invention. The information processing apparatus 11 includes a dialogue acquisition unit 101, a dialogue control unit 102, a response unit 103, a dialogue history storage unit 104, an intention estimation model storage unit 105, and a model storage unit 106. The information processing apparatus 11 functions as the dialogue acquisition unit 101, the dialogue control unit 102, and the response unit 103 by executing a program.
[0042] The dialogue acquisition unit 101 acquires dialogue data from the user terminal 12. Note that the dialogue data is not limited to voice data and may also be text data. The dialogue acquisition unit 101 converts the voice data acquired from the user terminal 12 into text data and stores it in the dialogue history storage unit 104, or stores the text data acquired from the user terminal 12 in the dialogue history storage unit 104.
[0043] The dialogue history storage unit 104 stores the history of dialogues in which the dialogue agent participates. The dialogue history storage unit 104 stores the history of dialogues (that is, the text data obtained by converting the voice data acquired by the dialogue acquisition unit 101 from the user terminal 12, or the text data acquired by the dialogue acquisition unit 101 from the user terminal 12). With reference to FIG. 10, an example of the history of dialogues stored in the dialogue history storage unit 104 will be described.
[0044] FIG. 10 is an example of the history of dialogues according to an embodiment of the present invention. As shown in FIG. 10, for each utterance (hereinafter also referred to as a message), the date and time when the utterance was made ("utterance date and time" in FIG. 10), the person who made the utterance ("utterer" in FIG. 10), and the content of the utterance ("message" in FIG. 10) are stored.
[0045] [History of Dialogues] Here, the history of dialogues will be described. The history of dialogues is either the history of dialogues between users or the history of dialogues between a user and a dialogue agent. The history of dialogues is not limited to the history of dialogues between two or more humans (for example, the user (salesperson) 2 and the dialogue partner (customer) 3 in FIGS. 1 and 2), and may also be the history of dialogues between a human and the dialogue system 1. · In the case of a dialogue between humans, the history of the dialogue is data on the "utterance date and time", "utterer", and "message (content of the utterance)" of each utterance by two or more humans. · In the case of a dialogue with the human - interaction system 1 (that is, when a human interacts with the information - processing device 11 via the user terminal 12), the dialogue history is data of the "utterance date and time", "utterer", and "message (content of the utterance)" of each utterance by one or more humans, and data of the "transmission date and time to the user terminal 12" and "content of the response message" of the response message created by the information - processing device 11.
[0046] Return to the description of FIG. 7. The dialogue control unit 102 selects one or more models from among a plurality of models based on the dialogue history, and creates a response message for the dialogue agent using the output data of the selected model. The dialogue control unit 102 selects one or more models from among two or more models stored in the model storage unit 106 based on the dialogue history stored in the dialogue - history storage unit 104, and creates a response message using the output data of the selected model.
[0047] [Estimation of the intention of the dialogue] The dialogue control unit 102 estimates the intention of the dialogue (for example, "request only recommendation information", "request both recommendation information and other information other than recommendation information", "request only other information other than recommendation information") based on the dialogue history, and can select one or more models from among two or more models stored in the model storage unit 106 based on the intention. For example, the dialogue control unit 102 can estimate the intention of the dialogue using the intention - estimation model stored in the intention - estimation model storage unit 105 or the large - language model 162 stored in the model storage unit 106 based on the dialogue or a summary of the dialogue. Note that the intention of the dialogue is not limited to the intention estimated from one utterance, and may be an intention estimated from two or more utterances.
[0048] The response unit 103 transmits the response message created by the dialogue control unit 102 to the user terminal 12. The response unit 103 may convert the text data of the response message created by the dialogue control unit 102 into voice data and transmit it to the user terminal 12, or may transmit the text data of the response message created by the dialogue control unit 102 to the user terminal 12.
[0049] When the dialogue history (dialogue or summary of the dialogue (in this case, the dialogue control unit 102 creates a summary of the dialogue)) is input to the intention estimation model storage unit 105, a machine learning model (intention estimation model) trained to output the intention of the dialogue is stored.
[0050] The model storage unit 106 stores a plurality of models including a large language model and a task-specific model that is machine-learned to be specialized for a specific task different from the large language model. Two or more models including one large language model 162 and one or more task-specific models 161 are stored in the model storage unit 106. Hereinafter, the large language model 162 and the task-specific model 161 will be described.
[0051] [Large Language Model] The large language model 162 is a general-purpose natural language processing model and is also called an LLM (Large Language Model). When the dialogue system 1 creates a response message using the output data of the large language model 162, in the response message, the needs of the dialogue partner (customer) 3 can be expressed using expressions not included in the dialogue between the user (salesperson) 2 and the dialogue partner (customer) 3. That is, the dialogue system 1 can provide at least one of a message indicating an understanding of the needs of the dialogue partner (customer) 3 (for example, potential needs) and a message indicating a question regarding the needs of the dialogue partner (customer) 3 (for example, a question to obtain additional information).
[0052] [Task-Specific Model] The task-specialized model 161 is a model that is machine-learned to be specialized for a specific task (for example, a specific task includes specific operations such as recommending a commercial product as a solution to the needs of the conversation partner (customer) 3. Note that the recommendation of a commercial product is just an example, and specific tasks also include tasks related to organizations such as companies and industry operations. In addition, tasks include not only recommendations but also actions such as providing information about commercial products and companies, providing information on search results, and providing commercial product information such as how to use commercial products).
[0053] For example, the task-specialized model 161 is a recommendation engine that searches for and recommends commercial products. For example, the task-specialized model 161 may be a recommendation engine for each type of commercial product to be recommended (for example, a recommendation engine for recommending products, a recommendation engine for recommending services). For example, the task-specialized model 161 may be a recommendation engine for each basis of recommendation (for example, a recommendation engine for recommending based on the conversation history, a recommendation engine for recommending based on past purchase history). When the dialogue system 1 creates a response message using the output data of the task-specialized model 161 such as a recommendation engine, the user (salesperson) 2 can provide a message indicating commercial products that can be proposed to the conversation partner (customer) 3.
[0054] Note that the task-specialized model 161 includes not only a recommendation engine but also a model for searching or referring to information used when the user (salesperson) 2 creates a business daily report.
[0055] [Selection of Model Based on Conversation History and Customer Information] The dialogue control unit 102 can select one or more models from among two or more models stored in the model storage unit 106 based on not only the conversation history but also the information of the conversation partner (customer) 3 (for example, past purchase history (note that a database indicating the past purchase history may be used, or the purchase history may be determined based on the conversation history)).
[0056] [Selection of a model for recommending products other than those recommended by the salesperson to the customer] The dialogue control unit 102 can select a recommendation engine that recommends products other than those recommended by the user (salesperson) 2 to the dialogue partner (customer) 3 from among two or more models stored in the model storage unit 106. In this case, the dialogue control unit 102 determines the products that the user (salesperson) 2 has already recommended to the dialogue partner (customer) 3 based on the dialogue history between the user (salesperson) 2 and the dialogue partner (customer) 3.
[0057] <Functional configuration> FIG. 8 is a functional block diagram of a communication support device and a terminal device according to an embodiment of the present invention.
[0058] (Functional configuration of the terminal device) The terminal device 100 realizes each functional configuration as shown in FIG. 8, for example, when the CPU 2001 included in the terminal device 100 executes a predetermined program. In the example of FIG. 8, the terminal device 100 includes a data transmission unit 421, a display unit 422, an audio output unit 423, and an operation reception unit 424, etc.
[0059] The data transmission unit 421 acquires the speech of the participants participating in the communication, the speech of the salesperson 2 participating in the business negotiation, and the speech of the customer 3. The data transmission unit 421 transmits the acquired speech (voice data) to the communication support device 10.
[0060] The display unit 422 executes display processing for displaying the video of the virtual salesperson 110 received from the communication support device 10, the recommendation information proposed to the customer 3, and the text of the speech of the salesperson 2 and the customer 3 participating in the business negotiation.
[0061] The audio output unit 423 executes audio output processing for outputting the speech of the dialogue agent received from the communication support device 10.
[0062] The operation reception unit 424 executes an operation reception process for receiving operations on the terminal device 100.
[0063] (Functional Configuration of Communication Support Device) The communication support device 10 realizes a functional configuration as shown in, for example, FIG. 8 by a computer 200 provided in the communication support device 10 executing a predetermined program stored in a storage medium. In the example of FIG. 8, the communication support device 10 has functional configurations such as an input unit 401, a voice recognition unit 402, a speaker identification unit 403, an identification unit 404, a response unit 405, a voice synthesis unit 410, a drawing unit 411, and an output unit 412. Note that at least a part of each of the above functional configurations may be realized by hardware.
[0064] Also, the communication support device 10 realizes a storage unit 413 by storage devices such as an HD 1004 and an HDD controller 1005, for example. Note that the storage unit 413 may be realized by, for example, a storage server provided outside the communication support device 10 or a cloud service.
[0065] The input unit 401 receives the input of the speech of the salesperson 2 and the speech of the customer 3 transmitted by the terminal device 100.
[0066] The voice recognition unit 402 executes known voice recognition processing on the speech received by the input unit 401 to convert the acquired speech into text. Note that when the speech in the communication input to the input unit 401 is text data, the communication support device 10 may not have the voice recognition unit 402.
[0067] The speaker identification unit 403 identifies the speech of the salesperson 2 by using known speaker recognition technology or the like, and identifies the speech other than that of the salesperson 2 as the speech of the customer 3. When the speech in the communication input to the input unit 401 is text data, the speaker identification unit 403 may determine the participant who made the speech based on the terminal device that input the text. For example, when text data is input from the terminal device 100 of the salesperson 2, the participant who made the speech may be determined as the salesperson 2, and when text data is input from another terminal device, the participant who made the speech may be determined as the customer 3.
[0068] The identification unit 404 executes an identification process for identifying the attributes of the participant identified by the speaker identification unit 403. Note that the identification unit 404 only needs to be able to identify the attributes of the participant who made the speech, and it is not essential for the speaker identification unit 403 to identify the participant.
[0069] It may be possible to set in advance in the communication support device 10 which attribute (role) the voice of the participant corresponds to, such as a customer, a salesperson, a support person, etc. In this case, the identification unit 404 can identify the attribute from the voice of the participant who made the speech. If it seems difficult to register the voice of the customer in advance, only the salesperson and the support person may be set in the communication support device 10, and the attributes of the other voices may be automatically set as those of the customer for the attributes (roles) of the participants who made the speech.
[0070] It may be possible to register in advance in the communication support device 10 the profile information of the participant (role, position, stance, specific industry, business type, occupation, or gender, etc.). In this case, the identification unit 404 may identify the attributes (such as the salesperson 2, the customer 3, etc.) of the participant who made the speech based on the profile information of the speaker identified by the speaker identification unit 403. If it seems difficult to register the profile information of the customer in advance, when the profile information of the participant who made the speech is not registered, the attributes of the participant who made the speech may be automatically set as those of the customer.
[0071] The communication support device 10 may collect voice data from various speakers in advance, analyze voice features such as voice tone, intonation, speed, and language accent from the collected voice data, and machine-learn a model for identifying attributes.
[0072] The response unit 405 executes response processing to respond to the utterance input to the input unit 401. The response unit 405 includes, for example, a determination unit 406, an intention interpretation unit 407, a generation unit 408, a control unit 409, and the like.
[0073] The determination unit 406 determines the actions of the virtual salesperson 110 according to the output result of the intention interpretation unit 407. The determination unit 406 determines the actions of the virtual salesperson 110, such as whether to propose a specific commercial material to the customer 3, whether to present candidate issues and then propose commercial materials related to the issues, or whether to engage in casual conversation. Specific examples of processing: Input: Output result of the intention interpretation unit 407. Any one of three types: "proposal of a specific commercial material", "proposal of a commercial material related to the issue after presenting candidate issues (latent needs)", "casual conversation". Processing: For example, the determination unit 406 selects an action defined by rules. · For example, when the output result of the intention interpretation unit 407 is "proposal of a specific commercial material", the determination unit 406 determines to use the model using the known recommendation logic of the model in the storage unit 413. · For example, when the output result of the intention interpretation unit 407 is "proposal of a commercial material related to the issue after presenting candidate issues (latent needs)", the determination unit 406 determines to use the general-purpose large-scale language model of the model in the storage unit 413, the model using the known recommendation logic, and the "Prompt for presenting candidate issues" in the Prompt of the storage unit 413. · For example, when the output result of the intention interpretation unit 407 is "casual conversation", the determination unit 406 determines to use the general-purpose large-scale language model of the model in the storage unit 413 and the "Prompt for casual conversation" in the Prompt of the storage unit 413. Note that the procedure for determining the process may be determined by rules, or may be determined using the general large-scale language model of the model in the storage unit 413 and the Prompt for determining the process from the Prompts in the storage unit 413.
[0074] The intention interpretation unit 407 analyzes the uttered speech texturized by the speech recognition unit 402 (hereinafter referred to as the uttered text) by a natural language processing (NLP) algorithm, and extracts, for example, the intention of the utterance, keywords, and areas of interest.
[0075] The intention interpretation unit 407 determines, for example, from the utterances of the customer 3 or the salesperson 2 whether what is required as the utterance of the virtual salesperson 110 is casual conversation, a proposal for a specific commercial material, or a dialogue for determining issues yet (a proposal for a commercial material related to the issue after presenting candidates for the issue). The intention interpretation unit 407 determines the actions of the virtual salesperson 110, such as whether to propose a specific commercial material to the customer 3, to propose a commercial material related to the issue after presenting candidates for the issue, or to engage in casual conversation, based on the uttered speech of the salesperson 2 or the uttered speech of the customer 3 texturized by the speech recognition unit 402. Examples of specific procedures · Input: Data of the dialogue history held in the storage unit 413. For example, although it is the texturized utterances of the salesperson 2 or the customer 3 or both, a dialogue related to the current dialogue is searched for and extracted from the dialogue history held in the format of FIG. 11. For example, it is searched for and extracted using the SessionID. · Process: Refer to the model and Prompt in the storage unit 413. For example, refer to the URL and API key for API connection of the general large-scale language model from the model in the storage unit 413 and the Prompt for intention interpretation from the Prompts in the storage unit 413, and input them into the model in consideration of the dialogue history searched for and extracted above in the Prompt. As the dialogue history to be considered in the Prompt, all the dialogue histories having the same Session ID may be included, or a part of them may be used. · Output: Output any one of the three types: "Proposal of specific commercial products", "Proposal of commercial products related to the issue after presenting candidates for the issue (potential needs)", and "Casual conversation".
[0076] The generation unit 408 executes the actions determined by the determination unit 406 to generate a response of the virtual salesperson 110. For example, the generation unit 408 refers to the data of the conversation history held in the storage unit 413 and generates a response with reference to the model in the storage unit 413 and the Prompt. For example, when the determination unit 406 determines to have a casual conversation, the generation unit 408 uses a large language model and a Prompt for casual conversation to generate a response to the casual conversation. For example, when the determination unit 406 determines to propose a specific commercial product, the generation unit 408 acquires information on the specific commercial product and generates a message proposing the specific commercial product based on the acquired result. As a method for acquiring information on the specific commercial product, for example, the result of searching the information in the storage unit 413 may be used, the output result of a model learned to respond to the specific commercial product may be used, or information may be acquired using the publicly known technology RAG. For example, when the determination unit 406 determines to propose a commercial product related to the issue after presenting candidates for the issue, the generation unit 408 uses the method described below to generate a message proposing the issue and the commercial product for the issue.
[0077] The control unit 409 executes a control process for controlling the dialogue agent according to the response content generated by the generation unit 408. For example, the control unit 409 outputs the language response included in the response content to the speech synthesis unit 410, outputs the non-verbal response to the drawing unit 411, and outputs the proposal materials to the output unit 412 to generate the video of the virtual salesperson 110.
[0078] The speech synthesis unit 410 executes a speech synthesis process for converting the input language response into speech by speech synthesis technology.
[0079] The drawing unit 411 executes a drawing process of drawing the dialogue agent according to the input non-verbal response. For example, the drawing unit 411 reflects expressions, gazes, postures, emotions, actions, and paralanguage, etc. in the drawing of the virtual salesperson 110 according to the non-verbal response. The drawing unit 411 also performs lip-sync drawing to move the mouth of the virtual salesperson 110 in accordance with the utterance of the virtual salesperson 110.
[0080] The output unit 412 executes an output process of outputting a video including the voice of the dialogue agent vocalized by the voice synthesis unit 410, the dialogue agent drawn by the drawing unit 411, and the proposal materials to the terminal device 100 or the like. For example, the output unit 412 transmits the video of the virtual salesperson 110 to the terminal device 100 via the communication network N.
[0081] Note that the system configuration of the communication support system 1 shown in FIG. 8 is an example. For example, the communication support system 1 may be configured by one terminal device 100 having the functional configuration of the communication support device 10 shown in FIG. 8. In this case, the terminal device 100 serves as the communication support device. Also, at least a part of each functional configuration of the communication support device 10 may be possessed by the terminal device 100. For example, the terminal device 100 may have a voice synthesis unit 410, a drawing unit 411, an output unit 412, etc.
[0082] FIG. 9 is an example of the storage unit 413 according to an embodiment of the present invention.
[0083] The "dialogue history" in FIG. 9 shows dialogue history data that holds the dialogue history. For example, the data is data held in the format of FIG. 11. For example, the data is in a format such as Json, txt, csv. For example, the data is assigned a UserID for identifying the system user and a SessionID for identifying a set of dialogues. The SessionID is, for example, a value that is assigned the same value from the start to the end of the business negotiation. It is possible to search for and extract the dialogue history related to the current dialogue using the SessionID or the like.
[0084] The "Model" in FIG. 9 indicates a model storage unit that holds various models used to implement the communication support device 10.
[0085] For example, the model is a rule-based program that performs processing according to specific rules.
[0086] For example, the model is a general-purpose large language model, such as a text generation language model called GPT-4 (Generative Pre-trained Transformer 4). For example, the feature of a general-purpose large language model is to input text information called Prompt and return the output according to the instructions written in the Prompt as text information.
[0087] For example, the model is a model specialized for making a specific response. Specific examples of models specialized for making a specific response are as follows. · A recommendation engine machine-learned to respond with commercial products · A model using known recommendation logic · For an input query for search, generate an expression for searching commercial products, and respond with commercial products containing text similar to the generated expression. Known RAG technology may be used. · A model learned to respond with commercial products highly relevant to the input query · A model obtained by fine-tuning a general-purpose large language model
[0088] Each model may be placed on the server of the communication support device 10 and can be directly called or called using API connection. When using API connection, the model storage unit may hold information such as URLs and keys necessary for the connection destination of the model.
[0089] The "Prompt" in FIG. 9 indicates a group of Prompts when operating a general-purpose large language model. Examples of Prompts that are inputs to the general-purpose large language model will be described later in FIG. 12.
[0090] [Recommendations other than the products recommended by the salesperson] Specific examples of recommending products other than those recommended by the salesperson to the customer will be described. For example, in the determination unit 406, when the output result of the intention interpretation unit 407 is "proposal of a specific product", the case where it is determined to use a model of a known recommendation logic of the model in the storage unit 413 will be described.
[0091] When generating a response using a model with a known recommendation logic in the generation unit 408 to obtain a response of the recommended product and then generating the response, it is possible to make recommendations other than the products recommended by the salesperson by following the following procedure.
[0092] Using a model with a known recommendation logic, output the recommended product, and use a general-purpose large language model, the conversation history, and a Prompt for excluding duplicate products to exclude the products already included in the conversation history from the previously output recommended products. In addition, in the process of accumulating the conversation history, at regular intervals, extract the products recommended by the salesperson from the conversation history, hold the information of the recommended products, and by referring to the information of the recommended products in the products output using a model with a known recommendation logic, it becomes possible to recommend products other than those recommended by the salesperson.
[0093] FIG. 11 is an example of the conversation history according to an embodiment of the present invention. As shown in FIG. 11, for each utterance (also referred to as a message), the date and time when the utterance was made ("utterance date and time" in FIG. 11), the person who made the utterance ("utterer" in FIG. 11), and the content of the utterance ("message" in FIG. 11) are stored. Search and extract using the UserID for identifying the system user and the SessionID for identifying the set of conversations. The SessionID is, for example, a value assigned the same from the start to the end of the business negotiation.
[0094] FIG. 12 shows an example of a Prompt (Prompt for casual conversation) of the storage unit 413 according to an embodiment of the present invention. It shows an example of a Prompt used when generating a response of the generation unit 408. For example, when the determination unit 406 determines to have a casual conversation and determines to use a general large-scale language model and a Prompt for casual conversation to generate a response, it shows an example of the Prompt to be used. In the column of the conversation history of the Prompt, "Salesperson: It's a nice day today, isn't it? Customer: Yes, it is." is the part where the conversation related to the SessionID of the current conversation in FIG. 11 is extracted and the text arranged in chronological order is input. By inputting this Prompt into the general large-scale language model, a response is generated. There are multiple Prompts. The response is generated in such a way that the Prompt selected by the determination unit 406 is input into the general large-scale language model selected by the determination unit 406.
[0095] <Method> FIG. 13 is a flowchart showing the dialogue control process according to an embodiment of the present invention.
[0096] In step 101 (S101), the dialogue control unit 102 acquires the dialogue history before a predetermined inquiry message (for example, "Alfred" in FIG. 2) from the dialogue history stored in the dialogue history storage unit 104.
[0097] In step 102 (S102), the dialogue control unit 102 selects a model corresponding to the intention of the dialogue acquired in S101 from the models (specifically, two or more models including one general large-scale language model 162 and one or more task-specific models 161) stored in the model storage unit 106.
[0098] In step 103 (S103), the dialogue control unit 102 estimates the needs of the dialogue partner (customer) 3. Specifically, the dialogue control unit 102 uses the model selected in S102 to output recommendation information (detailed information about the commercial product) of the commercial product to be recommended to the dialogue partner (customer) 3, understanding of the needs of the dialogue partner (customer) 3 (for example, potential needs), questions regarding the needs of the dialogue partner (customer) 3 (for example, questions to obtain additional information), etc.
[0099] In step 104 (S104), the dialogue control unit 102 creates a response message using the output data in S103.
[0100] FIG. 14 is a flowchart showing the dialogue control process (selection of a recommendation engine and a large language model) according to an embodiment of the present invention.
[0101] In step 201 (S201), the dialogue control unit 102 acquires the dialogue history before a predetermined inquiry message (for example, "Alfred" in FIG. 2) from the dialogue history stored in the dialogue history storage unit 104.
[0102] In step 202 (S202), the dialogue control unit 102 estimates the intention of the dialogue acquired in S201 (that is, the intention of the human to inquire of the dialogue system 1).
[0103] When the task-specific model 161 is a recommendation engine, the intention of the dialogue may be classified into three types: "requesting only recommendation information (for example, in the case of a dialogue requesting a response to a question about a commercial product ("question and answer" in FIG. 14))", "requesting both recommendation information and other information other than the recommendation information (for example, in the case of a dialogue about a recommendation but without specifying a commercial product ("problem solving" in FIG. 14))", and "requesting only other information other than the recommendation information (such as a response to casual conversation) ("other (casual conversation)" in FIG. 14)".
[0104] In step 203 (S203), the dialogue control unit 102 determines which of the intentions estimated in S202 is the case. If the intention of the dialogue is "Others (casual conversation)", the process proceeds to step 204. If the intention of the dialogue is "Question and answer", the process proceeds to step 206. If the intention of the dialogue is "Problem solving", the process proceeds to step 208. For example, the dialogue control unit 102 may determine whether the intention of the dialogue is an intention for which the task-specific model 161 should be used, and if it is not an intention for which the task-specific model 161 should be used, may use the large language model 162 instead.
[0105] For example, in S203, the dialogue control unit 102 can determine the intention by voice command recognition. For example, based on the voice of a salesperson, if it is "Let's have a casual conversation", the process proceeds to S204. If it is "Search for the product I'm about to mention", the process proceeds to S206. If it is "Are there any other good products?", the process proceeds to S208.
[0106] In step 204 (S204), the dialogue control unit 102 selects a model corresponding to the intention of the dialogue determined in S203 (that is, a large language model corresponding to the intention of the dialogue "Others (casual conversation)") from the models stored in the model storage unit 106.
[0107] In step 205 (S205), the dialogue control unit 102 uses the large language model selected in S204 to output at least one of potential needs and questions for obtaining additional information.
[0108] In step 206 (S206), the dialogue control unit 102 selects a model corresponding to the intention of the dialogue determined in S203 (that is, a recommendation engine corresponding to the intention of the dialogue "Question and answer") from the models stored in the model storage unit 106.
[0109] In step 207 (S207), the dialogue control unit 102 uses the recommendation engine selected in S206 to output recommendation information.
[0110] In step 208 (S208), the dialogue control unit 102 selects, from the models stored in the model storage unit 106, a model corresponding to the intention of the dialogue determined in S203 (that is, a recommendation engine and a large language model corresponding to the dialogue intention of "problem solving").
[0111] In step 209 (S209), the dialogue control unit 102 causes the recommendation engine selected in S208 to output recommendation information, and uses the large language model selected in S208 to output at least one of potential needs and questions for obtaining additional information.
[0112] In step 210 (S210), the dialogue control unit 102 creates a response message using the output data in S205 or S207 or S209. For example, the dialogue control unit 102 may create a response message using a template, may create a response message based on rules, or may create a response message using a large language model.
[0113] FIG. 15 is a flowchart showing the processing in the case of question and answer according to an embodiment of the present invention (S207 in FIG. 14).
[0114] In step 271 (S271), the dialogue control unit 102 acquires the dialogue history.
[0115] In step 272 (S272), the dialogue control unit 102 extracts the problems of customer 3 from the dialogue history acquired in S271.
[0116] In step 273 (S273), the dialogue control unit 102 searches for a product (specific merchandise) for solving the problem extracted in S272 using a recommendation engine.
[0117] In step 274 (S274), the dialogue control unit 102 determines the presence or absence of a product. If a product is found, it proceeds to step 275, and if no product is found, it returns to step 272.
[0118] In step 275 (S275), the dialogue control unit 102 outputs a product.
[0119] For example, the dialogue example in FIG. 15 is as follows. · Customer 3: I want to digitize invoices and delivery notes. · Salesperson 2: Isn't there a product that suits this customer? · Virtual Sales 110: How about "Document Electronic Storage Service"?
[0120] FIG. 16 is a flowchart showing the processing in the case of problem solving according to an embodiment of the present invention (S209 in FIG. 14).
[0121] In step 291 (S291), the dialogue control unit 102 acquires the dialogue history.
[0122] In step 292 (S292), the dialogue control unit 102 extracts the issues of Customer 3 from the dialogue history acquired in S291.
[0123] In step 293 (S293), the dialogue control unit 102 predicts potential issues from the dialogue history acquired in S291.
[0124] In step 294 (S294), the dialogue control unit 102 generates a solution for solving the potential issues in S293.
[0125] In step 295 (S295), the dialogue control unit 102 searches for a product (specific merchandise) for solving the issues extracted in S292 using the recommendation engine.
[0126] In step 296 (S296), the dialogue control unit 102 determines the availability of the product. If the product is found, it proceeds to step 297; if the product is not found, it returns to step 294.
[0127] In step 297 (S297), the dialogue control unit 102 outputs products and solutions.
[0128] For example, the dialogue example in FIG. 16 is as follows. · Salesperson 2: Are there any other good products besides Alfred? · Virtual Sales 110: As far as I can tell from what you've said so far, don't you have the problem of "the in-house training program is not sufficient and there are issues in the training of new employees"? · Customer 3: Come to think of it, I seem to remember the personnel saying something like that... · Virtual Sales 110: Then, may I propose this product. "[Sales Manual] Scrum P_Online Training Session Package_For Resellers". With this product, we expect to "introduce a more comprehensive in-house training program and improve the quality of trainers".
[0129] FIG. 17 is an example of the intention of a dialogue according to an embodiment of the present invention.
[0130] For example, when the message included in the dialogue history or the summary of the message is "Do you have any materials about Product A?" (that is, in the case of a dialogue asking for a response to a question about a product), the intention of the dialogue is "Question and Answer".
[0131] For example, when the message included in the dialogue history or the summary of the message is "I take pictures of the equipment when entering and leaving, and check them when leaving, but it takes time to search for the pictures." (that is, in the case of a dialogue about a recommendation but not specifying a product), the intention of the dialogue is "Problem Solving".
[0132] For example, when the message included in the dialogue history or the summary of the message is "Good morning." (that is, in the case of casual conversation), the intention of the dialogue is "Other (Casual Conversation)".
[0133] FIG. 18 is a sequence diagram showing the dialogue control process according to an embodiment of the present invention.
[0134] In step 301 (S301), the dialogue acquisition unit 101 receives a predetermined query message (for example, "Alfred" in FIG. 2) from the user terminal 12.
[0135] In step 302 (S302), the dialogue control unit 102 acquires the dialogue history before the query message in S301 from the dialogue history stored in the dialogue history storage unit 104.
[0136] In step 303 (S303), the dialogue control unit 102 estimates the intention of the dialogue acquired in S302. For example, the dialogue control unit 102 uses the intention estimation model stored in the intention estimation model storage unit 105 to output the intention of the dialogue.
[0137] In step 304 (S304), the dialogue control unit 102 selects a model corresponding to the intention of the dialogue estimated in S303 from the models stored in the model storage unit 106 (specifically, two or more models including one large language model and one or more task-specific models).
[0138] In step 305 (S305), the dialogue control unit 102 creates input data for input to the model selected in S304. For example, the dialogue control unit 102 creates a search query for input to the recommendation engine. For example, the dialogue control unit 102 creates a prompt for input to the large language model.
[0139] In step 306a (S306a), the dialogue control unit 102 inputs the input data (in this case, the search query) created in S305 to the model (in this case, the task-specific model (assumed to be the recommendation engine) 161) selected in S304 to output recommendation information.
[0140] In step 306b (S306b), the dialogue control unit 102 inputs the input data (in this case, the prompt) created in S305 into the model selected in S304 (in this case, the large language model 162) to output at least one of potential needs and questions for obtaining additional information.
[0141] In step 307 (S307), the dialogue control unit 102 creates a response message using the output data from S306a and S306b.
[0142] In step 308 (S308), the response unit 103 transmits the response message created in S307 to the user terminal 12.
[0143] <Effect> As described above, in one embodiment of the present invention, for general topics such as casual conversations, it is suitable to select and respond using a large language model. For topics related to specific tasks such as product recommendations, it is suitable to select and respond using a task-specific model specialized for the specific task. Therefore, according to the content of the dialogue between the user and the interlocutor, the dialogue agent can give an appropriate response. Further, in one embodiment of the present invention, by selecting both the large language model and the task-specific model and having the dialogue agent respond, the dialogue system 1 can not only provide information on the products proposed by the user (salesperson) to the interlocutor (customer), but also provide the potential needs of the interlocutor (customer). Therefore, the potential needs of the interlocutor (customer) can be fully explored to expand business negotiation opportunities.
[0144] Each function of the embodiment described above can be realized by one or more processing circuits. Here, the "processing circuit" in this specification refers to a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, an ASIC (Application Specific Integrated Circuit) designed to execute each function described above, a DSP (digital signal processor), an FPGA (field programmable gate array), and devices such as conventional circuit modules.
Explanation of Signs
[0145] 1 Interactive system, communication support system 2 User (salesperson, host) 3 Interaction partner (customer, guest) 11 Information processing device (server, communication support device) 12 User terminal 100 Terminal device 101 Interaction acquisition unit 102 Interaction control unit 103 Response unit 104 Interaction history storage unit 105 Intention estimation model storage unit 106 Model storage unit 161 Task specialization model 162 Large language model 401 Input unit 402 Speech recognition unit 403 Speaker identification unit 404 Identification unit 405 Response unit 406 Decision unit 407 Intention interpretation unit 408 Generation unit 409 Control unit 408 Generation unit 409 Control unit 410 Speech synthesis unit 411 Drawing unit 412 Output unit 421 Data transmission unit 422 Display unit 423 Audio output unit 424 Operation reception unit
Prior art documents
Patent documents
[0146]
Patent Document 1
Claims
1. a model storage unit that stores a plurality of models including a large-scale language model and a task-specific model that is machine-learned to specialize in a specific task different from the large-scale language model; a dialogue history storage unit for storing a history of dialogues in which the dialogue agent participates; a dialogue control unit that selects one or more models from the plurality of models based on the dialogue history, and creates a response message for the dialogue agent using output data of the selected model; An information processing device comprising:
2. The information processing apparatus according to claim 1 , wherein the dialogue control unit estimates an intention of the dialogue based on a history of the dialogue, and selects one or more models from the plurality of models based on the intention.
3. The information processing apparatus according to claim 1 , wherein the dialogue history is either a history of dialogue between users or a history of dialogue between the user and the dialogue agent.
4. The information processing device according to claim 1 , wherein the task specific model is a recommendation engine.
5. The information processing device according to claim 4 , wherein the dialogue agent is a virtual sales agent that supports business negotiations, and the recommendation engine is a recommendation engine that recommends products.
6. The interaction is a conversation between a sales representative and a customer, The information processing apparatus according to claim 1 , wherein the dialogue control unit selects one or more models from the plurality of models based on information about the customer in addition to the dialogue history.
7. The interaction is a conversation between a sales representative and a customer, The information processing apparatus according to claim 1 , wherein the dialogue control unit selects, from the plurality of models, a recommendation engine that recommends a product other than the product recommended by the sales representative to the customer.
8. A dialogue acquisition unit that receives voice data of a dialogue from a user terminal; a response unit that transmits voice data of the response message to the user terminal; The information processing device according to claim 1 , further comprising:
9. The dialogue control unit includes: Selecting one or more models from the plurality of models based on an intention of the dialogue, the intention being any one of a proposal for a specific product, a proposal for a product related to the problem after presenting a candidate for the problem, and casual conversation; If the intent is to propose a specific product, create a response message using the task-specific model, the response message including a product for solving the customer's problem; If the intention is to present a candidate problem and then propose a product related to the candidate problem, using the task specific model and the large-scale language model, create a response message including a product for solving the candidate problem; The information processing apparatus according to claim 1 , wherein, when the intention is chatting, a response message including chatting is created using the large-scale language model.
10. a model storage unit that stores a plurality of models including a large-scale language model and a task-specific model that is machine-learned to specialize in a specific task different from the large-scale language model; A method executed by an information processing device including a dialogue history storage unit that stores a history of dialogues in which the dialogue agent participates, The method includes selecting one or more models from the plurality of models based on the dialogue history, and generating a response message for the dialogue agent using output data of the selected models.
11. a model storage unit that stores a plurality of models including a large-scale language model and a task-specific model that is machine-learned to specialize in a specific task different from the large-scale language model; An information processing device including a dialogue history storage unit that stores a history of dialogues in which the dialogue agent participates, a program for executing a process of selecting one or more models from the plurality of models based on the dialogue history, and creating a response message for the dialogue agent using output data of the selected models;
12. An interactive system including an information processing device and a user terminal, The information processing device includes: a model storage unit that stores a plurality of models including a large-scale language model and a task-specific model that is machine-learned to specialize in a specific task different from the large-scale language model; a dialogue history storage unit for storing a history of dialogues in which the dialogue agent participates; A dialogue acquisition unit that receives dialogue data from the user terminal; a dialogue control unit that selects one or more models from the plurality of models based on the dialogue history, and creates a response message for the dialogue agent using output data of the selected model; a response unit for transmitting the response message to the user terminal; A dialogue system with
Citation Information
Patent Citations
Dialogue support system, dialogue support method, and dialogue support program
JP2018185561A