Information processing apparatus and method, computer program product, and dialog system
By designing the model storage unit, the dialogue history storage unit and the dialogue control unit in the dialogue system, and selecting the appropriate model to generate the response using the dialogue history, the problem of the existing dialogue system's response is solved and the response quality is improved.
Patent Information
- Application Number
- CN202411720864.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for existing dialogue systems to make consistent responses based on the dialogue content.
An information processing device is designed, including a model storage unit, a dialogue history storage unit, and a dialogue control unit. The dialogue history storage unit saves the dialogue history, and the dialogue control unit selects more than one model from a plurality of models based on the dialogue history, and creates a response message using the output data of the selected model.
It realizes appropriate responses based on the dialogue content, and improves the response quality of the dialogue system.
Smart Images

Figure CN120069882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, method, computer program product, and dialogue system. Background Art
[0002] As is well known, a dialogue system can be used for dialogue between a human and a computer. For example, information for computer responses is created using a machine learning model.
[0003] However, it has been difficult for a computer in a conventional dialogue system to make a response that matches the dialogue content. In view of this, an object of the present invention is to make an appropriate response according to the dialogue content. Summary of the Invention
[0004] An information processing apparatus according to an embodiment of the present invention includes:
[0005] a model storage unit that stores a plurality of models including a large language model and a task-specific model dedicated to a specific task different from the large language model and learned by machine learning;
[0006] a dialogue history storage unit that stores a dialogue history in which a dialogue agent participates; and
[0007] a dialogue control unit that selects one or more models from the plurality of models according to the dialogue history and creates a response message of the dialogue agent using output data of the selected model.
[0008] An effect of the present invention is that it is possible to make an appropriate response according to the dialogue content. Brief Description of the Drawings
[0009] Figure 1 is a schematic diagram of a usage scenario according to an embodiment of the present invention.
[0010] Figure 2 is an example of a dialogue schematic diagram according to an embodiment of the present invention.
[0011] Figure 3 is an example of a schematic diagram of an overall configuration according to an embodiment of the present invention.
[0012] Figure 4 is another example of a schematic diagram of an overall configuration according to an embodiment of the present invention.
[0013] Figure 5 is a hardware structure block diagram of an information processing apparatus according to an embodiment of the present invention.
[0014] Figure 6 is a hardware structure block diagram of a user terminal according to an embodiment of the present invention.
[0015] Figure 7 This is a functional block diagram of an information processing apparatus according to an embodiment of the present invention.
[0016] Figure 8 This is a functional block diagram of a sound output unit 423 and a terminal device according to an embodiment of the present invention.
[0017] Figure 9 This is a schematic diagram of an example of a storage unit according to an embodiment of the present invention.
[0018] Figure 10 This is a schematic diagram of an example of a conversation history according to an embodiment of the present invention.
[0019] Figure 11 This is a schematic diagram of an example of a conversation history according to an embodiment of the present invention.
[0020] Figure 12 This is a schematic diagram of an example of a prompt according to an embodiment of the present invention.
[0021] Figure 13 This is a flowchart of a conversation control process according to an embodiment of the present invention.
[0022] Figure 14 This is a flowchart of a conversation control process (selection of a recommendation engine and a large language model) according to an embodiment of the present invention.
[0023] Figure 15 This is a flowchart of a process during question and answer according to an embodiment of the present invention.
[0024] Figure 16 This is a flowchart of a process for solving problems according to an embodiment of the present invention.
[0025] Figure 17 This is a schematic diagram of an example of a conversation intention according to an embodiment of the present invention.
[0026] Figure 18 This is a timing diagram of a conversation control process according to an embodiment of the present invention.
[0027] Symbol Explanation
[0028] 1 Conversation system
[0029] 2 User (salesperson)
[0030] 3 Conversation partner (customer)
[0031] 11 Information processing apparatus (server, communication assistance apparatus)
[0032] 12 User terminal
[0033] 100 Terminal device
[0034] 101 Conversation acquisition unit
[0035] 102 Conversation control unit
[0036] 103 Response unit
[0037] 104 Conversation history storage unit
[0038] 105 Intention estimation model storage unit
[0039] 106 Model storage unit
[0040] 161 Task-specific model
[0041] 162 Large language model
[0042] 401 Input unit
[0043] 402 Speech recognition unit
[0044] 403 Speaker determination unit
[0045] 404 Recognition unit
[0046] 405 Response unit
[0047] 406 Decision unit
[0048] 407 Intention interpretation unit
[0049] 408 Creation unit
[0050] 409 Control unit
[0051] 410 Speech synthesis unit
[0052] 411 Drawing unit
[0053] 412 Output unit
[0054] 421 Data transmission unit
[0055] 422 Display unit
[0056] 423 Audio output unit
[0057] 424 Operation reception unit Detailed implementation mode
[0058] The embodiments of the present invention will be described below with reference to the accompanying drawings. Different from the prior art, the present invention can solve the problem of selecting a large language model and a machine learning model different from the large language model according to the conversation history and giving an appropriate response.
[0059] Figure 1It is a schematic diagram of the usage scenario involved in an embodiment of the present invention. Suppose the present invention is used in a negotiation (for example, when a salesperson 2 recommends a product to a customer 3). The overall process is described below.
[0060] In step 1 (S1), the user (salesperson) 2 asks the conversation partner (customer) 3 about the needs (also called issues) of the conversation partner 3.
[0061] In step 2 (S2), the conversation partner (customer) 3 points out the needs of the conversation partner (customer) 3 according to the question of the user (salesperson) 2 in S1.
[0062] In step 3 (S3), the user (salesperson) 2 asks the dialogue system 1 a question about the solution to the needs of the conversation partner (customer) 3.
[0063] In step 4 (S4), the dialogue system 1 outputs a response message according to the question in S3 (specifically, a message about the solution to the needs of the conversation partner (customer) 3, or the understanding of the needs (such as potential needs), or a question about the needs (such as a question for obtaining further information)).
[0064] Figure 2 It is a schematic diagram of an example conversation involved in an embodiment of the present invention.
[0065] An example is given in step 11 (S11) Figure 1 of what the user (salesperson) 2 said in S1 (for example, listening to the customer's needs, such as "What difficulties do you have?").
[0066] Step 12 (S12) is Figure 1 an example of what the conversation partner (customer) 3 said in S2 of
[0067] An example is given in step 13 (S13) Figure 1 of what the user (salesperson) 2 said in S3 (asking the dialogue system about the solution, such as "Mr. A, do you have any solutions?"). Here, it is set that the dialogue system 1 responds to a preset call message (for example, Figure 2 "Mr. A" in the example of
[0068] Step 14 (S14) shows an example of what the dialogue system 1 gave in Figure 1 S4 ("Your need is to reduce time-consuming and laborious operations and improve business efficiency. How about adopting the ABC system that can perform unified admission management?"). The dialogue system 1 can output the response message by playing sound, or by displaying text, or asFigure 2 Output sound and text simultaneously.
[0069] <Overall composition>
[0070] Figure 3 It is a schematic diagram of the overall composition (Example 1) related to an embodiment of the present invention. The dialogue system 1 includes an information processing device (server) 11 and a user terminal 12. In the description herein, the information processing device 11 and the user terminal 12 are regarded as different devices, but the information processing device 11 and the user terminal 12 can also be installed in the same device (that is, the information processing device 11 can have the functions of the user terminal 12).
[0071] 《Information processing device》
[0072] The information processing device 11 selects one or more machine learning models from two or more machine learning models according to the dialogue history, and uses the output data of the selected machine learning model to generate a response message. The information processing device 11 is composed of one or more computers. For example, the information processing device 11 is a server.
[0073] The device group described in the embodiments only represents one of the multiple computing environments for implementing the embodiments disclosed in the present invention. In some embodiments, the information processing device 11 includes multiple computing devices, such as a server cluster. The multiple computing devices are configured to communicate with each other through any type of communication link, including a network, shared memory, etc., to perform the processing disclosed herein.
[0074] 《User terminal》
[0075] The user terminal 12 sends the data of the sound emitted by a person to the information processing device 11, receives the response information generated by the information processing device 11 and outputs it (such as sound playback, text display, etc.). The user terminal 12 has a microphone function, a speaker function, and a display function. The user terminal 12 is, for example, a tablet, a smart phone, a personal computer, etc.
[0076] <System composition>
[0077] Figure 4 It is a schematic diagram of the overall composition (Example 2) related to an embodiment of the present invention. In Figure 4 this example, the dialogue system 1 includes a communication assistance device 10 and a terminal device 100 connected to a communication network N such as the Internet and a LAN (Local Area Network).
[0078] As an example of a usage scenario, during the negotiation between salesperson 2 and customer 3, the terminal device 100 that displays the virtual salesperson 110 is carried, and the negotiation is conducted with the virtual salesperson 110 interspersed. Here, the negotiation is an example of communication. The salesperson 2 is an example of a host participating in the communication, and the customer 3 is an example of a guest participating in the communication. The virtual salesperson 110 is an example of a dialogue agent. A dialogue agent is a virtual salesperson that assists in business negotiations.
[0079] The terminal device 100 is an information terminal such as a PC (Personal Computer), a tablet terminal, or a smartphone that is used by the salesperson 2. The terminal device 100 acquires the speech sounds of the salesperson 2 and the customer 3 who are participants in the negotiation, and sends the acquired speech sounds (voice data) to the communication assistance device 10. The speech sound is an example of a statement in communication. The statements in communication include, for example, statements of text data such as chat. In the following description, the statements in communication are set as speech.
[0080] The communication assistance device (server device) 10 is, for example, an information processing device having a computer configuration or a system including multiple computers. The communication assistance device 10 acquires the speech sounds sent by the terminal device 100, analyzes the acquired speech sounds, and generates a response corresponding to the needs of the customer 3. The response includes, for example, providing recommended information, providing specific products, responses for small talk, etc.
[0081] Based on the content of the response, the communication assistance device 10 controls the virtual salesperson 110 displayed on the terminal device 100. The communication assistance device 10 controls the posture, gestures, mannerisms, and speech.
[0082] The terminal device 100 displays the virtual salesperson 110 controlled by the communication assistance device 10, and at the same time outputs the speech of the virtual salesperson 110. In this way, the virtual salesperson 110 can, for example, provide information on specific products corresponding to the needs of the customer 3, or provide recommended information, or provide topics such as small talk to the customer 3 and the salesperson 2 in accordance with the process of the negotiation.
[0083] For example, if it is determined that the need of customer 3 is "electronic billing", the virtual salesperson 110 will propose product suggestions related to the electronic billing system, such as "Regarding electronic billing, how about product A?" etc.
[0084] Another specific example is that the need of customer 3 is to solve the vague problem of "AI and DX are very popular recently, and we don't know what to do". In response to this, the virtual salesperson 110 determines that it is a potential need, and thus proposes "For example, is there a problem of spending time on bill processing? For this situation, product A can be recommended." etc. Based on the problems that the customer is very likely to encounter, further product suggestions related to the problem are given.
[0085] As another specific example, when salesperson 2 asks casual questions such as "It's been really hot lately. What's going on?" to virtual salesperson 110, virtual salesperson 110 replies with casual responses like "Yeah, the temperature has risen to a certain degree today. It's really hot."
[0086] <Hardware Configuration>
[0087] Figure 5 It is a hardware configuration diagram of information processing device (server) 11 according to an embodiment of the present invention.
[0088] As Figure 5 shown, information processing machine (server) 11 is composed of a computer, which includes: central processing unit 1001, ROM 1002, RAM 1003, HD 1004, HDD (Hard Disk Drive) controller 1005, display 1006, peripheral connection I / F (Interface) 1007, network I / F 1008, bus 1009, keyboard 1010, pointing device 1011, DVD-RW (Digital Versatile Disk Rewritable) drive 1013, and medium I / F 1015.
[0089] Among them, CPU 1001 controls the overall operation of information processing device (server) 11. ROM 1002 stores programs such as IPL for driving CPU 1001. RAM 1003 is used as the working area of CPU 1001. HD 1004 stores various data such as programs. HDD controller 1005 controls the reading and writing of various data on HD 1004 according to the control of CPU 1001. Display 1006 displays various information such as cursor, menu, window, text, or image. Peripheral connection I / F 1007 is an interface for connecting various external devices. In this case, the peripherals are, for example, USB (Universal Serial Bus) memory or printer. Network I / F 1008 is an interface for data communication using a communication network. Bus 1009 is used for electrically connecting Figure 5 the address bus and data bus of each component such as CPU 1001 shown above.
[0090] The keyboard 1010 is an input device having a plurality of keys for inputting characters, numerical values, various instructions, etc. The pointing device 1011 is an input device for selecting and executing various instructions, selecting a processing object, moving a cursor, etc. The DVD-RW drive 1013 controls the reading and writing of various data to and from the DVD-RW 1012, which is a removable recording medium as an example. In addition to the DVD-RW, it may also be a DVD-R or the like. The medium I / F 1015 controls the reading and writing (storage) of data to and from the recording medium 1014 such as a flash memory.
[0091] Figure 6 It is a hardware structure block diagram of the user terminal 12 related to an embodiment of the present invention.
[0092] As Figure 6 shown, the user terminal 12 includes a CPU 2001, a ROM 2002, a RAM 2003, an EEPROM 2004, a CMOS sensor 2005, a camera element I / F 2006, an acceleration azimuth sensor 2007, a medium I / F 2009, and a GPS receiver 2011.
[0093] Among them, the CPU 2001 controls the overall operation of the user terminal 12. The ROM 2002 stores programs such as IPL for driving the CPU 2001. The RAM 2003 is used as the working area of the CPU 2001. The EEPROM 2004 reads and writes various data such as smartphone programs under the control of the CPU 2001. The CMOS (Complementary Metal Oxide Semiconductor) sensor 2005 is an in-built imaging device that obtains image data by photographing a subject (mainly its own image) under the control of the CPU 2001. In addition to the CMOS sensor, it may also be an imaging device such as a CCD (Charge Coupled Device) sensor. The camera element I / F 2006 is a circuit that controls the driving of the CMOS sensor 2005. The acceleration azimuth sensor 2007 is various sensors such as a magnetic compass for detecting geomagnetism, a gyrocompass, and an acceleration sensor. The medium I / F 2009 controls the reading and writing of data to and from the recording medium 2008 such as a flash memory. The GPS receiver 2011 receives GPS signals from GPS satellites.
[0094] The user terminal 12 further includes a long-distance communication circuit 2012, a CMOS sensor 2013, a camera element I / F 2014, a microphone 2015, a speaker 2016, a sound input / output I / F 2017, a display 2018, a peripheral connection I / F (Interface) 2019, a short-distance communication circuit 2020, an antenna 2020a of the short-distance communication circuit 2020, and a touch panel 2021.
[0095] Among them, the long-distance communication circuit 2012 is a circuit that communicates with other devices through a communication network. The CMOS sensor 2013 is an in-built imaging device that captures an object according to the control of the CPU 2001 to obtain image data. The imaging element I / F 2014 is a circuit that controls the driving of the CMOS sensor 2013. The microphone 2015 is an in-built circuit that converts sound into an electrical signal. The speaker 2016 is an in-built circuit that converts an electrical signal into physical vibration to generate sounds such as music or voice. The sound input / output I / F 2017 is a circuit that processes the input / output of sound signals between the microphone 2015 and the speaker 2016 according to the control of the CPU 2001. The display 2018 is a display device such as liquid crystal and organic EL (Electro Luminescence) that displays the image of the object and various icons, etc. The peripheral connection I / F 2019 is an interface for connecting various external devices. The short-distance communication circuit 2020 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 2021 is an input device that operates the user terminal 12 by the user touching the display 2018.
[0096] The user terminal 12 also has a bus 2010. The bus 2010 is used for electrically connecting Figure 6 the address bus, data bus, etc. of each component such as the CPU 2001 shown.
[0097] <Functional configuration>
[0098] Figure 7 is a functional block diagram of the information processing apparatus 11 according to an embodiment of the present invention. The information processing apparatus 11 includes a dialogue acquisition unit 101, a dialogue control unit 102, a response unit 103, a dialogue history storage unit 104, an intention estimation model storage unit 105, and a model storage unit 106. The information processing apparatus 11 functions as the dialogue acquisition unit 101, the dialogue control unit 102, and the response unit 103 by executing a program.
[0099] The dialogue acquisition unit 101 acquires dialogue data from the user terminal 12. The dialogue data is not limited to sound and can also be text. The dialogue acquisition unit 101 converts the sound data acquired from the user terminal 12 into text data and stores it in the dialogue history storage unit 104, or stores the text data acquired from the user terminal 12 in the dialogue history storage unit 104.
[0100] The dialogue history storage unit 104 stores the history of the dialogue in which the dialogue agent participates. The dialogue history storage unit 104 stores the dialogue history (that is, the text data obtained by converting the voice data obtained by the dialogue acquisition unit 101 from the user terminal 12, or the text data obtained by the dialogue acquisition unit 101 from the user terminal 12). Refer to Figure 10 , and a dialogue history example stored in the dialogue history storage unit 104 will be described.
[0101] [Dialogue History]
[0102] Figure 10 is an example of the dialogue history related to an embodiment of the present invention. As Figure 10 shown, the date and time of each utterance (hereinafter also referred to as a message) ( Figure 10 "utterance date and time"), the person who made the utterance ( Figure 10 "speaker"), and the content of the utterance ( Figure 10 "message") are stored.
[0103] The dialogue history will be described here. The dialogue history is one of the history of the dialogue between users and the history of the dialogue between the user and the dialogue agent. The dialogue history is not limited to the dialogue history between two or more people (for example, Figure 1 and Figure 2 user (salesperson) 2 and interlocutor (customer) 3), and can also be the dialogue history between a person and the dialogue system 1.
[0104] · In the case of a dialogue between people, the dialogue history is the data of "utterance date and time", "speaker", and "message (content of the utterance)" of each utterance by two or more people.
[0105] · In the case of a dialogue between a person and the dialogue system 1 (that is, when a person communicates with the information processing device 11 through the user terminal 12), the dialogue history is the data of "utterance date and time", "speaker", and "message (content of the utterance)" of each utterance by one or more people, and the data of "transmission date and time when the response message made by the information processing device 11 is sent to the user terminal 12" and "content of the response message".
[0106] Return Figure 7 to the description of. The dialogue control unit 102 selects one or more models from multiple models according to the dialogue history, and uses the output data of the selected model to generate the response message of the dialogue agent. The dialogue control unit 102 selects one or more models from two or more models stored in the model storage unit 106 according to the dialogue history stored in the dialogue history storage unit 104, and generates a response message using the output data of the selected model.
[0107] [Presumption of Dialogue Intention]
[0108] The dialogue control unit 102 infers the intention of the dialogue based on the dialogue history (for example, "only request recommendation information", "request recommendation information and other information other than recommendation information", "only request other information other than recommendation information"), and can select one or more models from two or more models stored in the model storage unit 106 according to this intention. For example, the dialogue control unit 102 can infer the dialogue intention using the intention inference model stored in the intention inference model storage unit 105 or the large language model 162 stored in the model storage unit 106 based on the dialogue or the dialogue summary. The intention of the dialogue is not limited to the intention inferred from the speech of one person, but can also be the intention inferred from the speech of two or more people.
[0109] The response unit 103 sends the response message created by the dialogue control unit 102 to the user terminal 12. The response unit 103 can either transform the text data of the response message created by the dialogue control unit 102 into voice data and send it to the user terminal 12, or send the text data of the response message created by the dialogue control unit 102 to the user terminal 12.
[0110] The intention inference model storage unit 105 stores a machine learning model (intention inference model), which is trained to output the intention of the dialogue after inputting the dialogue history (dialogue or dialogue summary (in this case, the dialogue control unit 102 creates the dialogue summary)).
[0111] The model storage unit 106 stores multiple models, including a large language model and a task-specific model that is different from the large language model and performs machine learning specifically for a specific task. The model storage unit 106 stores two or more models including one large language model 162 and one or more task-specific models 161. The large language model 162 and the task-specific model 161 are described below.
[0112] [Large Language Model]
[0113] The large language model 162 is a general natural language processing model, also known as an LLM (Large Language Model). When the dialogue system 1 creates a response message using the output information of the large language model 162, it can express the needs of the conversation partner (customer) 3 in the response message using expressions that are not included in the conversation between the user (salesperson) 2 and the conversation partner (customer) 3. That is, the dialogue system 1 can provide at least one of a message indicating an understanding of the needs of the conversation partner (customer) 3 (such as potential needs) and a message indicating a question about the needs of the conversation partner (customer) 3 (such as an additional question to obtain information).
[0114] [Task-Specific Model]
[0115] The task-specific model 161 is a model that has been specifically machine-learned for a specific task (for example, the specific task includes solutions to the needs of the conversation partner (customer) 3, specific tasks such as recommending products, etc.). The recommendation of products is just an example, and specific tasks also include tasks related to the operations of organizations such as enterprises and industries. The tasks include not only recommendations but also actions such as providing product and enterprise-related information, providing search result information, and providing product information such as the usage of products.
[0116] For example, the task-specific model 161 is a recommendation engine that can search for and recommend products. For example, the task-specific model 161 can be a recommendation engine for each recommended product (for example, a recommendation engine for recommended products, a recommendation engine for recommended services). For example, the task-specific model 161 can be a recommendation engine for each recommendation basis (for example, a recommendation engine that recommends based on the conversation history, a recommendation engine that recommends based on past purchase records). When the dialogue system 1 creates a response message using the output information of the task-specific model 161 such as a recommendation engine, it can provide information indicating products that the user (salesperson) 2 can recommend to the conversation partner (customer) 3.
[0117] In addition to the recommendation engine, the task-specific model 161 also includes models for retrieving or referring to information used by the user (salesperson) 2 when creating a business daily report.
[0118] [Model Selection Based on Conversation History and Customer Information]
[0119] In addition to the conversation history, the dialogue control unit 102 can also select one or more models from two or more models stored in the model storage unit 106 based on information about the conversation partner (customer) 3 (for example, past purchase records (which can be represented by a database of past purchase records or the purchase records can be judged based on the conversation history)).
[0120] [Selection of Recommendation Models Other than Products Recommended by the Salesperson to the Customer]
[0121] The dialogue control unit 102 can select a recommendation engine that recommends products other than those recommended by the user (salesperson) 2 to the conversation partner (customer) 3 from two or more models stored in the model storage unit 106. In this case, the dialogue control unit 102 judges the products that the user (salesperson) 2 has already recommended to the conversation partner (customer) 3 based on the conversation history between the user (salesperson) 2 and the conversation partner (customer) 3.
[0122] [Functional Configuration]
[0123] Figure 8 It is a functional block diagram of the voice output unit 423 and the terminal device related to an embodiment of the present invention.
[0124] [Functional Configuration of Terminal Device]
[0125] The terminal device 100 realizes each functional configuration as shown, for example, by the CPU 2001 included in the terminal device 100 executing a prescribed program. Figure 8 In the Figure 8 example, the terminal device 100 has a data transmission unit 421, a display unit 422, a sound output unit 423, an operation reception unit 424, etc.
[0126] The data transmission unit 421 acquires the speech sounds of the participants participating in the communication, the salesperson 2 participating in the negotiation, and the speech sounds of the customer 3. The data transmission unit 421 transmits the acquired speech sounds (sound data) to the sound output unit 42310.
[0127] The display unit 422 performs display processing for displaying the image of the virtual salesperson 110 received from the communication assistance device 10, the recommended information recommended to the customer 3, the salesperson 2 participating in the negotiation, and the content obtained by texturizing the speech sounds of the customer 3, etc.
[0128] The sound output unit 423 performs sound output processing and outputs the speech sounds of the dialogue agent received from the communication assistance device 10, etc.
[0129] The operation reception unit 424 performs operation reception processing and accepts operations on the terminal device 100.
[0130] [Functional Configuration of Communication Assistance Device]
[0131] The communication assistance device 10 realizes the functional configuration as shown, for example, by the computer included in the communication assistance device 10 executing a prescribed program stored in a storage medium. Figure 8 In the Figure 8 example, the communication assistance device 10 has various functional configurations such as an input unit 401, a sound recognition unit 402, a speaker determination unit 403, a recognition unit 404, a response unit 405, a speech synthesis unit 410, a drawing unit 411, and an output unit 412. At least a part of the above-mentioned various functional configurations can also be realized by hardware.
[0132] The communication assistance device 10 realizes the storage unit 413 through storage devices such as, for example, the HD 1004 and the HDD controller 1005. The storage unit 413 can also be realized through, for example, a storage server provided outside the communication assistance device 10 or a cloud service, etc.
[0133] The input unit 401 accepts the input of the speech sounds of the salesperson 2 and the speech sounds of the customer 3 sent by the terminal device 100.
[0134] The speech recognition unit 402 performs known speech recognition processing on the speech sound received by the input unit 401, and forms the obtained speech sound into text. When the speech in the communication input by the input unit 401 is text data, the communication assistance device 10 may not have the speech recognition unit 402 either.
[0135] The speaker determination unit 403 determines the speech of the salesperson 2 through known speaker recognition technology or the like, and determines the speech other than the speech of the salesperson 2 as the speech of the customer 3. When the speech in the communication input by the input unit 401 is text data, the speaker determination unit 403 may also determine the participants in the speech according to the terminal device that input the text. For example, when text data is input from the terminal device 100 of the salesperson 2, the participant in the speech may also be determined as the salesperson 2, and when text data is input from other terminal devices, the participant in the speech is determined as the customer 3.
[0136] The recognition unit 404 performs recognition processing to recognize the attributes of the participants determined by the speaker determination unit 403. The recognition unit 404 only needs to be able to recognize the attributes of the participants in the speech, and the speaker determination unit 403 does not necessarily have to determine the participants.
[0137] It is also possible to preset in the communication assistance device 10 which attribute (task) the voices of the participants conform to, such as customer, salesperson, assistant, etc. In this case, the recognition unit 404 can recognize the attributes according to the voices of the participants in the speech. If it is difficult to pre-register the customer voice, it is also possible to only set the salesperson and assistant in the communication assistance device 10, and automatically set the attribute of the participant in the speech as the customer attribute (task) for voices other than these.
[0138] It is also possible to pre-register the profile information (task, position, stance, specific industry, business type, occupation, or gender, etc.) of the participants in the communication assistance device 10. In this case, the recognition unit 404 can also recognize the attributes (such as the salesperson 2, the customer 3, etc.) of the participants in the speech according to the profile information of the speakers determined by the speaker determination unit 403. If it is difficult to pre-register the profile information of the customers, it is also possible to automatically set the attributes of the participants in the speech in the customer attribute (task) when the profile information of the participants in the speech is not logged.
[0139] The communication assistance device 10 can also pre-collect voice data from various speakers, analyze voice characteristics such as the pitch, intonation, speed, and stress of the language of the voice from the collected voice data, and let the model for recognizing attributes perform machine learning.
[0140] The response unit 405 performs response processing and gives a response to the speech input to the input unit 401. The response unit 405 includes, for example, a decision unit 406, an intent interpretation unit 407, a production unit 408, a control unit 409, and so on.
[0141] Based on the output result of the intent interpretation unit 407, the decision unit 406 decides the actions of the virtual salesperson 110. The decision unit 406 decides the actions of the virtual salesperson 110, such as suggesting a specific product to the customer 3, making a product recommendation related to the topic based on the presented candidate topics, or making small talk, etc. Specific examples of the processing:
[0142] Input: The output result of the intent interpretation unit 407. Any one of the three types: "suggestion of a specific product", "giving a suggestion related to the topic (potential demand) based on the presented candidate products", and "making small talk".
[0143] Processing: For example, the decision unit 406 selects an action determined by rules.
[0144] · For example, when the output result of the intent interpretation part 407 is "suggestion of a specific product", the decision unit 406 decides to use the model in the storage unit 413 that utilizes well-known recommendation logic.
[0145] · For example, when the output result of the intent interpretation part 407 is "giving a suggestion related to the topic (potential demand) based on the presented candidate topics", the decision unit 406 decides to adopt the model that uses a general large language model and well-known recommendation logic in the storage unit 413, as well as the "prompt for presenting candidate topics" in the prompts of the storage unit 413.
[0146] · For example, when the output result of the intent interpretation part 407 is "making small talk", the decision unit 406 decides to utilize the "prompt for making small talk" that uses a general large language model in the storage unit 413 and the prompts of the storage unit 413.
[0147] The steps of the decision-making process can be determined by rules, or by using the general large language model of the model in the storage unit 413 and the prompt for determining the process from the prompts in the storage unit 413.
[0148] The intent interpretation unit 407 analyzes the speech sound (hereinafter referred to as the speech text) converted into text by the voice recognition unit 402 through natural language processing (NLP: Natural Language Processing) algorithms, and extracts, for example, the speech intent, keywords, and areas of interest.
[0149] The intention interpretation unit 407 determines whether the content required for the speech of the virtual salesperson 110 is casual conversation, advice on specific products, or a conversation for determining issues (based on the extraction of candidate issues, providing advice on products related to the issues) according to the speeches of, for example, the customer 3 and the salesperson 2. The intention interpretation unit 407 decides whether the action of the virtual salesperson 110 is to recommend a specific product to the customer 3, or to provide advice on products related to the topic based on the extraction of candidate issues, or to have a casual conversation, etc., according to the speech voice of the salesperson 2 or the speech voice of the customer 3 converted into text by the voice recognition unit 402. An example of the specific steps is as follows.
[0150] · Input: It is the data of the conversation history stored in the storage unit 413. For example, although it is the speech of the salesperson 2 or the customer 3 or both converted into text, the conversation related to the current conversation is retrieved and extracted from the conversation history stored in the Figure 11 form. For example, it is retrieved and extracted using the Session ID.
[0151] · Processing: Refer to the model and prompts in the storage unit 413. For example, from the model in the storage unit 413, refer to the URL and API key for connecting to the general large-scale language model API, and from the prompts in the storage unit 413, refer to the prompts for intention interpretation. Add the above retrieved and extracted conversation history to the prompts and input them into the model. The conversation history added to the prompts can be either all the conversation history with the same Session ID or a part of it.
[0152] · Output: Output any one of the three types: "Advice on specific products", "Advice on products (potential needs) related to the topic based on the proposed candidate issues", and "Casual conversation".
[0153] The production unit 408 executes the action determined by the determination unit 406 and produces the response of the virtual salesperson 110. For example, the production unit 408 refers to the data of the conversation history stored in the storage unit 413 and refers to the model and prompts in the storage unit 413 to produce the response. For example, when the determination unit 406 decides to have a casual conversation, the production unit 408 uses the large-scale language model and casual conversation prompts to produce the response for the casual conversation. For example, when the determination unit 406 decides to give advice on specific products, the production unit 408 obtains the information of the specific products and generates a message for giving advice on specific products based on the obtained results. Regarding the method of obtaining information on specific products, it can utilize, for example, the result of information retrieval in the storage unit 413, or the output result of the model learned for responding to specific products, or the well-known RAG to obtain information. For example, when the determination unit 406 decides to provide advice on products related to the topic based on the proposed candidate issues, the production unit 408 uses the method described below to produce the message of the topic and the advice on products related to the topic.
[0154] Based on the response content produced by the production unit 408, the control unit 409 executes control processing for the dialogue agent. For example, the control unit 409 outputs the verbal response included in the response content to the speech synthesis unit 410, outputs the non-verbal response to the rendering unit 411, and outputs the recommended materials to the output unit 412 to produce an image of the virtual salesperson 110.
[0155] The speech synthesis unit 410 performs speech synthesis processing to convert the input verbal response into sound through speech synthesis technology.
[0156] Based on the input non-verbal response, the rendering unit 411 executes rendering processing for the dialogue agent. For example, based on the non-verbal response, the rendering unit 411 reflects expressions, gaze, postures, emotions, mannerisms, and paralanguage, etc. in the rendering of the virtual salesperson 110. The rendering unit 411 also performs lip-sync rendering, etc. in coordination with the voice output of the virtual salesperson 110 to make the mouth of the virtual salesperson 110 move.
[0157] The output unit 412 performs output processing to output the image of the dialogue agent whose voice has been converted into sound by the speech synthesis unit 410, the dialogue agent rendered by the rendering unit 411, and the recommended materials to the terminal device 100, etc. For example, the output unit 412 sends the image of the virtual salesperson 110 to the terminal device 100 via the communication network N.
[0158] Figure 8 An example of the system configuration of the dialogue system 1 shown, for example, the dialogue system 1 can also be composed of a single terminal device 100 having the Figure 8 functions of the communication assistance device 10 shown. In this case, the terminal device 100 becomes the communication assistance device. At least a part of each functional component of the communication assistance device 10 can also be provided by the terminal device 100. For example, the terminal device 100 can have a speech synthesis unit 410, a rendering unit 411, and an output unit 412, etc.
[0159] Figure 9 An example of the storage unit 413 related to an embodiment of the present invention.
[0160] Figure 9 The "dialogue history" represents dialogue history data that holds the dialogue history. For example, the data is in the Figure 11 form of data saved. The data can be in formats such as Json, txt, csv, etc. For example, a User ID for identifying the system user and a Session ID for identifying the dialogue set are assigned to the data. The Session ID is assigned the same value, for example, from the start to the end of the negotiation. Using the Session ID, etc., the dialogue history related to the current dialogue can be retrieved and extracted.
[0161] Figure 9 The "model" refers to the model storage unit, which holds various models used to implement the communication assistance device 10.
[0162] For example, the model is a rule-based program that performs operations according to specific rules.
[0163] For example, the model is a general large language model, such as the text creation language model called GPT-4 (Generative Pre-trained Transformer 4). For example, the feature of a general large language model is that it inputs text information called a prompt and returns the output as text information according to the instructions written in the prompt.
[0164] For example, the model is a model dedicated to performing specific responses. Specific examples of models that have been specially processed for dedicated responses are as follows.
[0165] · A recommendation engine that has undergone machine learning to respond to products
[0166] · A model that adopts a well-known recommendation logic
[0167] · Create an expression for retrieving products for the input query of the retrieval, and respond with products that contain text similar to the created expression. It is also possible to use the well-known RAG technology.
[0168] · A model that has been learned to respond to products with a high degree of relevance to the question for the input query
[0169] · A model obtained by fine-tuning a general large language model
[0170] Each model can be placed on the server of the communication assistance device 10 and can be directly called or called using API connection. When making an API connection, the model storage unit can also adopt a form that holds information such as the URL or key required for the connection object of the model.
[0171] Figure 9 The "prompt" refers to the group of prompts when running a general large-scale language model. Examples of prompts input into the general large-scale language model will be described with reference to Figure 12 description.
[0172] [Recommendations for products other than those already recommended by the salesperson]
[0173] Here, specific examples of recommendations for products other than those already recommended by the salesperson to the customer are described. For example, in the case where the output result of the intention interpretation unit 407 is "recommendation of specific products", an example when the decision unit 406 decides to use a model that uses a well-known recommendation logic in the models of the storage unit 413.
[0174] When the production department 408 creates a response based on the response of the recommended product obtained by using a model with a known recommendation logic, the following steps can be used to recommend products other than those already recommended by the salesperson when creating the response.
[0175] Using a model with a known recommendation logic, output recommended products, and use a general large-scale language model, conversation history, and prompts for excluding duplicate products to exclude products included in the conversation history from the previously output recommended products. In addition, during the process of accumulating the conversation history, every certain period of time, extract the products recommended by the salesperson from the conversation history, maintain the information of the recommended products, and refer to the information of the recommended products from the recommended products output by using a model with a known recommendation logic, so as to recommend products other than those already recommended by the salesperson.
[0176] Figure 11 This is an example of the conversation history related to an embodiment of the present invention. As Figure 11 shown, save the date and time of each speech (also called a message) ( Figure 11 "Speech Date and Time"), the person who made the speech ( Figure 11 "Speaker"), and the content of the speech ( Figure 11 "Message"). Use the user ID of the recognition system user and the Session ID for recognizing the conversation set for retrieval and extraction. The Session ID is given the same value from the start to the end of the negotiation, for example.
[0177] Figure 12 This is an example of the prompt (prompt during chatting) of the storage unit 413 related to an embodiment of the present invention. This figure shows an example of the prompt used when the production department 408 creates a response. For example, it shows an example of the prompt determined by the determination unit 406 to start chatting and used when determining the response by using a general large-scale language model and a prompt for chatting. In the conversation history column of the prompt, "Salesperson: The weather is nice today. Customer: Yes." is the part of the text input in chronological order of the date and time by extracting the conversation with the Session ID related to Figure 11 this conversation. This prompt is input into the general large-scale language model to create a response. There are multiple prompts. The response is created by inputting the prompt selected by the determination unit 406 into the general large-scale language model selected by the determination unit 406.
[0178] <Method>
[0179] Figure 13 This is a flowchart of the conversation control process related to an embodiment of the present invention.
[0180] In step 101 (S101), the dialogue control unit 102 obtains the dialogue history earlier than a predetermined inquiry message (for example Figure 2 Mr. A) from the dialogue history stored in the dialogue history storage unit 104.
[0181] In step 102 (S102), the dialogue control unit 102 selects a model corresponding to the dialogue intention obtained in S101 from the models stored in the model storage unit 106 (specifically, two or more models including one large language model 162 and one or more task-specific models 161).
[0182] In step 103 (S103), the dialogue control unit 102 estimates the needs of the conversation partner (customer) 3. Specifically, the dialogue control unit 102 uses the model selected in S102 to make the conversation partner (customer) 3 output recommendation information of the recommended product (detailed information about the product), understanding of the needs of the conversation partner (customer) 3 (for example, potential needs), questions about the needs of the conversation partner (customer) 3 (for example, questions for obtaining additional information), etc.
[0183] In step 104 (S104), the dialogue control unit 102 creates a response message using the output data of S103.
[0184] Figure 14 It is a flowchart of the dialogue control process (selection of the recommendation engine and the large language model) according to an embodiment of the present invention.
[0185] In step 201 (S201), the dialogue control unit 102 obtains the dialogue history earlier than a predetermined inquiry message (for example, Figure 2 Mr. A) from the dialogue history stored in the dialogue history storage unit 104.
[0186] In step 202 (S202), the dialogue control unit 102 estimates the dialogue intention obtained in S201 (that is, the intention of the person asking the dialogue system 1).
[0187] When the task-specific model 161 is a recommendation engine, the dialogue intention can be classified into the following three types: "request only recommendation information (for example, in the case of a dialogue that is a response to a question about a product ( Figure 14 Question and Answer")), "request recommendation information and other information other than recommendation information (for example, in the case of a dialogue about a recommendation, but a dialogue without specifying a product ( Figure 14 Problem Solving")), "request only other information other than recommendation information (response to small talk, etc.) ( Figure 14 Other (Small Talk)).
[0188] In step 203 (S203), the dialogue control unit 102 determines which type of intention is the intention inferred in S202. When the intention of the dialogue is "Other (idle talk)", it proceeds to step 204. When the dialogue intention is "Question and answer", it proceeds to step 206. When the intention of the dialogue is "Problem solving", it proceeds to step 208. For example, the dialogue control unit 102 may also determine whether the dialogue intention is an intention for which the task-specific model 161 should be used. If it is not an intention for which the task-specific model 161 should be used, the large language model 162 is used.
[0189] For example, in S203, the dialogue control unit 102 can determine the intention through voice command recognition. For example, based on the voice of the salesperson, the dialogue control unit 102 proceeds to S204 in the case of "idle talk", proceeds to S206 in the case of "retrieving the currently mentioned product", and proceeds to S208 in the case of "Are there no other good products?"
[0190] In step 204 (S204), the dialogue control unit 102 selects a model corresponding to the dialogue intention determined in S203 from the models stored in the model storage unit 106 (i.e., the large language model corresponding to the dialogue intention "Other (idle talk)").
[0191] In step 205 (S205), the dialogue control unit 102 uses the large language model selected in S204 to output at least one of the potential demand and the additional question for obtaining information.
[0192] In step 206 (S206), the dialogue control unit 102 selects a model corresponding to the dialogue intention determined in S203 from the models stored in the model storage unit 106 (i.e., the recommendation engine corresponding to the dialogue intention "Question and answer").
[0193] In step 207 (S207), the dialogue control unit 102 uses the recommendation engine selected in S206 to output recommendation information.
[0194] In step 208 (S208), the dialogue control unit 102 selects a model corresponding to the dialogue intention determined in S203 from the models stored in the model storage unit 106 (i.e., the recommendation engine and the large language model corresponding to the dialogue intention "Problem solving").
[0195] In step 209 (S209), the dialogue control unit 102 uses the recommendation engine selected in S208 to output recommendation information, and uses the large language model selected in S208 to output at least one of the potential demand and the additional question for obtaining information.
[0196] In step 210 (S210), the dialogue control unit 102 uses the output data of S205 or S207 or S209 to create a response message. For example, the dialogue control unit 102 creates a response message using a template, can also create a response message based on rules, or can utilize a large language model to create a response message.
[0197] Figure 15 This is the processing flowchart during question and answer according to an embodiment of the present invention ( Figure 14 for S207).
[0198] In step 271 (S271), the dialogue control unit 102 obtains the dialogue history.
[0199] In step 272 (S272), the dialogue control unit 102 extracts the issues of customer 3 from the dialogue history obtained in S271.
[0200] In step 273 (S273), the dialogue control unit 102 uses a recommendation engine to retrieve products (specific products) for solving the issues extracted in S272.
[0201] In step 274 (S274), the dialogue control unit 102 determines whether there are products. If a product is found, it proceeds to step 275; if no product is found, it returns to step 272.
[0202] In step 275 (S275), the dialogue control unit 102 outputs the product.
[0203] For example, Figure 15 the dialogue example in
[0204] · Customer 3: Wants to digitize the bill and delivery note
[0205] · Salesperson 2: Isn't there a product suitable for this customer?
[0206] · Virtual Sales 110: How about "Voucher Electronic Preservation Service"?
[0207] Figure 16 This is the flowchart of the processing during issue resolution according to an embodiment of the present invention ( Figure 14 for S209).
[0208] In step 291 (S291), the dialogue control unit 102 obtains the dialogue history.
[0209] In step 292 (S292), the dialogue control unit 102 extracts the issues of customer 3 from the dialogue history obtained in S291.
[0210] In step 293 (S293), the dialogue control unit 102 predicts potential issues based on the dialogue history obtained in S291.
[0211] In step 294 (S294), the dialogue control unit 102 creates a solution for resolving the potential issues in S293.
[0212] In step 295 (S295), the dialogue control unit 102 retrieves a product (specific product) for resolving the issues extracted in S292 through a recommendation engine.
[0213] In step 296 (S296), the dialogue control unit 102 determines whether there is a product. If a product is found, it proceeds to step 297; if no product is found, it returns to step 294.
[0214] In step 297 (S297), the dialogue control unit 102 outputs the product and the solution.
[0215] For example, Figure 16 the dialogue example in
[0216] · Salesperson 2: Mr. A, are there any other good products?
[0217] · Virtual Salesperson 110: From the current dialogue, is the customer troubled by "the internal training plan is not perfect enough and there are problems in the training of new employees"?
[0218] · Customer 3: Come to think of it, I think the personnel department said something like that...
[0219] · Virtual Salesperson 110: Then, I recommend this product here. "[Sales Manual] Scrum P_Online Training Course Package_For Sales Stores". This product is expected to "introduce a more substantial company internal training plan and improve the quality of coaches".
[0220] Figure 17 is an example of the dialogue intention according to an embodiment of the present invention.
[0221] For example, when the information included in the dialogue history or the summary of this information is "Is there any data about product A?" (i.e., a dialogue asking for an answer to a question about a product), the intention of the dialogue is "question and answer".
[0222] For example, when the message included in the dialogue history or the summary of this message may be shown as "Photographing equipment for taking pictures when entering and leaving, used for confirmation when exiting, but it takes a long time to find the photos." (i.e., a dialogue about a recommendation but without specifying a product), the intention of the dialogue is "solving issues".
[0223] For example, when the message included in the conversation history or the summary of the message is "Good morning." (i.e., small talk), the intention of the conversation is "Other (small talk).".
[0224] Figure 18 It is a sequence diagram of the conversation control process related to an embodiment of the present invention.
[0225] In step 301 (S301), the conversation acquisition unit 101 receives a predetermined inquiry message from the user terminal 12 (for example, Figure 2 Mr. A).
[0226] In step 302 (S302), the conversation control unit 102 acquires the conversation history before the inquiry message of S301 from the conversation history stored in the conversation history storage unit 104.
[0227] In step 303 (S303), the conversation control unit 102 infers the conversation intention obtained in S302. For example, the conversation control unit 102 uses the intention inference model stored in the intention inference model storage unit 105 to output the conversation intention.
[0228] In step 304 (S304), the conversation control unit 102 selects a model corresponding to the conversation intention inferred in S303 from the models stored in the model storage unit 106 (specifically, two or more models including one large language model and one or more task-specific models).
[0229] In step 305 (S305), the conversation control unit 102 creates input data to be input into the model selected in S304. For example, the conversation control unit 102 creates a retrieval query to be input into the recommendation engine. For example, the conversation control unit 102 generates a prompt for input into the large language model.
[0230] In step 306a (S306a), the conversation control unit 102 inputs the input data generated in S305 into the model selected in S304 (here, the task-specific model (as a recommendation engine) 161) (here, the retrieval query) to make it output recommendation information.
[0231] In step 306b (S306b), the conversation control unit 102 inputs the input data generated in S305 into the model selected in S304 (here, the large language model 162) (here, the prompt) and outputs at least one of the potential demand and the additional question for obtaining information.
[0232] In step 307 (S307), the conversation control unit 102 uses the output data of S306a and S306b to create a response message.
[0233] In step 308 (S308), the response unit 103 sends the response message created in S307 to the user terminal 12.
[0234] <Effect>
[0235] Thus, in one embodiment of the present invention, for example, for general topics such as small talk, it is suitable to select a large language model for response, and for topics related to specific tasks such as product recommendations, it is suitable to select a task-specific model dedicated to specific tasks for response. Therefore, according to the conversation content between the user and the conversation partner, the dialogue agent can give appropriate responses. Additionally, in one embodiment of the present invention, by simultaneously selecting a large language model and a task-specific model and having the dialogue agent make responses, the dialogue system 1 can not only provide information on the products recommended by the user (salesperson) to the conversation partner (customer), but also provide the potential needs of the conversation partner (customer). Therefore, it can fully explore the potential needs of the conversation partner (customer) and expand business opportunities.
[0236] Each function of the above-described embodiment can be implemented by one or more processing circuits. Here, the "processing circuit" includes: a processor programmed with software to execute each function, such as a processor installed through a circuit, and devices such as an ASIC (Application Specific Integrated Circuit), a DSP (digital signal processor), an FPGA (field programmable gate array), and existing circuit modules that are designed to execute the above-described functions.
Claims
1. An information processing device, comprising: A model storage unit for storing a plurality of models including a large-scale language model and a task-specific model dedicated to a specific task different from the large-scale language model and subjected to machine learning; A conversation record storage unit, used to store conversation records participated by the conversation agent; as well as The dialogue control unit is used to select one or more models from the multiple models based on the dialogue history, and use the output data of the selected model to create a response message of the dialogue agent.
2. The information processing device according to claim 1, wherein: The dialogue control unit estimates the intention of the dialogue based on the dialogue history, and selects one or more models from the plurality of models based on the intention.
3. The information processing device according to claim 1, wherein: The conversation record is any one of a record of conversations between users and a record of conversations between the user and the conversation agent.
4. The information processing device according to claim 1, wherein: The task-specific model is a recommendation engine.
5. The information processing device according to claim 4, wherein: The dialogue agent is a virtual salesperson that assists in negotiation, and the recommendation engine is a recommendation engine that recommends products.
6. The information processing device according to claim 1, wherein: The dialogue is between a salesperson and a customer. The interaction control unit selects one or more models from the plurality of models based on the customer information in addition to the interaction history.
7. The information processing device according to claim 1, wherein: The dialogue is between a salesperson and a customer. The dialogue control unit selects a recommendation engine from the plurality of models, the recommendation engine recommending a product other than a product that the salesperson has recommended to the customer.
8. The information processing device according to claim 1, further comprising: A conversation acquisition unit, configured to receive voice data of a conversation from a user terminal; as well as The response unit is used to send the sound data of the response message to the user terminal.
9. The information processing device according to claim 1, wherein: The dialogue control unit selects one or more models from the plurality of models according to the intention of the dialogue, wherein the intention is any one of suggestion of a specific product, suggestion of a product related to the topic based on the proposed candidate topic, and small talk. When the intention is to recommend a specific product, a response message is prepared using the task-specific model, and the response includes a product for solving the customer's problem. When the intention is to suggest a product related to the candidate problem based on the proposed problem, a response message is prepared using the task-specific model and the large-scale language model, and the response message includes a product for solving the candidate problem. When the intention is to chat, a response message is produced using the large-scale language model, and the response message includes chatting.
10. A method for execution by an information processing device, The information processing device comprises: a model storage unit for storing a plurality of models including a large-scale language model and a task-specific model dedicated to a specific task different from the large-scale language model and subjected to machine learning; and A conversation record storage unit is used to store conversation records participated by the conversation agent. The method is to select one or more models from the plurality of models based on the conversation history, and to create a response message of the conversation agent using output data of the selected model.
11. A computer program product capable of being executed by an information processing device, The information processing device comprises: a model storage unit for storing a plurality of models including a large-scale language model and a task-specific model dedicated to a specific task different from the large-scale language model and subjected to machine learning; and A conversation record storage unit is used to store conversation records participated by the conversation agent. The information processing device performs processing by selecting one or more models from the plurality of models based on the conversation history, and creating a response message of the conversation agent using output data of the selected model.
12. A dialogue system, comprising an information processing device and a user terminal, wherein the information processing device has: A model storage unit for storing a plurality of models including a large-scale language model and a task-specific model dedicated to a specific task different from the large-scale language model and subjected to machine learning; A conversation record storage unit, used to store conversation records participated by the conversation agent; A conversation acquisition unit, used for receiving conversation data from a user terminal; A dialogue control unit, configured to select one or more models from the plurality of models based on the dialogue history, and to create a response message of the dialogue agent using output data of the selected model; as well as A response unit is used to send the response message to the user terminal.