Communication system
The communication system uses generative AI to determine conversation context and emotions, addressing inappropriate remarks and enhancing spontaneity and creativity in user interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NOMURA RESEARCH INSTITUTE
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Existing AI systems lack the ability to accurately grasp the user's situation and emotions during conversations, leading to inappropriate remarks, lack of spontaneity, and uncreative responses, which can offend users.
A communication system that utilizes generative AI to determine the current position and scene of a conversation based on perceptual information, predict user emotions, and adjust its actions and responses accordingly, integrating long-term memory to support appropriate communication.
Enables AI to communicate appropriately by grasping the user's situation and emotions, enhancing spontaneity and creativity, thus improving user interaction.
Smart Images

Figure 2026085635000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the utilization technology of generative AI (Artificial Intelligence), and particularly to a technology effective when applied to a communication system that supports interactive communication with users such as customers.
Background Art
[0002] In recent years, generative AI (hereinafter sometimes simply abbreviated as "AI") has been used in information processing systems, applications, etc. in various fields. For example, with respect to conversations such as questions from users such as customers, many mechanisms have been studied to automatically and autonomously perform processes such as giving responses such as answers and proposals in the form of natural conversations using generative AI.
[0003] For example, in Japanese Patent Application Laid-Open No. 2018-81444 (Patent Document 1), a plurality of dialogue agent means for providing various services are provided for each service, and by specializing each for each service, a highly specialized service is provided. On the other hand, when each dialogue agent cannot handle it by itself, it is described that the user is guided to a more suitable dialogue agent by transferring the conversation sentence to another dialogue agent.
Prior Art Documents
Patent Documents
[0004] <
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] According to the prior art as described in Patent Document 1, by providing a plurality of highly specialized AIs specialized for each service and corresponding to the AI that can handle the problem, it is said that it is possible to handle various operations and tasks.
[0006] On the other hand, even if we can provide AI that can respond to user needs, there are cases where the AI fails to understand the situation or the other person's feelings during a conversation, making inappropriate remarks that "don't read the room" and causing offense. Similarly, while the AI may answer what the user asks, it often lacks spontaneity and interactivity, meaning it is uncreative, passive, and provides a uniform response.
[0007] Therefore, the object of the present invention is to provide a communication system that accurately grasps the user's situation and supports the realization of communication appropriate to the situation. The aforementioned and other objects and novel features of the present invention will become clear from the description herein and the accompanying drawings. [Means for solving the problem]
[0008] A brief overview of some of the representative inventions disclosed in this application is as follows:
[0009] A typical embodiment of the present invention is a communication system that supports communication between a user and another party, and generates AI (Artificial Intelligence) based on content including information related to the history of the communication obtained from a long-term memory unit. The system includes: a process determination unit that uses AI to determine the process indicating the current position in the communication and records it in the short-term memory; a scene determination unit that uses a generating AI to determine a scene indicating the current state of the communication based on perceptual information consisting of audio and video data related to the communication acquired from the outside at any time and records it in the short-term memory; an action determination unit that uses a generating AI to determine the next action to be taken in the communication based on the information including the process and the scene recorded in the short-term memory and records it in the medium-term memory; a specialized processing unit that causes the corresponding generating AI to execute processing related to one or more of the actions determined by the action determination unit and records the processing results in the medium-term memory; and an output determination unit that uses a generating AI to determine whether or not to output the processing results and what content to output if so, based on the information including the process, the scene related to the communication, and the processing results in the specialized processing unit recorded in the short-term memory, the medium-term memory, and the long-term memory, and outputs according to the determination result. [Effects of the Invention]
[0010] The effects obtained by some of the representative inventions disclosed in this application can be briefly explained as follows:
[0011] In other words, according to a typical embodiment of the present invention, it is possible to accurately grasp the user's situation and support the realization of communication appropriate to the situation. [Brief explanation of the drawing]
[0012] [Figure 1] This figure shows an overview of an example configuration of a communication system, which is one embodiment of the present invention. [Figure 2] This figure outlines an example of what a generative AI in one embodiment of the present invention does to engage in situationally appropriate communication. [Figure 3] This figure outlines an example of communication in one embodiment of the present invention. [Figure 4] This is a processing flow diagram that outlines an example of a processing flow that uses AI to support communication in one embodiment of the present invention. [Figure 5] This is a processing flow diagram that outlines an example of a processing flow that uses AI to support communication in one embodiment of the present invention. [Figure 6] This figure outlines an example of a process determination prompt in one embodiment of the present invention. [Figure 7] This figure outlines an example of output obtained by inputting a process determination prompt in one embodiment of the present invention into a process determination AI. [Figure 8] This figure outlines an example of a scene determination prompt in one embodiment of the present invention. [Figure 9] This figure outlines an example of output obtained by inputting a scene determination prompt in one embodiment of the present invention to a scene determination AI. [Figure 10] This figure provides an overview of an example of an action determination prompt in one embodiment of the present invention. [Figure 11] This figure outlines an example of output obtained by inputting an action determination prompt in one embodiment of the present invention into an action determination AI. [Figure 12] This figure provides an overview of an example of an output determination prompt in one embodiment of the present invention. [Figure 13] This figure outlines an example of an output obtained by inputting an output determination prompt in one embodiment of the present invention into an output determination AI. [Figure 14]This is a diagram showing an overview of an example of specific data in the process list in one embodiment of the present invention. [Figure 15] This is a diagram showing an overview of an example of specific data in the scene list in one embodiment of the present invention. [Figure 16] This is a diagram showing an overview of the data configuration and an example of specific data in the short-term memory DB in one embodiment of the present invention. [Figure 17] This is a diagram showing an overview of the data configuration and an example of specific data in the medium-term memory DB in one embodiment of the present invention. [Figure 18] (a) and (b) are diagrams showing an overview of the data configuration and an example of specific data in the long-term memory DB in one embodiment of the present invention.
Mode for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention will be described in detail based on the drawings. In all the drawings for explaining the embodiments, the same parts are generally denoted by the same reference numerals, and repeated explanations thereof are omitted. On the other hand, for the parts described with reference numerals in a certain drawing, they will not be shown again in the explanations of other drawings, but may be referred to with the same reference numerals.
[0014] <Overview> For example, sales staff such as insurance diplomats need to communicate effectively with customers, which is one of the important tasks. However, due to the high stress associated with communication, the turnover rate is high and recruitment is difficult. In addition, there are many problems surrounding sales staff, such as the easy occurrence of individual ability differences, the time-consuming nature of training, and the sophistication of roles due to the diversification of customer needs.
[0015] To address these challenges, one option to consider is whether generative AI can support or replace sales staff. AI does not experience stress, does not leave the company, and requires no recruitment; it only needs to be implemented. Furthermore, it can guarantee a certain level of competence and performance, and can quickly learn diverse knowledge, making it capable of addressing the aforementioned challenges.
[0016] On the other hand, unlike humans, AI lacks imagination, the ability to judge right from wrong, and the capacity for empathy. As a result, in conversations with customers, it may fail to understand the situation or the other person's feelings, making inappropriate remarks that "don't read the room" and causing offense. Similarly, lacking a sense of responsibility, mission, and creativity, and unable to act spontaneously, AI may answer questions asked by customers, but often in a passive and uniform manner. There are things that humans can do that AI cannot, or is not good at.
[0017] Humans are able to do these things because they can grasp the current situation from various external information, make judgments based on their own experience, and possess thoughts and emotions, in other words, they have "self-awareness." Therefore, in one embodiment of the present invention, a communication system enables the AI (hereinafter sometimes referred to as "conscious AI") to realize what self-awareness does in the process of a person perceiving and acting during communication with another party, thereby enabling the AI to communicate and act in a way that is appropriate to the situation.
[0018] For AI to communicate appropriately for a given situation, it first integrates information obtained from its visual and auditory senses (images and sounds), and then predicts the situation and the other person's emotions (i.e., what kind of "scene" it is) based on past experiences. For example, in a scenario where a customer comes to an insurance counter to discuss changing their life insurance contract, the AI can determine its current position ("process") in the workflow, such as the insurance proposal stage, based on the customer's visit history, and then determine the "scene" ("setting"), such as the current greeting stage, based on external observations. In addition, it needs to understand implicit rules, such as predicting the customer's emotions (e.g., if their facial expression is gloomy, they may be feeling unwell) by comparing it to accumulated memories related to the customer, and understanding what an insurance salesperson should say to the customer.
[0019] Based on the results of these "process" and "scene" assessments, the AI predicts the actions and words it can take and the consequences they will bring. For example, if it engages in small talk with the customer mentioned above before getting down to business, it predicts how much the customer's emotions will improve or worsen. Based on this prediction, it then determines the appropriate actions and words to take ("next action"). Based on these process, scene, and next action assessments, the AI determines its own "mode" of behavior and communicates and acts accordingly.
[0020] Figure 2 is a diagram illustrating an example of what a generating AI in one embodiment of the present invention does to communicate appropriately for a given situation. First, it receives input of information obtained from the outside through sight and hearing (i.e., video and audio data acquired by a camera or microphone) (S01). Then, as perceptual processing of this information, it first receives information related to the other party (customer) (S02), and records and stores the obtained information as short-term memory (storage of current perceptual information that changes moment by moment) (S03). After that, it organizes the information by selecting, integrating, and discarding the obtained information (S04).
[0021] Next, as a process to determine the next action the AI should take, it first extracts information from its accumulated long-term memory (an accumulation of "experiences" compiled from short-term memories, etc., at regular intervals) that is similar to or related to the current situation (S05). Then, based on the current perceptual information and the retrospective information from long-term memory, the conscious AI decides the next action to take and determines the mode in which the AI should behave (S06). Finally, it calls upon an AI that performs specialized processing corresponding to the mode of the next action (sometimes referred to as the "specialized AI" below) and instructs it to proceed (S07).
[0022] In specialized processing, the specialized AI first receives instructions from the conscious AI (S08), performs its own specialized processing based on those instructions (S09), and transmits the processing result to the conscious AI (S10). Note that the conscious AI may call on more than one specialized AI for specialized processing; depending on the content of the action, processing may be carried out by multiple specialized AIs. In this case, some or all of the processing by each specialized AI may be performed in parallel or asynchronously.
[0023] Subsequently, the conscious AI, as part of its output determination process, first receives processing results from each specialized AI (S11), and organizes the information by selecting, integrating, and discarding the obtained results (S12). Then, it discards information that should not be output or is unnecessary, and decides what to output (S13). Finally, it outputs this information to the other party (customer) as video and audio (S14). For example, by synchronizing audio output with the speech of an avatar with a human appearance on the screen, it is possible to create a situation where the other party feels as if they are conversing with the avatar.
[0024] Figure 3 is a diagram illustrating an example of communication in one embodiment of the present invention. Here, an example is shown in which a sales representative 5, such as an insurance agent, uses the communication system 1 of this embodiment when communicating with a customer 4. The diagram shows an example of communication in a scene where the sales representative 5 makes a second visit to the home of customer 4, who is considering purchasing insurance, in a top-to-bottom timeline. The communication process (current position in the flow of communication and work) starts with "small talk" (above the dashed line in the diagram) and progresses to an insurance "proposal" (below the dashed line in the diagram).
[0025] In the "small talk" process, communication system 1 uses AI to perform the processing shown in Figure 2 above. However, based on the content of short-term and long-term memories, the conscious AI determines that continuing the "small talk mode" is the appropriate next action in processing S06 (i.e., the action to be taken is "do nothing"). (As a result, the avatar of communication system 1 does not speak, nor does it call in a specialized AI to perform specialized processing.)
[0026] In the subsequent "proposal" process, communication system 1 uses AI to perform the processing shown in Figure 2 above. Here, we show a case where, in response to salesperson 5's initial proposal utterance ("I know you're interested in..."), the conscious AI determines in processing S06 that "proposal mode" is the appropriate next action mode and decides to perform a "compliance check" as the action to take. Then, in the specialized processing in S09, a specialized AI is called to perform the compliance check and the results are obtained. Since the processing result indicates that there are no compliance issues, the conscious AI decides in processing S13 to output nothing (as a result, the avatar in communication system 1 does not speak).
[0027] Furthermore, when salesperson 5 made their next suggestion ("Our educational insurance is better..."), the communication system 1 performed the same processing. As a result, in processing S11, the specialized AI issued a warning about a compliance violation. The system then examined whether or not the statement should be made and what the statement should be, and in processing S13, it decided to make the statement, and the avatar, etc., said, "Let me correct myself."
[0028] In this way, each time customer 4 or salesperson 5 speaks, communication system 1 rapidly repeats the series of processes shown in the example in Figure 2, appropriately grasping the "process" (current position in the flow of communication and work) and "scene" (external conditions judged from perceptual information) as the situation, determining the "mode" of action to take according to the situation (how the AI itself will behave), and speaking according to the mode to support communication by salesperson 5.
[0029] <System Configuration> Figure 1 is a diagram illustrating an example configuration of a communication system 1, which is one embodiment of the present invention. The communication system 1 is an information processing system that, for example, is composed of one or more server devices, virtual servers built on a cloud computing service, or information processing terminals, and uses a CPU (Central Processing Unit) (not shown) to execute middleware such as an OS (Operating System), DBMS (Database Management System), and Web server programs, as well as software running on them, which are deployed from a storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive) onto memory, thereby realizing various functions to support communication between users such as sales personnel and customers 4.
[0030] This communication system 1 includes, for example, a process determination unit 11, a scene determination unit 12, an action determination unit 13, a specialized processing unit 14, and an output determination unit 15, all implemented as software. It also includes data stores such as a process list 21, a scene list 22, a short-term memory database (DB) 23, a medium-term memory DB 24, and a long-term memory DB 25, all implemented using databases or file tables.
[0031] Each of the above components may be equipped with or utilize a generating AI in the form of a process determination AI, a scene determination AI, an action determination AI, a specialist AI, and an output determination AI. These generating AIs may be independently built on the communication system 1, or they may be configured to utilize services provided by external vendors, etc. Furthermore, each may be configured as an independent and separate AI system or service, or one or more generating AIs may be functionally divided and used on the same AI system or service.
[0032] The process determination unit 11 has the function of determining, using a conscious AI (process determination AI), where a conversation conducted by user 2 is positioned within the flow of communication and business operations, that is, which "process" it belongs to, based on instructions from user 2 (for example, the activation of communication system 1 or related applications). For example, it retrieves the content of past communications (history information) by referring to the long-term memory DB 25, and then uses the process determination AI to determine the process by referring to a list of processes that have been pre-input as a process list 21. The information of the determined process is recorded in the short-term memory DB 23.
[0033] The scene determination unit 12 acquires the communication status of user 2 as perceptual information and, based on the acquired information, has the function of determining what kind of situation the situation is, that is, which "scene" the communication corresponds to, using a conscious AI (scene determination AI). Perceptual information can be acquired, for example, by an input device 3 such as a video camera or microphone that can acquire and digitize the communication status as sound (auditory information) or video / images (visual information). The scene determination unit 12 may be configured to directly acquire perceptual information data from the input device 3, or the input device 3 may be connected to an information processing terminal connected to the communication system 1 via a network (not shown), and the perceptual information data acquired by the information processing terminal may be received via the network. The scene determination AI determines the scene by referring to the acquired perceptual information and a list of scenes that has been pre-input as a scene list 22. The information of the determined scene is also recorded in the short-term memory DB 23.
[0034] Furthermore, since the above-mentioned perceptual information is received in large quantities in real time without interruption, in this embodiment, for example, the processing is divided into two threads: one for acquiring and storing perceptual information and another for determining the scene. For example, the acquired perceptual information (sound, video, images, etc.) is stored in a database (not shown) as it is acquired, and the scene is determined in a separate thread based on the stored content and recorded in the short-term memory DB23. This allows for the continuous storage of large amounts of perceptual information in the short-term memory DB23, while the scene determination process can be performed asynchronously based on a snapshot at a certain point in time. In addition, multiple perceptual information, such as visual and auditory information, can be processed integrally.
[0035] The action determination unit 13 has the function of using a conscious AI (action determination AI) to determine the "mode" (how the AI itself will behave) of the next action that the AI should take, based on the contents of the short-term memory DB 23 recorded by the process determination unit 11 and the scene determination unit 12, i.e., the information of the current process and scene in the communication by user 2, and calling a specialized AI to perform the corresponding specialized processing. The content of the determined next action is recorded in the medium-term memory DB 24 as medium-term memory (content of the action and processing results by the specialized AI).
[0036] The specialized processing unit 14 has specialized AIs that perform various specialized processes, and has the function of performing processing using the corresponding specialized AI based on instructions from the action determination unit 13 and the content of the next action recorded in the medium-term storage DB 24. The processing results are passed to the output determination unit 15, which will be described later, by recording them in the medium-term storage DB 24 or the like.
[0037] The output determination unit 15 uses a conscious AI (output determination AI) to determine whether or not to output (speak) the results of processing by the specialized AI in the specialized processing unit 14 to the user 2. If it determines that output is appropriate, it has the function of outputting (speaking) to the user 2 (or on behalf of the user 2). When determining whether or not to output, the contents of the short-term memory DB 23, the medium-term memory DB 24, and the long-term memory DB 25 are referred to. The determination result is recorded and stored in the medium-term memory DB 24. The user interface for outputting (speaking) is not particularly limited, but for example, an avatar (digital alter ego) displayed on a display (not shown) can speak. The avatar may be configured to acquire the above-mentioned perceptual information by conversing with the user 2 or the other party.
[0038] Furthermore, the calling of individual consciousness AIs and specialized AIs, as well as their collaboration, can be implemented, for example, by using the API (Application Programming Interface) integration function provided by the AI platform.
[0039] <Processing flow> Figures 4 and 5 are processing flow diagrams that outline an example of a processing flow that uses AI to support communication in one embodiment of the present invention (the processing in Figure 4 and the processing in Figure 5 are continuous). Here, using an insurance sales scenario as an example, we show how to support communication between salesperson 5 and a customer who is considering educational insurance. For example, salesperson 5 and the customer reconnected at a class reunion or similar event, and salesperson 5 is visiting the customer's home to sell insurance. The first visit ended with salesperson 5 conducting an interview with the customer and beginning to propose insurance, and this is the second visit.
[0040] Prior to communication, the process determination unit 11 and the scene determination unit 12 perform pre-processing to input information about the processes and scenes set in the process list 21 and scene list 22, respectively (S20). For example, the information in the process list 21 and scene list 22 (for example, if used in insurance sales operations, a list of insurance sales operations processes and a list of scenes that occur in each process) can be pre-trained into each conscious AI, or appropriate methods such as inputting into the conscious AI using RAG (Retrieval-Augmented Generation) or prompts can be used.
[0041] To initiate communication, the sales representative 5 first activates the communication system 1 or a related application (S21). At this time, the sales representative 5 inputs information necessary to obtain prior information from the long-term memory DB 25 through the process described later (for example, "The customer's name is Mr. / Ms. XX"). After the application is activated, the process determination unit 11 obtains information about the customer's past communication history from the long-term memory DB 25 as prior information (S22). For example, it searches the long-term memory DB 23 to obtain information about the previous conversation with the customer (for example, "The previous conversation ended with the start of an insurance proposal"). It also searches the case DB (which is not shown and constitutes part of the long-term memory DB 25; its contents will be described later) to obtain information on similar successful and unsuccessful insurance sales cases (for example, "The case where a proposal was made on the second visit was successful," "The case where a proposal was made on the first visit was unsuccessful").
[0042] Subsequently, the process determination unit 11 generates a process determination prompt 31 based on information obtained from the long-term memory DB 25 and inputs this to the process determination AI to determine the current process in the communication (S23). The information of the determined process (for example, "insurance proposal") is recorded in the short-term memory DB 23. The process determination prompt 31 can be generated, for example, by setting information related to the communication with the customer obtained from the long-term memory DB 25 to a pre-prepared template (not shown) (the same applies to other prompts described later).
[0043] Figure 6 is a schematic diagram illustrating an example of a process determination prompt 31 in one embodiment of the present invention, and Figure 7 is a schematic diagram illustrating an example of an output obtained by inputting the process determination prompt 31 to the process determination AI.
[0044] In the example in Figure 6, the tags are set with content obtained from the long-term memory DB 25. The tags also contain the results obtained by performing a vector search on the process list 21 based on the tag content (i.e., information on possible processes). Furthermore, the tags contain the results obtained by performing a vector search on the case DB (which is not shown in the diagram and constitutes part of the long-term memory DB 25; its contents will be described later) based on the tag content. Finally, the tags indicate an example of the AI output format.
[0045] In response to such a process determination prompt 31, the output example in Figure 7 shows that the current communication status ("This is the second visit and an insurance proposal is planned"), the reason for the process determination ("The hearing was completed during the previous visit"), and the determined process ("Proposal") are set.
[0046] Returning to Figure 4, salesperson 5 then begins a conversation (for example, "Hello. It's getting cold, isn't it?") (S24). At this time, the scene determination unit 12 acquires perceptual information, such as visual information (video) and auditory information (sound), from the outside via the input device 3 and records this in the short-term memory DB 23 (S25). This recording is done continuously at short time intervals, such as every second, in order to grasp the situation that is changing moment by moment. Separately, in a different thread, a scene determination prompt 32 is generated based on the acquired perceptual information, and this is input to the scene determination AI to determine the current scene in the communication (S26). Information on the determined scene (for example, "icebreaker") is recorded in the short-term memory DB 23.
[0047] Figure 8 is a schematic diagram illustrating an example of a scene determination prompt 32 in one embodiment of the present invention, and Figure 9 is a schematic diagram illustrating an example of an output obtained by inputting the scene determination prompt 32 to the scene determination AI.
[0048] In the example in Figure 8, the tags are set with the tag and the perceptual information that the tag will input. The tag is set with encoded image data (for example, image data of a conversation scene), and the tag is set with the audio data converted to text (for example, "Hello. It's getting cold."). The tag is also set with the result obtained by vector searching the scene list 22 based on the content of the tag (i.e., the input perceptual information). And the tag indicates an example of the output format by the AI.
[0049] In response to such a scene determination prompt 32, the output example in Figure 9 shows that the current communication situation (salesperson 5 is greeting the customer at their home), the reason for determining the scene (greeting the customer at their home at the beginning of the business negotiation), and the determined scene ("Greeting / Icebreaker") have been set. There can be multiple possible scenes, but they can be represented by "5W1H" information, prior information about the person being spoken to, and information about changes from the past. By inputting this information into the AI, the scene can be determined.
[0050] Returning to Figure 4, the action determination unit 13 then retrieves the process and scene determination results from the short-term memory DB 23 (S27), generates an action determination prompt 33 based on the retrieved information, and inputs this to the action determination AI to determine the next action (S28). The information of the determined action (for example, calling specialized AIs for "small talk," "insurance proposal," and "business etiquette check" to process the information) is recorded in the medium-term memory DB 24.
[0051] Figure 10 is a schematic diagram illustrating an example of an action determination prompt 33 in one embodiment of the present invention, and Figure 11 is a schematic diagram illustrating an example of an output obtained by inputting the action determination prompt 33 to the action determination AI.
[0052] In the example shown in Figure 10, the tags contain the content of the process ("proposal") and scene ("greeting / icebreaker") retrieved from the short-term memory DB23. The tags also indicate an example of the output format used by the AI.
[0053] In response to such an action determination prompt 33, the output example in Figure 11 shows the reason for determining to call the specialized AI (e.g., "A proposal is about to be made, so it is necessary to bring up a good topic to get the person to open up") and that the specialized AI to be called (e.g., "Conversation AI") is set for each specialized AI to be called.
[0054] Returning to Figure 5, the specialized processing unit 14 then retrieves the content of the next action determined and called by the action determination unit 13 from the medium-term memory DB 24 (S29), generates an action prompt 34 based on the retrieved information, and inputs this to the called specialized AI to execute the action as a specialized process (S30). The execution results of the action (for example, a business etiquette violation warning output as a result of processing by the "business etiquette check AI", the content of the statement "Did you watch yesterday's baseball game?" by the "casual conversation AI", the content of the suggestion "I recommend the following educational insurance plan..." by the "insurance suggestion AI") are recorded in the medium-term memory DB 24. These processes are performed for each called specialized AI, but they may be performed sequentially or one or more specialized processes may be performed in parallel. Note that the content of the action prompt 34 differs for each specialized AI, so an explanation is omitted.
[0055] Subsequently, the output determination unit 15 collects and integrates information obtained from the short-term memory DB 23, medium-term memory DB 24, and long-term memory DB 25, such as the process determination result (e.g., "proposal"), the scene determination result (e.g., "greeting / icebreaker"), and the results of specialized processing by the specialized AI (e.g., "small talk AI"), which are obtained from the series of processes so far (S31). Based on this information, it generates an output determination prompt 35 and inputs it to the output determination AI to determine whether or not to actually output, and what content and method of output to use (S32). The content of the output (e.g., "continue small talk") is recorded in the medium-term memory DB 24.
[0056] Figure 12 is a diagram illustrating an example of an output determination prompt 35 in one embodiment of the present invention, and Figure 13 is a diagram illustrating an example of an output obtained by inputting the output determination prompt 35 to the output determination AI.
[0057] In the example in Figure 12, the tags contain the content of the process ("proposal") and scene ("greeting / icebreaker") obtained from the short-term memory DB23. The tags also contain the results of each specialized AI ("small talk AI," etc.) and their respective specialized processing ("Did you watch yesterday's baseball game?", etc.) obtained from the medium-term memory DB24. Furthermore, the tags indicate an example of the AI's output format.
[0058] In response to such an output decision prompt 35, the output example in Figure 13 shows that the reason for the output decision ("The customer is smiling," "It would be better to chat a little longer to help them relax") and the content of the output ("Did you watch yesterday's baseball game?") are set.
[0059] By repeating the above series of processes each time salesperson 5 or customer 4 speaks, the system supports salesperson 5 in communicating in a way that is appropriate to the situation.
[0060] <Data Structure> Figure 14 is a diagram illustrating an example of specific data in a process list 21 in one embodiment of the present invention. The process list 21 holds master information for each process that constitutes communication, and includes items such as process name and process summary. The process name item holds the name of the target process. The process summary item holds summary information that allows the user and the generating AI to understand the content of the target process.
[0061] Figure 15 is a diagram illustrating an example of specific data in a scene list 22 according to one embodiment of the present invention. The scene list 22 holds master information for each scene that appears in communication, and includes items such as scene name and scene summary. The scene name item holds the name of the target scene. The scene summary item holds summary information that allows the user and the generating AI to understand the content of the target scene.
[0062] Figure 16 is a diagram illustrating the data structure and specific data examples of the short-term memory DB23 in one embodiment of the present invention. The short-term memory DB23 is a table that holds current perceptual information (short-term memory) related to constantly changing external circumstances in the form of scenes, and includes items such as session ID, timestamp, data format, data content, and scene.
[0063] The Session ID field holds identification information such as a number that uniquely identifies each communication. The Timestamp field holds information about the timestamp when the target's perceptual information was acquired (or recorded in the short-term memory DB23). The Data Format field holds information about the data format of the target's perceptual information (scene) (e.g., whether it is content determined by the generating AI based on the perceptual information ("memory") or input content based on the perceptual information ("input"). The Data Content field holds information about the content determined by the generating AI as the content of the target's perceptual information (scene). The Scene field holds information about the scene determined by the generating AI based on the content of the target's perceptual information.
[0064] Figure 17 is a diagram illustrating the data structure and specific data examples of the mid-term memory DB24 in one embodiment of the present invention. The mid-term memory DB24 is a table that holds information related to the actions and processing results (mid-term memory) performed by the specialized AI in the specialized processing unit 14, and has items such as timestamp, specialized AI, action, adoption flag, and evaluation. The timestamp item holds information about the timestamp when the target specialized processing was performed (or recorded in the mid-term memory DB24). The specialized AI item holds information about the specialized AI that performed the target specialized processing. The action item holds information about the specific content of the target specialized processing (statement). The adoption flag item holds information about whether the target action was adopted or not. The evaluation item holds information about the evaluation of the adopted action. This evaluation information can be obtained, for example, as a result of feedback from user 2 or the AI itself after communication.
[0065] Figure 18 is a diagram illustrating the data structure and specific data examples of the long-term memory DB25 in one embodiment of the present invention. The long-term memory DB25 is a table that stores information aggregated from short-term memory at regular intervals, as well as information related to the history (long-term memory) of actions taken by the user. It consists of the long-term memory DB25a shown in Figure 18(a) and the case DB25b shown in Figure 18(b). The long-term memory information stored in the long-term memory DB25 is accumulated, for example, as feedback information from user 2 or the AI itself after communication.
[0066] The Long-Term Memory DB25a is a table that holds summary information about long-term memory in each communication, and has items such as timestamp, summary, and process. The timestamp item holds the timestamp information when the long-term memory information of the target was recorded. The summary item holds the information summarized by the generating AI as the content of the long-term memory of the target. The process item holds information about the process related to the long-term memory of the target.
[0067] Case DB25b is a table that separates and records historical information for each individual communication (case) from the contents of long-term memory, and has items such as Case ID, Process, Case, and Evaluation. The Case ID item holds identification information such as a number that uniquely identifies the target case. The Process item holds information about the history of the process in the target case. The Case item holds information determined by the generating AI as the situation and content of the target case. The Evaluation item holds information about the evaluation of the communication in the target case (for example, in this embodiment, whether or not an insurance contract was concluded).
[0068] As described above, according to the communication system 1, which is one embodiment of the present invention, each time a customer 4 or salesperson 5 speaks, the communication system 1 uses conscious AI to appropriately grasp the process and scene as the situation, determines the mode of action to be taken according to the situation, performs processing by specialized AI according to the mode, and after determining the output content based on the processing result, speaks, thereby effectively supporting the communication by the salesperson 5.
[0069] The present inventors have described the invention in detail based on embodiments above, but it goes without saying that the present invention is not limited to the above embodiments and can be modified in various ways without departing from its essence. Furthermore, the above embodiments are described in detail for the purpose of explaining the present invention in an easy-to-understand manner and are not necessarily limited to those having all the configurations described. In addition, it is possible to add, delete, or replace some of the configurations of the above embodiments with other configurations.
[0070] Furthermore, each of the above configurations, functions, processing units, and processing means may be implemented in hardware, in whole or in part, for example, by designing them as integrated circuits. Alternatively, each of the above configurations, functions, and means may be implemented in software by having the processor interpret and execute programs that implement each function. Information such as programs, tables, and files that implement each function can be stored in memory, hard disks, SSDs, or other recording devices, or in recording media such as IC cards, SD cards, or DVDs.
[0071] Furthermore, in the diagrams above, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines that would be present in the actual implementation. In reality, it can be assumed that almost all components are interconnected. [Industrial applicability]
[0072] This invention can be used in communication systems that support interactive communication with users such as customers. [Explanation of Symbols]
[0073] 1...Communication system, 2...User, 3...Input device, 4...Customer, 5...Sales staff 11...Process determination unit, 12...Scene determination unit, 13...Action determination unit, 14...Specialized processing unit, 15...Output determination unit, 21...Process list, 22...Scene list, 23...Short-term memory DB, 24...Medium-term memory DB, 25...Long-term memory DB, 25a...Long-term memory DB, 25b...Case DB, 31...Process determination prompt, 32...Scene determination prompt, 33...Action determination prompt, 34...Action prompt, 35...Output determination prompt
Claims
1. A communication system that supports communication between users and other parties, A process determination unit, based on content including information related to the history of the communication obtained from the long-term memory unit, uses generating AI (Artificial Intelligence) to determine the process indicating the current position in the communication and records it in the short-term memory unit. A scene determination unit, which uses generating AI to determine a scene representing the current state of the communication based on perceptual information consisting of audio and video data related to the communication acquired from an external source at any time, and records it in the short-term memory unit, Based on the information including the process and the scene recorded in the short-term memory, the generating AI determines the next action to be taken in the communication and records it in the medium-term memory; A specialized processing unit which causes the corresponding generation AI to execute processing related to one or more of the actions determined by the action determination unit and records the processing results in the medium-term storage unit, A communication system comprising: an output determination unit that, based on information recorded in the short-term memory unit, the medium-term memory unit, and the long-term memory unit, including the process, the scene, and the processing results in the specialized processing unit related to the communication, generates an AI to determine whether or not to output the processing results and, if so, what the content of the output should be, and performs output according to the determination result.
2. In the communication system described in claim 1, The scene determination unit is a communication system that asynchronously performs the process of recording the perceptual information acquired from an external source at any time and the process of determining the scene based on the recorded perceptual information.