Intervention device, intervention method, and program
The intervention device addresses communication gaps by estimating and presenting tailored information to enhance mutual understanding between parties, improving communication quality.
Patent Information
- Application Number
- PCT/JP2024/026027
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-22
AI Technical Summary
Conventional communication technologies fail to bridge gaps in perception, understanding, and emotional differences between communication participants, leading to misunderstandings and dissatisfaction.
An intervention device that acquires information from communicating parties, estimates communication gaps, and presents tailored information to alleviate these gaps through multimodal means such as text, audio, and visual cues.
Facilitates true mutual understanding by addressing discrepancies in perception, language, and emotional differences, enhancing communication quality.
Smart Images

Figure JP2024026027_22012026_PF_FP_ABST
Abstract
Description
Intervention device, intervention method, and program
[0001] The present disclosure relates to an intervention device, an intervention method, and a program.
[0002] In communication (including not only communication between people but also communication between people and machines and between machines), it is important to accurately understand the intentions of the other party. A related technology is known, which is a technology that can clarify ambiguous questions in question answering called machine reading comprehension (Non-Patent Document 1).
[0003] Atsushi Otsuka, Kyosuke Nishida, Itsumi Saito, Hisako Asano, Junji Tomita, and Tetsuji Sato, "Proposal of a Question Answering Method Using Machine Reading Comprehension Focusing on Clarifying Question Intent," Transactions of the Japanese Society for Artificial Intelligence 34(5) A-J14_1-12, September 2019.
[0004] However, conventional technologies including the technology described in Non-Patent Document 1 have not been able to bridge gaps such as discrepancies in perception, differences in perception, and differences in understanding between communication participants.
[0005] The present disclosure has been made in consideration of the above points, and aims to provide a technology that can support communication.
[0006] An intervention device according to one aspect of the present disclosure is an intervention device that intervenes in communication between one or more subjects, and includes: an acquisition unit that acquires information about at least one of the one or more subjects; an estimation unit that estimates an index value for evaluating a communication gap based on the information about the subject; and a presentation unit that presents information to at least one of the one or more subjects to assist in eliminating the gap based on the index value.
[0007] It can support communication.
[0008] FIG. 1 is a diagram (1) showing an example of intervention. FIG. 2 is a diagram (2) showing an example of intervention. FIG. 3 is a diagram (3) showing an example of intervention. FIG. 4 is a diagram (4) showing an example of intervention. FIG. 5 is a diagram (5) showing an example of intervention. FIG. 1 is a diagram showing an example of the overall configuration of an intervention system including an intervention device according to the present embodiment. FIG. 2 is a diagram showing an example of the hardware configuration of an intervention device according to the present embodiment. FIG. 3 is a diagram showing an example of the functional configuration of an intervention device according to the present embodiment. FIG. 4 is a diagram showing an example of the detailed functional configuration of a gap estimator according to the present embodiment. FIG. 5 is a diagram schematically showing an example of a multimodal recognition model. FIG. 6 is a flowchart showing intervention processing in Example 1. FIG. 7 is a flowchart showing intervention processing in Example 2. FIG. 8 is a flowchart showing intervention processing in Example 3. FIG. 9 is a flowchart showing intervention processing in Example 4. FIG. 10 is a flowchart showing intervention processing in Example 5. FIG. 11 is a diagram showing a modified example of the functional configuration of the intervention device according to the present embodiment. FIG. 12 is a flowchart showing intervention processing in a modified example. FIG. 13 is a diagram showing an example of a UI. FIG. 14 is a diagram showing an example of a UI. FIG. 15 is a diagram showing an example of a UI. FIG. 16 is a diagram showing an example of a UI. FIG. 7 is a diagram showing an example of a UI (No. 7); FIG. 8 is a diagram showing an example of a UI (No. 8); FIG. 9 is a diagram showing an example of the functional configuration of a learning device that learns a gap recognition model; FIG. 10 is a diagram showing an example of first learning data; FIG. 11 is a flowchart showing an example of a learning process for learning a gap recognition model; FIG. 12 is a diagram showing an example of the functional configuration of a learning device that learns a presentation information generation model; FIG. 13 is a diagram showing an example of second learning data; FIG. 14 is a flowchart showing an example of a learning process for learning a presentation information generation model.
[0009] An embodiment of the present invention will be described in detail below with reference to the drawings. The following describes an intervention device 10 that can intervene to bridge a gap that occurs between two or more people (hereinafter referred to as "users") who are communicating with each other. Use of the intervention device 10 according to this embodiment helps to bridge a gap that occurs between users, thereby enabling communication that achieves true mutual understanding.
[0010] <Definitions of Terms, etc.> A user is the subject of communication. A user is not limited to a person, but may also be a machine. Furthermore, a machine is not limited to a device, equipment, terminal, etc., but may be, for example, a program or module that realizes artificial intelligence (AI), or a program or module that realizes a character or avatar that simulates a past or future self, another person's personality, a real or virtual creature, etc.
[0011] Communication is the act of one user transmitting some information to another user between two or more users.
[0012] Intervention means presenting information (hereinafter also referred to as "presented information") to at least one of the users who are communicating in order to fill the gap. Note that, although the presented information is assumed to be mainly text below, the presented information is not limited to text. The presented information may be, for example, audio, images, videos, vibrations, blinking lights, or other information, or may be information such as a program or script for executing a predetermined process, or may be a combination of two or more of these types of information.
[0013] A gap is any difference that occurs between users who are communicating. Specific examples of gaps include discrepancies in perception, differences in perception, differences in understanding, changes in emotions, differences in emotions with the other person, differences between expectations of the other person and their actual response, differences in satisfaction, differences in trust, differences in interest, differences in concentration, differences in fun, differences in excitement, the emergence of awe or fear, etc. However, these are all just examples, and gaps can include a variety of differences that occur between users who are communicating.
[0014] The type of gap refers to the type or category of difference that has occurred between users who are communicating. For example, if "discrepancy in perception" is used as the gap, the type of gap will be "discrepancy in perception."
[0015] <Example of Intervention> As an example, the following describes a case where an intervention is performed when a gap occurs in communication between a customer and an operator in a call center (which may also be called a "contact center"). In the following, the customer is referred to as a "first user" and the operator is referred to as a "second user." Furthermore, a device such as a PC (personal computer) used by the first user is referred to as a "first user device 20," and a device such as a PC used by the second user is referred to as a "second user device 30."
[0016] - When there is a gap due to the "order of speaking" As an example of the difference between expectations of the other person and their actual response, we will explain the case where there is a gap due to the order of speaking.
[0017] As shown in Figure 1, suppose a first user inquires at a call center about whether a certain product is available for purchase, and a second user gives a lengthy explanation of the product, and finally informs the user that the product is still available for purchase. In this case, the first user, who simply wanted to know whether the product was still available for purchase, becomes dissatisfied, feeling that "the final conclusion alone was enough," and a gap has arisen between the first and second users due to the "order of speaking." Note that "AAA" and "BBB" in Figure 1 refer to letters or character strings (e.g., proper nouns) that refer to a specific subject being discussed. This is also true for other figures.
[0018] Therefore, in this case, the interventional device 10 according to the present embodiment presents presentation information to at least one of the first user device 20 and the second user device 30 to fill the gap. For example, when presenting presentation information to the first user device 20, the interventional device 10 according to the present embodiment presents presentation information such as, "Is all the customer wanting to know whether or not they can apply?" as shown in FIG. 1 . When presenting presentation information to the second user device 30, the interventional device 10 according to the present embodiment presents presentation information such as, "It seems that the customer only wants to know whether or not they can apply. If the explanation is too long, they may complain." As a result, when the presentation information is presented to the first user device 20, the first user can, for example, tell the second user that they only want to know whether or not they can apply. When the presentation information is presented to the second user device 30, the second user can, for example, apologize for the long explanation.
[0019] The intervention device 10 according to the present embodiment uses various information acquired from the first user device 20 and the second user device 30 to determine whether gaps need to be filled and to acquire presentation information. A typical example of information acquired from the first user device 20 and the second user device 30 is the user's voice. However, the information acquired from the first user device 20 and the second user device 30 is not limited to this. Examples of information acquired from the first user device 20 and the second user device 30 include, in addition to voice, images and videos of the user, text entered by the user, and sensor information measured by a sensor on the user or their surroundings (e.g., vital signs such as the user's blood pressure and body temperature, and environmental information such as the temperature and humidity of the user's environment). In other words, user information is multimodal information including voice, images, videos, text, sensor information, and the like. Hereinafter, information acquired from the first user device 20 will be referred to as "first user information," and information acquired from the second user device 30 will be referred to as "second user information."
[0020] - When there is a gap due to "language usage" As an example of the difference between expectations of the other person and their actual response, we will explain the case where there is a gap due to the other person's language usage.
[0021] As shown in Figure 2, suppose a second user says to a first user, "Please hurry up and order a certain product." In this case, the first user feels dissatisfied, thinking, "It feels like I'm being told to hurry up and order, and I feel uncomfortable. I think it would be better to add a phrase like, 'I'm very sorry, but'." This can be said to be a gap between the first and second users due to their "language usage."
[0022] Therefore, in this case, the interventional device 10 according to the present embodiment presents presentation information to at least one of the first user device 20 and the second user device 30 to fill the gap. For example, when presenting presentation information to the first user device 20, the interventional device 10 according to the present embodiment presents presentation information such as, "We apologize for the inappropriate language used by the operator, as a new employee," as shown in FIG. 2 . When presenting presentation information to the second user device 30, the interventional device 10 according to the present embodiment presents presentation information such as, "We used language that may be perceived by the customer as a complaint. Please apologize later." This can be expected to alleviate the first user's dissatisfaction when the presentation information is presented to the first user device 20. When presenting presentation information to the second user device 30, the second user can take action, such as apologizing for the inappropriate language used.
[0023] - When there is a gap due to the "time it takes to explain" As an example of the difference between expectations of the other party and their actual response, we will explain the case where there is a gap due to the time it takes to explain.
[0024] 3, when a first user inquires at a call center about how to change the address of an account, the second user begins to explain how to change the address without informing the call center of the time required for the operation, the number of steps, etc. In this case, the first user, who wants to know how long it will take and how many steps it will take to change the address, becomes dissatisfied and says, "I wish they would tell me first how long it will take and how many steps there are," and a gap has arisen between the first user and the second user due to the "time required for the explanation."
[0025] Therefore, in this case, the interventional device 10 according to the present embodiment presents presentation information to at least one of the first user device 20 and the second user device 30 to fill the gap. For example, when presenting presentation information to the first user device 20, as shown in FIG. 3, the interventional device 10 according to the present embodiment presents presentation information such as, "Changing your address will take approximately 10 minutes. The address change procedure involves three steps. We will send you reference materials." Furthermore, when presenting presentation information to the second user device 30, as shown in FIG. 3, the interventional device 10 according to the present embodiment presents presentation information such as, "Please first explain to me the time required to change your address and the overall operation." This allows the first user to know the time and number of steps required to change their address, which is expected to alleviate their dissatisfaction. Furthermore, when presenting presentation information to the second user device 30, the second user can take action, such as explaining to the first user the time required to change their address and the overall operation required.
[0026] - When there is a gap caused by "content that is difficult to convey in words" As an example of the difference between expectations of the other party and their actual response, we will explain the case where there is a gap caused by content that is difficult to convey in words.
[0027] 4, when a first user inquires at a call center about how to set up the system initially, a second user may say something that is difficult to convey in words, such as "Please open the control panel at the top of the computer screen." In this case, the first user, who does not understand what the term "control panel" means, becomes dissatisfied and says, "I wish there was a visual explanation using an image." This can be said to create a gap between the first and second users due to "something that is difficult to convey in words."
[0028] Therefore, in this case, the interventional device 10 according to the present embodiment presents presentation information to at least one of the first user device 20 and the second user device 30 to fill the gap. For example, when presenting presentation information to the first user device 20, as shown in FIG. 4, the interventional device 10 according to the present embodiment presents an image pointing to a control panel on a personal computer screen and text such as "The control panel is here." Furthermore, when presenting presentation information to the second user device 30, as shown in FIG. 4, the interventional device 10 according to the present embodiment presents presentation information such as "It seems you haven't reached the control panel yet. Please explain it to me in detail." This allows the first user to visually grasp the location of the control panel when the presentation information is presented to the first user device 20, which is expected to alleviate their dissatisfaction. Furthermore, when presenting presentation information to the second user device 30, the second user can respond by, for example, carefully explaining the location of the control panel.
[0029] - When there is a gap due to "level of understanding" As an example of differences in level of understanding, we will explain the case where there is a gap due to each other's level of understanding.
[0030] 5, suppose that the second user says, "Is there anything you're unsure about so far?", and the first user does not say anything about it, even though he or she has questions. In this case, the first user's level of understanding differs from that of the second user, and a gap in "level of understanding" has arisen between the first user and the second user.
[0031] Therefore, in this case, the interventional device 10 according to the present embodiment presents presentation information to at least one of the first user device 20 and the second user device 30 to fill the gap. For example, when presenting presentation information to the first user device 20, the interventional device 10 according to the present embodiment presents presentation information such as, "Was there a lack of explanation about DDD?" as shown in FIG. 5 . Also, when presenting presentation information to the second user device 30, the interventional device 10 according to the present embodiment presents presentation information such as, "You may not understand DDD. Please listen carefully." as shown in FIG. 5 . As a result, when the presentation information is presented to the first user device 20, the first user can take action, such as requesting an explanation from the second user. Also, when the presentation information is presented to the second user device 30, the second user can take action, such as listening to the first user.
[0032] As described above, the intervention device 10 according to the present embodiment intervenes in the communication between the first user and the second user, thereby helping to eliminate the gap that exists between the first user and the second user. Therefore, by using the intervention device 10 according to the present embodiment, communication that realizes true mutual understanding between the first user and the second user can be expected.
[0033] <Example of Overall Configuration of Intervention System 1 Including Intervention Device 10> Fig. 6 shows an example of the overall configuration of the intervention system 1 including the intervention device 10 according to the present embodiment. As shown in Fig. 6, the intervention device 10 according to the present embodiment includes the intervention device 10, a first user device 20, and a second user device 30. The intervention device 10, the first user device 20, and the second user device 30 are communicatively connected via an arbitrary communication network 40. Note that the communication network 40 includes various communication networks and communication means, such as the Internet, a local area network (LAN), a wide area network (WAN), a telephone network (public switched telephone network, mobile communication network, etc.), etc.
[0034] The intervention device 10 is a computer or computer system that intervenes in communication between a first user and a second user in accordance with a gap between the first user and the second user. The intervention device 10 is realized, for example, by a PC, a general-purpose server, or a system configured thereof.
[0035] The first user device 20 is a computer or computer system used by or serving as the first user. The first user device 20 is realized, for example, by a terminal 21 such as a PC used by the first user. However, this is just one example, and the first user device 20 may be realized by various devices 22 including, for example, a server, a smartphone, a wearable device, a robot, industrial machinery, home appliances, a vending machine, a cash register terminal, a sensor device, etc. Note that these are examples of the device 22, and the device 22 may also include various devices such as an in-vehicle device, a game device, a tablet terminal, medical equipment, nursing care equipment, digital signage, an electronic whiteboard, etc.
[0036] The second user device 30 is a computer or computer system used by or serving as a second user. The second user device 30 is realized, for example, by a terminal 31 such as a PC used by the second user. However, this is just one example, and the second user device 30 may be realized by various devices 32 including, for example, a server, a smartphone, a wearable device, a robot, industrial machinery, home appliances, a vending machine, a cash register terminal, a sensor device, etc. Note that these are examples of the device 32, and the device 33 may include various devices such as an in-vehicle device, a game device, a tablet terminal, medical equipment, nursing care equipment, digital signage, an electronic whiteboard, etc.
[0037] 6 is merely an example, and the overall configuration of the intervention system 1 is not limited thereto. For example, the intervention device 10 and the first user device 20 or the second user device 30 may be integrally configured.
[0038] <Example of Hardware Configuration of Intervention Device 10> An example of the hardware configuration of the intervention device 10 according to this embodiment is shown in Fig. 7. As shown in Fig. 7, the intervention device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a random access memory (RAM) 105, a read only memory (ROM) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is communicatively connected via a bus 109.
[0039] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the intervention device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.
[0040] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0041] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is one of various arithmetic devices such as a central processing unit (CPU) or a graphics processing unit (GPU).
[0042] 7 is an example, and the hardware configuration of the intervention device 10 is not limited to this. The intervention device 10 may have, for example, multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.
[0043] <Example of Functional Configuration of Intervention Device 10> Fig. 8 shows an example of the functional configuration of the intervention device 10 according to this embodiment. As shown in Fig. 8 , the intervention device 10 according to this embodiment includes a first user information acquisition unit 201, a second user information acquisition unit 202, a gap estimation unit 203, a presentation information acquisition unit 204, a first user presentation unit 205, and a second user presentation unit 206. These units are realized, for example, by processing in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like. The intervention device 10 according to this embodiment also includes a first user history DB 207, a second user history DB 208, and an intervention plan DB 209. Each of these DBs (databases) is realized, for example, by a storage area of the auxiliary storage device 107 or the like. However, at least one of these DBs may be realized, for example, by a storage area of a storage device (e.g., a storage device included in a database server) connected to the intervention device 10 via the communication network 40.
[0044] The first user information acquisition unit 201 acquires first user information from the first user device 20. The first user information acquisition unit 201 also stores the first user information acquired from the first user device 20 in the first user history DB 207. The first user information acquisition unit 201 acquires the first user information from the first user device 20 at predetermined intervals (e.g., every few seconds to every few tens of seconds). The first user information is multimodal information including at least one of the following: the voice of the first user during communication with the second user; images or videos of the first user (particularly, the face of the first user, etc.); text entered by the first user; and sensor information obtained by measuring the first user and their surroundings using a sensor.
[0045] The second user information acquisition unit 202 acquires second user information from the second user device 30. The second user information acquisition unit 202 also stores the second user information acquired from the second user device 30 in the second user history DB 208. The second user information acquisition unit 202 acquires the second user information from the second user device 30 at predetermined time intervals (e.g., every few seconds to every few tens of seconds). The second user information is multimodal information that includes at least one of the following: the voice of the second user during communication with the first user; images or videos of the second user (particularly, the face of the second user); text entered by the second user; and sensor information obtained by measuring the second user and their surroundings using a sensor.
[0046] The gap estimation unit 203 estimates the gap between the first user and the second user using at least one of the first user information and the second user information. Specifically, the gap estimation unit 203 estimates the type of gap between the first user and the second user (hereinafter also referred to as the "gap type") and the presence or absence of a gap or a gap degree score. A detailed functional configuration example of the gap estimation unit 203 will be described later. The gap degree score is an index value for evaluating the gap between the first user and the second user, and is, for example, a value that quantitatively evaluates the degree or level of the gap. Note that the gap degree score may be a continuous value, a discrete value, a percentage, or the like. Furthermore, for example, if the gap degree score is a discrete value taking 0 or 1, the gap degree score may be regarded as the presence or absence of a gap by determining that a gap degree score of 0 indicates no gap and a gap degree score of 1 indicates a gap.
[0047] The presented information acquisition unit 204 determines whether or not it is necessary to fill the gap. If it determines that it is necessary to fill the gap, the presented information acquisition unit 204 acquires at least one of presented information for the first user and presented information for the second user from the intervention DB 209 according to the gap type, the presence or absence of a gap, or the gap degree score.
[0048] The first user presentation unit 205 presents presentation information for the first user to the first user device 20. Hereinafter, the presentation information for the first user will also be referred to as “first presentation information.” Note that the first user presentation unit 205 can present the first presentation information by outputting the first presentation information to the first user device 20.
[0049] The second user presentation unit 206 presents presentation information for the second user to the second user device 30. Hereinafter, the presentation information for the second user will also be referred to as “second presentation information.” Note that the second user presentation unit 206 can present the second presentation information by outputting the second presentation information to the second user device 30.
[0050] The first user history DB 207 stores the history of the first user information (that is, time-series data of the first user information).
[0051] The second user history DB 208 stores the history of second user information (that is, time-series data of second user information).
[0052] The intervention DB 209 stores first presented information and second presented information for each gap type (or for each gap type and score range of the gap degree score). That is, the intervention DB 209 stores information in a format such as (gap type, first presented information, second presented information) or (gap type, score range of the gap degree score, first presented information, second presented information).
[0053] Note that the functional configuration example of the interventional device 10 shown in FIG. 8 is just an example, and the functional configuration of the interventional device 10 is not limited thereto. For example, when only the first user information is used in the gap estimation unit 203, the interventional device 10 may not have the second user information acquisition unit 202 and the second user history DB 208. Similarly, when only the second user information is used in the gap estimation unit 203, the interventional device 10 may not have the first user information acquisition unit 201 and the first user history DB 207. Furthermore, for example, when the first presentation information is presented only to the first user device 20, the interventional device 10 may not have the second user presentation unit 206. Similarly, for example, when the second presentation information is presented only to the second user device 30, the interventional device 10 may not have the first user presentation unit 205.
[0054] <Detailed Functional Configuration Example of Gap Estimation Unit 203> A detailed functional configuration example of the gap estimation unit 203 according to this embodiment is shown in Fig. 9. As shown in Fig. 9, the gap estimation unit 203 according to this embodiment includes a multimodal recognition unit 211 and a gap recognition unit 212.
[0055] The multimodal recognition unit 211 receives first user information as input and outputs recognition information (hereinafter also referred to as "first recognition information") including linguistic information, non-linguistic information, paralinguistic information, etc. recognized or detected from the audio, images, videos, text, sensor information, etc. included in the first user information. Similarly, the multimodal recognition unit 211 receives second user information as input and outputs recognition information (hereinafter also referred to as "second recognition information") including linguistic information, non-linguistic information, paralinguistic information, etc. recognized or detected from the audio, images, videos, text, sensor information, etc. included in the second user information. Note that linguistic information refers to information that represents the content of a user's utterance, and non-linguistic information refers to non-voluntary information such as gender, age, etc., among non-linguistic information. Furthermore, paralinguistic information refers to voluntary information such as emotions, attitudes, intentions, etc., among non-linguistic information.
[0056] Here, the multimodal recognition unit 211 is realized by a known multimodal recognition model (which may also be called "multimodal AI" or "multimodal LLMs (Large Language Models)"). An example of a multimodal recognition model is shown in FIG. 10. The multimodal recognition model 1000 shown in FIG. 10 is composed of one or more neural network layers (generally multiple neural network layers). It receives multimodal information as input, recognizes various information from the multimodal information, and outputs recognition information as a result. The multimodal information includes various information such as the user's voice, images, text, sensor information, etc. The recognition information is information that recognizes the content of the user's utterance, the user's state, actions, etc., and may include, for example, speech-recognition text, the user's emotions (e.g., joy, anger, sadness, happiness, positivity / negativity, etc.), whether or not the user is smiling, the way they speak, the direction of their face, their level of alertness, their level of empathy, the presence or absence of negative words, their gender, their age, the presence or absence of a certain illness, the number of fillers, etc. Emotions may include, for example, feelings, emotions, feelings, moods, thoughts, feelings, emotions, impressions, feelings, sentiments, values, thoughts, etc. Furthermore, speech-recognized text may be called, for example, "spoken text" or simply "text."
[0057] Other examples that may be included in the recognition information include, for example, whether or not there is crying, whether or not there is screaming, whether or not there is coughing, speaking style, speech interval, speaker, speaker diarization, speaking rate, speech recognition text after adding punctuation, speech recognition text converted from spoken language into written language, understandability, funniness, whether or not there is harassment, detection results of objects including people, detection results of faces with or without masks, text area in the image, key points on the face, gaze direction, clothing, belongings, behavior, gestures, poker face level, amount of gaze movement, average amount of gaze movement, participation level, number of nods, number of blinks, text representing explanatory text for the image, personality traits, speaking behavior, communication skill level, likability level, sales impression level, willingness to engage in conversation, smoothness of conversation, activeness of conversation, excitement of conversation, concentration level, understanding level, satisfaction level, motivation, interest level, etc. The recognition information recognized by the multimodal recognition unit 211 may be information recognized from one type of information (e.g., audio only, image only) or information recognized from multiple types of information (e.g., both audio and image).
[0058] Note that the above recognition information is an example, and the multimodal recognition unit 211 is capable of using first user information or second user information as input to output various recognition information that can be recognized, detected, or estimated using a known multimodal recognition model.
[0059] To obtain the above recognition information, the multimodal recognition unit 211 realizes, for example, the following functions. However, it goes without saying that these functions are merely examples. It also goes without saying that these functions may be combined as appropriate. Furthermore, although terms such as "recognition," "detection," and "estimation" are used below, these terms are not strictly distinct and may be interchangeable terms.
[0060] - Speech period detection: Using audio data as input, the speaker's speech period is detected and output.
[0061] ・Speech recognition: Takes voice data as input and outputs voice-recognized text.
[0062] - Multi-speaker speech recognition: Speech data is input and speech recognition text for each speaker is output.
[0063] Voice gender recognition: Voice data is input and the speaker's gender is recognized and output.
[0064] - Voice age recognition: Voice data is used as input to recognize and output the speaker's age.
[0065] - Voice emotion recognition: Voice data is used as input to recognize and output the speaker's emotions.
[0066] Crying / screaming detection: Audio data is input and a pair of crying / screaming labels and their confidence scores is output.
[0067] Laughter / cough detection: Audio data is input and a pair of laugh / cough labels and their confidence scores is output.
[0068] Speaking style estimation: Speech data is input and a pair of speaking style labels and their confidence scores is output.
[0069] Speaker vector extraction: Inputs speech data and outputs speaker vectors.
[0070] Speaker diarization recognition: Speech recognition texts of multiple speakers and speaker vectors of those multiple speakers are input, and pairs of utterance start times, utterance end times, and speaker IDs are output. Note that speaker diarization recognition may be performed after speech recognition and speaker vector extraction, or may use the outputs of speech recognition and speaker vector extraction as input.
[0071] Adding punctuation: Text data is input and text data with punctuation added is output. Note that adding punctuation may be performed after speech recognition, or the output of speech recognition may be used as input.
[0072] Spoken-to-written conversion: Text data is input, and text data obtained by converting spoken data into written data is output. Note that the spoken-to-written conversion may be performed after speech recognition, or the speech recognition output may be used as input.
[0073] Positive / negative estimation: Text data is input, and pairs of positive / negative labels and their confidence scores are output. Note that positive / negative estimation may be performed in a later stage of speech recognition processing, or the speech recognition output may be used as input.
[0074] - Comprehensibility estimation: Text data is input, and a pair of comprehensibility labels and their reliability is output. Note that comprehensibility estimation may be performed after speech recognition, or may use the speech recognition output as input.
[0075] Interestingness estimation: Text data is input, and a pair of interestingness labels and their reliability is output. Interestingness estimation may be performed after speech recognition, or the speech recognition output may be used as input.
[0076] Harassment detection: Text data is input, and a pair of harassment labels and their reliability is output. Harassment detection may be performed in a later stage of speech recognition, or the speech recognition output may be used as input.
[0077] Filler count estimation: Text data is input, and the number of fillers is estimated and output. Note that filler count estimation may be performed in a later stage of speech recognition processing, or the output of speech recognition may be used as input.
[0078] Face detection: Image data is input and the coordinates of the face areas for the number of people in the image are output.
[0079] - Object detection: Image data is input and the coordinates of the object area for the object in the image are output.
[0080] - Person detection: Image data is input and the coordinates of the person areas for the number of people in the image are output.
[0081] Face detection with or without mask: Image data is input, and pairs of coordinates of face areas for the number of people in the image and labels indicating whether or not a face is masked are output.
[0082] Character detection: Image data is input and the coordinates of a rectangular area surrounding the character area in the image are output.
[0083] Facial emotion recognition: Image data of the facial region is input, and a pair of an emotion label and its reliability is output. Note that facial emotion recognition may be performed after face detection or face detection with or without mask, and may use image data of the facial region represented by the coordinates output by face detection or face detection with or without mask as input.
[0084] Face / gender estimation: Image data of the face area is input, and a pair of a gender label and its reliability is output. Note that face / gender estimation may be performed after face detection or face detection with or without a mask, and image data of the face area represented by the coordinates output by face detection or face detection with or without a mask may be input.
[0085] Face age estimation: Image data of the face area is input, and the age is estimated and output. Note that face age estimation may be performed after face detection or face detection with or without mask, and image data of the face area represented by coordinates output by face detection or face detection with or without mask may be input.
[0086] Facial emotional arousal estimation: Image data representing a face image is input, and the arousal level is estimated and output. Note that the facial emotional arousal estimation may be performed after face detection or face detection with or without a mask, and may use image data of the face area represented by the coordinates output by face detection or face detection with or without a mask as input.
[0087] Face direction estimation: Using image data of the face region as input, estimate and output the face direction. Note that face direction estimation may be performed in a subsequent process after face detection or face detection with or without a mask, and may use image data of the face region represented by coordinates output by face detection or face detection with or without a mask as input.
[0088] Facial keypoint estimation: Image data representing a facial image is input, and facial keypoints are estimated and output. Note that facial keypoint estimation may be performed after face detection or face detection with or without mask, and may use image data of the facial area represented by the coordinates output by face detection or face detection with or without mask as input.
[0089] Gaze estimation: Using image data of the face region as input, the vertical and horizontal gaze angles of the right and left eyes are estimated and output. Note that gaze estimation may be performed in a subsequent process after face detection or face detection with or without a mask, and image data of the face region represented by coordinates output by face detection or face detection with or without a mask may be used as input.
[0090] Clothing recognition: Image data of a person area is input, and a set of clothing labels and their reliability is output. Note that clothing recognition may be performed after person detection, and the image data of the person area represented by the coordinates output by person detection may be used as input.
[0091] Personal item recognition: Image data of a person area is input, and a pair of personal item labels and their reliability is output. Note that personal item recognition may be performed after person detection, and the image data of the person area represented by the coordinates output by person detection may be used as input.
[0092] Action recognition: Image data of a human region is input, and a pair of an action label and its reliability is output. Note that action recognition may be performed after human detection, and the image data of the human region represented by the coordinates output by human detection may be input.
[0093] Gesture recognition: A time series of image data is input, and a pair of gesture labels and their reliability is output.
[0094] Frame average of face and gender estimation: A time series of face and gender estimation results for the same person is input, and a pair of the average gender label and its reliability is output.
[0095] - Frame average of facial age estimation: The time series of facial age estimation results for the same person is input and the average age is output.
[0096] Poker face degree estimation: The time series of facial emotion recognition results for the same person is used as input, and the poker face degree is output. The poker face degree is defined as the proportion of facial expressions that are not specified in advance.
[0097] - Gaze movement amount estimation: The time series of gaze estimation results is input and the gaze movement amount is output.
[0098] Average gaze movement estimation: The gaze estimation results for two consecutive frames of the same person are used as the average gaze movement amount, and the average gaze movement amount is output.
[0099] - Participation estimation: The time series of face direction estimation results is input, and the percentage of people facing the direction of the camera that took the photo is output.
[0100] Nodding frequency estimation: The time series of face direction estimation results is input, and the number of noddings is output.
[0101] - Blink frequency estimation: The time series of facial keypoint estimation results is used as input, and the blink frequency is output.
[0102] Image description generation: Image data is input and text describing the image is output.
[0103] Character recognition: Image data and the coordinates of a rectangular area surrounding a character area in the image are input, and the text of the character area is output. Note that character recognition may be performed after character detection, and the coordinates output by character detection may be used as input.
[0104] Personality trait estimation: Voice data and a time series of image data of the face region are input, and a pair of personality trait labels and their reliability is output. Note that personality trait estimation may be performed in a later stage of face detection or face detection with or without a mask, and may use as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask.
[0105] Speech behavior recognition: It takes voice data and a time series of image data of the face region as input, and outputs a pair of a speech behavior label and its reliability. Note that speech behavior recognition may be positioned after face detection or face detection with or without mask, and may take as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without mask.
[0106] Communication skill estimation: Voice data and a time series of image data of the face region are input, and a set of communication skill labels and their reliability is output. Note that communication skill estimation may be performed in a later process after face detection or face detection with or without a mask, and may use as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask.
[0107] Likeability estimation: Voice data and a time series of image data of the face region are input, and a pair of likeability labels and their reliability is output. Note that likeability estimation may be performed in a later process after face detection or face detection with or without a mask, and the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask may be input.
[0108] Conversational sales impression estimation: Voice data and image data are input, and a sales impression label is output.
[0109] Conversation Positive / Negative Degree Estimation: The facial emotion recognition results of all conversation participants are used as input to output the conversation positive / negative degree. The conversation positive / negative degree is defined as the difference between the positive and negative proportions of all participants.
[0110] Conversation empathy estimation: The facial emotion recognition results of all conversation participants are input, and the conversation empathy is output. The conversation empathy is defined as the similarity of the facial emotion recognition results of all conversation participants.
[0111] Conversational engagement estimation: The conversational engagement estimation results for all conversation participants are input, and the conversational engagement is output. The conversational engagement is defined as the average conversational empathy of all conversation participants.
[0112] - Estimation of conversational fluency: The result of voice activity detection is used as input and conversational fluency is output. The conversational fluency is defined as the proportion of non-silence periods.
[0113] Conversational enthusiasm estimation: The results of the conversational positive / negative degree estimation and the conversational fluency estimation are input, and the conversational enthusiasm degree is output. The conversational enthusiasm degree is defined as the sum of the conversational positive / negative degree and the conversational fluency degree.
[0114] The gap recognition unit 212 receives at least one of the first recognition information and the second recognition information as input, and recognizes the type of gap between at least one of the first user and the second user and the other, and the presence or absence of a gap or a gap degree score. The gap recognition unit 212 is realized, for example, by a trained machine learning model that receives recognition information as input and outputs the type of gap and the presence or absence of a gap or a gap degree score. However, the gap recognition unit 212 may also be realized, for example, by a program that implements a rule base or a knowledge graph, or a program that performs simple calculations. Hereinafter, the machine learning model that realizes the gap recognition unit 212 will be referred to as a "gap recognition model." A learning method for the gap recognition model will be described later.
[0115] The multimodal recognition unit 211 and the gap recognition unit 212 may recognize the presence or absence of a gap or the gap degree score by processing the multimodal information (i.e., the first user information, the second user information, or both) in an end-to-end manner. In this case, the multimodal recognition unit 211 and the gap recognition unit 212 may process each type of information (e.g., image, audio, etc.) in parallel, or may process each type of information sequentially or collectively. For details about end-to-end processing, see, for example, References 1-4.
[0116] Furthermore, a model that processes multimodal information in an end-to-end manner (i.e., a model that realizes the multimodal recognition unit 211 and the gap recognition unit 212 in an end-to-end manner) can be configured in various ways. For example, it can be configured with a first input layer that inputs linguistic information included in the multimodal information, a second input layer that inputs paralinguistic information included in the multimodal information, an intermediate layer, and an output layer. In this case, the output layer may be, for example, a layer that outputs a binary value indicating whether or not there is a gap, or a layer that outputs a gap degree score. Alternatively, the output layer may be, for example, a layer that outputs at least one of linguistic information and paralinguistic information.
[0117] <Example of Intervention Process> An example of the intervention process according to this embodiment will be described below. In the following Examples 1 to 5, as an example, the index value representing emotions is assumed to be a positive / negative degree, and the gap is assumed to be a change in one's own emotions or a difference in emotions between oneself and the other person.
[0118] Example 1 An example of an intervention process in which a gap is recognized using first user information and, if the gap needs to be filled, first presentation information is presented to the first user will be described below with reference to Fig. 11. The intervention process shown in Fig. 11 is repeatedly executed at predetermined time intervals (e.g., every few seconds to every few tens of seconds).
[0119] The first user information acquisition unit 201 acquires first user information from the first user device 20 (step S101). The first user information is stored in the first user history DB 207.
[0120] The multimodal recognition unit 211 of the gap estimation unit 203 recognizes first recognition information using the first user information stored in the first user history DB 207 (step S102). Hereinafter, as an example, it is assumed that the first recognition information is a first text representing a speech-recognized text of the voice included in the first user information acquired in step S101 and the current emotion of the first user. It is also assumed that the emotion of the first user is represented by an index value called a positive / negative degree. The positive / negative degree is a value that quantifies the positive degree (or negative degree) of the user's emotion. Hereinafter, it is assumed that the positive / negative degree is represented as a higher numerical value, a more positive emotion, and a lower numerical value, a more negative emotion. However, the positive / negative degree is merely an example, and other index values that can represent the user's emotion (e.g., joy, anger, sadness, happiness, positive / negative degree, etc.) may be used.
[0121] The gap recognition unit 212 of the gap estimation unit 203 recognizes a gap using the first text and the emotion of the first user (step S103). That is, the gap recognition unit 212 calculates the type of gap between the first text and the second user and the presence or absence of a gap or a gap degree score using the first text and the emotion of the first user.
[0122] For example, the gap type may be calculated based on the type of index value representing emotion, and the presence or absence of a gap or the gap degree score may be calculated based on a change in emotion. For example, in this embodiment, the gap type may be "positive / negative degree," and the gap degree score may be calculated as |current positive / negative degree-previous positive / negative degree|. When calculating the presence or absence of a gap, if the gap degree score is greater than a threshold value, it is determined that there is a gap; otherwise, it is determined that there is no gap. The threshold value is a preset value.
[0123] The presented information acquisition unit 204 determines whether or not it is necessary to fill the gap (step S104). For example, when the gap degree score is calculated in the above step S103, the presented information acquisition unit 204 determines that it is necessary to fill the gap if the gap degree score exceeds a preset threshold and the current positive / negative degree - the previous positive / negative degree < 0, and determines that it is not necessary to fill the gap if this is not the case. Also, for example, when the presence or absence of a gap is calculated in the above step S103, the presented information acquisition unit 204 may determine that it is necessary to fill the gap if there is a gap, and that it is not necessary to fill the gap if this is not the case.
[0124] If it is determined in step S104 that the gap needs to be filled, the presented information acquisition unit 204 acquires first presented information from the intervention DB 209 according to the gap type and the presence or absence of a gap or the gap degree score (step S105). For example, if the gap degree score is calculated in step S103, the presented information acquisition unit 204 acquires first presented information corresponding to the gap type "positive / negative degree" and the gap degree score calculated in step S103 from the intervention DB 209. On the other hand, for example, if the presence or absence of a gap is calculated in step S103, the presented information acquisition unit 204 acquires first presented information corresponding to the gap type "positive / negative degree" from the intervention DB 209.
[0125] If it is determined in step S104 that there is no need to fill the gap, the presentation information acquisition unit 204 ends the intervention process.
[0126] The first user presentation unit 205 presents the first presentation information acquired in step S105 to the first user device 20 (step S106).
[0127] Example 2 An example of an intervention process in which a gap is recognized using second user information and, if the gap needs to be filled, second presentation information is presented to the second user will be described below with reference to Fig. 12. The intervention process shown in Fig. 12 is repeatedly executed at predetermined time intervals (e.g., every few seconds to every few tens of seconds).
[0128] The second user information acquisition unit 202 acquires second user information from the second user device 30 (step S201). The second user information is stored in the second user history DB 208.
[0129] The multimodal recognition unit 211 of the gap estimation unit 203 recognizes second recognition information using the second user information stored in the second user history DB 208 (step S202). Hereinafter, as an example, it is assumed that the second recognition information is a combination of a second text representing a speech-recognized text of the voice included in the second user information acquired in step S201 and the current emotion of the second user. It is also assumed that the emotion of the second user is expressed as a positive / negative degree.
[0130] The gap recognition unit 212 of the gap estimation unit 203 recognizes a gap using the second text and the emotion of the second user (step S203). That is, the gap recognition unit 212 calculates the type of gap between the first user and the second user, and the presence or absence of a gap or a gap degree score, using the second text and the emotion of the second user.
[0131] For example, the gap type may be calculated based on the type of index value representing emotion, and the presence or absence of a gap or the gap degree score may be calculated based on a change in emotion. For example, in this embodiment, the gap type may be "positive / negative degree," and the gap degree score may be calculated as |current positive / negative degree-previous positive / negative degree|. When calculating the presence or absence of a gap, if the gap degree score is greater than a threshold value, it is determined that there is a gap; otherwise, it is determined that there is no gap. The threshold value is a preset value.
[0132] The presented information acquisition unit 204 determines whether or not it is necessary to fill the gap (step S204). For example, when the gap degree score is calculated in the above step S203, the presented information acquisition unit 204 determines that it is necessary to fill the gap if the gap degree score exceeds a preset threshold and the current positive / negative degree - the previous positive / negative degree < 0, and determines that it is not necessary to fill the gap if this is not the case. Also, for example, when the presence or absence of a gap is calculated in the above step S203, the presented information acquisition unit 204 may determine that it is necessary to fill the gap if there is a gap, and that it is not necessary to fill the gap if this is not the case.
[0133] If it is determined in step S204 that the gap needs to be filled, the presented information acquisition unit 204 acquires first presented information from the intervention DB 209 according to the gap type and the presence or absence of a gap or the gap degree score (step S205). For example, if the gap degree score is calculated in step S203, the presented information acquisition unit 204 acquires second presented information corresponding to the gap type "positive / negative degree" and the gap degree score calculated in step S203 from the intervention DB 209. On the other hand, for example, if the presence or absence of a gap is calculated in step S203, the presented information acquisition unit 204 acquires second presented information corresponding to the gap type "positive / negative degree" from the intervention DB 209.
[0134] If it is determined in step S204 that there is no need to fill the gap, the presentation information acquisition unit 204 ends the intervention process.
[0135] The second user presentation unit 206 presents the second presentation information acquired in step S205 to the second user device 30 (step S206).
[0136] Example 3 An example of an intervention process in which a gap is recognized using first user information and second user information, and if the gap needs to be filled, first presentation information is presented to the first user will be described below with reference to Fig. 13. The intervention process shown in Fig. 13 is repeatedly executed at predetermined time intervals (e.g., every few seconds to every few tens of seconds).
[0137] The first user information acquisition unit 201 acquires first user information from the first user device 20 (step S301). The first user information is stored in the first user history DB 207.
[0138] The multimodal recognition unit 211 of the gap estimation unit 203 recognizes first recognition information using the first user information stored in the first user history DB 207 (step S302). Hereinafter, as an example, it is assumed that the first recognition information is a first text representing a speech-recognized text of the voice included in the first user information acquired in step S301 and the current emotion of the first user. It is also assumed that the emotion of the first user is expressed as a positive / negative degree.
[0139] The second user information acquisition unit 202 acquires the second user information from the second user device 30 (step S303). The second user information is stored in the second user history DB 208.
[0140] The multimodal recognition unit 211 of the gap estimation unit 203 recognizes second recognition information using the second user information stored in the second user history DB 208 (step S304). As in step S302, the second recognition information is assumed to be a combination of a second text representing a speech-recognized text of the speech included in the second user information acquired in step S303 and the current emotion of the second user. The emotion of the second user is assumed to be expressed as a positive / negative degree.
[0141] The above steps S301 to S302 and steps S303 to S304 may be performed in any order. That is, steps S301 to S302 may be performed after steps S303 to S304. Furthermore, steps S301 to S302 and steps S303 to S304 may be performed in parallel.
[0142] The gap recognition unit 212 of the gap estimation unit 203 recognizes a gap using the first text, the first user's emotion, and the second text, and the second user's emotion (step S305). That is, the gap recognition unit 212 calculates the type of gap between the first user and the second user, and the presence or absence of a gap or a gap degree score, using the first text, the first user's emotion, and the second text, and the second user's emotion.
[0143] For example, the gap type may be the type of index value representing emotion, and the presence or absence of a gap or gap degree score may be calculated from the difference between the emotion of the first user and the emotion of the second user. For example, in this embodiment, the gap type may be "positive / negative degree," and the gap degree score may be |current positive / negative degree of the first user--current positive / negative degree of the second user|. When calculating the presence or absence of a gap, if the gap degree score is greater than a threshold value, it is determined that there is a gap; otherwise, it is determined that there is no gap. The threshold value is a preset value.
[0144] The presented information acquisition unit 204 determines whether or not it is necessary to fill the gap (step S306). For example, when the gap degree score is calculated in step S305, the presented information acquisition unit 204 determines that it is necessary to fill the gap if the gap degree score exceeds a preset threshold and the current positive / negative degree of the first user minus the current positive / negative degree of the second user is less than 0, and determines that it is not necessary to fill the gap otherwise. Also, for example, when the presence or absence of a gap is calculated in step S305, the presented information acquisition unit 204 determines that it is necessary to fill the gap if it is present, and determines that it is not necessary to fill the gap otherwise.
[0145] If it is determined in step S306 that the gap needs to be filled, the presented information acquisition unit 204 acquires first presented information from the intervention DB 209 according to the gap type and the presence or absence of a gap or the gap degree score (step S307). For example, if the gap degree score is calculated in step S305, the presented information acquisition unit 204 acquires first presented information corresponding to the gap type "positive / negative degree" and the gap degree score calculated in step S305 from the intervention DB 209. On the other hand, for example, if the presence or absence of a gap is calculated in step S305, the presented information acquisition unit 204 acquires first presented information corresponding to the gap type "positive / negative degree" from the intervention DB 209.
[0146] If it is determined in step S306 above that there is no need to fill the gap, the presentation information acquisition unit 204 ends the intervention process.
[0147] The first user presentation unit 205 presents the first presentation information acquired in step S307 to the first user device 20 (step S308).
[0148] Example 4 An example of an intervention process in which a gap is recognized using first user information and second user information, and second presentation information is presented to a second user if the gap needs to be filled will be described below with reference to Fig. 14. The intervention process shown in Fig. 14 is repeatedly executed at predetermined time intervals (e.g., every few seconds to every few tens of seconds).
[0149] Steps S401 to S406 in FIG. 14 may be similar to steps S301 to S306 in FIG. 13, respectively, and therefore a description thereof will be omitted.
[0150] If it is determined in step S406 of Fig. 14 that the gap needs to be filled, the presented information acquisition unit 204 acquires second presented information from the intervention DB 209 according to the gap type and the presence or absence of a gap or the gap degree score (step S407). For example, if the gap degree score is calculated in step S405 of Fig. 14, the presented information acquisition unit 204 acquires second presented information corresponding to the gap type and gap degree score calculated in step S405 of Fig. 14 from the intervention DB 209. On the other hand, for example, if the presence or absence of a gap is calculated in step S405 of Fig. 14, the presented information acquisition unit 204 acquires second presented information corresponding to the gap type calculated in step S405 of Fig. 14 from the intervention DB 209.
[0151] If it is determined in step S406 of FIG. 14 that the gap does not need to be filled, the presentation information acquiring unit 204 ends the intervention process.
[0152] The second user presentation unit 206 presents the second presentation information acquired in step S407 to the second user device 30 (step S408).
[0153] Example 5: Hereinafter, an example of an intervention process will be described with reference to FIG. 15 in which a gap is recognized using first user information and second user information, and if it is necessary to fill the gap, first presented information is presented to the first user and second presented information is presented to the second user.
[0154] Steps S501 to S506 in FIG. 15 may be similar to steps S401 to S406 in FIG. 14, respectively, and therefore a description thereof will be omitted.
[0155] 15 , when it is determined that the gap needs to be filled, the presented information acquiring unit 204 acquires first presented information and second presented information from the intervention DB 209 according to the gap type and the presence or absence of a gap or the gap degree score (step S507). For example, when the gap degree score is calculated in step S505 of Fig. 15 , the presented information acquiring unit 204 acquires first presented information and second presented information corresponding to the gap type and gap degree score calculated in step S505 of Fig. 15 from the intervention DB 209. On the other hand, when the presence or absence of a gap is calculated in step S505 of Fig. 15 , the presented information acquiring unit 204 acquires first presented information and second presented information corresponding to the gap type calculated in step S505 of Fig. 15 from the intervention DB 209.
[0156] If it is determined in step S506 of FIG. 15 that the gap does not need to be filled, the presentation information acquisition unit 204 ends the intervention process.
[0157] The first user presentation unit 205 presents the first presentation information acquired in step S507 to the first user device 20 (step S508).
[0158] The second user presentation unit 206 presents the second presentation information acquired in step S507 to the second user device 30 (step S509).
[0159] The order of step S508 and step S509 is not limited to this. That is, step S508 may be executed after step S509. Furthermore, step S508 and step S509 may be executed in parallel.
[0160] <Modification of Functional Configuration of Intervention Device 10> Fig. 16 shows a modification of the functional configuration of the intervention device 10 according to the present embodiment. As shown in Fig. 16, the intervention device 10 in this modification has a presentation information acquisition unit 204A instead of the presentation information acquisition unit 204. The presentation information acquisition unit 204A is realized, for example, by a process in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like. Furthermore, the intervention device 10 in this modification does not have the intervention plan DB 209.
[0161] The presentation information acquisition unit 204A determines whether the gap needs to be filled, and if it determines that the gap needs to be filled, generates at least one of first presentation information and second presentation information according to the gap type, the presence or absence of a gap, or the gap degree score. Here, the presentation information acquisition unit 204A includes a trained machine learning model that receives the gap type and the presence or absence of a gap or the gap degree score as input and outputs at least one of the first presentation information and the second presentation information, and the presentation information acquisition unit 204A generates at least one of the first presentation information and the second presentation information using this trained machine learning model. Note that such a machine learning model is realized using a large-scale language model (LLM) or the like, and may also be referred to as, for example, a "generative model," "generative AI," or "generative AI." Hereinafter, a trained machine learning model that receives the gap type and the presence or absence of a gap or the gap degree score as input and outputs at least one of the first presentation information and the second presentation information will be referred to as a "presentation information generation model."
[0162] <Modification of Intervention Process> Hereinafter, as a modification, an example of the intervention process in the case where the presentation information acquisition unit 204A generates the first presentation information in the above-described first embodiment will be described with reference to Fig. 17. The intervention process shown in Fig. 17 is repeatedly executed at predetermined time intervals (e.g., every few seconds to every few tens of seconds).
[0163] Steps S601 to S604 and step S606 in FIG. 17 may be similar to steps S101 to S104 and step S106 in FIG. 11, respectively, and therefore a description thereof will be omitted.
[0164] 17 , when it is determined that the gap needs to be filled, the presentation information acquiring unit 204A generates first presentation information according to the gap type and the presence or absence of a gap or the gap degree score (step S605). For example, when the gap degree score is calculated in step S603 of Fig. 17 , the presentation information acquiring unit 204A inputs the gap type and the gap degree score into the presentation information generation model and generates the first presentation information as its output. On the other hand, when the presence or absence of a gap is calculated in step S603 of Fig. 17 , the presentation information acquiring unit 204A inputs the gap type and the presence or absence of a gap into the presentation information generation model and generates the first presentation information as its output.
[0165] In the intervention process shown in FIG. 17, the case where the first presentation information in Example 1 is generated by the presentation information acquisition unit 204A has been described. However, the presentation information acquisition unit 204A can also generate the presentation information (the first presentation information, the second presentation information, or both) in Examples 2 to 5 in the same manner.
[0166] <Other Modifications of the Interventional Device 10> The following describes other modifications of the above-described interventional device 10. The following modifications can be combined as appropriate as long as they are not inconsistent with each other.
[0167] Variation 1-1: Instead of the first presentation information and the second presentation information, templates thereof may be stored in the intervention DB 209. In this case, if a user's label (e.g., emotion label, attitude label, etc.) can be recognized as recognition information by the multimodal recognition unit 211, the presentation information acquisition unit 204 may create at least one of the first presentation information and the second presentation information using the label and the template stored in the intervention DB 209.
[0168] For example, assume that a template of second presentation information "The customer's emotion is *" is stored in the intervention DB 209 and the label "anger" is recognized. In this case, the presentation information acquisition unit 204 may create second presentation information "The customer's emotion is 'anger'" by replacing the "*" part of the template with the label "anger."
[0169] Variation 1-2: Because the presentation information generation model is realized by a large-scale language model or the like, the presentation information acquisition unit 204A may generate, in addition to the presentation information, the basis or reason for generating the presentation information (hereinafter, the basis and reason are also collectively referred to as "basis, etc."). Note that the basis, etc. for generating the presentation information can be generated, for example, by providing an instruction (also referred to as a "prompt, etc.") for generating the basis, etc. of the presentation information to the presentation information generation model.
[0170] Variation 1-3: Because a multimodal recognition model can also be realized by a large-scale language model or the like, the multimodal recognition unit 211 may generate, in addition to the recognition information, a basis for recognizing the recognition information, etc. Note that the basis for recognizing the recognition information, etc. can be generated, for example, by providing an instruction (prompt) to the multimodal recognition model to generate a basis for the recognition information, etc.
[0171] Variation 1-4: The gap estimation unit 203 may calculate an index value (hereinafter also referred to as a "first overall evaluation value") that comprehensively evaluates the degree of mutual understanding in communication between the first user and the second user. The first overall evaluation value may be calculated from the first recognition information, the second recognition information, and the gap degree score. For example, the first overall evaluation value may be calculated by calculating the average value or weighted average value of the gap degree scores calculated for each gap type, and then assigning a sign determined from the first recognition information and the second recognition information to the average value or weighted average value. As a specific example, if there is only one gap type, "positive / negative degree," the first overall evaluation value may be calculated by assigning a positive sign to the gap degree score calculated from the positive / negative degree if the positive / negative degree of the second user minus the positive / negative degree of the first user is ≧0, or a negative sign otherwise.
[0172] Modification 1-5 The gap estimation unit 203 may use the user information of each user to calculate an index value (hereinafter also referred to as a "second overall evaluation value") that comprehensively evaluates the emotions, etc. of the user. The second overall evaluation value may be calculated from the recognition information of the user. Specifically, when the recognition information includes index values that quantify emotions for each type of emotion, the average value, weighted average value, etc. of these index values may be calculated as the second overall evaluation value.
[0173] In the above embodiment, the case where communication is performed between a first user and a second user has been described, but the present invention can also be applied to a case where communication is performed between three or more users. Note that when communication is performed between N users, the intervention system 1 includes N user devices.
[0174] Modification 1-7 The gap estimation unit 203 and the presentation information acquisition unit 204A may be realized by a machine learning model such as a large-scale language model.
[0175] Modification 1-8 When a functional unit is realized by a machine learning model, the machine learning model may be available via the communication network 40, for example, via a Web API or the like.
[0176] In the first embodiment, the first user information is acquired and then the first presentation information is presented to the first user. However, for example, the second presentation information may be presented to the second user after the first user information is acquired. That is, the second presentation information may be acquired from the intervention DB 209 in step S105.
[0177] In the second embodiment, the second user information is acquired and then the second presentation information is presented to the second user. Alternatively, the first presentation information may be presented to the first user after the second user information is acquired. That is, the first presentation information may be acquired from the intervention DB 209 in step S205.
[0178] Modification 1-11 In the above embodiment, presentation information for eliminating the gap (i.e., "good presentation information") is presented to the user. However, in addition to this, presentation information that is thought to cause a greater gap (i.e., "bad presentation information") may also be presented. This allows the user to know both the good presentation information and the bad presentation information, which is expected to lead to better communication. Note that bad presentation information may be associated with good presentation information and stored in the intervention DB 209, and when good presentation information is acquired, bad presentation information related to the good presentation information may also be acquired from the intervention DB 209.
[0179] Modification 1-12 The functional configurations of the intervention device 10 shown in FIGS. 8 and 16 are merely examples, and the functions and roles of each unit may be modified as appropriate. For example, in the intervention device 10 shown in FIG. 8, the presentation information acquisition unit 204 exists separately from the gap estimation unit 203, but the presentation information acquisition unit 204 may be included in the gap estimation unit 203. Similarly, in the intervention device 10 shown in FIG. 16, the presentation information acquisition unit 204A exists separately from the gap estimation unit 203, but the presentation information acquisition unit 204A may be included in the gap estimation unit 203. As another example, for example, the gap estimation unit 203 and the presentation information acquisition unit 204 (or the presentation information acquisition unit 204A) may be collectively referred to as an "intervention estimation unit" or the like. As another example, for example, the gap estimation unit 203 shown in FIG. 9 includes a multimodal recognition unit 211, but the multimodal recognition unit 211 may be located outside the gap estimation unit 203. In addition to these, it is possible to appropriately combine the functional units or divide each functional unit into a plurality of functional units as needed.
[0180] Modification 1-13 All or some of the functional units of the interventional device 10 shown in FIGS. 8 and 16 may be included in, for example, a server (e.g., a cloud server) communicatively connected to the interventional device 10.
[0181] <Example of UI> Below, examples of UIs (user interfaces) displayed on displays, etc. of the first user device 20 and the second user device 30 will be described. The UIs described below are all displayed on displays, etc. of the first user device 20 and the second user device 30 by the first user presentation unit 205 and the second user presentation unit 206, respectively. However, for example, the intervention device 10 may have functional units such as a "UI control unit" and an "output unit," and these UI control units and output units may display the UIs on displays, etc. of the first user device 20 and the second user device 30. Note that all or part of the UIs displayed on displays, etc. of the first user device 20 and the second user device 30 may be different. In particular, the UIs displayed on displays, etc. of the user devices used by the users may differ depending on the type of user (e.g., whether the user is a customer or an operator). However, the customer and the operator are examples of user types, and the types of users are not limited thereto. Other examples of user types include types based on the user's position (e.g., job title, etc.), role (e.g., job content, etc.), position (e.g., sales or customer, etc.), and types based on the verbal interaction situation (e.g., the one who made the gaffe or the one who received the gaffe).
[0182] 18 includes a first user display field 1101 in which an image or video captured of a first user is displayed, and a second user display field 1102 in which an image or video captured of a second user is displayed. Note that a character (a so-called avatar) representing the first user may be displayed in the first user display field 1101. Similarly, a character (a so-called avatar) representing the second user may be displayed in the second user display field 1102.
[0183] 18 includes an utterance display field 1103 displaying a first text representing an utterance of a first user and a second text representing a second utterance, a gap degree score display field 1104 displaying a gap degree score, and a presented information display field 1105 displaying presented information. In the UI 1100 shown in FIG. 18 , presented information 1111 and presented information 1112 are displayed in the presented information display field 1105. Note that only the first presented information is displayed in the presented information display field 1105 of the UI 1100 displayed by the first user device 20, and only the second presented information is displayed in the presented information display field 1105 of the UI 1100 displayed by the second user device 30. However, both the first presented information and the second presented information may be displayed in the presented information display field 1105.
[0184] 18 includes a first facial emotion display field 1106 that displays information about emotions recognized from the facial image of the first user, and a first voice emotion display field 1107 that displays information about emotions recognized from the voice of the first user. Similarly, the UI 1100 shown in FIG. 18 includes a second facial emotion display field 1108 that displays information about emotions recognized from the facial image of the second user, and a second voice emotion display field 1109 that displays information about emotions recognized from the voice of the second user.
[0185] UI Example (Part 2) The UI 1200 shown in Fig. 19 includes a gap presence / absence display field 1201 that displays whether or not there is a gap, a gap type display field 1202 that displays the gap type, and a recognition information display field 1203 that displays recognition information recognized from the user's own user information. The UI 1200 shown in Fig. 19 also includes a basis etc. display field 1204 that displays the basis etc. for the recognition information being recognized, and a presented information display field 1205 that displays presented information to the user. The basis etc. for generating the presented information can be generated by the presented information acquisition unit 204A in the above-mentioned modified example 1-2.
[0186] UI Example (3) When the basis etc. display field 1204 included in the UI 1200 shown in Fig. 19 is selected by the user, a UI 1300 shown in Fig. 20 may be displayed. The UI 1300 shown in Fig. 20 displays instructions (prompts) given to the multimodal recognition model when generating basis etc. for recognition of the recognition information displayed in the recognition information display field 1203, and responses to the prompts. The UI 1300 shown in Fig. 20 includes an instruction 1301, a response 1302 to the instruction 1301, an instruction 1303, and a response 1304 to the instruction 1303.
[0187] UI Example (4) The UI 1400 shown in Fig. 21 displays a line graph 1401 that represents a time series of the first overall index value calculated by the gap estimation unit 203 in the above-described modified example 1-4. Here, in the UI 1400 shown in Fig. 21, the closer the graph 1401 is to the center in the horizontal direction (in other words, the closer the first overall index value is to 0), the higher the level of mutual understanding that is achieved in communication between the first user and the second user. This enables the first user and the second user to communicate in a way that further enhances their level of mutual understanding.
[0188] Note that displaying the time series of the first overall index value as a line graph is just one example, and is not limited to line graphs. For example, the time series of the first overall index value may be displayed as a bar graph, a heat map, or the like.
[0189] UI Example (5) The UI 1500 shown in FIG. 22 displays line graphs 1501 and 1502 representing the time series of the second overall index values calculated by the gap estimation unit 203 in the above-described variants 1-5. Graph 1501 represents the time series of the second overall index values of the first user, and graph 1502 represents the time series of the second overall index values of the second user. In the UI 1500 shown in FIG. 22, the closer the graphs 1501 and 1502 are to the center in the horizontal direction, the greater the emotional compatibility between the first and second users. This allows the first and second users to communicate in a way that enhances their mutual understanding while taking into account the degree of emotional compatibility between them.
[0190] Note that displaying the time series of the second overall index value as a line graph is just one example, and is not limited to line graphs. For example, the time series of the second overall index value may be displayed as a bar graph, a heat map, or the like.
[0191] UI Example (No. 6) The UI 1600 shown in Fig. 23 displays a polygonal graph 1601 showing the current second overall index values of each of the first to sixth users. Each vertex of the graph 1601 represents the current second overall index value of each user. In the UI 1600 shown in Fig. 23, the closer the shape of the graph 1601 is to a circle, the greater the emotional compatibility between the users. This allows each user to communicate in a way that enhances mutual understanding while taking into account the degree of emotional compatibility between them.
[0192] 23 displays the current second overall index value of each user as graph 1601, but may also display, for example, a graph representing each user's past second overall index value. In this case, the graph representing each user's past second overall index value may be displayed in a manner different from graph 1601 (e.g., a manner in which the color of the graph representing the older second overall index value is lighter, etc.).
[0193] UI Example (7) The UI 1700 shown in FIG. 24 is a UI displayed on the second user device 30 in the above-described variant 1-11. It includes a face image display field 1701 displaying a face image of the other party (first user) and an AI avatar display field 1702 displaying an AI avatar that anthropomorphizes the intervention device 10. The UI 1700 shown in FIG. 24 also includes a comment display field 1703 displaying at least one of good and bad presentation information for the second user as a comment from the AI avatar, and a gap degree display field 1704 displaying a gap degree score in a predetermined format (e.g., numerical value, icon, gauge, graph, etc.). The comment display field 1703 may display information such as FAQs in addition to at least one of good and bad presentation information for the second user. This allows the second user to know the good and bad presentation information, thereby enabling communication with the first user that enhances mutual understanding.
[0194] UI Example (No. 8) The UI 1800 shown in FIG. 25 is a UI displayed on the first user device 20 in the above-described variation 1-11. It includes a face image display field 1801 displaying a face image of the other party (second user) and an AI avatar display field 1802 displaying an AI avatar that anthropomorphizes the intervention device 10. The UI 1800 shown in FIG. 25 also includes a comment display field 1803 displaying at least one of good and bad presentation information for the first user as a comment from the AI avatar, and a gap degree display field 1804 displaying a gap degree score in a predetermined format (e.g., numerical value, icon, gauge, graph, etc.). The comment display field 1803 may display information such as FAQs in addition to at least one of good and bad presentation information for the first user. This allows the first user to know the good and bad presentation information, thereby enabling communication with the second user that enhances mutual understanding.
[0195] <Modifications Regarding UI> Modifications regarding the above UI will be described below. Note that the following modifications can be combined as appropriate as long as they do not contradict each other.
[0196] Modification 2-1 The UI may be generated by the interventional device 10, or may be generated by a user device (such as the first user device 20 or the second user device 30), and the display of the UI may be updated based on information (such as presentation information) received from the interventional device 10. When the interventional device 10 generates the UI, the interventional device 10 may have a functional unit such as a "UI generation unit."
[0197] Modification 2-2 A UI may be displayed as a pop-up on another UI that is already displayed on the display of a user device (such as the first user device 20 or the second user device 30).
[0198] Variation 2-3: In addition to the presented information, the UI displayed on the display of a user device used by an operator or the like in a call center may also display information related to the presented information (e.g., Q&A related to the presented information, customer attribute information, etc.).
[0199] <Application Examples> The above embodiment has been described mainly with reference to a call center, but the above embodiment and its modifications can be applied to various fields. Some application examples are given below.
[0200] Application Example 1: This application can be applied to the field of education, such as schools and cram schools (including schools and cram schools in virtual spaces), with the first user being a student and the second user being a teacher. This is expected to eliminate a gap in understanding between students and teachers, for example, and realize communication that improves the students' learning effectiveness. A specific example of a gap in understanding is when a teacher thinks they have "taught something clearly," but the students "find it difficult."
[0201] The number of students is not limited to one, and the present invention can be applied to cases where there are multiple students for one teacher. Indicators for evaluating gaps in the field of education, in addition to level of understanding, include, for example, level of assent, level of interest, concentration, and motivation of students and teachers. Other indicators that can be used include, for example, level of poker face, level of communication skills, smoothness of conversation, level of empathy, level of enthusiasm, level of positive / negative, and level of initiative. A specific example of a gap in assent may be when a teacher believes that he or she has "taught the truth," while the student perceives it as "empty theory." A specific example of a gap in interest may be when a teacher believes that the student is "listening with interest," while the student perceives it as "uninterested." A specific example of a gap in concentration may be when a teacher believes that the student is "listening attentively," while the student "feels sleepy." A specific example of a gap in student motivation is when a teacher thinks, "The students are quiet so they are listening to me," but the students themselves feel bored.
[0202] Application Example 2: When applying the present invention to the field of education in schools, cram schools, etc. (including schools, cram schools, etc. in a virtual space), with the first user as a student and the second user as a teacher, it is possible to visualize index values (e.g., weighted sums, etc.) that comprehensively evaluate each student's level of understanding, satisfaction, interest, motivation, smoothness of conversation, empathy, enthusiasm, positive / negative, proactiveness, etc., using groups, classes, etc. categorized by classroom, seating area, student attributes, etc. In this case, any visualization method can be used, and examples include visualization using a heat map.
[0203] Application Example 3: While Application Example 2 above targets schools, cram schools, etc., it is also conceivable to target, for example, concerts and live venues (including concerts and live venues in virtual spaces) and visualize index values (e.g., weighted sums) that comprehensively evaluate the level of excitement, etc., as a heat map.
[0204] Application Example 4: This application can be applied to fields such as employee education and training, with the first user acting as an employee and the second user acting as an instructor. As with Application Example 1, this can be expected to eliminate any gaps in understanding between employees and instructors, thereby realizing communication that enhances the effectiveness of employee training. This application can also be applied to group work-style training, for example, where multiple employees are grouped together. In this case, in addition to the gap in understanding, other factors that can be used include the level of enthusiasm in the training or group work within it, the level of smoothness of conversation, the level of empathy, the level of positive / negative, and the level of initiative. Furthermore, in this case, for example, an index value (e.g., a weighted sum) that comprehensively evaluates the level of enthusiasm, smoothness, empathy, positive / negative, and initiative may be used, and the time series of the index value may be visualized as a graph.
[0205] Application Example 5: This can be applied to corporate meetings, with each user being an employee. This can be expected to realize communication that enhances mutual understanding between employees, leading to more productive meetings.
[0206] Application Example 6: With the first user as a customer and the second user as a salesperson, this can be applied to fields such as over-the-counter sales at a store or online sales (e.g., banking, insurance, real estate, etc.). This is expected to lead to communication that enhances mutual understanding between the customer and the salesperson, which can lead to more effective sales and sales that provide higher customer satisfaction.
[0207] The gaps that can be used include, for example, a gap in understanding, a gap in satisfaction, a gap in interest, and whether or not the customer is focused. In addition to these, a gap in sales impression can also be used. The sales impression is a label that represents the impression of a salesperson, and can take values such as gentlemanly, sincere, powerful, calm, driven, and amiable.
[0208] Furthermore, different index values may be used for evaluating gaps depending on the sales phase (e.g., first approach, interview, presentation, closing, etc.). For example, in the first approach, the gap may be evaluated based on the sales impression, in the interview, the gap may be evaluated based on the level of understanding, and in the presentation, the gap may be evaluated based on the level of satisfaction and interest. Similarly, different index values may be used for evaluating gaps depending on the type of salesperson or the type of customer. Here, the type of salesperson is a category in which the salesperson is classified based on their sales techniques, and examples include analytical, driving, amiable, and expressive. On the other hand, the type of customer is a category in which the customer is classified based on their words and actions, and for example, a category called social style is similarly classified into four types: analytical, driving, amiable, and expressive.
[0209] Application Example 7: The present invention can be applied to fields such as career counseling, with the first user being an employee, subordinate, or job seeker, and the second user being a human resources officer, supervisor, career consultant, counselor, etc. This can be expected to enhance communication between, for example, the employee, subordinate, or job seeker and the human resources officer, supervisor, career consultant, or counselor, thereby enabling more appropriate career counseling. Career counseling may be conducted one-on-one between the first user and the second user, between multiple first users and the second user, or between multiple first users and multiple second users. Various indicators can be used to evaluate the gap, such as trust, personality traits, poker face level, presence or absence of harassment, number of fillers, gaze direction, gaze movement amount, average gaze movement amount, gestures, clothing, belongings, etc.
[0210] In application example 7, a specific example of a gap could be when the HR department "thinks an employee looks timid based on their appearance and conducts an emotional consultation," but the employee is dissatisfied because the employee's personality is "cheerful, positive, and logical," resulting in a gap. Another example could be when a career consultant intends to "provide sympathetic advice," but the job seeker feels "untrustworthy because the consultant's gaze, blinking, and gestures bother him," resulting in a gap. Another example could be when a supervisor intends to "conduct an appropriate personnel interview," but the subordinate feels that "this is power harassment or sexual harassment," resulting in a gap.
[0211] Application Example 8: The medical field can be applied to the nursing care field, etc., with the first user being a patient or a care recipient and the second user being a doctor, nurse, pharmacist, caregiver, etc. For example, by using the level of understanding or the level of assent as indicators for evaluating the gap, it is possible to support the formation of informed consent between a doctor and a patient. Specifically, for example, when a doctor "explains bluntly and without facing the patient," the patient's dissatisfaction that "the content of the explanation is unreliable" can be detected as a gap. As another example, it is also possible to support the determination of whether or not a second opinion is necessary. Specifically, for example, it is possible to detect a gap when a doctor thinks that he or she "explained properly," but the patient feels that "the opinion of another doctor or expert, etc., is also considered."
[0212] Other Application Examples In addition to the above application examples, the above embodiments and their variations can be similarly applied to any field in which communication occurs between people, people, or machines. Examples of communication between people and machines include communication between an e-commerce site and its users, communication between a person and an in-vehicle device when driving a car, and communication for controlling a machine using human voice. Examples of communication between machines include communication between a vehicle and a roadside unit in vehicle control using autonomous driving technology, and communication between autonomously operating robots.
[0213] Examples of other application fields to which the intervention device 10 according to this embodiment can be applied include document creation in government and local governments, meetings (e.g., brainstorming for inventions) and the creation and evaluation of various documents (e.g., creating and evaluating patent specifications, creating contracts and financial statements) in the legal profession (e.g., patent attorneys, accountants, etc.), new product development activities and promotions in the food and beverage industry, real estate asset value calculation work in the real estate industry, analysis of cases and drugs in the medical field, X-ray checks, content creation and effect analysis in web marketing, website construction, various surveys and document creation in the consulting industry, curriculum consideration in the education and training field, interactions with customers in the finance and insurance industry, and analysis of surveillance cameras and investigation information in the police.
[0214] In addition to the above, this technology can also be applied to the field of education, where questions at an appropriate level are assigned to each student. For example, in a class where many students are taking the same class, it is possible to detect gaps between each student and the question (e.g., gaps in understanding), and by deterministically or probabilistically changing the next question to be applied depending on those gaps, it is possible to assign questions at an appropriate level to each student.
[0215] In addition to the above, the present invention can also be applied to situations such as dialogue with so-called AI (i.e., robots that apply artificial intelligence technology, digital content such as animation, etc.). Note that dialogue with AI may include not only dialogue between humans and AI, but also dialogue between AIs. Note that the facial expressions and movements of digital content (avatars and characters) that visually represent AI may be anthropomorphized according to the content and degree of the gap. Similarly, the facial expressions and movements of robots that use AI may be anthropomorphized.
[0216] As a specific example, the intervention device 10 according to the present embodiment can be applied in a situation where a person and an AI interact via a smartphone, digital signage, etc. Similarly, the intervention device 10 according to the present embodiment can be applied in a situation where AIs interact and cooperate with each other to perform some processing.
[0217] <Example of Functional Configuration of Learning Device 50 that Learns a Gap Recognition Model> FIG. 26 shows an example of the functional configuration of a learning device 50 that learns a gap recognition model that implements the gap recognition unit 212. As shown in FIG. 26 , the learning device 50 has an input unit 301, the gap recognition unit 212, and a parameter update unit 302. These units are implemented, for example, by a processor or the like executing one or more programs installed in the learning device 50. The learning device 50 also has a learning data DB 303. The learning data DB 303 is implemented, for example, by a storage area of an auxiliary storage device or the like. However, the learning data DB 303 may also be implemented by a storage area of a storage device (e.g., a storage device included in a database server) connected to the learning device 50 via a communication network. The learning device 50 may be implemented by the same device as the intervention device 10, or by a device different from the intervention device 10.
[0218] The input unit 301 inputs learning data (hereinafter also referred to as "first learning data") stored in the learning data DB 303. The first learning data is data for training the gap recognition model, and includes input data to be input to the gap recognition model and training data for the input data. The training data is the correct answer for the presence or absence of a gap or the gap degree score output from the gap recognition model when input data is input.
[0219] As an example, let us assume the above-mentioned Example 2, where the input data is second recognition information, and refer to the gap degree score output from the gap recognition model when the input data is input as the "gap degree score estimate."
[0220] The parameter update unit 302 updates the learnable parameters of the gap recognition model using the error between the gap degree score estimate recognized by the gap recognition unit 212 and the training data included in the first training data.
[0221] The learning data DB 303 stores first learning data. Note that the first learning data is created in advance by, for example, the creator of the gap recognition model.
[0222] <<Example of First Learning Data>> An example of the first learning data is shown in FIG. 27. As shown in FIG. 27, the first learning data includes input data and training data. The input data is data representing recognition information to be input to the gap recognition model. The input data included in the first learning data shown in FIG. 27 represents second recognition information consisting of text representing a slip of the tongue of a second user and the second user's emotions at that time. The training data is data representing the correct answer to the gap degree score. The value of the gap degree score used as training data is determined, for example, by the first user who made the slip of the tongue or a third party who heard it.
[0223] 27, the input data is the second recognition information consisting of text representing a slip of the second user's tongue and the first user's emotion at that time. However, this is merely an example. For example, the input data may be only text representing a certain utterance of the second user, or the input data may be text representing a certain utterance of the second user and some identification information of the second user at that time (e.g., level of understanding, level of interest, level of concentration, etc.). As another example, the input may be a group of text representing an utterance of the second user at a certain position and a change in some identification information of the second user at that time (e.g., change in emotion, change in level of understanding, change in level of interest, change in level of concentration, etc.).
[0224] <Example of Learning Process for Learning Gap Recognition Model> An example of the learning process for learning the gap recognition model will be described below with reference to Fig. 28. The learning process shown in Fig. 28 is repeatedly executed until a predetermined condition is satisfied. Here, the predetermined condition may be, for example, that the number of repetitions has reached a certain number or more, or that the learnable parameters of the gap recognition model have converged, etc.
[0225] The input unit 301 inputs the first learning data stored in the learning data DB 303 (step S701).
[0226] The gap recognition unit 212 recognizes a gap using the input data included in the first learning data input in step S701 (step S702). That is, the gap recognition unit 212 inputs the input data to a gap recognition model and obtains a gap degree score estimate as its output.
[0227] The parameter update unit 302 updates the learnable parameters of the gap recognition model using the error between the gap degree score estimate obtained in step S702 and the training data included in the first training data input in step S701 (step S703). Note that the parameter update unit 302 may use a known optimization method to update the learnable parameters so as to minimize the error as an objective function. For example, the mean square error may be used as the error.
[0228] <Example of Functional Configuration of a Learning Device 50 for Learning a Presentation Information Generation Model> FIG. 29 shows an example of the functional configuration of a learning device 50 for learning a presentation information generation model that implements a presentation information acquisition unit 204A. As shown in FIG. 29 , the learning device 50 includes an input unit 301A, a gap estimation unit 203, a presentation information acquisition unit 204A, and a parameter update unit 302A. Each of these units is implemented, for example, by a processor or the like executing one or more programs installed in the learning device 50. The learning device 50 also includes a learning data DB 303A. The learning data DB 303A is implemented, for example, by a storage area of an auxiliary storage device or the like. However, the learning data DB 303A may also be implemented by a storage area of a storage device (e.g., a storage device included in a database server) connected to the learning device 50 via a communication network. The learning device 50 may be implemented by the same device as the intervention device 10, or by a device different from the intervention device 10.
[0229] The input unit 301A inputs learning data (hereinafter also referred to as "second learning data") stored in the learning data DB 303A. The second learning data is data for learning the presentation information generation model, and includes input data input to the gap estimation unit 203 and training data for the input data. The training data is the correct answer of the presentation information output from the presentation information generation model when the input data is input to the gap estimation unit 203. Hereinafter, assuming a case where the presentation information generation model generates second presentation information in the above-mentioned second embodiment, the presentation information output from the presentation information generation model will be referred to as "estimated presentation information."
[0230] The parameter update unit 302A updates the learning target parameters of the presentation information generation model using the error between the generation probability of the text representing the estimated presentation information generated by the presentation information acquisition unit 204A and the training data included in the second learning data.
[0231] The training data DB 303A stores second training data. Note that the second training data is created in advance by, for example, the creator of the presentation information generation model.
[0232] <<Example of Second Training Data>> An example of the second training data is shown in FIG. 30 . As shown in FIG. 30 , the second training data includes input data and training data. The input data is data representing multimodal information input to the gap estimation unit 203. The input data included in the second training data shown in FIG. 30 is data representing multimodal information when the second user makes a slip of the tongue. The training data included in the second training data shown in FIG. 30 is text data representing the second user's apology for the slip of the tongue. Note that the text used as training data may be text actually uttered or input by the second user, or may be text edited by the second user or the creator of the presentation information generation model, or may be text created by a third party or the creator of the presentation information generation model, etc.
[0233] <Learning Process for Learning Presentation Information Generation Model> An example of the learning process for learning the presentation information generation model will be described below with reference to Fig. 31. The learning process shown in Fig. 31 is repeatedly executed until a predetermined condition is satisfied. Here, examples of the predetermined condition include the number of repetitions reaching a certain number or more, or the trainable parameters of the presentation information generation model converging, etc.
[0234] The input unit 301A inputs the second learning data stored in the learning data DB 303A (step S801).
[0235] The gap estimation unit 203 uses the input data included in the second learning data input in step S801 above to estimate the gap (gap type and presence or absence of a gap or gap degree score) between the first user and the second user (step S802).
[0236] The presentation information acquiring unit 204A calculates the generation probability of text representing the second presentation information from the gap estimated in step S802 (the gap type and the presence / absence of a gap or the gap degree score) (step S803). That is, the presentation information acquiring unit 204 inputs the gap type and the presence / absence of a gap or the gap degree score estimated in step S802 into a presentation information generation model and calculates the generation probability of text representing the estimated presentation information that is output from the model. Note that the generation probability of text refers to a series of posterior probabilities of each token constituting the text, for example, when the text is composed of tokens. Furthermore, when the number of token types is K and the text is composed of M tokens, the posterior probability of the m (1≦m≦M)th token refers to data represented by a K-dimensional vector whose k-th element represents the probability that a k (1≦k≦K)-th token is generated as the m-th token when a token sequence up to the m-1-th token is obtained, and whose sum is 1.
[0237] The parameter update unit 302A updates the training parameters of the presentation information generation model using the error between the generation probability calculated in step S803 and the generation probability of the training data included in the second training data input in step S801 (step S804). The generation probability of the training data is, for example, a series of correct answer probabilities for each token constituting the text represented by the training data. The correct answer probability of a token is data represented by a K-dimensional vector in which, if the token is the kth type of token, only the kth element is 1 and the other elements are 0. The parameter update unit 302A may use the error as an objective function and update the trainable parameters using a known optimization method so as to minimize the objective function. For example, a cross-entropy error or the like can be used as the error.
[0238] <Modifications of the Learning Device 50> The following describes modifications of the learning device 50. The following modifications can be combined as appropriate as long as they do not contradict each other.
[0239] Modification 3-1 The learning device 50 may simultaneously learn the gap recognition model and the presentation information generation model by using learning data that combines both the first learning data and the second learning data.
[0240] Modification 3-2: When the gap estimation unit 203 and the presentation information acquisition unit 204A are realized by a machine learning model such as a large-scale language model, the learning device 50 may train this machine learning model. In this case, the learning device 50 may train the machine learning model that realizes the gap estimation unit 203 and the presentation information acquisition unit 204A by fine-tuning the parameters of the existing large-scale language model.
[0241] The following supplementary notes are further disclosed regarding the above embodiments. (Supplementary Item 1) An intervention device for intervening in communication between one or more subjects, comprising: a memory; and at least one processor connected to the memory, wherein the processor acquires information about at least one of the one or more subjects, estimates an index value for evaluating a communication gap based on the information about the subject, and presents information to at least one of the one or more subjects to help close the gap based on the index value. (Supplementary Item 2) The intervention device according to Supplementary Item 1, wherein the processor determines whether there is a communication gap based on the index value, and if it is determined that there is a communication gap, presents information to help close the gap to at least one of the one or more subjects. (Supplementary Item 3) The intervention device according to Supplementary Item 1 or 2, wherein the processor estimates recognition information that recognizes the speech content, appearance, state, or action of the subject from information about the subject, and estimates the index value based on the recognition information. (Supplementary Item 4) The intervention device according to Supplementary Item 3, wherein the recognition information includes information recognizing an emotion of the subject, and the processor estimates the index value from a change in the emotion of the subject. (Supplementary Item 5) The intervention device according to Supplementary Item 3, wherein the recognition information includes information recognizing an emotion of the subject, and the processor estimates the index value from a difference in the emotion between two or more subjects performing the communication. (Supplementary Item 6) The intervention device according to Supplementary Item 3, wherein the processor further estimates a basis for the recognition, and further presents the index value and the basis.(Supplementary Item 7) A non-transitory storage medium storing a program that causes a computer to execute a process of intervening in communication between one or more entities, wherein the process of intervening in communication between the one or more entities includes the following processes: acquiring information about at least one of the one or more entities; estimating an index value for evaluating a gap in communication based on the information about the entity; and presenting information to at least one of the one or more entities to assist in resolving the gap based on the index value.
[0242] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0243] [References] Reference 1: Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando, "End-to-End Joint Target and Non-Target Speakers ASR", arXiv:2306.02273 [cs.CL]. Reference 2: Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura, "Leveraging Large Text Corpora for End-to-End Speech Summarization", arXiv:2303.00978 [cs.CL]. Reference 3: Kaori Kumagai, Motohiro Takagi, Shigekuni Kondo, Aono Yuji, Kobayashi Ichiro. "Controllable End-to-End Image Captioning Using Verb-Specific Semantic Role Labels," Transactions of the Information Processing Society of Japan, vol. 63, No. 12, pp. 1884-1894, Dec. 2022. Reference 4: Yamazaki Yoshihiro, Orihashi Shota, Masumura Ryo, Uchida Mihiro, Takashima Eihiko, "End-to-End Multimodal Dialogue Response Generation Based on Video Understanding Mechanism Using Space-Time Attention," 2022 Annual Conference of the Japanese Society for Artificial Intelligence (36th).
[0244] 1 Intervention system 10 Intervention device 20 First user device 30 Second user device 40 Communication network 50 Learning device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 First user information acquisition unit 202 Second user information acquisition unit 203 Gap estimation unit 204, 204A Presentation information acquisition unit 205 First user presentation unit 206 Second user presentation unit 207 First user history DB 208 Second user history DB 209 Intervention plan DB 211 Multimodal recognition unit 212 Gap recognition unit 301, 301A Input unit 302, 302A Parameter update unit 303, 303A Learning data DB
Claims
An intervention device for intervening in communication between one or more subjects, comprising: an acquisition unit that acquires information about at least one subject among the one or more subjects; an estimation unit that estimates an index value for evaluating the communication gap based on information about the subject; a presentation unit that presents information to at least one of the one or more subjects based on the index value to support elimination of the gap; An interventional device having: The presentation unit determining whether there is a gap in the communication based on the index value; The intervention device of claim 1 , wherein if it is determined that there is a gap in communication, information is presented to at least one of the one or more subjects to assist in closing the gap. The estimation unit Estimating recognition information that recognizes the subject's speech content, appearance, state, or action from information about the subject; The intervention device according to claim 1 or 2, wherein the index value is estimated based on the recognition information. The recognition information includes information that recognizes the emotion of the subject, The estimation unit The intervention device according to claim 3 , wherein the index value is estimated from a change in the subject's emotion. The recognition information includes information that recognizes the emotion of the subject, The estimation unit The intervention device according to claim 3 , wherein the index value is estimated from a difference in the emotions between two or more subjects performing the communication. The estimation unit Further infer the basis for said knowledge, The presentation unit The interventional device according to claim 3 , further comprising: a display of the index value and the reason. A computer that intervenes in communication between one or more entities, an acquisition step of acquiring information about at least one of the one or more entities; an estimation step of estimating an index value for evaluating the communication gap based on information about the subject; a presentation step of presenting information to at least one of the one or more subjects based on the index value to assist in eliminating the gap; Intervention methods to implement. A computer that intervenes in communication between one or more subjects, an acquisition step of acquiring information about at least one of the one or more entities; an estimation step of estimating an index value for assessing the communication gap based on information about the subject; a presentation step of presenting information to at least one of the one or more subjects based on the index value to assist in eliminating the gap; A program that executes the following.
Citation Information
Patent Citations
Emotional state estimation program, emotional state estimation method, emotional state estimation device and emotional state estimation system
JP2017029386A
Automated decision making for selecting scaffolds after a partially correct answer in conversational intelligent tutor systems (ITS)
US20200402414A1