Output device, output method, and program

The intervention device addresses the challenge of inappropriate communication in remote interactions by analyzing and intervening with inappropriate content, ensuring smooth communication through appropriate information replacement.

WO2026018445A1PCT designated stage Publication Date: 2026-01-22NT T INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/026031
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Conventional communication technologies struggle to handle multimodal information effectively, leading to inappropriate communication that can evoke negative emotions or discomfort in remote interactions, such as video calls, due to issues like slips of the tongue, rude behavior, or unpleasant facial expressions.

Method used

An intervention device that analyzes multimodal information from users during communication, determining inappropriate content and either blocks or replaces it with appropriate information, using a combination of rule-based systems, knowledge graphs, and machine learning models to generate presentation information.

Benefits of technology

Enables smooth communication by preventing the transmission of inappropriate information and replacing it with appropriate content, thereby enhancing the communication experience between users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024026031_22012026_PF_FP_ABST
    Figure JP2024026031_22012026_PF_FP_ABST
Patent Text Reader

Abstract

An output device according to one aspect of the present invention is for outputting a UI to a display unit that is used by at least one entity among one or more entities engaged in communication. The UI includes information expressing an intervention with respect to inappropriate information transfer in the communication.
Need to check novelty before this filing date? Find Prior Art

Description

Output device, output method, and program

[0001] The present disclosure relates to an output device, an output method, and a program.

[0002] For smooth communication, it is important to avoid inappropriate communication, such as facial expressions or behavior that may cause discomfort to the other party. For this reason, for example, when people communicate remotely using video calls or the like, a mechanism is expected to be developed that blocks inappropriate communication or replaces it with an appropriate one when inappropriate communication occurs. A conventional technique that may realize such a mechanism is, for example, the technique described in Non-Patent Document 1.

[0003] Non-Patent Document 1 describes a technology that understands the content of input text information and expresses emotions using images and voices that show changes in a person's facial expression.

[0004] R. Anderson, B. Stenger, V. Wan, and R. Cipolla, "Expressive visual text-to-speech using active appearance models," in CVPR, 2013.

[0005] However, conventional technologies, including the technology described in Non-Patent Document 1, have not been able to handle multimodal information, and therefore may not be able to fully support smooth communication between people.

[0006] The present disclosure has been made in consideration of the above points, and aims to provide a technology that can support the realization of smooth communication.

[0007] An output device according to one aspect of the present disclosure is an output device that outputs a UI to a display unit used by at least one of one or more subjects engaged in communication, wherein the UI includes information representing an intervention in inappropriate information transmission in the communication.

[0008] It can help to achieve smooth communication.

[0009] FIG. 1 is a diagram (part 1) showing an example of intervention. FIG. 2 is a diagram (part 3) showing an example of intervention. FIG. 1 is a diagram showing an example of an intervention. FIG. 2 is a diagram (part 3) showing an example of intervention. FIG. 3 is a diagram showing an example of the overall configuration of an intervention system including an intervention device according to the present embodiment. FIG. 4 is a diagram showing an example of the hardware configuration of an intervention device according to the present embodiment. FIG. 4 is a diagram showing an example of the functional configuration of an intervention device according to the present embodiment. FIG. 5 is a diagram showing an example of a detailed functional configuration of an intervention estimation unit according to the present embodiment. FIG. 5 is a diagram schematically showing an example of a multimodal recognition model. FIG. 6 is a flowchart (part 1) showing intervention processing in an example. FIG. 7 is a flowchart (part 2) showing intervention processing in an example. FIG. 7 is a diagram showing a first modification of the functional configuration of the intervention device according to the present embodiment. FIG. 8 is a flowchart (part 1) showing intervention processing in the first modification. FIG. 8 is a flowchart (part 2) showing intervention processing in the first modification. FIG. 9 is a diagram showing a second modification of the functional configuration of the intervention device according to the present embodiment. FIG. 9 is a flowchart (part 1) showing intervention processing in the second modification. FIG. 10 is a flowchart (part 2) showing intervention processing in the second modification. FIG. 11 is a diagram showing a first modification of the functional configuration of the intervention device according to the present embodiment. FIG. 4 is a diagram (part 4) showing an example of a UI. FIG. 5 is a diagram (part 5) showing an example of a UI. FIG. 6 is a diagram showing an example of the functional configuration of a learning device that learns an intervention necessity recognition model. FIG. 7 is a diagram schematically showing an example of first learning data. FIG. 8 is a flowchart showing an example of a learning process for learning an intervention necessity recognition model. FIG. 9 is a diagram showing an example of the functional configuration of a learning device that learns a presentation information generation model. FIG. 10 is a diagram schematically showing an example of second learning data. FIG. 11 is a flowchart showing an example of a learning process for learning a presentation information generation model.

[0010] An embodiment of the present invention will be described in detail below with reference to the drawings. The following describes an intervention device 10 that can intervene in inappropriate communication when inappropriate communication occurs, assuming a situation in which two or more people (hereinafter referred to as "users") are communicating. Examples of intervention in inappropriate communication include blocking the communication or replacing it with appropriate communication. By using the intervention device 10 according to this embodiment, when inappropriate communication occurs between users, the communication can be blocked or replaced with appropriate communication, thereby enabling smooth communication.

[0011] <Definitions of Terms, etc.> A user is the subject of communication. While a user is primarily assumed to be a human, it is sufficient that at least one user who is the subject of communication is a human, and other users may also be machines. Furthermore, a machine is not limited to a device, equipment, terminal, etc., but may also be, for example, a program or module that realizes artificial intelligence (AI), or a program or module that realizes a character or avatar that simulates your past or future self, another person's personality, or a real or virtual creature.

[0012] Communication is the transmission of some information between two or more users, where one user transmits some information to another user via video, audio, images (including video), text, etc. Information transmission is the transmission of some information from one user to another user. Information transmission includes not only linguistic information (e.g., speech content, text, etc.), but also non-linguistic information (e.g., facial expressions, gestures, posture, voice volume, speaking speed, etc.).

[0013] Examples of communication via video include video calls and online conferences. Examples of communication via audio include voice calls. Examples of communication via images include chats using images. Examples of communication via text include text chats. However, these are merely examples, and communication may involve one user transmitting some information to another user between two or more users via two or more of video, audio, images, text, etc.

[0014] Inappropriate communication refers to the transmission of information (hereinafter also referred to as inappropriate information) that may evoke negative emotions (e.g., discomfort, anger, sadness, humiliation, anxiety, fear, loss of psychological safety, etc.) in the recipient. However, inappropriate communication may also include, for example, the transmission of information that may result in unnecessary communication in the recipient. Specific examples of inappropriate communication include, for example, gaffes, rude behavior, offensive facial expressions, and rude or offensive words and actions. Other examples of inappropriate communication include the transmission of information that violates public order and morals or compliance. On the other hand, appropriate communication refers to the transmission of information that is unlikely or unlikely to evoke negative emotions in the recipient (hereinafter also referred to as appropriate information). Note that words and actions include not only words but also facial expressions, gestures, and other actions.

[0015] Intervention refers to, for example, blocking the transmission of inappropriate information when at least one of the users communicating transmits inappropriate information or replacing the transmission with appropriate information. Blocking the transmission of inappropriate information is achieved by transmitting information for blocking the inappropriate information instead of the inappropriate information. Similarly, replacing the inappropriate information with appropriate information is achieved by transmitting appropriate information instead of the inappropriate information. Hereinafter, information for blocking inappropriate information and appropriate information transmitted in place of inappropriate information will be collectively referred to as "presented information." Note that presented information can be expressed, for example, as video, images, audio, text, etc.

[0016] <Example of Intervention> As an example, the following describes a case where communication between a customer and an operator in a call center (which may also be called a "contact center") is assumed, and intervention in the communication occurs when inappropriate information is transmitted during the communication. In the following, the customer is referred to as a "first user," and the operator is referred to as a "second user." Furthermore, a device such as a PC (personal computer) used by the first user is referred to as a "first user device 20," and a device such as a PC used by the second user is referred to as a "second user device 30."

[0017] - When there is a "slip of the tongue" As an example of inappropriate communication of information, we will explain the case of a slip of the tongue.

[0018] As shown in FIG. 1 , it is assumed that a first user and a second user are communicating via video (images and audio). While the first user and the second user are communicating, the intervention device 10 transmits video (images and audio) of the second user to the first user device 20 and video (images and audio) of the first user to the second user device 30. As a result, an image (video) of the second user is displayed on the display of the first user device 20 and the audio of the second user is output from the speaker. Similarly, an image (video) of the first user is displayed on the display of the second user device 30 and the audio of the first user is output from the speaker.

[0019] At this time, it is assumed that the second user makes a rude utterance (a slip of the tongue) to the first user, and in this case, it can be said that the second user has communicated inappropriate information to the first user.

[0020] Therefore, in this case, the intervention device 10 according to the present embodiment presents presentation information for blocking the video of the second user when the gaffe was made, instead of the video of the second user, thereby blocking the transmission of inappropriate information to the first user.

[0021] An example of information to be presented to block the video of the second user when a gaffe is made is, for example, "Communication has been blocked because an inappropriate expression was used," when it is acceptable to inform the first user that inappropriate information has been transmitted. On the other hand, for example, when it is not acceptable to inform the first user that inappropriate information has been transmitted, it is acceptable to say, "The communication situation seems to be poor at the moment. Please wait a moment," or simply "Please wait a moment," or to display an error screen.

[0022] The intervention device 10 according to the present embodiment determines whether intervention in information transmission is necessary using various information (hereinafter also referred to as user information) acquired from the first user device 20 and the second user device 30. Typical examples of information acquired from the first user device 20 and the second user device 30 include user video (images and audio). However, the information acquired from the first user device 20 and the second user device 30 is not limited to video. Examples of information acquired from the first user device 20 and the second user device 30 include, in addition to video, text input by the user, sensor information measured by a sensor on the user or their surroundings (e.g., vital signs such as the user's blood pressure and body temperature, and environmental information such as the temperature and humidity of the user's environment), and the like. In other words, user information is multimodal information including audio, images, videos, text, sensor information, and the like. Hereinafter, the user information acquired from the first user device 20 will be referred to as "first user information," and the user information acquired from the second user device 30 will be referred to as "second user information."

[0023] - When there is a "rude attitude" As an example of inappropriate information transmission, we will explain when there is a rude attitude.

[0024] 2 , it is assumed that a first user and a second user are communicating via video (images and audio). During the communication between the first user and the second user, the intervention device 10 transmits the video (images and audio) of the second user to the first user device 20 and transmits the video (images and audio) of the first user to the second user device 30.

[0025] In this case, if the second user responds to the first user in a rude manner, it can be said that the second user has communicated inappropriate information to the first user.

[0026] In this case, the intervention device 10 according to the present embodiment presents an image of the second user with an appropriate attitude as presentation information, instead of an image of the second user in a video of the second user behaving rudely, thereby replacing inappropriate communication of information with appropriate communication of information to the first user.

[0027] - When there is an "unpleasant facial expression" As an example of inappropriate communication, we will explain the case where there is an unpleasant facial expression.

[0028] 3 , it is assumed that a first user and a second user are communicating via video (images and audio). During the communication between the first user and the second user, the intervention device 10 transmits the video (images and audio) of the second user to the first user device 20 and the video (images and audio) of the first user to the second user device 30.

[0029] In this case, if the second user responds to the first user with an unpleasant facial expression, it can be said that the second user has communicated inappropriate information to the first user.

[0030] Therefore, in this case, instead of an image included in a video of the second user responding with an unpleasant facial expression, the intervention device 10 according to the present embodiment processes the image to have an appropriate facial expression and presents the processed image as presentation information. This allows inappropriate communication of information to the first user to be replaced with appropriate communication. As described above, the intervention device 10 according to the present embodiment intervenes in inappropriate communication of information during communication between the first user and the second user, thereby preventing the inappropriate communication from being presented to the other party. Therefore, use of the intervention device 10 according to the present embodiment is expected to realize smooth communication between the first user and the second user.

[0031] 1 to 3, the examples of intervention are described in relation to cases where the second user transmits inappropriate information, but conversely, intervention can also be performed in relation to cases where the first user transmits inappropriate information. Furthermore, intervention can also be performed in relation to cases where both the first user and the second user transmit inappropriate information.

[0032] <Example of Overall Configuration of Intervention System 1 Including Intervention Device 10> Fig. 4 shows an example of the overall configuration of the intervention system 1 including the intervention device 10 according to the present embodiment. As shown in Fig. 4, the intervention device 10 according to the present embodiment includes the intervention device 10, a first user device 20, and a second user device 30. The intervention device 10, the first user device 20, and the second user device 30 are communicatively connected via an arbitrary communication network 40. Note that the communication network 40 includes various communication networks and communication means, such as the Internet, a local area network (LAN), a wide area network (WAN), a telephone network (public switched telephone network, mobile communication network, etc.), etc.

[0033] The intervention device 10 is a computer or a computer system that intervenes in inappropriate information transmission during communication between a first user and a second user. The intervention device 10 is realized, for example, by a PC, a general-purpose server, or a system configured thereof.

[0034] The first user device 20 is a computer or computer system used by a first user or that serves as the first user. The first user device 20 may be implemented, for example, as a PC, smartphone, tablet, wearable device, or other device used by the first user. However, this is merely an example, and the first user device 20 may be implemented as a variety of devices, including a server, robot, industrial machinery, home appliances, vending machines, cash registers, sensor devices, and the like. Note that these are merely examples, and the first user device 20 may be implemented as a variety of devices, such as an in-vehicle device, a game device, medical equipment, nursing care equipment, digital signage, an electronic whiteboard, and the like.

[0035] The second user device 30 is a computer or computer system used by or serving as a second user. The second user device 30 may be implemented, for example, by a PC, smartphone, tablet, wearable device, or other device used by the second user. However, this is merely an example, and the second user device 30 may be implemented by various devices including, for example, a server, robot, industrial machinery, home appliances, vending machines, cash registers, sensor devices, and the like. Note that these are merely examples, and the second user device 30 may be implemented by various devices such as an in-vehicle device, a game device, medical equipment, nursing care equipment, digital signage, an electronic whiteboard, and the like.

[0036] 4 is an example, and the overall configuration of the intervention system 1 is not limited to this. For example, the intervention device 10 and the first user device 20 or the second user device 30 may be integrally configured.

[0037] <Example of Hardware Configuration of Intervention Device 10> An example of the hardware configuration of the intervention device 10 according to this embodiment is shown in Fig. 5. As shown in Fig. 5, the intervention device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is communicatively connected via a bus 109.

[0038] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the intervention device 10 does not necessarily have to include at least one of the input device 101 and the display device 102, for example.

[0039] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0040] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is one of various arithmetic devices such as a central processing unit (CPU) or a graphics processing unit (GPU).

[0041] 5 is an example, and the hardware configuration of the intervention device 10 is not limited to this. The intervention device 10 may have, for example, multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.

[0042] <Example of Functional Configuration of Intervention Device 10> Fig. 6 shows an example of the functional configuration of the intervention device 10 according to this embodiment. As shown in Fig. 6, the intervention device 10 according to this embodiment includes a first user information acquisition unit 201, a second user information acquisition unit 202, an intervention estimation unit 203, a presentation information generation unit 204, a first user presentation unit 205, and a second user presentation unit 206. These units are realized, for example, by a process in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like. The intervention device 10 according to this embodiment also includes a first user history DB 207 and a second user history DB 208. Each of these DBs (databases) is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, at least one of these DBs may be realized, for example, by a storage area of ​​a storage device (e.g., a storage device included in a database server) connected to the intervention device 10 via the communication network 40.

[0043] The first user information acquisition unit 201 acquires first user information from the first user device 20. Furthermore, the first user information acquisition unit 201 stores the first user information acquired from the first user device 20 in the first user history DB 207. Note that the first user information acquisition unit 201 acquires the first user information from the first user device 20 at predetermined time intervals (e.g., every few seconds). The first user information is multimodal information including at least one of the following: the voice of the first user during communication with the second user; images or videos of the first user (particularly, the face of the first user, etc.); text input by the first user; and sensor information obtained by measuring the first user and their surroundings using a sensor.

[0044] The second user information acquisition unit 202 acquires second user information from the second user device 30. The second user information acquisition unit 202 also stores the second user information acquired from the second user device 30 in the second user history DB 208. The second user information acquisition unit 202 acquires the second user information from the second user device 30 at predetermined time intervals (e.g., every few seconds). The second user information is multimodal information that includes at least one of the following: the voice of the second user during communication with the first user; images or videos of the second user (particularly, the face of the second user, etc.); text entered by the second user; and sensor information obtained by measuring the second user and their surroundings using a sensor.

[0045] The intervention estimation unit 203 estimates whether intervention is necessary for information transmission in communication between the first user and the second user, using at least one of the first user information and the second user information. Specifically, the intervention estimation unit 203 estimates the type of information transmission in communication between the first user and the second user (hereinafter also referred to as the "information transmission type") and a value that quantitatively evaluates the need for intervention for that information transmission (hereinafter also referred to as the "intervention necessity score"). A detailed example of the functional configuration of the intervention estimation unit 203 will be described later. Note that the intervention necessity score may be a continuous value, a discrete value, a percentage, or the like. Furthermore, for example, if the intervention necessity score is a discrete value that takes the values ​​of 0 or 1, the intervention necessity score can be regarded as whether intervention is necessary or not, by determining that intervention is not necessary when the intervention necessity score is 0 and that intervention is necessary when the intervention necessity score is 1.

[0046] The presentation information generation unit 204 determines whether to intervene in the information transmission using the type of information transmission in the communication between the first user and the second user and its intervention necessity score. Furthermore, if the presentation information generation unit 204 determines to intervene in the information transmission, it generates at least one of presentation information for the first user and presentation information for the second user using the information transmission type and intervention necessity score of the information transmission. The presentation information generation unit 204 may generate the presentation information using, for example, a rule base, a knowledge graph, a trained machine learning model, etc.

[0047] For example, when generating presentation information based on a rule base, the presentation information generation unit 204 may refer to a table or the like in which presentation information (e.g., past videos, images, audio, etc. of each user) is stored for each type of information transmission (or for each type of information transmission and a range of the intervention necessity score), and acquire from the table or the like presentation information corresponding to the type of information transmission (or the type of information transmission and the intervention necessity score) estimated by the intervention estimation unit 203. Alternatively, for example, the presentation information generation unit 204 may refer to a table in which methods of processing and correcting images and audio for each type of information transmission (or for each type of information transmission and a range of the intervention necessity score), and process and correct images, audio, etc. that represent inappropriate information transmission using the processing and correction method corresponding to the type of information transmission (or the type of information transmission and the intervention necessity score) estimated by the intervention estimation unit 203, to generate presentation information.

[0048] Furthermore, for example, when generating presentation information using a knowledge graph, the presentation information generation unit 204 may refer to a knowledge graph for generating presentation information (e.g., each user's past video, image, audio, etc.) from the information transmission type (or the information transmission type and the intervention necessity score) and generate the presentation information from the information transmission type (or the information transmission type and the intervention necessity score) estimated by the intervention estimation unit 203. Alternatively, for example, the presentation information generation unit 204 may refer to a knowledge graph for identifying a method of processing or correcting images or audio from the information transmission type (or the information transmission type and the intervention necessity score), identify the processing or correction method from the information transmission type (or the information transmission type and the intervention necessity score) estimated by the intervention estimation unit 203, and generate presentation information by processing or correcting images, audio, etc. that represent inappropriate information transmission using the processing or correction method.

[0049] Furthermore, for example, when generating presentation information using a trained machine learning model, the presentation information generation unit 204 may generate presentation information by inputting the type of information transmission (or the type of information transmission and the intervention necessity score) estimated by the intervention estimation unit 203 into the machine learning model. Alternatively, for example, the presentation information generation unit 204 may input an image or audio, or both, representing inappropriate information transmission into the machine learning model in addition to the type of information transmission (or the type of information transmission and the intervention necessity score) estimated by the intervention estimation unit 203, and generate presentation information in which the image or audio, or both, are processed or modified. Hereinafter, a machine learning model that outputs presentation information using at least the type of information transmission (or the type of information transmission and the intervention necessity score) as input will be referred to as a "presentation information generation model." Note that a learning method for the presentation information generation model will be described later.

[0050] The first user presentation unit 205 presents presentation information for the first user to the first user device 20. Hereinafter, the presentation information for the first user will also be referred to as “first presentation information.” Note that the first user presentation unit 205 can present the first presentation information by outputting (transmitting) the first presentation information to the first user device 20.

[0051] In addition, while communication is taking place between the first user and the second user, the first user presentation unit 205 outputs (transmits), for example, video (or only image, only audio, only text, etc.) included in the second user information to the first user device 20.

[0052] The second user presentation unit 206 presents presentation information for the second user to the second user device 30. Hereinafter, the presentation information for the second user will also be referred to as “second presentation information.” Note that the second user presentation unit 206 can present the second presentation information by outputting (transmitting) the second presentation information to the second user device 30.

[0053] In addition, while communication is taking place between the first user and the second user, the second user presentation unit 206 outputs (transmits), for example, video (or only image, only audio, only text, etc.) included in the first user information to the second user device 30.

[0054] The first user history DB 207 stores the history of the first user information (i.e., time-series data of the first user information). Note that the history of the first user information may include various information in addition to the information included in the first user information (e.g., video, image, audio, text, etc.). For example, the history of the first user information may include identification information such as an ID that identifies the first user, recognition information (e.g., emotions, etc.) by the multimodal recognition unit 251 described later, an intervention necessity score recognized by the intervention necessity recognition unit 252 described later, information indicating whether an intervention has actually been performed, video after the intervention if an intervention has actually been performed (or link information to the video, etc.), and actual video corresponding to the video (i.e., actual video included in the first user information).

[0055] The second user history DB 208 stores the history of the second user information (i.e., time-series data of the second user information). Note that the history of the second user information may include various information in addition to the information included in the second user information (e.g., video, image, audio, text, etc.). For example, the history of the second user information may include identification information such as an ID that identifies the second user, recognition information (e.g., emotions, etc.) by the multimodal recognition unit 251 described later, an intervention necessity score recognized by the intervention necessity recognition unit 252 described later, information indicating whether an intervention has actually been performed, video after the intervention if an intervention has actually been performed (or link information to the video, etc.) and actual video corresponding to the video (i.e., actual video included in the second user information), etc.

[0056] The functional configuration of the interventional device 10 shown in FIG. 6 is an example, and the functional configuration of the interventional device 10 is not limited to this.

[0057] <Detailed Functional Configuration Example of Intervention Estimation Unit 203> A detailed functional configuration example of the intervention estimation unit 203 according to this embodiment is shown in Fig. 7. As shown in Fig. 7, the intervention estimation unit 203 according to this embodiment includes a multimodal recognition unit 251 and an intervention necessity recognition unit 252.

[0058] The multimodal recognition unit 251 receives first user information as input and outputs recognition information (hereinafter also referred to as "first recognition information") including linguistic information, non-linguistic information, paralinguistic information, etc. recognized or detected from the audio, images, videos, text, sensor information, etc. included in the first user information. Similarly, the multimodal recognition unit 251 receives second user information as input and outputs recognition information (hereinafter also referred to as "second recognition information") including linguistic information, non-linguistic information, paralinguistic information, etc. recognized or detected from the audio, images, videos, text, sensor information, etc. included in the second user information. Note that linguistic information refers to information that represents the content of a user's utterance, and non-linguistic information refers to non-voluntary information such as gender, age, etc., among non-linguistic information. Furthermore, paralinguistic information refers to voluntary information such as emotions, attitudes, intentions, etc., among non-linguistic information.

[0059] Here, the multimodal recognition unit 251 is realized by a known multimodal recognition model (which may also be called "multimodal AI" or "multimodal LLMs (Large Language Models)"). An example of a multimodal recognition model is shown in FIG. 8. The multimodal recognition model 1000 shown in FIG. 8 is composed of one or more neural network layers (generally multiple neural network layers). It receives multimodal information as input, recognizes various information through multimodal recognition, and outputs recognition information as a result. The multimodal information includes various information such as the user's voice, images, text, sensor information, etc. The recognition information is information obtained by recognizing the content of the user's utterance, the user's state, actions, etc., and may include, for example, speech recognition text, the user's emotions, whether or not they smile, their speaking style, their facial direction, their level of alertness, their empathy level, the presence or absence of negative words, their gender, their age, the presence or absence of a certain illness, the number of fillers, their facial expressions, their tone of voice, their gestures, and whether or not their words and actions are inconsistent. The mismatch between words and actions refers to a mismatch between the content of the utterance (e.g., the content of the speech-recognition text) and the expression of emotions (e.g., facial expressions, tone of voice, gestures, etc.). The speech-recognition text may be called, for example, "speech text" or simply "text."

[0060] Other examples that may be included in the recognition information include, for example, whether or not there is crying, whether or not there is screaming, whether or not there is coughing, speaking style, speaker, speaker diarization, speaking rate, speech recognition text after adding punctuation, speech recognition text converted from spoken language into written language, understandability, funniness, whether or not there is harassment, detection results of objects including people, detection results of faces with or without masks, text area in the image, key points on the face, gaze direction, clothing, belongings, behavior, gestures, poker face level, amount of gaze movement, average amount of gaze movement, participation level, number of nods, number of blinks, text representing explanatory text for the image, personality traits, speaking behavior, communication skill level, likability level, sales impression level, willingness to engage in conversation, smoothness of conversation, activeness of conversation, excitement of conversation, concentration level, understanding level, satisfaction level, motivation, interest level, etc. The recognition information recognized by the multimodal recognition unit 251 may be information recognized from one type of information (e.g., audio only, image only) or information recognized from multiple types of information (e.g., both audio and image).

[0061] Note that the above recognition information is an example, and the multimodal recognition unit 251 is capable of using first user information or second user information as input to output various recognition information that can be recognized, detected, or estimated using a known multimodal recognition model.

[0062] To obtain the above recognition information, the multimodal recognition unit 211 realizes, for example, the following functions. However, it goes without saying that these functions are merely examples. It also goes without saying that these functions may be combined as appropriate. Furthermore, although terms such as "recognition," "detection," and "estimation" are used below, these terms are not strictly distinct and may be interchangeable terms.

[0063] - Speech period detection: Using audio data as input, the speaker's speech period is detected and output.

[0064] ・Speech recognition: Takes voice data as input and outputs voice-recognized text.

[0065] - Multi-speaker speech recognition: Speech data is input and speech recognition text for each speaker is output.

[0066] Voice gender recognition: Voice data is used as input to recognize and output the speaker's gender.

[0067] - Voice age recognition: Voice data is used as input to recognize and output the speaker's age.

[0068] - Voice emotion recognition: Voice data is used as input to recognize and output the speaker's emotions.

[0069] Crying / screaming detection: Audio data is input and a pair of crying / screaming labels and their confidence scores is output.

[0070] Laughter / cough detection: Audio data is input and pairs of laugh / cough labels and their confidence scores are output.

[0071] Speaking style estimation: Speech data is input and a pair of speaking style labels and their confidence scores is output.

[0072] Speaker vector extraction: Inputs speech data and outputs speaker vectors.

[0073] Speaker diarization recognition: Speech recognition texts of multiple speakers and speaker vectors of those multiple speakers are input, and pairs of utterance start times, utterance end times, and speaker IDs are output. Note that speaker diarization recognition may be performed after speech recognition and speaker vector extraction, or may use the outputs of speech recognition and speaker vector extraction as input.

[0074] Adding punctuation: Text data is input and text data with punctuation added is output. Note that adding punctuation may be performed after speech recognition, or the output of speech recognition may be used as input.

[0075] Spoken-to-written conversion: Text data is input, and text data obtained by converting spoken data into written data is output. Note that the spoken-to-written conversion may be performed after speech recognition, or the speech recognition output may be used as input.

[0076] Positive / negative estimation: Text data is input, and pairs of positive / negative labels and their confidence scores are output. Note that positive / negative estimation may be performed in a later stage of speech recognition processing, or the speech recognition output may be used as input.

[0077] - Comprehensibility estimation: Text data is input, and a pair of comprehensibility labels and their reliability is output. Note that comprehensibility estimation may be performed after speech recognition, or may use the speech recognition output as input.

[0078] Interestingness estimation: Text data is input, and a set of interestingness labels and their reliability is output. Interestingness estimation may be performed after speech recognition, or the speech recognition output may be used as input.

[0079] Harassment detection: Text data is input, and a pair of harassment labels and their reliability is output. Harassment detection may be performed after speech recognition, or the speech recognition output may be used as input.

[0080] Filler count estimation: Text data is input, and the number of fillers is estimated and output. Note that filler count estimation may be performed in a later stage of speech recognition processing, or the output of speech recognition may be used as input.

[0081] Face detection: Image data is input and the coordinates of the face areas for the number of people in the image are output.

[0082] - Object detection: Image data is input and the coordinates of the object area for the object in the image are output.

[0083] - Person detection: Image data is input and the coordinates of the person areas for the number of people in the image are output.

[0084] Face detection with or without mask: Image data is input, and pairs of coordinates of face areas for the number of people in the image and labels indicating whether or not a face is masked are output.

[0085] Character detection: Image data is input and the coordinates of a rectangular area surrounding the character area in the image are output.

[0086] Facial emotion recognition: Image data of the facial region is input, and a pair of an emotion label and its reliability is output. Note that facial emotion recognition may be performed after face detection or face detection with or without mask, and may use image data of the facial region represented by the coordinates output by face detection or face detection with or without mask as input.

[0087] Face / gender estimation: Image data of the face area is input, and a pair of a gender label and its reliability is output. Note that face / gender estimation may be performed after face detection or face detection with or without a mask, and image data of the face area represented by the coordinates output by face detection or face detection with or without a mask may be input.

[0088] Face age estimation: Image data of the face area is input, and the age is estimated and output. Note that face age estimation may be performed after face detection or face detection with or without mask, and image data of the face area represented by coordinates output by face detection or face detection with or without mask may be input.

[0089] Facial emotional arousal estimation: Image data representing a face image is input, and the arousal level is estimated and output. Note that the facial emotional arousal estimation may be performed after face detection or face detection with or without a mask, and may use image data of the face area represented by the coordinates output by face detection or face detection with or without a mask as input.

[0090] Face direction estimation: Using image data of the face region as input, estimate and output the face direction. Note that face direction estimation may be performed in a subsequent process after face detection or face detection with or without a mask, and may use image data of the face region represented by coordinates output by face detection or face detection with or without a mask as input.

[0091] Facial keypoint estimation: Image data representing a facial image is input, and facial keypoints are estimated and output. Note that facial keypoint estimation may be performed after face detection or face detection with or without mask, and may use image data of the facial area represented by the coordinates output by face detection or face detection with or without mask as input.

[0092] Gaze estimation: Using image data of the face region as input, the vertical and horizontal gaze angles of the right and left eyes are estimated and output. Note that gaze estimation may be performed in a subsequent process after face detection or face detection with or without a mask, and image data of the face region represented by coordinates output by face detection or face detection with or without a mask may be used as input.

[0093] Clothing recognition: Image data of a person area is input, and a set of clothing labels and their reliability is output. Note that clothing recognition may be performed after person detection, and the image data of the person area represented by the coordinates output by person detection may be used as input.

[0094] Personal item recognition: Image data of a person area is input, and a pair of personal item labels and their reliability is output. Note that personal item recognition may be performed after person detection, and the image data of the person area represented by the coordinates output by person detection may be used as input.

[0095] Action recognition: Image data of a human region is input, and a pair of an action label and its reliability is output. Note that action recognition may be performed after human detection, and the image data of the human region represented by the coordinates output by human detection may be input.

[0096] Gesture recognition: A time series of image data is input, and a pair of gesture labels and their reliability is output.

[0097] Frame average of face and gender estimation: A time series of face and gender estimation results for the same person is input, and a pair of the average gender label and its reliability is output.

[0098] - Frame average of facial age estimation: The time series of facial age estimation results for the same person is input and the average age is output.

[0099] Poker face degree estimation: The time series of facial emotion recognition results for the same person is used as input, and the poker face degree is output. The poker face degree is defined as the proportion of facial expressions that are not specified in advance.

[0100] - Gaze movement amount estimation: The time series of gaze estimation results is input and the gaze movement amount is output.

[0101] Average gaze movement estimation: The gaze estimation results for two consecutive frames of the same person are used as the average gaze movement amount, and the average gaze movement amount is output.

[0102] - Participation estimation: The time series of face direction estimation results is input, and the percentage of people facing the direction of the camera that took the photo is output.

[0103] Nodding frequency estimation: The time series of face direction estimation results is input, and the number of noddings is output.

[0104] - Blink frequency estimation: The time series of facial keypoint estimation results is used as input, and the blink frequency is output.

[0105] Image description generation: Image data is input and text describing the image is output.

[0106] Character recognition: Image data and the coordinates of a rectangular area surrounding a character area in the image are input, and the text of the character area is output. Note that character recognition may be performed after character detection, and the coordinates output by character detection may be used as input.

[0107] Personality trait estimation: Voice data and a time series of image data of the face region are input, and a pair of personality trait labels and their reliability is output. Note that personality trait estimation may be performed in a later stage of face detection or face detection with or without a mask, and may use as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask.

[0108] Speech behavior recognition: It takes voice data and a time series of image data of the face region as input, and outputs a pair of a speech behavior label and its reliability. Note that speech behavior recognition may be positioned after face detection or face detection with or without mask, and may take as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without mask.

[0109] Communication skill estimation: Voice data and a time series of image data of the face region are input, and a set of communication skill labels and their reliability is output. Note that communication skill estimation may be performed in a later process after face detection or face detection with or without a mask, and may use as input the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask.

[0110] Likeability estimation: Voice data and a time series of image data of the face region are input, and a pair of likeability labels and their reliability is output. Note that likeability estimation may be performed in a later process after face detection or face detection with or without a mask, and the time series of image data of the face region represented by the coordinates output by face detection or face detection with or without a mask may be input.

[0111] Conversational sales impression estimation: Voice data and image data are input, and a sales impression label is output.

[0112] Conversation Positive / Negative Degree Estimation: The facial emotion recognition results of all conversation participants are used as input to output the conversation positive / negative degree. The conversation positive / negative degree is defined as the difference between the positive and negative proportions of all participants.

[0113] Conversation empathy estimation: The facial emotion recognition results of all conversation participants are input, and the conversation empathy is output. The conversation empathy is defined as the similarity of the facial emotion recognition results of all conversation participants.

[0114] Conversational engagement estimation: The conversational engagement estimation results for all conversation participants are input, and the conversational engagement is output. The conversational engagement is defined as the average conversational empathy of all conversation participants.

[0115] - Estimation of conversational fluency: The result of voice activity detection is used as input and conversational fluency is output. The conversational fluency is defined as the proportion of non-silence periods.

[0116] Conversation enthusiasm estimation: The results of the conversation positive / negative estimation and the conversation fluency estimation are input, and the conversation enthusiasm is output. The conversation enthusiasm is defined as the sum of the conversation positive / negative and conversation fluency.

[0117] The intervention necessity recognition unit 252 receives at least one of the first recognition information and the second recognition information as input, and recognizes the type of information transmission from one of the first user and the second user to the other, and the intervention necessity score for that information transmission. The intervention necessity recognition unit 252 is realized, for example, by a trained machine learning model that receives recognition information as input and outputs the type of information transmission and the intervention necessity score. However, the intervention necessity recognition unit 252 may also be realized, for example, by a program that implements a rule base or a knowledge graph, or a program that performs simple calculations. Hereinafter, the machine learning model that realizes the intervention necessity recognition unit 252 will be referred to as the "intervention necessity recognition model." The learning method for the intervention necessity recognition model will be described later.

[0118] The multimodal recognition unit 251 and the intervention necessity recognition unit 252 may recognize the intervention necessity score by processing the multimodal information (i.e., the first user information, the second user information, or both) in an end-to-end manner. In this case, the multimodal recognition unit 251 and the intervention necessity recognition unit 252 may process information of each type (e.g., image, audio, etc.) in parallel, or may process each type of information sequentially or collectively. For details about end-to-end processing, see, for example, References 1-4.

[0119] Furthermore, a model that processes multimodal information in an end-to-end manner (i.e., a model that realizes the multimodal recognition unit 251 and the intervention necessity recognition unit 252 in an end-to-end manner) can be configured in various ways. For example, it can be configured with a first input layer that inputs linguistic information included in the multimodal information, a second input layer that inputs paralinguistic information included in the multimodal information, an intermediate layer, and an output layer. In this case, the output layer may be, for example, a layer that outputs an intervention necessity score as a binary value, or a layer that outputs an intervention necessity score as a continuous value. Alternatively, the output layer may be, for example, a layer that outputs at least one of linguistic information and paralinguistic information.

[0120] <Example of Intervention Processing> An example of the intervention processing according to this embodiment will be described below.

[0121] Example of Intervention Process (Part 1) Hereinafter, an intervention process will be described with reference to Fig. 9, in which a determination is made as to whether or not intervention is necessary in the information transmission of a second user using second user information, and if intervention in the information transmission is necessary, first presentation information is presented to a first user. The intervention process shown in Fig. 9 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0122] The second user information acquisition unit 202 acquires second user information from the second user device 30 (step S101). The second user information is stored in the second user history DB 208.

[0123] The multimodal recognition unit 251 of the intervention estimation unit 203 recognizes second recognition information using the second user information stored in the second user history DB 208 (step S102). Hereinafter, as an example, it is assumed that the second recognition information is recognized as a second text representing a speech recognition text of the voice included in the second user information acquired in step S101, an expression of the second user's emotions (e.g., facial expressions, tone of voice, gestures, etc.), and whether or not there is a discrepancy between words and actions.

[0124] The intervention necessity recognition unit 252 of the intervention estimation unit 203 receives the second recognition information as input and recognizes whether intervention is necessary for the second user's communication (step S103). That is, the intervention necessity recognition unit 252 receives the second recognition information as input and recognizes the type of communication by the second user and an intervention necessity score for that communication. Examples of communication types include "speech content," "emotional expression (facial expression)," "emotional expression (tone of voice)," "emotional expression (gestures)," and "inconsistency between words and actions." The intervention necessity score is calculated such that the more inappropriate the speech content or emotional expression, or the more inconsistent the words and actions, the higher the degree of intervention necessity, and the lower the degree of intervention necessity. Note that when multiple communication types are recognized, the intervention necessity recognition unit 252 may calculate a single intervention necessity score for the multiple communication types, or may recognize an intervention necessity score for each communication type. For simplicity, it is assumed below that the intervention necessity recognition unit 252 recognizes one intervention necessity score for one information transmission type, and that one or more information transmission types and their corresponding intervention necessity scores are obtained.

[0125] The presentation information generator 204 determines whether intervention is necessary in the information transmission of the second user (step S104). For example, the presentation information generator 204 determines that intervention is necessary if at least one intervention need score exceeds a predetermined threshold, and that intervention is not necessary if not. Here, the threshold may be determined commonly for all information transmission types, or may be determined for each information transmission type. If a threshold is determined for each information transmission type, the presentation information generator 204 determines, for each intervention need score, whether the intervention need score exceeds the threshold, using a threshold determined in advance for the information transmission type corresponding to the intervention need score.

[0126] For example, if the intervention necessity score corresponding to the information transmission type "utterance content" exceeds the threshold, it corresponds to, for example, a case where the second user made a slip of the tongue in communicating information. Furthermore, if the intervention necessity score corresponding to the information transmission type "inconsistency between words and actions" exceeds the threshold, it corresponds to, for example, a case where the second user communicated information in a rude manner. Furthermore, if the intervention necessity score corresponding to the information transmission type "expression of emotion (facial expression)" exceeds the threshold, it corresponds to, for example, a case where the second user communicated information in an unpleasant manner.

[0127] If it is determined in step S104 that intervention is necessary, the presentation information generator 204 generates first presentation information using the information transmission type and the intervention necessity score (step S105). That is, the presentation information generator 204 generates the first presentation information using, for example, the intervention necessity score determined in step S104 to exceed the threshold and the information transmission type corresponding to the intervention necessity score, based on a rule base, a knowledge graph, or a presentation information generation model.

[0128] If it is determined in step S104 that no intervention is necessary, the presentation information generation unit 204 ends the intervention process.

[0129] The first user presentation unit 205 presents the first presentation information generated in step S105 above instead of the video of the second user (or the image or audio included in the video) (step S106).

[0130] In addition, by switching the "first user" and "second user" in the intervention process shown in Figure 9, it is possible to similarly realize an intervention process in which the first user information is used to determine whether intervention is necessary in the information transmission of the first user, and if intervention in the information transmission is necessary, the second presentation information is presented to the second user.

[0131] Example of Intervention Process (Part 2) Hereinafter, an intervention process will be described with reference to Fig. 10 , in which a first user and a second user are used to determine whether or not an intervention is necessary in an information transmission by a second user, and if an intervention in the information transmission is necessary, the first presentation information is presented to the first user. The intervention process shown in Fig. 10 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0132] The second user information acquisition unit 202 acquires second user information from the second user device 30 (step S201). The second user information is stored in the second user history DB 208.

[0133] The multimodal recognition unit 251 of the intervention estimation unit 203 recognizes second recognition information using the second user information stored in the second user history DB 208 (step S202). Hereinafter, as an example, it is assumed that the second recognition information is recognized as a second text representing a speech recognition text of the voice included in the second user information acquired in step S201, an expression of the second user's emotions (e.g., facial expressions, tone of voice, gestures, etc.), and whether or not there is a discrepancy between words and actions.

[0134] The first user information acquisition unit 201 acquires the first user information from the first user device 20 (step S203). The first user information is stored in the first user history DB 207.

[0135] The multimodal recognition unit 251 of the intervention estimation unit 203 recognizes the first recognition information (step S204) using the first user information stored in the first user history DB 207. Hereinafter, as an example, it is assumed that the first text representing the speech recognition text of the speech included in the first user information acquired in step S203, the expression of the first user's emotions (e.g., facial expressions, tone of voice, gestures, etc.), and the presence or absence of inconsistency between words and actions are recognized as the second recognition information.

[0136] The above steps S201 to S202 and steps S203 to S204 may be performed in any order. That is, steps S201 to S202 may be performed after steps S203 to S204. Furthermore, steps S201 to S202 and steps S203 to S204 may be performed in parallel.

[0137] The intervention necessity recognition unit 252 of the intervention estimation unit 203 receives the first recognition information and the second recognition information as input and recognizes whether intervention is necessary for the second user's communication of information (step S205). That is, the intervention necessity recognition unit 252 receives the first recognition information and the second recognition information as input and recognizes the type of communication of the second user and an intervention necessity score for that communication of information. Examples of the type of communication of information include "speech content," "emotional expression (facial expression)," "emotional expression (tone of voice)," "emotional expression (gestures)," and "inconsistency between words and actions." The intervention necessity score is calculated so that the more inappropriate the second user's speech content or emotional expression, or the more inconsistent the second user's speech and actions, the higher the degree of intervention necessity, and the lower the degree of intervention necessity. Furthermore, the intervention necessity score is calculated such that the more the first user expresses negative emotions (e.g., discomfort, anger, sadness, etc.) or the more inconsistency there is between words and actions, the higher the degree of intervention necessity, and the less so, the lower the degree of intervention necessity. Note that when multiple types of information communication are recognized, the intervention necessity recognition unit 252 may calculate one intervention necessity score for the multiple types of information communication, or may recognize an intervention necessity score for each of the types of information communication. For simplicity, it is assumed below that the intervention necessity recognition unit 252 recognizes one intervention necessity score for one type of information communication, and that one or more types of information communication and their corresponding intervention necessity scores have been obtained.

[0138] The presentation information generator 204 determines whether intervention is necessary in the information transmission of the second user (step S206). For example, the presentation information generator 204 determines that intervention is necessary if at least one intervention need score exceeds a predetermined threshold, and that intervention is not necessary if not. Here, the threshold may be determined commonly for all information transmission types, or may be determined for each information transmission type. If a threshold is determined for each information transmission type, the presentation information generator 204 determines, for each intervention need score, whether the intervention need score exceeds the threshold, using a threshold determined in advance for the information transmission type corresponding to the intervention need score.

[0139] For example, if the intervention necessity score corresponding to the information transmission type "utterance content" exceeds the threshold, it corresponds to, for example, a case where the second user made a slip of the tongue in communicating information. Furthermore, if the intervention necessity score corresponding to the information transmission type "inconsistency between words and actions" exceeds the threshold, it corresponds to, for example, a case where the second user communicated information in a rude manner. Furthermore, if the intervention necessity score corresponding to the information transmission type "expression of emotion (facial expression)" exceeds the threshold, it corresponds to, for example, a case where the second user communicated information in an unpleasant manner.

[0140] If it is determined in step S206 that intervention is necessary, the presentation information generator 204 generates first presentation information using the information transmission type and the intervention necessity score (step S207). That is, the presentation information generator 204 generates the first presentation information using, for example, the intervention necessity score determined in step S206 to exceed the threshold and the corresponding information transmission type, based on a rule base, a knowledge graph, or a presentation information generation model.

[0141] If it is determined in step S206 above that intervention is not necessary, the presentation information generation unit 204 ends the intervention process.

[0142] The first user presentation unit 205 presents the first presentation information generated in step S207 above instead of the video of the second user (or the image or audio included in the video) (step S208).

[0143] In addition, by switching the "first user" and "second user" in the intervention process shown in Figure 10, it is possible to similarly realize an intervention process in which the first user information and the second user information are used to determine whether intervention is necessary in the information transmission of the first user, and if intervention in the information transmission is necessary, the second presentation information is presented to the second user.

[0144] <Variation 1 of Functional Configuration of Intervention Device 10> Fig. 11 shows Variation 1 of the functional configuration of the intervention device 10 according to the present embodiment. As shown in Fig. 11 , the intervention device 10 in Variation 1 further includes a first advice information generator 209 and a second advice information generator 210. These units are realized, for example, by processing in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like.

[0145] When it is determined that intervention is necessary in the communication of the first user (i.e., when the first user has communicated inappropriate information), the first advice information generation unit 209 generates advice information (hereinafter also referred to as "first advice information") that represents advice to the first user. Examples of the first advice information include information representing the occurrence of inappropriate communication of information (e.g., a slip of the tongue, an inappropriate attitude, an unpleasant facial expression, etc.), and information representing advice to suppress the communication of inappropriate information.

[0146] When it is determined that intervention is necessary in the communication of the second user (i.e., when the second user has communicated inappropriate information), the second advice information generation unit 210 generates advice information (hereinafter also referred to as "second advice information") that represents advice to the second user. Examples of the second advice information include information representing the occurrence of inappropriate communication of information (e.g., a slip of the tongue, an inappropriate attitude, an unpleasant facial expression, etc.), and information representing advice to suppress the communication of inappropriate information.

[0147] Here, the first advice information generation unit 209 and the second advice information generation unit 210 may generate advice information using, for example, a rule base, a knowledge graph, a trained machine learning model, or the like.

[0148] For example, when generating the first advice information based on a rule base, the first advice information generation unit 209 may refer to a table or the like in which the first advice information is stored for each information transmission type (or for each information transmission type and the range of the intervention necessity score), and acquire from the table or the like the first advice information corresponding to the information transmission type (or the information transmission type and the intervention necessity score) estimated by the intervention estimation unit 203. The same applies to the case in which the second advice information generation unit 210 generates the second advice information.

[0149] Furthermore, for example, when generating the first advice information using a knowledge graph, the first advice information generation unit 209 may refer to the knowledge graph for generating the first advice information from the information transmission type (or the information transmission type and the intervention necessity score), and generate the first advice information from the information transmission type (or the information transmission type and the intervention necessity score) estimated by the intervention estimation unit 203. The same applies to the case where the second advice information generation unit 210 generates the second advice information.

[0150] Furthermore, for example, when generating the first advice information using a trained machine learning model, the first advice information generation unit 209 may generate the first advice information by inputting the information transmission type (or the information transmission type and the intervention necessity score) estimated by the intervention estimation unit 203 into the machine learning model. Hereinafter, a machine learning model that takes the information transmission type (or the information transmission type and the intervention necessity score) as input and outputs advice information will be referred to as an "advice information generation model." Note that the advice information generation model can be trained in the same manner as the presentation information generation model that takes the information transmission type (or the information transmission type and the intervention necessity score) as input and outputs presentation information (however, the correct answer for the advice information is used as training data instead of the correct answer for the presentation information), and therefore the learning method will be omitted.

[0151] <Modification 1 of Intervention Processing> Modification 1 of the intervention processing according to this embodiment will be described below.

[0152] Variation 1 of Intervention Processing (Part 1) Hereinafter, an intervention processing will be described with reference to Fig. 12 , in which second user information is used to determine whether intervention in information transmission by a second user is necessary, and if intervention in the information transmission is necessary, first presentation information is presented to the first user and second advice information is presented to the second user. The intervention processing shown in Fig. 12 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0153] Steps S301 to S306 may be similar to steps S101 to S106 in FIG. 9, respectively, and therefore a description thereof will be omitted.

[0154] If it is determined in step S304 that intervention is necessary, the second advice information generation unit 210 generates second advice information using the information transmission type and the intervention necessity score (step S307). That is, the second advice information generation unit 210 generates the second advice information by a rule base, a knowledge graph, or an advice information generation model, for example, using the intervention necessity score determined in step S304 to exceed the threshold and the information transmission type corresponding to the intervention necessity score.

[0155] The second user presenting unit 206 presents the second advice information generated in step S307 to the second user (step S308), which allows the second user to know that he or she has transmitted inappropriate information and receive advice on how to suppress such information transmission.

[0156] In addition, by swapping the "first user" and the "second user" in the intervention process shown in Figure 14 and generating and presenting "first advice information" instead of "second advice information" in steps S307 to S308, it is possible to similarly realize an intervention process in which the first user information is used to determine whether intervention is necessary in the information transmission of the first user, and if intervention in the information transmission is necessary, the second presentation information is presented to the second user and the first advice information is presented to the first user.

[0157] Variation 1 (Part 2) of Intervention Processing Hereinafter, an intervention processing will be described with reference to Fig. 13 , in which first user information and second user information are used to determine whether intervention in information communication by a second user is necessary, and if intervention in the information communication is necessary, first presentation information is presented to the first user and second advice information is presented to the second user. The intervention processing shown in Fig. 13 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0158] Steps S401 to S408 may be similar to steps S201 to S208 in FIG. 10, respectively, and therefore a description thereof will be omitted.

[0159] If it is determined in step S406 that intervention is necessary, the second advice information generation unit 210 generates second advice information using the information transmission type and the intervention necessity score (step S409). That is, the second advice information generation unit 210 generates the second advice information by a rule base, a knowledge graph, or an advice information generation model, for example, using the intervention necessity score determined in step S406 to exceed the threshold value and the information transmission type corresponding to the intervention necessity score.

[0160] The second user presenting unit 206 presents the second advice information generated in step S409 to the second user (step S410). This allows the second user to know that he or she has transmitted inappropriate information and receive advice on how to suppress such information transmission.

[0161] In addition, by swapping the "first user" and the "second user" in the intervention process shown in Figure 13 and generating and presenting "first advice information" instead of "second advice information" in steps S409 to S410, it is possible to similarly realize an intervention process in which the first user information and the second user information are used to determine whether intervention is necessary in the information transmission of the first user, and if intervention in the information transmission is necessary, the second presentation information is presented to the second user and the first advice information is presented to the first user.

[0162] <Modification 2 of Functional Configuration of Intervention Device 10> Fig. 14 shows Modification 2 of the functional configuration of the intervention device 10 according to the present embodiment. As shown in Fig. 14, the intervention device 10 in Modification 2 further includes a permission information acquisition unit 211. Furthermore, the intervention device 10 in Modification 2 includes a presentation information generation unit 204A instead of the presentation information generation unit 204. These units are realized, for example, by processing in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like.

[0163] The permission information acquisition unit 211 acquires permission information of at least one of the first user and the second user. The permission information is information indicating whether the user permits intervention in inappropriate information transmission when that user transmits the information. The first user can use the first user device 20 to set, in advance or in real time, whether to permit intervention in inappropriate information transmission when that user transmits the information. Similarly, the second user can use the second user device 30 to set, in advance or in real time, whether to permit intervention in inappropriate information transmission when that user transmits the information. Note that the permission information may take various forms, and may, for example, be a binary value, where 0 indicates no permission and 1 indicates permission.

[0164] The presentation information generation unit 204A further determines whether intervention in inappropriate information transmission is permitted, using the permission information acquired by the permission information acquisition unit 211.

[0165] <Modification 2 of Intervention Processing> Modification 2 of the intervention processing according to this embodiment will be described below.

[0166] Variation 2 (Part 1) of Intervention Processing Below, we will explain, with reference to Fig. 15 , an intervention processing in which, after determining whether or not intervention in information communication by a second user is necessary using second user information, intervention in the information communication is necessary and if such intervention is permitted, first presentation information is presented to a first user. The intervention processing shown in Fig. 15 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0167] Steps S501 to S503 may be similar to steps S101 to S103 in FIG. 9, respectively, and therefore a description thereof will be omitted.

[0168] The permission information acquisition unit 211 acquires the permission information of the second user (step S504). That is, the permission information acquisition unit 211 acquires permission information indicating whether or not the second user permits intervention in inappropriate information transmission when the second user transmits inappropriate information. The permission information acquisition unit 211 may acquire the permission information of the second user from the second user device 30, or, if the permission information of the second user is stored in a storage area such as a memory, may acquire the permission information of the second user from that storage area.

[0169] The presentation information generation unit 204A determines whether the permission information acquired in step S504 above represents permission and whether intervention in the information communication of the second user is necessary (step S505). Note that the presentation information generation unit 204A may determine whether intervention in the information communication of the second user is necessary, for example, by a method similar to that of step S104 in FIG. 9.

[0170] If it is determined in step S505 above that the permission information represents permission and that intervention is necessary, the presentation information generation unit 204A generates first presentation information (step S506), similar to step S105 in Figure 9.

[0171] If it is determined in step S505 above that the permission information indicates non-permission or that intervention is not required, the presentation information generation unit 204A ends the intervention process.

[0172] Similar to step S106 in FIG. 9, the first user presentation unit 205 presents the first presentation information generated in step S506 above instead of the video of the second user (or the image or audio contained in the video) (step S507).

[0173] In addition, by switching the "first user" and "second user" in the intervention process shown in Figure 15, it is possible to similarly realize an intervention process in which the first user information is used to determine whether intervention is necessary in the information transmission of the first user, and then intervention in the information transmission is required, and if such intervention is permitted, the second presentation information is presented to the second user.

[0174] Variation 2 (Part 2) of Intervention Processing Hereinafter, an intervention processing will be described with reference to FIG. 16 , in which first user information and second user information are used to determine whether intervention in information communication by a second user is necessary, and if intervention in the information communication is necessary and permitted, first presentation information is presented to the first user. The intervention processing shown in FIG. 16 is repeatedly executed at predetermined time intervals (e.g., every few seconds). In the following, it is assumed, as an example, that the first user and the second user are communicating via video.

[0175] Steps S601 to S605 may be similar to steps S201 to S205 in FIG. 10, respectively, and therefore a description thereof will be omitted.

[0176] The permission information acquisition unit 211 acquires the permission information of the second user (step S606). That is, the permission information acquisition unit 211 acquires permission information indicating whether or not the second user permits intervention in inappropriate information transmission when the second user transmits inappropriate information. The permission information acquisition unit 211 may acquire the permission information of the second user from the second user device 30, or, if the permission information of the second user is stored in a storage area such as a memory, may acquire the permission information of the second user from that storage area.

[0177] The presentation information generation unit 204A determines whether the permission information acquired in step S606 above represents permission and whether intervention in the information communication of the second user is necessary (step S607). Note that the presentation information generation unit 204A may determine whether intervention in the information communication of the second user is necessary, for example, by a method similar to that of step S206 in FIG. 10.

[0178] If it is determined in step S607 above that the permission information represents permission and that intervention is necessary, the presentation information generation unit 204A generates first presentation information (step S608), similar to step S207 in Figure 10.

[0179] If it is determined in step S607 above that the permission information indicates non-permission or that intervention is not required, the presentation information generation unit 204A ends the intervention process.

[0180] Similar to step S208 in FIG. 10, the first user presentation unit 205 presents the first presentation information generated in step S608 above instead of the video of the second user (or the image or audio contained in the video) (step S609).

[0181] In addition, by switching the "first user" and the "second user" in the intervention process shown in FIG. 16, it is possible to similarly realize an intervention process in which it is necessary to determine whether or not intervention is necessary in the information transmission of the first user using the first user information and the second user information, and then to intervene in the information transmission, and if such intervention is permitted, to present the second presentation information to the second user.

[0182] <Modification 3 of Functional Configuration of Intervention Device 10> Fig. 17 shows Modification 3 of the functional configuration of the intervention device 10 according to the present embodiment. As shown in Fig. 17 , the intervention device 10 in Modification 3 further includes a third user presentation unit 212. The third user presentation unit 212 is realized, for example, by a process in which one or more programs installed in the intervention device 10 are executed by the processor 108 or the like. The intervention device 10 in Modification 3 also includes a log history DB 213. The log history DB 213 is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, the log history DB 213 may also be realized, for example, by a storage area of ​​a storage device (e.g., a storage device included in a database server) connected to the intervention device 10 via the communication network 40.

[0183] The third user presentation unit 212 presents the log stored in the log history DB 213 to the third user device. The log history DB 213 stores, as a log, at least one of the video and first presentation information output to the first user device 20 and the video and second presentation information output to the second user device 30. The third user device refers to various terminals, such as a PC, smartphone, tablet terminal, or wearable device, used by a third user other than the first user and the third user. Various types of third users are conceivable. For example, if the second user is a call center operator, the third user may be a supervisor who supervises the operator.

[0184] The log history DB 213 stores, as a log, at least one of the video and first presentation information output to the first user device 20 and the video and second presentation information output to the second user device 30. In particular, when targeting a call center, it is preferable that the log history DB 213 stores, as a log, the video and first presentation information output to the first user device 20. Among the logs stored in the log history DB 213, logs representing presentation information may be associated with, for example, a predetermined flag value. This allows, for example, a third user to easily check the presentation information when checking the log. In addition to the method of associating flag values, flag values ​​may be embedded in the logs using, for example, digital watermarking technology.

[0185] By using the intervention device 10 in the third modification example, a third user, such as a call center supervisor, can view a log of video and presented information representing past interactions between an operator and a customer. Therefore, the third user can know whether, for example, a second user, such as an operator, received inappropriate information and whether the first presented information was presented to the first user as a result of the inappropriate information transmission. When the first presented information is presented to the first user, the third user may simultaneously view not only the first presented information but also a video of the second user when the first presented information is presented to the first user.

[0186] <Other Modifications of the Interventional Device 10> The following describes other modifications of the above-described interventional device 10. The following modifications can be combined as appropriate as long as they are not inconsistent with each other.

[0187] Variation 1-1: In the above embodiment, the video is a video of the user, and the presented information blocks inappropriate information transmission in the video or replaces it with appropriate information transmission. However, the video does not necessarily have to be a video of the user. For example, the video may be a video in which the user is replaced by an avatar, such as a predetermined character. In this case, the presented information also blocks inappropriate information transmission by the avatar or replaces it with appropriate information transmission.

[0188] Variation 1-2: When a user's inappropriate communication is blocked or replaced with an appropriate communication, the user may be notified of this. This allows the user to know that their communication has been blocked or replaced with an appropriate communication.

[0189] Modification 1-3: When communication between users is initiated, the value of the permission information of each user may be determined depending on the degree of accuracy required for that communication, etc. Specifically, when accurate communication is required, the permission information of each user may be information indicating that permission is not granted, whereas when accurate communication is not necessarily required, the permission information of each user may be information indicating that permission is granted.

[0190] Variation 1-4 When generating the first presentation information in the above embodiment for a call center, for example, a video, image, audio, or the like of a similar response by a second user in the past may be generated as the first presentation information.

[0191] Modification 1-5: When generating the first presentation information in Modification 1-4 above, for example, a video, image, or audio of another operator performing a similar service in the past may be generated as the first presentation information. In addition, at this time, for example, information in which the face or voice of the other operator is changed to the face or voice of the second user may be generated as the first presentation information.

[0192] Variation 1-6: When replacing inappropriate information transmission with appropriate information transmission, it is difficult to present the presentation information in real time if it takes time to generate the presentation information. Therefore, for example, video of each user communicating with another user may be buffered for a certain period of time, and then the video may be output to each user with a certain time delay from real time. This makes it possible to achieve real-time performance with a certain time delay from real time, as long as the time required to generate the presentation information is within a certain period of time.

[0193] Modifications 1-7 The permission information for each user in Modification 2 above may be information representing other users who, when the user transmits inappropriate information, authorize the presentation of presentation information to block or replace the inappropriate information. This makes it possible, when a user transmits inappropriate information, to output the presentation information only to other users authorized by the permission information. In other words, it is possible to output the video of the user who transmitted inappropriate information as is, without presenting the presentation information to other users not authorized by the permission information. Therefore, for example, by authorizing only users such as customers and not users such as supervisors or one's own superiors in the permission information, it is possible to output the presentation information to customers and output the video of the user who transmitted inappropriate information as is to users such as supervisors or one's own superiors.

[0194] Modification 1-8: In accordance with a user's selection, for example, in addition to the information presented to the user, a video of the user who has transmitted inappropriate information may be output as is. In other words, in accordance with a user's selection, both the information presented to the user and a video of the user who has transmitted inappropriate information may be output to the user.

[0195] Modification 1-9 In the above-described modification 1, for example, advisory information (particularly, advisory information for mitigating the impact of inappropriate information transmission) may also be presented to a user to whom the presented information is presented. To give a specific example, when the first presented information is presented to a first user due to a slip of the tongue by a second user, the first advisory information may be presented to the first user. This makes it possible to present the first advisory information (e.g., advisory information such as "Please calm down") to the first user in order to calm the first user's anger, for example, when the first user becomes angry due to a slip of the tongue by the second user.

[0196] In this case, for example, if a second user makes a slip of the tongue due to a first user's past words or actions (e.g., provocation, etc.), the first user may be presented with first advice information (e.g., advice information such as ``Please do not provoke'') to inhibit the second user from making similar or related words or actions to the first user's past words or actions.

[0197] In the above-described modification 3, the third user who is a supervisor checks the log of the communication between the first user and the second user. However, the third user may also be able to check the communication between the first user and the second user in real time. In this case, the third user may also be able to set the permission information of the second user in real time.

[0198] Variation 1-11 All or part of the functional units of the intervention device 10 shown in Figures 6, 11, 14, and 17 may be possessed, for example, by a server (e.g., a cloud server) that is communicatively connected to the intervention device 10.

[0199] Modification 1-12 In the above embodiment, user information is mainly assumed to include the user's voice during communication, images and videos of the user, text entered by the user, and sensor information measured by sensors on the user and their surroundings, but user information may also include information related to the user. Examples of user-related information include background images of images and videos of the user, information displayed in conjunction with the user's own image or avatar (e.g., dialogue in speech bubbles, avatar skin and clothing, etc.), background music, etc. Other user-related information includes information that is not related to the user's own words and actions but that the user can set or bring in for communication.

[0200] At this time, examples of the presented information include information for correcting unpleasant images (images of clothing, background images, etc.) and information for blocking or changing unpleasant background music.

[0201] Modification 1-13 In the above embodiment, the multimodal recognition unit 251 is implemented by the multimodal recognition model 1000. However, for example, the multimodal recognition unit 251 may be implemented by a model (hereinafter also referred to as a subject imitation model) that imitates a real or virtual living thing (including a person, character, animal, etc.). Here, the subject imitation model is a model that receives relevant user information as input and outputs recognition information of the subject. The subject imitation model is, for example, a model (e.g., a machine learning model) created using DTC (digital twin computing) technology. A specific example of a subject imitation model is, for example, a machine learning model trained to understand the words and actions of the subject and output the subject's reaction to those words and actions. Note that the multimodal recognition unit 251 and the intervention necessity recognition unit 252 may be configured by a subject imitation model, and the subject imitation model may output an intervention necessity score (including whether or not to intervene).

[0202] However, DTC is just one example, and the subject imitation model may be created by a technology other than DTC as long as it is a model that takes the relevant user information as input and outputs the subject's recognition information. Specifically, the subject imitation model may be a model created by analyzing the subject's communication habits and reactions at that time, for example.

[0203] Although it is preferable that a subject imitation model be created for each subject, for example, subjects that show the same or similar reactions to the same words and actions may be categorized, and a subject imitation model may be created for each category. Furthermore, data (learning data) for creating a subject imitation model may be collected by sensing the subject in question, or may be collected using social media such as SNS.

[0204] Modification 1-14 In the above embodiment, inappropriate communication may include cases where the information communicated by the communication is inappropriate in relation to the other party. For example, if the other party belongs to a certain company and the logo of a company that is a competitor of that company is displayed, the other party may feel uncomfortable. For this reason, the display of such a logo may be included in inappropriate communication.

[0205] Other specific examples of inappropriate communication in a relationship with another person include communication of political information, communication of information about a team other than the one the other person supports in a sport, etc. Communication that the other person has previously felt to be inappropriate also counts as inappropriate communication in a relationship with another person.

[0206] <Example of UI> Below, examples of UIs (user interfaces) displayed on displays, etc. of the first user device 20 and the second user device 30 will be described. All of the UIs described below are displayed on displays, etc. of the first user device 20 and the second user device 30 by the first user presentation unit 205 and the second user presentation unit 206, respectively. However, for example, the intervention device 10 may have functional units such as a "UI control unit" or an "output unit," and these UI control units or output units may display the UIs on displays, etc. of the first user device 20 and the second user device 30. Note that all or part of the UIs displayed on displays, etc. of the first user device 20 and the second user device 30 may be different. In particular, the UIs displayed on displays, etc. of the user devices used by the users may differ depending on the type of user (e.g., whether the user is a customer or an operator, etc.).

[0207] 18 is a UI displayed on the display of the first user device 20, and includes a first user display field 1101 in which an image or video taken of the first user (first user U1) is displayed, and a second user display field 1102 in which an image or video taken of the second user U2 is displayed. Note that the first user display field 1101 may display an avatar or the like of the first user U1, and similarly, the second user display field 1102 may display an avatar or the like of the second user U2.

[0208] In this case, for example, if the second user U2 behaves rudely, first presentation information is output to the first user device 20, and the first presentation information is displayed in the second user display field 1102. In the second user display field 1102 of the UI 1100 shown in Fig. 18, first presentation information 1111 is displayed, which displays the text "The communication status seems to be poor at the moment. Please wait a moment." This temporarily interrupts communication between the first user and the second user, and therefore, for example, the second user can prevent a situation in which their relationship with the first user becomes even worse.

[0209] 19 is a UI displayed on the display of the first user device 20, and includes a first user display field 1201 in which an image or video taken of the first user (first user U1) is displayed, and a second user display field 1202 in which an image or video taken of the second user U2 is displayed. Note that the first user display field 1201 may display an avatar or the like of the first user U1, and similarly, the second user display field 1202 may display an avatar or the like of the second user U2.

[0210] In this case, for example, if the second user U2 behaves rudely, first presentation information is output to the first user device 20 and the first presentation information is displayed in the second user display field 1202. An image 1211 indicating that the second user U2 is apologizing is displayed in the second user display field 1202 of the UI 1200 shown in Fig. 19. This is expected to prevent a situation in which the relationship between the first user and the second user deteriorates.

[0211] UI Example (Part 3) It is assumed that four users, a first user U1, a second user U2, a third user U3, and a fourth user U4, are communicating via video.

[0212] 20 is a UI displayed on the display of the first user device 20, and includes a first user display field 1301 in which an image or video of the first user (U1) is displayed, a second user display field 1302 in which an image or video of the second user U2 is displayed, a third user display field 1303 in which an image or video of the third user U3 is displayed, and a fourth user display field 1304 in which an image or video of the fourth user U4 is displayed. Note that the first user display field 1301 may display an avatar or the like of the first user U1, and similarly, the second user display field 1302 may display an avatar or the like of the second user U2. The same applies to the third user display field 1303 and the fourth user display field 1304.

[0213] Furthermore, the second user display field 1302, the third user display field 1303, and the fourth user display field 1304 each include setting buttons 1312, 1313, and 1314 for setting the consent information of the first user in the above-described variant 1-7. The first user can set the consent to presenting the presentation information to the second user by turning the setting button 1312 "ON," and can set the consent to not presenting the presentation information to the second user by turning the setting button 1312 "OFF." Similarly, the first user can set the consent to presenting the presentation information to the third user by turning the setting button 1313 "ON," and can set the consent to not presenting the presentation information to the third user by turning the setting button 1313 "OFF." Similarly, the first user can set the setting to "ON" to allow the fourth user to receive the presentation information by setting the setting button 1314, and can set the setting to "OFF" to not allow the fourth user to receive the presentation information. In this way, the setting buttons 1312, 1313, and 1314 allow the first user to set in real time whether or not to allow other users to receive the presentation information.

[0214] 20, the setting buttons 1312, 1313, and 1314 are included in the second user display field 1302, the third user display field 1303, and the fourth user display field 1304, respectively. However, this is merely an example, and the setting buttons 1312, 1313, and 1314 may be displayed in a list, for example. The setting buttons themselves are also an example, and, for example, a check box or the like may be used to set whether or not other users have permission to present information. Furthermore, a case in which permission is set in real time is also an example, and permission may be set in advance. In this case, information indicating permission (e.g., "ON / OFF" or an icon indicating "ON / OFF") that has been set in advance may be displayed on the UI 1300.

[0215] 21 is a UI displayed on the display of the first user device 20, and includes a first user display field 1401 in which images or videos taken of the first user (first user U1) are displayed, and second user display fields 1402 and 1403 in which images or videos taken of the second user U2 are displayed. The second user display field 1403 also includes a setting button 1421 for setting whether or not to display images or videos taken of the second user U2. When the setting button 1421 is "ON," images or videos taken of the second user U2 are displayed in the second user display field 1403, and when the setting button 1421 is "OFF," nothing is displayed in the second user display field 1403. In addition, the first user display field 1401 may display an avatar or the like of the first user U1, and similarly, the second user display fields 1402 and 1403 may display an avatar or the like of the second user U2.

[0216] In this case, for example, if the second user U2 expresses discomfort due to the rude attitude of the first user U1, first presentation information is output to the first user device 20, and the first presentation information is displayed in the second user display field 1402. An image 1411 representing the smiling second user U2 is displayed in the second user display field 1402 of the UI 1400 shown in FIG. 21 . Meanwhile, an image or video captured of the second user U2 is displayed in the second user display field 1403. This allows the first user U1 to know, in addition to the presentation information about himself, for example, that the second user U2 is actually expressing discomfort.

[0217] For example, when the subject imitation model described in Modification Example 1-13 is used to estimate information indicating discomfort in the subject imitation model of the second user U2 even though the second user U2 does not show discomfort, first presentation information indicating that the second user U2 is actually uncomfortable may be output to the first user device 20. The first presentation information indicating that the second user U2 is actually uncomfortable may be, for example, information such as a character string indicating that fact, or information for changing the facial expression of the second user U2 to one that indicates discomfort.

[0218] 22 is a UI displayed on the display of the second user device 30, and includes a first user display field 1501 in which an image or video captured of the first user U1 is displayed, and a second user display field 1502 in which an image or video captured of the user (second user U2) is displayed. The UI 1500 shown in FIG. 22 also includes a first presented information display field 1503 in which the first presented information is displayed when the first presented information is presented on the first user device 20.

[0219] For example, if the first presented information is presented to the first user device 20 due to inappropriate information transmission by the second user U2, the first presented information is displayed in the first presented information display field 1503. An image 1511 indicating that the second user U2 is apologizing is displayed in the first presented information display field 1503 of the UI 1500 shown in FIG. 22 . At this time, information indicating that the first presented information is being presented to the first user is displayed in the first user display field 1501. Text 1521 indicating "Replacement Occurring" is displayed in the first user display field 1501 of the UI 1500 shown in FIG. 22 . This allows the second user U2 to know the content of the first presented information being presented to the first user and also to know that the first presented information is being presented to the first user.

[0220] 22 may be displayed on a display of a third user device used by a third user such as a supervisor, etc. This allows the third user to similarly know the content of the first presented information being presented to the first user and to know that the first presented information is being presented to the first user.

[0221] <Modifications Regarding UI> Modifications regarding the above UI will be described below. Note that the following modifications can be combined as appropriate as long as they do not contradict each other.

[0222] Modification 2-1 The UI may be generated by the interventional device 10, or may be generated by a user device (such as the first user device 20 or the second user device 30), and the display of the UI may be updated based on information (such as presentation information) received from the interventional device 10. When the interventional device 10 generates the UI, the interventional device 10 may have a functional unit such as a "UI generation unit."

[0223] Modification 2-2 A UI may be displayed as a pop-up on another UI that is already displayed on the display of a user device (such as the first user device 20 or the second user device 30).

[0224] <Application Examples> The above embodiment has been described mainly with reference to a call center, but the above embodiment and its modifications can be applied to various fields. Some application examples are given below.

[0225] Application Example 1: This application can be applied to the field of education, such as schools and cram schools, with the first user being a student and the second user being a teacher. This is expected to facilitate smooth communication between students and teachers, for example.

[0226] The number of students is not limited to one, and the present invention can be similarly applied to cases where there are multiple students for one teacher, for example.

[0227] Application Example 2 This application can be applied to fields such as employee education and training, with the first user acting as an employee and the second user acting as an instructor. As a result, similar to Application Example 1, smooth communication between employees and instructors can be expected. Note that this application can also be applied to group work-style training where multiple employees are grouped together.

[0228] Application Example 3: This can be applied to corporate meetings, with each user being an employee. This can be expected to facilitate smooth communication between employees, leading to more productive meetings.

[0229] Application Example 4: With the first user as a customer and the second user as a salesperson, this can be applied to fields such as store counter sales and online sales. This is expected to facilitate smooth communication between the customer and the salesperson, leading to more effective sales and sales with higher customer satisfaction.

[0230] Other Application Examples In addition to the above application examples, the above embodiments and their modifications can be similarly applied to any field in which communication occurs between people (or between people and machines). For example, the above embodiments and their modifications can be similarly applied to communication such as interviews with human resources personnel, advisors, consultants, etc., and sales with customers. Other examples include communication such as interviews with medical professionals such as doctors and nurses, and interviews with professionals such as legal professionals.

[0231] <Example of Functional Configuration of Learning Device 50 That Learns Intervention Need Recognition Model> FIG. 23 shows an example of the functional configuration of a learning device 50 that learns an intervention need recognition model that implements the intervention need recognition unit 252. As shown in FIG. 23 , the learning device 50 has an input unit 301, an intervention need recognition unit 252, and a parameter update unit 302. These units are implemented, for example, by a processor or the like executing one or more programs installed in the learning device 50. The learning device 50 also has a learning data DB 303. The learning data DB 303 is implemented, for example, by a storage area of ​​an auxiliary storage device or the like. However, the learning data DB 303 may also be implemented by a storage area of ​​a storage device (e.g., a storage device included in a database server) connected to the learning device 50 via a communication network. The learning device 50 may be implemented by the same device as the intervention device 10, or by a device different from the intervention device 10.

[0232] The input unit 301 inputs learning data (hereinafter also referred to as "first learning data") stored in the learning data DB 303. The first learning data is data for training the intervention necessity recognition model, and includes input data to be input to the intervention necessity recognition model and teacher data for the input data. The teacher data is the correct answer to the necessity of intervention (i.e., the type of information transmission and the intervention degree score) output from the intervention necessity recognition model when input data is input. Hereinafter, the necessity of intervention output from the intervention necessity recognition model when input data is input will be referred to as the "estimated necessity of intervention value."

[0233] The parameter update unit 302 updates the learnable parameters of the intervention necessity recognition model using the error between the intervention necessity estimated value recognized by the intervention necessity recognition unit 252 and the teacher data included in the first learning data.

[0234] The learning data DB 303 stores first learning data. The first learning data is created in advance by, for example, a creator of the intervention necessity recognition model.

[0235] <<Example of First Learning Data>> FIG. 24 shows an example of the first learning data. As shown in FIG. 24, the first learning data includes input data and teacher data. The input data is data representing recognition information input to the intervention necessity recognition model. The input data included in the first learning data shown in FIG. 25 is composed of first recognition information and second recognition information. However, for example, only one of the first recognition information and the second recognition information may be used as input data. The teacher data is data representing the correct answer to the intervention necessity estimation value. The teacher data included in the first learning data shown in FIG. 25 is composed of an information transmission type and an intervention level score. However, for example, only the intervention level score may be used as the teacher data. The teacher data is created, for example, by a user who received inappropriate information transmission or a third party who references it.

[0236] <Example of Learning Process for Learning Intervention Necessity Recognition Model> An example of the learning process for learning the intervention necessity recognition model will be described below with reference to Fig. 25. The learning process shown in Fig. 25 is repeatedly executed until a predetermined condition is satisfied. Here, examples of the predetermined condition include that the number of repetitions is equal to or greater than a certain number, that the learnable parameters of the intervention necessity recognition model have converged, etc.

[0237] The input unit 301 inputs the first learning data stored in the learning data DB 303 (step S701).

[0238] The intervention necessity recognizing unit 252 recognizes the necessity of intervention using the input data included in the first learning data input in the above step S701 (step S702). That is, the intervention necessity recognizing unit 252 inputs the input data to an intervention necessity recognition model and obtains an intervention necessity estimated value as an output thereof.

[0239] The parameter update unit 302 updates the trainable parameters of the intervention necessity recognition model using the error between the intervention necessity estimation value obtained in step S702 and the teacher data included in the first learning data input in step S701 (step S703). That is, the parameter update unit 302 updates the trainable parameters of the intervention necessity recognition model using a first error representing the error between the information transmission type included in the intervention necessity estimation value and the information transmission type included in the teacher data, and a second error representing the error between the intervention level score included in the intervention necessity estimation value and the intervention level score included in the teacher data. The parameter update unit 302 may update the trainable parameters using a known optimization method, for example, to minimize an objective function, which is the sum of the first error and the second error. The first error may be, for example, a cross-entropy error. The second error may be, for example, a mean square error.

[0240] <Example of Functional Configuration of Learning Device 50 for Learning Presentation Information Model> FIG. 26 shows an example of the functional configuration of a learning device 50 for learning a presentation information generation model that implements the presentation information generation unit 204. As shown in FIG. 26 , the learning device 50 includes an input unit 301A, an intervention estimation unit 203, a presentation information generation unit 204, and a parameter update unit 302A. Each of these units is implemented, for example, by a processor or the like executing one or more programs installed in the learning device 50. The learning device 50 also includes a learning data DB 303A. The learning data DB 303A is implemented, for example, by a storage area of ​​an auxiliary storage device or the like. However, the learning data DB 303A may also be implemented by a storage area of ​​a storage device (e.g., a storage device included in a database server) connected to the learning device 50 via a communication network. The learning device 50 may be implemented by the same device as the intervention device 10, or by a device different from the intervention device 10.

[0241] The input unit 301A inputs learning data (hereinafter also referred to as "second learning data") stored in the learning data DB 303A. The second learning data is data for training the presentation information generation model, and includes input data input to the intervention estimation unit 203 and training data for the input data. The training data is the correct answer of the presentation information output from the presentation information generation model when the input data is input to the intervention estimation unit 203. Hereinafter, the presentation information output from the presentation information generation model will be referred to as "estimated presentation information."

[0242] The parameter update unit 302A updates the learning target parameters of the presentation information generation model using the error between the estimated presentation information generated by the presentation information generation unit 204 and the training data included in the second training data.

[0243] The training data DB 303A stores second training data. Note that the second training data is created in advance by, for example, the creator of the presentation information generation model.

[0244] <<Example of Second Learning Data>> FIG. 27 shows an example of the second learning data. As shown in FIG. 27, the second learning data includes input data and training data. The input data is data representing multimodal information input to the intervention estimation unit 203. The second learning data shown in FIG. 27 is composed of first user information and second user information. However, for example, only one of the first user information and the second user information may be used as input data. The input data may include an image or audio, or both, representing inappropriate communication input to the presentation information generation unit 204. The training data is data representing the correct answer to the estimated presentation information. The training data is created, for example, by a user who received inappropriate communication or a third party who references it.

[0245] <Learning Process for Learning Presentation Information Generation Model> An example of the learning process for learning the presentation information generation model will be described below with reference to Fig. 28. The learning process shown in Fig. 28 is repeatedly executed until a predetermined condition is satisfied. Here, examples of the predetermined condition include the number of repetitions reaching a certain number or more, or the trainable parameters of the presentation information generation model converging, etc.

[0246] The input unit 301A inputs the second learning data stored in the learning data DB 303A (step S801).

[0247] The intervention estimation unit 203 estimates the necessity of intervention (type of information transmission and intervention level score) using input data included in the second learning data input in step S801 (step S802).

[0248] The presentation information generation unit 204 generates presentation information using the intervention necessity estimated in step S802 (step S803). That is, the presentation information generation unit 204 inputs the information transmission type estimated in step S802 (or the information transmission type and the intervention necessity score) into a presentation information generation model and generates presentation information as its output. Note that, for example, if an image or audio, or both, representing inappropriate information transmission is further included in the input data of the second learning data input in step S801, the image or audio, or both, are also input into the presentation information generation model.

[0249] The parameter update unit 302A updates the learning parameters of the presentation information generation model using the error between the presentation information generated in step S803 and the teacher data included in the second learning data input in step S801 (step S804). Note that the parameter update unit 302A may update the learnable parameters using a known optimization method so as to minimize the error as an objective function. Depending on the type of information to be generated as presentation information (e.g., video, image, audio, text, etc.), a known error used in the task of generating that information may be used as the error.

[0250] <Modifications of the Learning Device 50> The following describes modifications of the learning device 50. The following modifications can be combined as appropriate as long as they do not contradict each other.

[0251] Modification 3-1 The learning device 50 may simultaneously learn the intervention necessity recognition model and the presentation information generation model by using learning data that combines both the first learning data and the second learning data.

[0252] Modification 3-2: A presentation information generation model may be learned by fine-tuning a machine learning model called a large-scale language model or the like that realizes generative AI or generative AI or the like.

[0253] The following supplementary notes are further disclosed with respect to the above embodiments. (Supplementary Item 1) An output device including: a memory; and at least one processor connected to the memory, wherein the processor outputs a UI to a display unit used by at least one of one or more subjects communicating, the UI including information representing an intervention in response to inappropriate communication in the communication. (Supplementary Item 2) The output device according to Supplementary Item 1, wherein the UI includes information representing the intervention instead of the inappropriate communication. (Supplementary Item 3) The output device according to Supplementary Item 2, wherein the UI includes at least one of a predetermined image and sound as information representing the intervention instead of at least one of an image and sound representing the inappropriate communication. (Supplementary Item 4) The output device according to Supplementary Item 3, wherein the predetermined image and sound are, respectively, an image and sound of a time when the subject who transmitted the inappropriate communication transmitted appropriate information in the past. (Supplementary Item 5) The output device according to Supplementary Item 3 or 4, wherein the UI includes at least one of an image and a sound generated by processing or modifying at least one of an image and a sound representing the inappropriate communication. (Supplementary Item 6) The output device according to Supplementary Item 5, wherein the UI includes a component for setting whether to generate at least one of an image and a sound resulting from processing or modifying at least one of an image and a sound representing the inappropriate communication. (Supplementary Item 7) The output device according to Supplementary Item 1, wherein the UI includes advice information including at least one of information representing that the communication is inappropriate to the subject who has transmitted the inappropriate communication and information for suppressing the inappropriate communication. (Supplementary Item 8) The output device according to Supplementary Item 1, wherein the UI is also output to a display unit used by a person other than the one or more subjects performing the communication. (Supplementary Item 9) The output device according to Supplementary Item 8, wherein the UI includes information representing the intervention in accordance with permission information representing whether to present the information representing the intervention. (Supplementary Item 10) The output device according to Supplementary Item 9, wherein the permission information is set by a person other than the one or more subjects performing the communication.(Appendix 11) A non-transitory storage medium storing a program that causes a computer to execute an output process, wherein the output process outputs a UI to a display unit used by at least one of one or more subjects communicating, and the UI includes information representing an intervention in inappropriate information transmission in the communication.

[0254] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0255] [References] Reference 1: Ryo Masumura, Naoki Makishima, Taiga Yamane, Yoshihiko Yamazaki, Saki Mizuno, Mana Ihori, Mihiro Uchida, Keita Suzuki, Hiroshi Sato, Tomohiro Tanaka, Akihiko Takashima, Satoshi Suzuki, Takafumi Moriya, Nobukatsu Hojo, Atsushi Ando, ​​"End-to-End Joint Target and Non-Target Speakers ASR", arXiv:2306.02273 [cs.CL]. Reference 2: Kohei Matsuura, Takanori Ashihara, Takafumi Moriya, Tomohiro Tanaka, Atsunori Ogawa, Marc Delcroix, Ryo Masumura, "Leveraging Large Text Corpora for End-to-End Speech Summarization", arXiv:2303.00978 [cs.CL]. Reference 3: Kaori Kumagai, Motohiro Takagi, Shigekuni Kondo, Aono Yuji, Ichiro Kobayashi. "Controllable End-to-End Image Captioning Using Verb-Specific Semantic Role Labels," Transactions of the Information Processing Society of Japan, vol. 63, No. 12, pp. 1884-1894, Dec. 2022. Reference 4: Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida, Eihiko Takashima, "End-to-End Multimodal Dialogue Response Generation Based on Video Understanding Mechanism Using Space-Time Attention," 2022 Annual Conference of the Japanese Society for Artificial Intelligence (36th).

[0256] 1 Intervention system 10 Intervention device 20 First user device 30 Second user device 40 Communication network 50 Learning device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 First user information acquisition unit 202 Second user information acquisition unit 203 Intervention estimation unit 204, 204A Presentation information generation unit 205 First user presentation unit 206 Second user presentation unit 207 First user history DB 208 Second user history DB 209 First advice information generation unit 210 Second advice information generation unit 211 Permission information acquisition unit 251 Multimodal recognition unit 252 Intervention necessity recognition unit 301 Input unit 302, 302A Parameter update unit 303, 303A Learning data DB

Claims

1. An output device that outputs a UI to a display unit used by at least one of one or more subjects engaged in communication, wherein the UI includes information representing an intervention in inappropriate information transmission in the communication.

2. The output device of claim 1, wherein the UI includes information representing the intervention in place of the inappropriate communication of information.

3. The output device according to claim 2, wherein the UI includes at least one of a predetermined image and sound as information representing the intervention, instead of at least one of an image and sound representing the inappropriate information transmission.

4. The output device according to claim 3, wherein the predetermined image and sound are images and sounds of the subject who transmitted the inappropriate information transmitting in the past transmitting appropriate information.

5. An output device according to claim 3 or 4, wherein the UI includes at least one of an image and a sound generated by processing or modifying at least one of an image and a sound representing the inappropriate communication of information.

6. The output device according to claim 5, wherein the UI includes a component for setting whether or not to generate at least one of an image and an audio that has been processed or modified from at least one of an image and an audio that expresses the inappropriate communication of information.

7. The output device of claim 1, wherein the UI includes advisory information including at least one of information indicating that the inappropriate information transmission is inappropriate to the entity that transmitted the inappropriate information and information for suppressing the inappropriate information transmission.

8. The output device according to claim 1, wherein the UI is also output to a display unit used by a person other than the one or more subjects performing the communication.

9. The output device according to claim 8, wherein the UI includes information representing the intervention in accordance with permission information representing whether or not to present the information representing the intervention.

10. The output device according to claim 9, wherein the permission information is set by a party other than the one or more subjects engaging in the communication.

11. An output method in which a computer outputs a UI to a display unit used by at least one of one or more subjects engaged in communication, wherein the UI includes information representing an intervention in inappropriate information transmission in the communication.

12. A program that causes a computer to output a UI to a display unit used by at least one of one or more subjects engaged in communication, wherein the UI includes information indicating an intervention in inappropriate information transmission in the communication.

Citation Information

Patent Citations

  • Communication device

    CN111131742A

  • Conference system, information processing apparatus, and information processing method

    JP2012054897A

  • Output apparatus, output method, and output program

    JP2021149664A