Conversation inconsistency detection device, conversation inconsistency detection method, and program

The conversation inconsistency detection system uses a generation model to assess conversation consistency and notify users of deviations, addressing the limitations of existing technologies by enhancing ease of use and accuracy in detecting actual conversation deviations.

JP2026054241APending Publication Date: 2026-03-26DAI NIPPON PRINTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing conversation inconsistency detection technologies only detect the presence of pronouns and ambiguous expressions, failing to accurately assess conversation consistency and requiring manual algorithm creation, thus lacking ease of use in detecting actual deviations.

Method used

A conversation inconsistency detection system utilizing a generation model to sequentially input conversation content, detect consistency, and notify inconsistencies through a notification control unit.

Benefits of technology

Facilitates easy detection of conversation consistency, providing real-time alerts for inconsistencies, applicable to various communication platforms and settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026054241000001_ABST
    Figure 2026054241000001_ABST
Patent Text Reader

Abstract

This invention provides a conversation inconsistency detection device, a conversation inconsistency detection method, and a program that can easily detect the consistency of conversations. [Solution] One aspect of the conversation inconsistency detection device of the present disclosure comprises: a consistency detection control unit that sequentially inputs the content of a conversation between multiple users into a generation model and obtains a detection result from the generation model that detects the consistency of the conversation with the content input up to that point; and a notification control unit that performs notification control when the detection result indicates an inconsistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a conversation inconsistency detection device, a conversation inconsistency detection method, and a program.

Background Art

[0002] Conventionally, various techniques for supporting conversations have been known.

[0003] For example, Patent Document 1 discloses a technique for detecting and accumulating a deviation in a conversation conducted among a plurality of users, determining that the conversation has broken down when the accumulated value of the conversation deviation exceeds a threshold value, and displaying the transition of the accumulated value of the conversation deviation. According to the technique disclosed in Patent Document 1, each user can grasp at which stage the conversation has become mismatched by checking the transition of the accumulated value of the conversation deviation, and it is said that the conversation can be efficiently restarted from the point when the conversation has become mismatched.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, the detection of conversation deviation in the technique disclosed in Patent Document 1 described above only detects whether pronouns, ambiguous expressions, etc. are included in the utterances within the conversation. That is, the technique disclosed in Patent Document 1 is only a technique for determining that there is a high possibility that the conversation has broken down when pronouns, ambiguous expressions, etc., are used a certain number of times or more within the conversation, and does not detect whether there is actually a deviation in the conversation or whether the conversation has actually broken down. Therefore, the technique disclosed in Patent Document 1 cannot detect the consistency of the conversation.

[0006] Furthermore, with the technology disclosed in Patent Document 1 mentioned above, it is necessary to create the algorithm for performing the processing described above oneself, and it is not possible to easily detect the consistency of the conversation.

[0007] This disclosure is made in view of the above circumstances and aims to provide a conversation inconsistency detection device, a conversation inconsistency detection method, and a program that can easily detect the consistency of conversations. [Means for solving the problem]

[0008] One embodiment of the conversation inconsistency detection device of this disclosure comprises: a consistency detection control unit that sequentially inputs the content of a conversation between multiple users into a generation model and obtains a detection result from the generation model that detects the consistency of the conversation with the content input up to that point; and a notification control unit that performs notification control when the detection result indicates an inconsistency.

[0009] One aspect of the conversation inconsistency detection method of the present disclosure includes a consistency detection control step in which a consistency detection control unit sequentially inputs the content of a conversation between multiple users into a generation model and obtains a detection result from the generation model that detects the consistency of the conversation with the content input up to that point; and a notification control step in which a notification control unit performs notification control of the occurrence of an inconsistency if the detection result indicates an inconsistency.

[0010] One aspect of the program of this disclosure causes a computer to sequentially input the content of a conversation between multiple users into a generation model and obtain a consistency detection control step from the generation model that detects the consistency of the conversation with the content input up to that point; and a notification control step that, if the detection result indicates inconsistency, performs notification control to indicate that an inconsistency has occurred. [Effects of the Invention]

[0011] This disclosure provides a conversation inconsistency detection device, a conversation inconsistency detection method, and a program that can easily detect the consistency of conversations. [Brief explanation of the drawing]

[0012] [Figure 1] Figure 1 is a block diagram showing an example of the system configuration of the conversation inconsistency detection system of this embodiment. [Figure 2] Figure 2 is a block diagram showing an example of the hardware configuration of the conversation inconsistency detection device and user terminal according to this embodiment. [Figure 3] Figure 3 is a block diagram showing an example of the functional configuration of the conversation inconsistency detection system of this embodiment. [Figure 4] Figure 4 shows an example of the first prompt in this embodiment. [Figure 5] Figure 5 shows an example of the first prompt in this embodiment. [Figure 6] Figure 6 shows an example of the first prompt in this embodiment. [Figure 7] Figure 7 shows an example of the format of the second prompt in this embodiment. [Figure 8] Figure 8 shows an example of a conversation screen in this embodiment. [Figure 9] Figure 9 shows an example of the content of a conversation in this embodiment. [Figure 10] Figure 10 shows an example of the interaction between the conversation inconsistency detection device and the generation model in this embodiment. [Figure 11] Figure 11 shows an example of a conversation screen in this embodiment. [Figure 12] Figure 12 shows an example of the interaction between the conversation inconsistency detection device and the generation model of this embodiment. [Figure 13] Figure 13 shows an example of a conversation screen in this embodiment. [Figure 14] Figure 14 shows an example of the interaction between the conversation inconsistency detection device and the generation model in this embodiment. [Figure 15] Figure 15 shows an example of a conversation screen in this embodiment. [Figure 16]FIG. 16 is a diagram showing an example of the interaction between the conversation inconsistency detection device and the generation model of the present embodiment. [Figure 17] FIG. 17 is a diagram showing an example of the conversation screen of the present embodiment. [Figure 18] FIG. 18 is a diagram showing an example of the conversation screen of the present embodiment. [Figure 19] FIG. 19 is a diagram showing an example of the conversation screen of the present embodiment. [Figure 20] FIG. 20 is a flowchart showing an example of the preprocessing performed by the conversation inconsistency detection device of the present embodiment. [Figure 21] FIG. 21 is a flowchart showing an example of the inconsistency detection process performed by the conversation inconsistency detection device of the present embodiment. [Figure 22] FIG. 22 is a diagram showing an example of the first prompt of Modification 1. [Figure 23] FIG. 23 is a diagram showing an example of the interaction between the conversation inconsistency detection device and the generation model of Modification 1. [Figure 24] FIG. 24 is a diagram showing an example of the conversation screen of Modification 1. [Figure 25] FIG. 25 is a diagram showing an example of the conversation screen of Modification 1. [Figure 26] FIG. 26 is a diagram showing an example of the first prompt of Modification 2. [Figure 27] FIG. 27 is a diagram showing an example of the conversation screen of Modification 2. [Figure 28] FIG. 28 is a diagram showing an example of the interaction between the conversation inconsistency detection device and the generation model of Modification 2. [Figure 29] FIG. 29 is a diagram showing an example of the conversation screen of Modification 2. [Figure 30] FIG. 30 is a diagram showing an example of the conversation screen of Modification 2. [Figure 31] FIG. 31 is a diagram showing an example of the first prompt of Modification 3. [Figure 32] FIG. 32 is a diagram showing an example of the conversation screen of Modification 3. [Figure 33]Figure 33 shows an example of the interaction between the conversation inconsistency detection device and the generative model in Modification 3. [Figure 34] Figure 34 shows an example of the conversation screen for Modification Example 3. [Figure 35] Figure 35 shows an example of the conversation screen for Modification Example 3. [Modes for carrying out the invention]

[0013] Hereinafter, embodiments of the present disclosure (hereinafter simply referred to as "these embodiments") will be described in detail with reference to the drawings. However, this disclosure is not limited to the following embodiments. Furthermore, the following embodiments and their variations can be combined as appropriate.

[0014] In the following, the conversation inconsistency detection system of this embodiment will be described using the example of each user engaging in conversation using a communication application (hereinafter, the application will simply be referred to as "app"). Examples of communication apps include metaverse apps, video conferencing apps, and chat apps, but are not limited to these; any app with communication functionality may be used.

[0015] However, the conversations targeted by the conversation inconsistency detection system of this embodiment are not limited to conversations using a communication application, but may also be conversations that users have in person. Furthermore, the conversations targeted by the conversation inconsistency detection system of this embodiment may be conversations between people with any kind of relationship, conversations that take place in any situation, and conversations that contain any content.

[0016] The following describes the conversation inconsistency detection system of this embodiment, using the example of a conversation involving two users. However, the conversation inconsistency detection system of this embodiment is not limited to this case and can naturally be applied to conversations involving three or more users.

[0017] First, the configuration of the conversation inconsistency detection system of this embodiment will be described.

[0018] Figure 1 is a block diagram showing an example of the system configuration of the conversation inconsistency detection system 1 of this embodiment. As shown in Figure 1, the conversation inconsistency detection system 1 comprises a conversation inconsistency detection device 10, user terminals 20-1 and 20-2, and a generation model 30.

[0019] The conversation inconsistency detection device 10, user terminals 20-1 and 20-2, and the generation model 30 are connected via network 2. Network 2 can be implemented by at least one of the following: the Internet and / or a LAN (Local Area Network). Network 2 may be a wired network, a wireless network, or a mixture of wired and wireless networks.

[0020] In the following explanation, if it is not necessary to distinguish between user terminals 20-1 and 20-2, they may simply be referred to as user terminal 20.

[0021] The conversation inconsistency detection device 10 controls user-to-user conversations conducted by each user using the user terminal 20, and uses the generation model 30 to detect the consistency of the content of those conversations. The conversation inconsistency detection device 10 can be implemented, for example, by a server device or a cloud service.

[0022] User terminal 20 is a terminal device with a communication application installed that each user uses to have a conversation. Examples include smartphones, tablet devices, or PCs (Personal Computers). In this embodiment, we will explain using the case where there are two users having a conversation as an example, so two user terminals 20-1 and 20-2 are given as examples of user terminals 20. However, the number of user terminals 20 is not limited to this and may be three or more.

[0023] The generative model 30 is a natural language processing model, also known as generative AI (Generative Artificial Intelligence), and provides services using natural language processing. In this embodiment, the generative model 30 is described as an AI chat service that has been fine-tuned from a large language model (LLM), but it is not limited to this.

[0024] Figure 2 is a block diagram showing an example of the hardware configuration of the conversation inconsistency detection device 10 and user terminal 20 in this embodiment.

[0025] First, the hardware configuration of the conversation inconsistency detection device 10 will be described. As shown in Figure 2, the conversation inconsistency detection device 10 comprises a control device 11, a main memory 12, an auxiliary memory 13, a communication device 14, and various buses 19. The control device 11, the main memory 12, the auxiliary memory 13, and the communication device 14 are connected via the various buses 19. Thus, the conversation inconsistency detection device 10 of this embodiment has a general hardware configuration that utilizes a normal computer.

[0026] The control device 11 controls the overall operation of the conversation inconsistency detection device 10. The control device 11 may be, but is not limited to, at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). There may be one or more CPUs and GPUs, and they may be single-core or multi-core.

[0027] Examples of main memory 12 include, but are not limited to, ROM (Read Only Memory) and RAM (Random Access Memory). ROM stores various programs, such as a program for controlling the conversation inconsistency detection device 10 and a program for the generation model 30 to detect the consistency of the conversation content. RAM is used as a work area when the control device 11 performs various controls based on the programs stored in ROM.

[0028] The auxiliary storage device 13 stores the various programs mentioned above and various data used to allow the generation model 30 to detect the consistency of the conversation content. The various programs mentioned above only need to be stored in at least one of the main storage device 12 and the auxiliary storage device 13. The auxiliary storage device 13 may be, but is not limited to, at least one of existing storage devices that can store magnetically, electrically, or optically, such as an HDD (Hard Disk Drive), SSD (Solid State Drive), and DVD (Digital Versatile Disc). The auxiliary storage device 13 may be built into the conversation inconsistency detection device 10 or attached externally to the conversation inconsistency detection device 10 via an interface such as USB (Universal Serial Bus). The auxiliary storage device 13 may also be a NAS (Network Attached Storage) connected via a network such as a LAN or WAN (Wide Area Network).

[0029] The communication device 14 is used to communicate with the user terminal 20 and the generation model 30, etc., via the network 2. Examples of the communication device 14 include, but are not limited to, a communication device for a wired LAN or a wireless communication device for a wireless LAN.

[0030] In addition to the above configuration, the conversation mismatch detection device 10 may further include hardwired circuits such as ICs (Integrated Circuits), ASICs (Application Specific Integrated Circuits), and FPGAs (Field-Programmable Gate Arrays) that are specific to the conversation mismatch detection device 10.

[0031] Next, the hardware configuration of the user terminal 20 will be described. As shown in Figure 2, the user terminal 20 comprises a control device 21, a main memory 22, an auxiliary memory 23, a communication device 24, an input device 25, a display device 26, an audio input device 27, an audio output device 28, and various buses 29. The control device 21, main memory 22, auxiliary memory 23, communication device 24, input device 25, display device 26, audio input device 27, and audio output device 28 are connected via various buses 29. Thus, the user terminal 20 of this embodiment has a general hardware configuration, such as that of a PC or tablet terminal.

[0032] The control device 21 controls the overall operation of the user terminal 20. The implementation method for the control device 21 is the same as that for the control device 11.

[0033] The implementation method for main memory 22 is the same as that for main memory 12, so a detailed explanation will be omitted. The ROM of main memory 22 stores various programs, such as programs for controlling the user terminal 20 and programs for running communication applications.

[0034] The auxiliary storage device 23 stores the various programs and data described above. Note that the various programs described above only need to be stored in at least one of the main storage device 22 and the auxiliary storage device 23. The implementation method for the auxiliary storage device 23 is the same as that for the auxiliary storage device 13, so a detailed explanation is omitted.

[0035] The communication device 24 is used to communicate with the conversation inconsistency detection device 10 via the network 2. Since the implementation method for the communication device 24 is the same as that for the communication device 14, a detailed explanation is omitted.

[0036] The input device 25 is used for various inputs, selections, and specifications for conversation, and serves as a user interface between the user and the input device. Examples of input devices 25 include, but are not limited to, keyboards, mice, and touch panels. The input device 25 may be built into the user terminal 20 or it may be externally connected via an interface such as USB.

[0037] The display device 26 displays various screens for confirming the content of a conversation and whether or not inconsistencies have occurred in the conversation, and serves as a user interface between the user and the device. Examples of the display device 26 include, but are not limited to, liquid crystal displays, organic electro-luminescence (OLED) displays, and touch panel displays. The display device 26 may be an internal display built into the user terminal 20, or an external display connected to the user terminal 20 via a display interface such as HDMI®.

[0038] The voice input device 27 inputs the user's spoken voice during a conversation, and can be implemented, for example, by a sound collection device such as a microphone.

[0039] The voice output device 28 outputs the speech of other users during a conversation, and can be implemented, for example, by a speaker.

[0040] In addition to the above configuration, the user terminal 20 may further include hardwired circuits such as ICs, ASICs, and FPGAs for speech recognition of the spoken voice input by the voice input device 27.

[0041] Figure 3 is a block diagram showing an example of the functional configuration of the conversation inconsistency detection system 1 of this embodiment, and mainly shows an example of the functional configuration of the conversation inconsistency detection device 10 and user terminals 20-1 and 20-2 included in the conversation inconsistency detection system 1 of this embodiment.

[0042] As shown in Figure 3, the conversation inconsistency detection device 10 includes a prompt creation unit 101, a consistency detection control unit 103, an acquisition unit 105, a speech recognition unit 106, a screen generation unit 107, and an output control unit 109 (an example of a notification control unit). The prompt creation unit 101, consistency detection control unit 103, acquisition unit 105, speech recognition unit 106, screen generation unit 107, and output control unit 109 can be realized, for example, by the control device 11, main memory 12, and communication device 14 described in Figure 2. For example, the control device 11 reads a program that causes the generation model 30 to detect the consistency of the conversation content stored in the main memory 12 (ROM) or auxiliary memory 13, and expands it into the main memory 12 (RAM). The control device 11 realizes each of the above-mentioned functional units by executing various processes according to the expanded program. Here, the case in which each of the above-mentioned functional units is realized as software has been explained as an example, but at least a part of each of the above-mentioned functional units may be realized as hardware. In this case, the functional parts to be implemented as hardware can be implemented, for example, by the hardwired circuits described above. Alternatively, any of the above-mentioned functional parts may be implemented through the cooperation of software and hardware.

[0043] In this embodiment, the conversation inconsistency detection device 10 does not necessarily need to have all of the above-described functional units as essential components; at least some of the functional units may be omitted. For example, if the communication application is a chat application, the speech recognition unit 106 can be omitted.

[0044] Furthermore, as shown in Figure 3, user terminal 20-1 includes a reception unit 201-1, a notification unit 205-1, and an output unit 207-1. User terminal 20-2 includes a reception unit 201-2, a notification unit 205-2, and an output unit 207-2.

[0045] Note that reception units 201-1 and 201-2 have common functions, and therefore, in the following explanation, they may be simply referred to as reception unit 201 unless it is necessary to distinguish between them. Similarly, notification units 205-1 and 205-2 have common functions, and therefore, in the following explanation, they may be simply referred to as notification unit 205 unless it is necessary to distinguish between them. Similarly, output units 207-1 and 207-2 have common functions, and therefore, in the following explanation, they may be simply referred to as output unit 207 unless it is necessary to distinguish between them.

[0046] The reception unit 201, notification unit 205, and output unit 207 can be implemented, for example, by the control device 21, main memory 22, and communication device 24 described in Figure 2. For example, the control device 21 reads a program for executing a communication application stored in the main memory 22 (ROM) or auxiliary memory 23 and loads it into the main memory 22 (RAM). The control device 21 implements each of the above-mentioned functional units within the communication application by executing various processes according to the loaded program. Here, the above-mentioned functional units have been explained using the example of implementation as software, but at least a part of each of the above-mentioned functional units may be implemented as hardware. In this case, the functional unit to be implemented as hardware can be implemented, for example, by the hardwired circuit described above. Alternatively, any of the above-mentioned functional units may be implemented through the cooperation of software and hardware.

[0047] In this embodiment, the user terminal 20 does not need to have all of the above-mentioned functional units as essential components, and at least some of the functional units can be omitted.

[0048] The following describes in detail the functions of the conversation inconsistency detection device 10 and the user terminal 20. In the following, we will use as an example the case where user A and user B use user terminals 20-1 and 20-2, each with a communication application installed, to have a conversation on the conversation inconsistency detection system 1. However, as mentioned above, the number of people who can participate in the conversation is not limited to this.

[0049] First, we will explain the preprocessing performed on the conversation inconsistency detection system 1 before the conversation between user A and user B begins.

[0050] The prompt creation unit 101 creates a first prompt that instructs the generation model 30 to detect the consistency of the conversation with the content entered so far when the content of a conversation between multiple users is input to the generation model 30. For example, the first prompt instructs the generation model 30 to detect the consistency of the conversation based on inconsistency definition information that defines patterns in which inconsistencies are considered to have occurred in the conversation, when the content of the conversation is input. More specifically, the first prompt instructs the generation model 30 to detect the consistency of the conversation based on the inconsistency definition information each time the content of the conversation is input.

[0051] Figures 4 and 5 show an example of the first prompt in this embodiment. In the example shown in Figures 4 and 5, the first prompt consists of two prompts: the prompt shown in Figure 4 and the prompt shown in Figure 5.

[0052] The prompts shown in Figure 4 correspond to inconsistency definition information and define patterns in which a conversation is considered to have occurred (conversation inconsistency). In the prompts shown in Figure 4, four patterns are defined as patterns in which a conversation is considered to have occurred: interruption of conversation, sudden change of topic, unnatural flow, and lack of back-and-forth conversation, and further, a description of the situation for each pattern is provided. However, the patterns in which a conversation is considered to have occurred are not limited to these.

[0053] The prompt shown in Figure 5 is a prompt that instructs the system to detect the consistency of the conversation with the previous content (whether or not there are conversation inconsistencies), as the content of the conversation is input for each user statement. The prompt shown in Figure 5 is input to the generative model 30 after the prompt shown in Figure 4, instructing the generative model 30 to detect the consistency of the conversation based on the prompt shown in Figure 4 (inconsistency definition information).

[0054] In this embodiment, as described above, an example in which the first prompt is composed of two prompts has been explained, but the embodiment is not limited to this, and the two prompts may be integrated into one prompt, so that the first prompt is realized with a single prompt.

[0055] Figure 6 shows an example of the first prompt of this embodiment, and illustrates another example of a prompt corresponding to the inconsistency definition information shown in Figure 4. The prompts shown in Figure 6 are obtained by applying Few-shot Learning to the prompts shown in Figure 4. Specifically, the prompts shown in Figure 6 present concrete example sentences as demonstrations for each pattern in which an inconsistency in the conversation is considered to have occurred. By using the prompts shown in Figure 6, it is expected that the accuracy of the generative model 30 in detecting the consistency of the conversation (presence or absence of conversational inconsistency) will be improved.

[0056] The prompts shown in Figures 4 to 6 may be stored in advance in the auxiliary storage device 23, for example. In this case, the prompt creation unit 101 creates the first prompt by reading these prompts from the auxiliary storage device 23. Alternatively, for example, text data of sentences corresponding to the prompts shown in Figures 4 to 6 may be stored in advance in the auxiliary storage device 23, and the prompt creation unit 101 may create the first prompt by reading this text data from the auxiliary storage device 23 and formatting it into the prompt format.

[0057] The consistency detection control unit 103 inputs the first prompt created by the prompt creation unit 101 into the generation model 30. Specifically, the consistency detection control unit 103 inputs the first prompt into the generation model 30 prior to the start of a conversation between multiple users. In this embodiment, the consistency detection control unit 103 inputs the prompt shown in Figure 4 or Figure 6 into the generation model 30, and then inputs the prompt shown in Figure 5 into the generation model 30.

[0058] Through the above preprocessing, by inputting the content of the conversation into the generative model 30 for each user statement, it becomes possible to detect consistency between the conversation and the content up to that point.

[0059] Next, we will explain the inconsistency detection process performed during the conversation between user A and user B on the conversation inconsistency detection system 1.

[0060] The reception unit 201 receives the content of conversations between multiple users. Specifically, in conversations between multiple users, the reception unit 201 receives the content of user statements input by voice from the voice input device 27. For example, in a conversation between user A and user B, reception unit 201-1 receives user A's statements by voice, and reception unit 201-2 receives user B's statements by voice.

[0061] The notification unit 205 notifies the conversation inconsistency detection device 10 of the content of the conversation received by the reception unit 201. For example, in the case of a conversation between user A and user B, notification unit 205-1 notifies the conversation inconsistency detection device 10 of the audio data of user A's statements. In addition, notification unit 205-2 notifies the conversation inconsistency detection device 10 of the audio data of user B's statements.

[0062] The acquisition unit 105 acquires the content of conversations between multiple users. Specifically, the acquisition unit 105 acquires the content of each user's statement each time they speak in the conversation. For example, when user A speaks in the conversation and audio data of user A's statement is notified from user terminal 20-1, the acquisition unit 105 acquires the audio data of user A's statement as user A's statement. Similarly, when user B speaks in the conversation and audio data of user B's statement is notified from user terminal 20-2, the acquisition unit 105 acquires the audio data of user B's statement as user B's statement.

[0063] The speech recognition unit 106 performs speech recognition on the audio content of the conversation acquired by the acquisition unit 105 and generates conversation information that shows the content of the conversation in text format. For example, in the case of a conversation between user A and user B, the speech recognition unit 106 performs speech recognition on the audio data of user A's statements and generates conversation information for user A that shows the content of user A's statements in text format. The speech recognition unit 106 also performs speech recognition on the audio data of user B's statements and generates conversation information for user B that shows the content of user B's statements in text format.

[0064] In cases where text data is used instead of audio data of a conversation, such as in a chat application, the speech recognition unit 106 can be omitted, and the notification unit 205 can simply notify the conversation inconsistency detection device 10 of the text data of the conversation content received by the reception unit 201 as conversation information.

[0065] The prompt creation unit 101 creates a second prompt indicating the content of a user's statement each time that content is acquired on a user-by-user basis. Specifically, the second prompt indicates the speech recognition result of the speech content by the speech recognition unit 106. For example, when the acquisition unit 105 acquires conversation information of user A as the content of user A's statement, the prompt creation unit 101 uses the conversation information of user A to create a second prompt indicating the content of user A's statement. Similarly, when the acquisition unit 105 acquires conversation information of user B as the content of user B's statement, the prompt creation unit 101 uses the conversation information of user B to create a second prompt indicating the content of user B's statement.

[0066] Figure 7 shows an example of the format of the second prompt in this embodiment. In the example shown in Figure 7, the format of the second prompt consists of the speaker and the content of the statement. For example, when the acquisition unit 105 acquires conversation information of user A, the prompt creation unit 101 creates the second prompt by inserting the name of user A (for example, Mr. / Ms. A) into the speaker part of the format shown in Figure 7 and inserting the content of the statement of user A as indicated by the conversation information of user A into the content of the statement. Similarly, when the acquisition unit 105 acquires conversation information of user B, the prompt creation unit 101 creates the second prompt by inserting the name of user B (for example, Mr. / Ms. B) into the speaker part of the format shown in Figure 7 and inserting the content of the statement of user B as indicated by the conversation information of user B into the content of the statement.

[0067] The consistency detection control unit 103 sequentially inputs the content of conversations between multiple users into the generation model 30 and obtains a detection result from the generation model 30 that detects the consistency of the conversation with the content input up to that point. Specifically, each time a second prompt is generated by the prompt creation unit 101, the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result. Because the above-described preprocessing has been performed, the consistency detection control unit 103 can detect the consistency of the conversation with the content up to that point by inputting the content of the conversation into the generation model 30 for each user statement.

[0068] The screen generation unit 107 generates conversation screens for conversations between multiple users. Specifically, each time the acquisition unit 105 acquires user-specific statements, the screen generation unit 107 generates a conversation screen that associates the user who made the statement with the statement content. Furthermore, if the detection result acquired by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that includes information to inform each user that an inconsistency has occurred in the conversation.

[0069] The information used to inform each user that an inconsistency has occurred in the conversation may be textual information, non-textual information such as graphics, an effect on the conversation screen, or audio information output along with the conversation screen, and at least one of these types of information is included. The effect on the conversation screen may be any special effect that occurs on the conversation screen, such as the blinking of an object on the conversation screen (for example, a blinking object such as a lamp), the appearance of an object indicating that an inconsistency has occurred in the conversation, or the conversation screen going dark.

[0070] If the detection result obtained by the consistency detection control unit 103 indicates a mismatch, the output control unit 109 performs notification control to indicate that a mismatch has occurred. Specifically, the output control unit 109 outputs to each user terminal 20 that a mismatch has occurred and causes each user terminal 20 to perform notification control to indicate that a mismatch has occurred. The notification control may be screen-based control, voice-based control, or control of a notification device (not shown), such as the illumination of a lamp, and at least one of these controls is included.

[0071] For example, the output control unit 109 outputs the conversation screen generated by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20. As a result, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, notification control of the inconsistency is performed on the conversation screen.

[0072] When the conversation screen is output from the conversation inconsistency detection device 10, the output unit 207 outputs the conversation screen to the display device 26 and displays the conversation screen on the display device 26.

[0073] The following explains the inconsistency detection process when a communication app is a metaverse app, using specific examples.

[0074] Figure 8 shows an example of the conversation screen 401 of this embodiment, and shows the conversation screen 401 displayed on the display device 26 of each user terminal 20 before the conversation between user A and user B starts on the conversation inconsistency detection system 1. The conversation screen 401 shows the state before Avatar A, which is the avatar of user A, and Avatar B, which is the avatar of user B, start a conversation in a virtual chat room. When a conversation between users starts via the avatars, the content of the users' statements is displayed in association with the avatars. Figure 9 shows an example of the content of a conversation of this embodiment. In the following, we will explain using as an example the case in which the conversation between user A and user B takes place with the statements 501 to 507 shown in Figure 9.

[0075] First, when User A makes the statement "Thank you for your hard work at yesterday's event" (hereinafter referred to as "Statement 501"), the acquisition unit 105 acquires User A's voice data indicating Statement 501 from User Terminal 20-1, and the voice recognition unit 106 performs voice recognition on the voice data to generate User A's conversation information indicating Statement 501. The prompt creation unit 101 creates a second prompt by inserting User A's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0076] Figure 10 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 in this embodiment. In this embodiment, the chat history, which is the interaction between the conversation inconsistency detection device 10 and the generation model 30, is assumed to be held by the consistency detection control unit 103. In the example shown in Figure 10, the chat history displays the second prompt 511, which shows the content of user A's statement 501, and the consistency detection result 513 of the conversation content up to the second prompt 511, in chronological order. The detection result 513 shows no inconsistency.

[0077] The screen generation unit 107 generates a conversation screen that associates the utterance content 501, indicated by the conversation information of user A generated by the speech recognition unit 106, with avatar A speaking. However, since the detection result 513 obtained by the consistency detection control unit 103 does not indicate an inconsistency, a conversation screen containing information to notify each user that an inconsistency has occurred in the conversation is not generated.

[0078] Figure 11 shows an example of the conversation screen 411 of this embodiment. In the conversation screen 411 shown in Figure 11, a speech bubble 415 indicating the content of the statement 501 is associated with avatar A, and the screen shows avatar A speaking the content of the statement 501.

[0079] The output control unit 109 outputs the conversation screen 411 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 411 on each user terminal 20. This allows user B to confirm what user A has said, and user A to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 501 to user terminal 20-2, and the spoken content 501 may be synthesized into speech and output as audio on user terminal 20-2.

[0080] Next, when User B says "Thanks for your hard work, that was fun." (hereinafter referred to as "Statement 503"), the acquisition unit 105 acquires User B's voice data indicating Statement 503 from User Terminal 20-2, and the voice recognition unit 106 performs voice recognition on the voice data to generate User B's conversation information indicating Statement 503. The prompt creation unit 101 creates a second prompt by inserting User B's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0081] Figure 12 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of this embodiment. The chat history shown in Figure 12 is the same as the chat history shown in Figure 10, but with the addition of a second prompt 521 indicating the content of user B's statement 503 and a detection result 523 of the consistency of the conversation content up to the second prompt 521, and is displayed in chronological order. The detection result 523 shows no inconsistencies.

[0082] The screen generation unit 107 generates a conversation screen that associates the utterance content 503, indicated by the conversation information of user B generated by the speech recognition unit 106, with avatar B speaking. However, since the detection result 523 obtained by the consistency detection control unit 103 does not indicate an inconsistency, a conversation screen containing information to notify each user that an inconsistency has occurred in the conversation is not generated.

[0083] Figure 13 shows an example of the conversation screen 421 of this embodiment. In the conversation screen 421 shown in Figure 13, a speech bubble 425 indicating the content of the statement 503 is associated with avatar B, and the screen shows avatar B speaking the content of the statement 503.

[0084] The output control unit 109 outputs the conversation screen 421 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 421 on each user terminal 20. This allows user A to confirm what user B has said, and user B to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 503 to user terminal 20-1, and then synthesize the spoken content 503 into speech and output it as audio on user terminal 20-1.

[0085] Next, when User A says, "I wonder what happened to Suzuki and the others after that?" (hereinafter referred to as "Statement 505"), the acquisition unit 105 acquires User A's voice data indicating Statement 505 from User Terminal 20-1, and the speech recognition unit 106 performs speech recognition on the voice data to generate User A's conversation information indicating Statement 505. The prompt creation unit 101 creates a second prompt by inserting User A's conversation information acquired by the acquisition unit 105 into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0086] Figure 14 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of this embodiment. The chat history shown in Figure 14 is the same as the chat history shown in Figure 12, but with the addition of a second prompt 531 indicating the content of user A's statement 505 and a detection result 533 of the consistency of the conversation content up to the second prompt 531, and is displayed in chronological order. The detection result 533 shows no inconsistencies.

[0087] The screen generation unit 107 generates a conversation screen that associates the utterance content 505, indicated by the conversation information of user A generated by the speech recognition unit 106, with avatar A speaking. However, since the detection result 533 obtained by the consistency detection control unit 103 does not indicate an inconsistency, a conversation screen containing information to notify each user that an inconsistency has occurred in the conversation is not generated.

[0088] Figure 15 shows an example of the conversation screen 431 of this embodiment. In the conversation screen 431 shown in Figure 15, a speech bubble 435 indicating the content of the statement 505 is associated with avatar A, and the screen shows avatar A speaking the content of the statement 505.

[0089] The output control unit 109 outputs the conversation screen 421 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 421 on each user terminal 20. This allows user B to confirm what user A said, and user A to confirm what they said. Alternatively, the output control unit 109 may output the spoken content 505 to user terminal 20-2, and the spoken content 505 may be synthesized into speech and output as audio on user terminal 20-2.

[0090] Next, when User B says, "By the way, I'm hungry," (hereinafter referred to as "Statement 507"), the acquisition unit 105 acquires User B's voice data indicating Statement 507 from User Terminal 20-2, and the voice recognition unit 106 performs voice recognition on the voice data to generate User B's conversation information indicating Statement 507. The prompt creation unit 101 creates a second prompt by inserting User B's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0091] Figure 16 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of this embodiment. The chat history shown in Figure 16 is the same as the chat history shown in Figure 14, but with the addition of a second prompt 541 indicating the content of User B's statement 507 and a detection result 543 indicating the consistency of the conversation content up to the second prompt 541, and is displayed in chronological order. Note that User B's statement 507 is a sudden change of topic, so the detection result 543 is an inconsistency detection.

[0092] The screen generation unit 107 generates a conversation screen that associates the utterance content 507, indicated by the conversation information of user B generated by the speech recognition unit 106, with avatar B speaking. Furthermore, since the detection result 543 obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that includes information to inform each user that an inconsistency has occurred in the conversation.

[0093] Figure 17 shows an example of a conversation screen 441 in this embodiment. In the conversation screen 441 shown in Figure 17, a speech bubble 445 indicating the content of the statement 507 is associated with avatar B, and the screen shows avatar B speaking the content of the statement 507. Figure 18 shows an example of a conversation screen 451 in this embodiment. In the conversation screen 451 shown in Figure 18, as a result of avatar B speaking the content of the statement 507, an inconsistency occurs in the conversation, and the conversation screen includes information 455 to inform each user that an inconsistency has occurred in the conversation.

[0094] The output control unit 109 outputs the conversation screen 441 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 441 on each user terminal 20. This allows user A to confirm what user B has said, and user B to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 507 to user terminal 20-1, and then synthesize the spoken content 507 on user terminal 20-1 and output it as speech.

[0095] The output control unit 109 further outputs the conversation screen 451 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 451 on each user terminal 20. This allows each user to confirm that an inconsistency has occurred in the conversation. Alternatively, the output control unit 109 may output information 455 to each user terminal 20 and output a speech synthesis message on each user terminal 20 indicating that an inconsistency has occurred in the conversation (information 455).

[0096] The above explains the inconsistency detection process when a communication app is a metaverse app, using specific examples. However, the above explanation can also be applied when the communication app is a chat app, except for the UI (User Interface) of the conversation screen.

[0097] Figure 19 shows an example of a conversation screen 601 in this embodiment. The conversation screen 601 shown in Figure 19 is an example of a conversation screen (chat screen) when the communication application is a chat application.

[0098] If the communication application is a chat application, when the conversation screen 411 shown in Figure 11 of the above explanation is displayed on the user terminal 20, a text message 611 indicating the content of the message 501 is displayed on the conversation screen 601 shown in Figure 19. Also, when the conversation screen 421 shown in Figure 13 of the above explanation is displayed on the user terminal 20, a text message 621 indicating the content of the message 503 is displayed on the conversation screen 601 shown in Figure 19. Also, when the conversation screen 431 shown in Figure 15 of the above explanation is displayed on the user terminal 20, a text message 631 indicating the content of the message 505 is displayed on the conversation screen 601 shown in Figure 19. Also, when the conversation screen 441 shown in Figure 17 of the above explanation is displayed on the user terminal 20, a text message 641 indicating the content of the message 507 is displayed on the conversation screen 601 shown in Figure 19. Also, when the conversation screen 451 shown in Figure 18 of the above explanation is displayed on the user terminal 20, a text message 651 indicating information 455 is displayed on the conversation screen 601 shown in Figure 19.

[0099] Next, the operation of the conversation inconsistency detection system of this embodiment will be described.

[0100] Figure 20 is a flowchart showing an example of preprocessing performed by the conversation inconsistency detection device 10 of this embodiment.

[0101] First, the prompt generation unit 101 creates a prompt that defines a pattern in which an inconsistency in the conversation is considered to have occurred (step S101).

[0102] Next, the consistency detection control unit 103 inputs the prompt created in step S101 by the prompt creation unit 101 to the generation model 30 (step S103).

[0103] Next, the prompt generation unit 101 creates a prompt that instructs the system to detect the consistency of the conversation with the previously entered content each time new conversation content is input (step S105).

[0104] Next, the consistency detection control unit 103 inputs the prompt created in step S105 by the prompt creation unit 101 to the generation model 30 (step S107).

[0105] Figure 21 is a flowchart showing an example of the inconsistency detection process performed by the conversation inconsistency detection device 10 of this embodiment.

[0106] First, the acquisition unit 105 acquires audio data of each user's speech each time they speak during a conversation (step S201).

[0107] Next, when the speech recognition unit 106 acquires audio data of the user's speech, it performs speech recognition on the audio data and generates conversation information indicating the content of the user's speech (step S202).

[0108] Next, when conversation information indicating the content of the user's statement is generated, the prompt creation unit 101 uses that conversation information to create a prompt indicating the content of that statement (step S203).

[0109] Next, the consistency detection control unit 103 inputs the second prompt created by the prompt creation unit 101 to the generation model 30 (step S205), and obtains the result of detecting the consistency of the conversation with the content of the conversation so far from the generation model 30 (step S207).

[0110] Next, the screen generation unit 107 generates a conversation screen that associates the spoken content acquired by the acquisition unit 105 with the user who made the spoken word (step S209).

[0111] Next, the output control unit 109 outputs the conversation screen generated in step S209 by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20 (step S211).

[0112] Next, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency (Yes in step S213), the screen generation unit 107 generates a conversation screen containing information to inform each user that an inconsistency has occurred in the conversation (step S215).

[0113] Next, the output control unit 109 outputs the conversation screen generated in step S215 by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20 (step S217).

[0114] If the detection result obtained by the consistency detection control unit 103 does not indicate a mismatch (No in step S213), the processing in steps S215 to S217 is not performed.

[0115] The conversation inconsistency detection device 10 repeats the processes in steps S201 to S219 as long as the conversation between multiple users continues (No in step S219), and terminates the process when the conversation between multiple users ends (Yes in step S219). Whether or not the conversation between multiple users has ended can be determined, for example, by detecting whether or not the user has closed the communication application, or whether or not the conversation has been interrupted for a certain period of time.

[0116] As described above, in this embodiment, since a generative model (generative AI) is used to detect the consistency of a conversation, there is no need to develop a function to detect conversation consistency oneself, and the detection of conversation consistency can be easily implemented. This reduces development costs and also reduces the cost of introducing a function to detect conversation consistency.

[0117] Furthermore, in this embodiment, the consistency of the conversation with previous content is detected for each user statement in the conversation, and if an inconsistency is detected, the system notifies the user that an inconsistency has occurred. As a result, users participating in the conversation can understand in real time when an inconsistency has occurred. Therefore, according to this embodiment, it is possible to prompt the user who made the inconsistent statement to supplement the content of their statement early on, or to prompt other users to supplement the inconsistent statement, and it is expected that this will prevent the conversation from developing into a breakdown. In particular, this embodiment is suitable for supporting conversations of users who have characteristics that cause inconsistencies in conversations unintentionally, such as speaking only what they want to say.

[0118] Furthermore, in this embodiment, as a preprocessing step before the conversation begins, patterns that are considered to indicate a conversation inconsistency are presented (input) to the generation model 30. Therefore, during the conversation, the generation model 30 can detect the consistency of the conversation simply by presenting (inputting) the content of each user's statements to the generation model 30. Accordingly, this embodiment is expected to shorten the time required for the generation model 30 to detect the consistency of the conversation, and it is possible to notify in real time when a conversation inconsistency occurs.

[0119] (Variation 1) In the above embodiment, an example was described in which a notification is given when an inconsistency in the conversation is detected. In Modification 1, an example is described in which supplementary content that corrects the inconsistency in the conversation is reported as a statement made by the user who made the inconsistent statement. In Modification 1, the parts that differ from the above embodiment will be mainly described, and the parts that are the same as the above embodiment will not be explained.

[0120] First, we will explain the differences in pretreatment compared to the above embodiment.

[0121] In Modification 1, the prompt creation unit 101 creates a prompt as a first prompt that instructs the generation model 30 to detect the consistency of the conversation based on the inconsistency definition information when conversation content is input, and to generate supplementary content if an inconsistency is detected. More specifically, the first prompt instructs the generation model 30 to detect the consistency of the conversation based on the inconsistency definition information each time conversation content is input, and to generate supplementary content if an inconsistency is detected.

[0122] In Modification 1, the prompts corresponding to the inconsistency definition information that constitute the first prompt are the same as the prompts in the above embodiment (for example, the prompts shown in Figures 4 and 6), but the prompts that instruct the detection of inconsistencies are different from the prompts in the above embodiment (for example, the prompt shown in Figure 5).

[0123] Figure 22 shows an example of the first prompt in Modification 1. The prompt shown in Figure 22 is a prompt that instructs the generation model 30 to detect the consistency of the conversation with the content so far (whether or not there is a conversation inconsistency), and if an inconsistency is detected, to generate supplementary content from the speaker's perspective, as the content of the conversation is input for each user statement. The prompt shown in Figure 22 is input to the generation model 30 after the prompts shown in Figures 4 and 6, instructing the generation model 30 to detect the consistency of the conversation based on the prompts (inconsistency definition information) shown in Figures 4 and 6, and to generate supplementary content from the speaker's perspective if an inconsistency is detected.

[0124] In the modified example 1, the consistency detection control unit 103 inputs the prompt shown in Figure 4 or Figure 6 to the generation model 30, and then inputs the prompt shown in Figure 22 to the generation model 30.

[0125] Next, we will explain the differences between the inconsistency detection process and the above embodiment.

[0126] In Modification 1, if the consistency detection result of the speech obtained from the generation model 30 indicates an inconsistency, the consistency detection control unit 103 further obtains supplementary content from the generation model 30 to supplement the inconsistency that occurred in the conversation. In Modification 1, the supplementary content is content that supplements the inconsistency from the perspective of the user who spoke the content that caused the inconsistency in the conversation.

[0127] In the modified example 1, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that associates the supplementary content obtained by the consistency detection control unit 103 with the user who made the inconsistent statement, generated by the generation model 30.

[0128] In Modification 1, the output control unit 109 controls the notification of supplementary content if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency. Specifically, the output control unit 109 controls the output of at least one of the image and audio of the user who spoke the content causing the inconsistency in the conversation speaking the supplementary content. For example, the output control unit 109 outputs the conversation screen generated by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20. As a result, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the notification of supplementary content is controlled on the conversation screen.

[0129] Next, we will explain the differences from the above embodiment regarding the inconsistency detection process when the communication app is a metaverse app. Note that the difference from the above embodiment occurs after User B says, "By the way, I'm hungry." Therefore, Modification 1 will explain the process after User B makes the above statement.

[0130] When User B makes the statement "By the way, I'm hungry" (hereinafter referred to as "Statement 507"), the acquisition unit 105 acquires User B's voice data indicating Statement 507 from User Terminal 20-2, and the voice recognition unit 106 performs voice recognition on the voice data to generate User B's conversation information indicating Statement 507. The prompt creation unit 101 creates a second prompt by inserting User B's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0131] Figure 23 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of Modification 1. The chat history shown in Figure 23 is the same as the chat history shown in Figure 14, but with the addition of a second prompt 541 indicating the content 507 spoken by user B, and supplementary content 743 generated by the generation model 30 because the detection result for the consistency of the conversation content up to the second prompt 541 was found to be inconsistent, and these are displayed in chronological order. Note that the content 507 spoken by user B is a sudden change of topic, so the detection result is found to be inconsistent.

[0132] The screen generation unit 107 generates a conversation screen that associates the utterance content 507, which is indicated by the conversation information of user B generated by the speech recognition unit 106, with avatar B speaking. Furthermore, because the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that associates the supplementary content 743 obtained by the consistency detection control unit 103 with avatar B speaking.

[0133] Figure 24 shows an example of the conversation screen 751 of Modification 1. In the conversation screen 751 shown in Figure 24, an inconsistency occurs in the conversation as a result of Avatar B uttering statement 507. To correct this inconsistency, a speech bubble 745 showing supplementary content 743 is associated with Avatar B, and the screen shows Avatar B uttering supplementary content 743.

[0134] The output control unit 109 outputs the conversation screen 441 shown in Figure 17, which is generated by the screen generation unit 107, to user terminals 20-1 and 20-2, and displays the conversation screen 441 on each user terminal 20. This allows user A to confirm what user B has said, and user B to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 507 to user terminal 20-1, and then synthesize the spoken content 507 on user terminal 20-1 and output it as speech.

[0135] The output control unit 109 further outputs the conversation screen 751 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 751 on each user terminal 20. This allows each user to check the supplementary information regarding any inconsistencies that occurred in the conversation. Alternatively, the output control unit 109 may output the supplementary information 743 to each user terminal 20, and each user terminal 20 may synthesize the supplementary information 743 into speech and output it as audio.

[0136] Next, we will explain the differences between the conversation screen when the communication app is a chat app and the embodiment described above.

[0137] Figure 25 shows an example of the conversation screen 801 in Modification Example 1. The conversation screen 801 shown in Figure 25 is an example of a conversation screen (chat screen) when the communication application is a chat application.

[0138] If the communication application is a chat application, at the same time that the conversation screen 441 shown in Figure 17 of the above explanation is displayed on the user terminal 20, a text message 641 indicating the content of the statement 507 is displayed on the conversation screen 801 shown in Figure 25. Also, at the same time that the conversation screen 751 shown in Figure 24 of the above explanation is displayed on the user terminal 20, a text message 851 indicating supplementary content 743 is displayed on the conversation screen 801 shown in Figure 25.

[0139] In Modification 1, the system detects the consistency of each user's statement in a conversation with the content up to that point. If an inconsistency is detected, the system notifies the user who made the inconsistent statement of the supplementary information that corrects the inconsistency. As a result, users participating in the conversation can grasp the correction of any inconsistencies in the conversation in real time. According to Modification 1, the flow of the conversation can be automatically corrected based on the supplementary information, allowing each user participating in the meeting to continue the conversation naturally.

[0140] (Modification 2) In the first variation described above, supplementary information that corrects inconsistencies in a conversation is reported as being made by the user who made the inconsistent statement. In the second variation, however, supplementary information is reported as being made by the intermediary. In the second variation, the differences from the first variation will be explained, and the parts that are the same as in the first variation will be omitted.

[0141] First, I will explain the differences in the pretreatment compared to Modification Example 1.

[0142] Figure 26 shows an example of the first prompt in Modification 2. The prompt shown in Figure 26 is a prompt that instructs the generation model 30 to detect the consistency of the conversation with the content so far (whether or not there is a conversation inconsistency), and if an inconsistency is detected, to generate supplementary content from the perspective of an intermediary. The intermediary is a third party different from the users participating in the conversation. The prompt shown in Figure 26 is input to the generation model 30 after the prompts shown in Figures 4 and 6, instructing the generation model 30 to detect the consistency of the conversation based on the prompts (inconsistency definition information) shown in Figures 4 and 6, and to generate supplementary content from the perspective of an intermediary if an inconsistency is detected.

[0143] In the modified example 2, the consistency detection control unit 103 inputs the prompt shown in Figure 4 or Figure 6 to the generation model 30, and then inputs the prompt shown in Figure 26 to the generation model 30.

[0144] Next, we will explain the differences in the inconsistency detection process compared to Modification Example 1.

[0145] In Modification 2, if the consistency detection result of the speech obtained from the generation model 30 indicates inconsistency, the consistency detection control unit 103 further obtains supplementary content from the generation model 30 to supplement the inconsistency that occurred in the conversation. In Modification 1, the supplementary content is content that supplements the inconsistency from the perspective of the mediator.

[0146] In the modified example 2, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that is generated by the generation model 30 and associates the supplementary content obtained by the consistency detection control unit 103 with the intermediary.

[0147] In Modification 2, the output control unit 109 controls the notification of supplementary content if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency. Specifically, the output control unit 109 controls the output of at least one of the image and audio in which the intermediary speaks the supplementary content. For example, the output control unit 109 outputs the conversation screen generated by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20. As a result, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the notification of supplementary content is controlled on the conversation screen.

[0148] Next, we will explain the differences between the inconsistency detection process when the communication app is a metaverse app and Modification 1.

[0149] Figure 27 shows an example of the conversation screen 901 of Modification 2, and shows the conversation screen 901 displayed on the display device 26 of each user terminal 20 before the conversation between user A and user B begins on the conversation inconsistency detection system 1. The conversation screen 901 shows the state before Avatar A, which is the avatar of user A, and Avatar B, which is the avatar of user B, begin a conversation in a virtual chat room. When a conversation between users begins via the avatars, the content of the users' statements is displayed in association with the avatars. The conversation screen 901 also includes an intermediary 991, which is an NPC (Non-Player Character) not associated with the users. If an inconsistency in the conversation is detected, the intermediary 991 will correct the inconsistency. Below, the conversation between user A and user B takes place, with the statements 501 to 507 shown in Figure 9.

[0150] The difference from Modification 1 described below occurs after User B says, "By the way, I'm hungry." Therefore, Modification 2 will explain the process that takes place after User B makes the aforementioned statement.

[0151] When User B makes the statement "By the way, I'm hungry" (hereinafter referred to as "Statement 507"), the acquisition unit 105 acquires User B's voice data indicating Statement 507 from User Terminal 20-2, and the voice recognition unit 106 performs voice recognition on the voice data to generate User B's conversation information indicating Statement 507. The prompt creation unit 101 creates a second prompt by inserting User B's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0152] Figure 28 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of Modification 2. The chat history shown in Figure 28 is the same as the chat history shown in Figure 14, but with the addition of a second prompt 541 indicating the content 507 spoken by user B, and supplementary content 943 generated by the generation model 30 because the detection result for the consistency of the conversation content up to the second prompt 541 was found to be inconsistent, and these are displayed in chronological order. Note that the content 507 spoken by user B is a sudden change of topic, so the detection result is found to be inconsistent.

[0153] The screen generation unit 107 generates a conversation screen that associates the utterance content 507, which is indicated by the conversation information of user B generated by the speech recognition unit 106, with avatar B speaking. Furthermore, because the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that associates the supplementary content 943 obtained by the consistency detection control unit 103 with mediator 991 speaking.

[0154] Figure 29 shows an example of the conversation screen 951 of Modification 2. In the conversation screen 951 shown in Figure 29, an inconsistency occurs in the conversation as a result of Avatar B making a statement 507. To resolve this inconsistency, a speech bubble 945 showing supplementary content 943 is associated with the mediator 991, and the screen shows the mediator 991 making the supplementary content 943.

[0155] The output control unit 109 outputs the conversation screen 441 shown in Figure 17, which is generated by the screen generation unit 107, to user terminals 20-1 and 20-2, and displays the conversation screen 441 on each user terminal 20. This allows user A to confirm what user B has said, and user B to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 507 to user terminal 20-1, and then synthesize the spoken content 507 on user terminal 20-1 and output it as speech.

[0156] The output control unit 109 further outputs the conversation screen 951 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 951 on each user terminal 20. This allows each user to check the supplementary information regarding any inconsistencies that occurred in the conversation. Alternatively, the output control unit 109 may output the supplementary information 943 to each user terminal 20, and each user terminal 20 may synthesize the supplementary information 943 into speech and output it as speech.

[0157] Next, we will explain the differences between the conversation screen when the communication app is a chat app and the modified example 1.

[0158] Figure 30 shows an example of the conversation screen 1001 of Modification 2. The conversation screen 1001 shown in Figure 30 is an example of a conversation screen (chat screen) when the communication application is a chat application.

[0159] If the communication application is a chat application, at the time the conversation screen 441 shown in Figure 17 of the above explanation is displayed on the user terminal 20, a text message 641 indicating the content of the statement 507 is displayed on the conversation screen 1001 shown in Figure 30. Also, at the time the conversation screen 951 shown in Figure 29 of the above explanation is displayed on the user terminal 20, a text message 1051 indicating supplementary content 943 is displayed on the conversation screen 1001 shown in Figure 30.

[0160] In Modification 2, the system detects the consistency of each user's statement in the conversation with the content up to that point. If an inconsistency is detected, the system, acting as an intermediary, notifies the user of supplementary information to address the inconsistency. As a result, users participating in the conversation can grasp the correction of any inconsistencies in the conversation in real time. According to Modification 2, the flow of the conversation can be automatically corrected based on the supplementary information provided by the intermediary, who acts as a third party, allowing each user participating in the meeting to continue the conversation more naturally.

[0161] (Variation 3) In the above Modification 1, an example was described in which supplementary information that corrects inconsistencies in a conversation is reported as being made by the user who made the inconsistent statement. In Modification 3, an example is described in which the supplementary information is reported as being made by the user's supporter. In Modification 3, the differences from Modification 1 will be mainly explained, and the parts that are the same as in Modification 1 will be omitted.

[0162] First, I will explain the differences in the pretreatment compared to Modification Example 1.

[0163] Figure 31 shows an example of the first prompt in Modification 3. The prompt shown in Figure 31 is a prompt that instructs the generation model 30 to detect the consistency of the conversation with the content so far (whether or not there is a conversation inconsistency), and if an inconsistency is detected, to generate supplementary content from the perspective of a supporter of the user, as the content of the conversation is input for each statement made by the user. The prompt shown in Figure 31 is input to the generation model 30 after the prompts shown in Figures 4 and 6, instructing the generation model 30 to detect the consistency of the conversation based on the prompts (inconsistency definition information) shown in Figures 4 and 6, and to generate supplementary content from the perspective of a supporter of the user if an inconsistency is detected.

[0164] In the third modified example, the consistency detection control unit 103 inputs the prompt shown in Figure 4 or Figure 6 to the generation model 30, and then inputs the prompt shown in Figure 31 to the generation model 30.

[0165] Next, we will explain the differences in the inconsistency detection process compared to Modification Example 1.

[0166] In Modification 3, if the consistency detection result of the speech obtained from the generation model 30 indicates inconsistency, the consistency detection control unit 103 further obtains supplementary content from the generation model 30 to supplement the inconsistency that occurred in the conversation. In Modification 3, the supplementary content is content that supplements the inconsistency from the perspective of a user's supporter.

[0167] In the modified example 3, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the screen generation unit 107 generates a conversation screen that is generated by the generation model 30 and associates the supplementary content obtained by the consistency detection control unit 103 with the user's supporter.

[0168] In Modification 3, the output control unit 109 controls the notification of supplementary content if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency. Specifically, the output control unit 109 controls the output of at least one of the image and audio in which the user's supporter speaks the supplementary content. For example, the output control unit 109 outputs the conversation screen generated by the screen generation unit 107 to each user terminal 20 and displays it on each user terminal 20. As a result, if the detection result obtained by the consistency detection control unit 103 indicates an inconsistency, the notification of supplementary content is controlled on the conversation screen.

[0169] Next, we will explain the differences between the inconsistency detection process when the communication app is a metaverse app and Modification 1.

[0170] Figure 32 shows an example of the conversation screen 1101 of Modification 3, and shows the conversation screen 1101 displayed on the display device 26 of each user terminal 20 before the conversation between user A and user B begins on the conversation inconsistency detection system 1. The conversation screen 1101 shows the state before Avatar A, which is the avatar of user A, and Avatar B, which is the avatar of user B, begin a conversation in a virtual chat room. When a conversation between users begins via their avatars, the content of each user's statements is displayed in association with their avatars. In addition, in the conversation screen 1101, User A's supporter 1191A, which is an NPC associated with user A, is displayed in association with Avatar A, and User B's supporter 1191B, which is an NPC associated with user B, is displayed in association with Avatar B. If a conversation inconsistency is detected, supporter 1191A will supplement the inconsistency as User A's supporter, and if a conversation inconsistency is detected, supporter 1191B will supplement the inconsistency as User B's supporter. In the following, a conversation takes place between User A and User B, consisting of statements 501-507 as shown in Figure 9.

[0171] The difference from Modification 1 described below arises after User B says, "By the way, I'm hungry." Therefore, Modification 3 will explain the process that takes place after User B makes the aforementioned statement.

[0172] When User B makes the statement "By the way, I'm hungry" (hereinafter referred to as "Statement 507"), the acquisition unit 105 acquires User B's voice data indicating Statement 507 from User Terminal 20-2, and the voice recognition unit 106 performs voice recognition on the voice data to generate User B's conversation information indicating Statement 507. The prompt creation unit 101 creates a second prompt by inserting User B's conversation information into the format shown in Figure 7, and the consistency detection control unit 103 inputs the second prompt into the generation model 30 and obtains the detection result.

[0173] Figure 33 shows an example of the interaction between the conversation inconsistency detection device 10 and the generation model 30 of Modification 3. The chat history shown in Figure 33 is the same as the chat history shown in Figure 14, but with the addition of the second prompt 541 indicating the content 507 of user B's statement, and supplementary content 1143B and 1143A generated by the generation model 30 because the detection result for the consistency of the conversation content up to the second prompt 541 was found to be inconsistent, and is displayed in chronological order. Note that the content 507 of user B's statement is a sudden change of topic, so the detection result is found to be inconsistent.

[0174] The screen generation unit 107 generates a conversation screen that associates the utterance content 507, indicated by the conversation information of user B generated by the speech recognition unit 106, with avatar B speaking. Furthermore, because the detection results obtained by the consistency detection control unit 103 show inconsistencies, the screen generation unit 107 generates a conversation screen that associates the supplementary content 1143B and 1143A obtained by the consistency detection control unit 103 with supporters 1191B and 1191A speaking, respectively.

[0175] Figure 34 shows an example of the conversation screen 1151 of Modification 3. In the conversation screen 1151 shown in Figure 34, an inconsistency occurs in the conversation as a result of Avatar B making a statement 507. To resolve this inconsistency, a speech bubble 1145B indicating supplementary content 1143B is associated with supporter 1191B, and the screen shows supporter 1191B making a statement about supplementary content 1143B. Also in the conversation screen 1151 shown in Figure 34, a speech bubble 1145A indicating supplementary content 1143A is associated with supporter 1191A, and the screen shows supporter 1191A making a statement about supplementary content 1143B by supporter 1191B and responding with supplementary content 1143A.

[0176] The output control unit 109 outputs the conversation screen 441 shown in Figure 17, which is generated by the screen generation unit 107, to user terminals 20-1 and 20-2, and displays the conversation screen 441 on each user terminal 20. This allows user A to confirm what user B has said, and user B to confirm what they have said. Alternatively, the output control unit 109 may output the spoken content 507 to user terminal 20-1, and then synthesize the spoken content 507 on user terminal 20-1 and output it as speech.

[0177] The output control unit 109 further outputs the conversation screen 1151 generated by the screen generation unit 107 to user terminals 20-1 and 20-2, and displays the conversation screen 1151 on each user terminal 20. This allows each user to check the supplementary information regarding any inconsistencies that occurred in the conversation. Alternatively, the output control unit 109 may output the supplementary information 1143A and 1143B to each user terminal 20, and each user terminal 20 may synthesize the supplementary information 1143A and 1143B into speech and output it as speech.

[0178] Furthermore, the user's supporter and any supplementary comments made by the supporter may be displayed only to the user being supported. In this case, the screen generation unit 107 generates a conversation screen for user A and a conversation screen for user B, and the output control unit 109 outputs the conversation screen for user A to user terminal 20-1 and the conversation screen for user B to user terminal 20-2. The conversation screen for user A is a conversation screen in which user B's supporter 1191B and supporter 1191B's speech bubble 1145B are not displayed, but user A's supporter 1191A and supporter 1191A's speech bubble 1145A are displayed. The conversation screen for user B is a conversation screen in which user A's supporter 1191A and supporter 1191A's speech bubble 1145A are not displayed, but user B's supporter 1191B and supporter 1191B's speech bubble 1145B are displayed.

[0179] Next, we will explain the differences between the conversation screen when the communication app is a chat app and the modified example 1.

[0180] Figure 35 shows an example of the conversation screen 1201 of Modification 3. The conversation screen 1201 shown in Figure 35 is an example of a conversation screen (chat screen) when the communication application is a chat application.

[0181] If the communication app is a chat app, at the time the conversation screen 441 shown in Figure 17 of the above explanation is displayed on the user terminal 20, a text message 641 indicating the content of the statement 507 is displayed on the conversation screen 1201 shown in Figure 35. Also, at the time the conversation screen 1151 shown in Figure 34 of the above explanation is displayed on the user terminal 20, a text message 1251B indicating supplementary content 1143B and a text message 1251A indicating supplementary content 1143A are displayed on the conversation screen 1201 shown in Figure 35.

[0182] In Modification 3, the system detects the consistency of each user's statement in the conversation with the content up to that point. If an inconsistency is detected, the system provides supplementary information to address the inconsistency from the perspective of a supporter of the user. As a result, users participating in the conversation can grasp the correction of any inconsistencies in the conversation in real time. According to Modification 3, the flow of the conversation can be automatically corrected based on the supplementary information provided by the supporter of the user, allowing each user participating in the meeting to continue the conversation more naturally.

[0183] (Modification 4) In the above embodiments and their variations, the conversation inconsistency detection system was described using the example of a conversation using a communication application. However, it may also be used for conversations that users have in person. In this case, the functions of the user terminal 20 can be integrated into the conversation inconsistency detection device 10, and the conversation inconsistency detection device 10 can be implemented using a smartphone, tablet terminal, or PC.

[0184] (Variation 5) In the above embodiments and their variations, the example of implementing the generation model 30 as a cloud service was used for explanation, but the generation model 30 may also be included in the conversation inconsistency detection device 10. In this case, the generation model 30 can be implemented, for example, as an AI chat service that fine-tunes a small language model (SLM).

[0185] (Experimental variation 6) In the above embodiments and their variations, an example was described in which, as a preprocessing step before the conversation begins, a first prompt is input to the generation model 30, and during the conversation, the content of each user's statement is input to the generation model 30 as a second prompt. However, the content of the first prompt, which is a prompt corresponding to inconsistency definition information or a prompt instructing the detection of consistency, may be incorporated into the second prompt to form a single prompt, and this prompt may be input to the generation model 30 each time the user makes a statement to detect the consistency of the conversation.

[0186] (Example 7) In the above embodiments and their variations, the case where speech recognition is performed on the conversation inconsistency detection device 10 side was described as an example. However, the speech recognition unit 106 may be included on the user terminal 20 side, speech recognition may be performed on the user terminal 20 side, and the conversation information generated by the speech recognition may be notified to the conversation inconsistency detection device 10.

[0187] (program) The programs executed by the conversation inconsistency detection devices of the above embodiments and each of the above modifications are provided as installable or executable files stored on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD).

[0188] Furthermore, the programs executed by the conversation inconsistency detection devices of the above embodiments and their respective modifications may be stored on a computer connected to a network such as the Internet and provided by allowing downloads via the network. Alternatively, the programs executed by the conversation inconsistency detection devices of the above embodiments and their respective modifications may be provided or distributed via a network such as the Internet. Furthermore, the programs executed by the conversation inconsistency detection devices of the above embodiments and their respective modifications may be pre-installed in ROM or the like and provided.

[0189] The programs executed in the conversation inconsistency detection devices of the above embodiments and their respective modifications are configured as modules for implementing the above-described parts on a computer. In actual hardware, for example, the CPU reads the program from the HDD into RAM and executes it, thereby enabling the implementation of the above-described parts on the computer.

[0190] The above embodiments and their respective modifications are merely examples of how this disclosure may be implemented, and they do not restrict the technical scope of this disclosure. Therefore, this disclosure can be implemented in various ways without departing from its essence or its main features. For example, the above embodiments and their respective modifications may be combined as appropriate on a component basis. Also, for example, some components may be removed from the total components in the above embodiments and their respective modifications.

[0191] This disclosure includes the following aspects:

[0192] (1) A consistency detection control unit that sequentially inputs the content of a conversation between multiple users into a generative model and obtains a detection result from the generative model that detects the consistency of the conversation with the content input up to that point, If the detection result indicates an inconsistency, the notification control unit performs notification control to indicate that an inconsistency has occurred. A conversation inconsistency detection device equipped with the following features.

[0193] (2) If the consistency detection control unit indicates an inconsistency in the detection result, it further obtains supplementary content from the generation model to supplement the inconsistency that occurred in the conversation, The notification control unit performs notification control of the supplementary content. The conversation inconsistency detection device described in (1) above.

[0194] (3) The system further comprises a prompt creation unit that, upon inputting the content of the conversation, creates a first prompt that instructs the generation model to detect the consistency of the conversation based on inconsistency definition information that defines patterns in which inconsistencies are considered to have occurred in the conversation, and to generate the supplementary content if an inconsistency is detected, The consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device described in (2) above.

[0195] (4) The system further comprises a prompt generation unit that, upon inputting the content of the conversation, generates a first prompt that instructs the generation model to detect the consistency of the conversation based on inconsistency definition information that defines patterns in which inconsistencies are considered to have occurred in the conversation, The consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device described in (1) above.

[0196] (5) The system further includes an acquisition unit that acquires the content of each of the multiple users who speak in the conversation, The prompt generation unit generates a second prompt indicating the content of a user's statement each time the content of that statement is obtained. Each time the second prompt is generated, the consistency detection control unit inputs the second prompt to the generation model and obtains the detection result. A conversation inconsistency detection device as described in (3) or (4) above.

[0197] (6) The first prompt instructs the generative model to detect the coherence of the conversation each time the content of the conversation is input, Prior to the start of the conversation, the consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device described in (5) above.

[0198] (7) The acquisition unit acquires audio data of the spoken content as the spoken content, The system further includes a speech recognition unit that performs speech recognition on the acquired speech data, The second prompt indicates the speech recognition result of the spoken content. The conversation inconsistency detection device described in (5) above.

[0199] (8) The supplementary information is provided from the perspective of the user who made the statement causing the inconsistency in the conversation, the user's supporter, or the mediator of the conversation, and is intended to address the inconsistency. A conversation inconsistency detection device as described in (2) or (3) above.

[0200] (9) The notification control unit controls the output of at least one of the image and sound of the person in the position speaking the supplementary content. The conversation inconsistency detection device described in (8) above.

[0201] (10) A consistency detection control step in which the consistency detection control unit sequentially inputs the content of a conversation between multiple users into a generation model and obtains a detection result from the generation model that detects the consistency of the conversation with the content input up to that point, If the notification control unit indicates an inconsistency in the detection result, the notification control step includes performing a notification control to indicate that an inconsistency has occurred, A method for detecting conversational inconsistencies, including the following.

[0202] (11) A consistency detection control step in which the content of a conversation between multiple users is sequentially input into a generative model, and a detection result is obtained from the generative model that detects the consistency of the conversation with the content that has been input up to that point, If the detection result indicates an inconsistency, a notification control step is performed to notify that an inconsistency has occurred. A program that causes a computer to execute something. [Explanation of symbols]

[0203] 1. Conversation Inconsistency Detection System 10 Conversation Inconsistency Detection Device 20, 20-1, 20-2 User terminals 30 Generative Models 101 Prompt Creation Section 103 Consistency Detection Control Unit 105 Acquisition Department 106 Voice Recognition Unit 107 Screen generation section 109 Output Control Unit Rooms 201, 201-1, 201-2: Reception area 205, 205-1, 205-2 Notification section 207, 207-1, 207-2 Output Section

Claims

1. A consistency detection control unit sequentially inputs the content of conversations between multiple users into a generative model and obtains a detection result from the generative model that detects the consistency of the conversation with the content input up to that point. If the detection result indicates an inconsistency, the notification control unit performs notification control to indicate that an inconsistency has occurred. A conversation inconsistency detection device equipped with the following features.

2. If the consistency detection control unit indicates an inconsistency in the detection result, it further obtains supplementary content from the generation model to supplement the inconsistency that occurred in the conversation. The notification control unit performs notification control of the supplementary content. The conversation inconsistency detection device according to claim 1.

3. The system further includes a prompt creation unit that, upon inputting the content of the aforementioned conversation, creates a first prompt instructing the generation model to detect the consistency of the conversation based on inconsistency definition information that defines patterns in which inconsistencies are considered to have occurred in the conversation, and to generate the supplementary content if an inconsistency is detected. The consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device according to claim 2.

4. The system further includes a prompt generation unit that, upon inputting the content of the aforementioned conversation, creates a first prompt instructing the generation model to detect the consistency of the conversation based on inconsistency definition information that defines patterns in which inconsistencies are considered to have occurred in the conversation. The consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device according to claim 1.

5. The system further includes an acquisition unit that acquires the content of each of the multiple users who speak in the aforementioned conversation. The prompt generation unit generates a second prompt indicating the content of a user's statement each time the content of that statement is acquired. Each time the second prompt is generated, the consistency detection control unit inputs the second prompt to the generation model and obtains the detection result. The conversation inconsistency detection device according to claim 3 or 4.

6. The first prompt instructs the generative model to detect the coherence of the conversation each time the content of the conversation is input. Prior to the start of the conversation, the consistency detection control unit inputs the first prompt to the generation model. The conversation inconsistency detection device according to claim 5.

7. The acquisition unit acquires audio data of the spoken content as the spoken content, The system further includes a speech recognition unit that performs speech recognition on the acquired speech data, The second prompt indicates the speech recognition result of the spoken content. The conversation inconsistency detection device according to claim 5.

8. The aforementioned supplementary information is provided from the perspective of the user who made the statement causing the inconsistency in the aforementioned conversation, the user's supporter, or the mediator of the conversation, and aims to clarify the inconsistency. The conversation inconsistency detection device according to claim 2 or 3.

9. The notification control unit controls the output of at least one of the image and audio of the person in the position speaking the supplementary content. The conversation inconsistency detection device according to claim 8.

10. A consistency detection control step involves a consistency detection control unit sequentially inputting the content of a conversation between multiple users into a generation model, and obtaining a detection result from the generation model that detects the consistency of the conversation with the content input up to that point. If the notification control unit indicates an inconsistency in the detection result, the notification control step includes performing a notification control to indicate that an inconsistency has occurred, A method for detecting conversational inconsistencies, including the following.

11. A consistency detection control step involves sequentially inputting the content of a conversation between multiple users into a generative model, and obtaining a detection result from the generative model that detects the consistency of the conversation with the content input up to that point. If the detection result indicates an inconsistency, a notification control step is performed to notify that an inconsistency has occurred. A program that causes a computer to execute something.

Citation Information

Patent Citations

  • Information processing device, and information processing program

    JP2022094027A