Response language expression output device and method
The system classifies user dialogue interaction states using non-semantic features to generate contextually appropriate responses, addressing the limitations of conventional AI by enhancing user engagement through empathetic and contextually aware interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional conversational AI systems struggle to accurately classify the conversational interaction state of users, leading to inappropriate or over-responsive outputs, particularly when users seek a 'presence' rather than direct responses, and fail to recognize non-verbal cues.
A system that classifies user dialogue interaction states based on non-semantic features such as intonation, word endings, and speaking intervals, and generates responses that match these states, including the ability to withhold responses when appropriate.
Enables empathetic and contextually appropriate responses by classifying user interaction states, allowing the system to resonate with the user's dialogue state, even in silence, thus enhancing user engagement.
Smart Images

Figure 0007843093000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to outputting a response language expression that responds to a language expression given by a user (including language expressions such as words spoken by the user, language expressions by voice input from a microphone like a sentence, language expressions input by the user using an input device such as a keyboard, words and sentences, language expressions obtained by reading strings written on paper, etc.).
[0002] In particular, this invention takes the user's speech in natural language as input, classifies the user's dialogue interaction state as an index band based on non-semantic feature quantities (including non-semantic feature quantities that can be understood from external features recognized from the language expression given by the user) such as intonation, word endings, and speaking intervals, and generates a non-semantic natural language response according to the result. Terms such as "dialogue interaction state", "fatigue", and "introspection" in this specification are all treated as pattern classifications of non-semantic feature quantities (external feature quantities) such as word endings, intonation, word length, speaking intervals, and speaking frequencies included in the user's speech, and do not estimate and analyze the internal states of an individual's emotions, psychology, spirit, etc. This invention does not aim to estimate and judge the user's inner self, and is a technology for performing dialogue control such as response timing and response intensity based on non-semantic feature quantities (external feature quantities). Therefore, this invention does not fall under mental state inference.
Background Art
[0003] In recent years, dialogue-type AI (Artificial Intelligence) using large language models (LLMs :Large Language Models) such as ChatGPT (Generative Pre-trained Transformer) has become popular and enables highly accurate natural language responses in information provision, casual conversation, advice, etc. Also there are attempts to give machines emotions through emotion AI. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2025-47448 [Overview of the project] [Problems that the invention aims to solve]
[0005] Conventional conversational AI primarily focuses on semantic processing and informational responses, making it difficult to classify the conversational interaction state behind user utterances (e.g., low activity / active system) and construct "responses that respond to the atmosphere," such as not speaking, not responding, or responding only based on intuition. Furthermore, in situations where the user wants to be understood without speaking, or is seeking a "presence" rather than a direct response, conventional AI tended to over-respond or give incorrect responses. In addition, the invention described in Patent Document 1, as described in Claim 1 of Patent Document 1, involves sharing the joy of travel with a partner.
[0006] This invention aims to enable responses that correspond to the user's dialogue interaction state (such as the outwardly observable state when the user interacts with the response language expression output device, for example, the user's appearance, attitude, behavior, and activity state). In particular, it aims to enable the output of the response language expression to be stopped when it is determined, based on the user's dialogue interaction state, that it is better not to respond to the user. [Means for solving the problem]
[0007] The response language expression output device according to this invention includes a language expression input means for inputting language expressions provided by the user, and a language expression portion of the language expression input from the language expression input means, excluding meaningful content words (the language expression portion excluding meaningful content words includes not only the words themselves, such as the endings of the language expression portion, but also the length of the language expression, the feeling derived from the language expression, such as rhythm, presence or absence of long vowels, etc., the interval between utterances, the frequency of utterances, the intonation of the voice, and, if the language expression is input from a keyboard, the speed of the language input from the keyboard, the interval of input time, etc.) which responds to the user. The system includes a first classification means for classifying the state of a conversational interaction; a modification means for modifying a response language expression (for example, at least one of the ending, length, nuance, interval, and frequency of the response language expression) which is a response to a language expression input from the language expression input means, to match the user's conversational interaction state inferred by the first classification means; an output means for outputting the response language expression modified by the modification means; a first response determination means for determining whether or not to respond based on the conversational interaction state classified by the first classification means; and a first output control means for controlling the output means to stop outputting the response language expression in response to the determination by the first response determination means that no response is expected.
[0008] This invention also provides a method for outputting a response language expression. Specifically, in this method, a language expression input means inputs a language expression provided by a user, a first classification means classifies the user's dialogue interaction state based on the language expression portion of the language expression input from the language expression input means, excluding meaningful content words, a modification means modifies the response language expression, which is a response to the language expression input from the language expression input means, to match the user's dialogue interaction state classified by the first classification means, an output means outputs the response language expression modified by the modification means, a first response determination means determines whether or not to respond based on the dialogue interaction state classified by the first classification means, and a first output control means controls the output means to stop outputting the response language expression in accordance with the determination by the first response determination means that no response is expected.
[0009] Furthermore, this invention also provides a program for controlling the computer of a response language expression output device, and a recording medium storing that program.
[0010] The above modification means may also involve modifying the response language expression portion of the response language expression, which is a response to the language expression input from the above language expression input means, excluding the meaningful content words, so as to match the user's dialogue interaction state classified by the first classification means.
[0011] Preferably, the system includes a first response determination means for determining whether or not to respond based on the dialogue interaction state classified by the first classification means, and a first output control means for controlling the output means to stop outputting the response language expression in response to the determination by the first response determination means that no response is expected.
[0012] The first output control means may, for example, output the modified response language expression by the modification means after a predetermined time has elapsed since the first response determination means determined that there is no response.
[0013] The above modification means includes, for example, a word ending dictionary storing candidates for word endings or word nuances, a word nuance adjustment unit that modifies the selected word ending by vowel extension or intonation change, and an interval control unit that adjusts the timing of the response output. In this case, it is preferable to include a word nuance response generation unit that combines the word ending dictionary, the word nuance adjustment unit, and the interval control unit to generate a word nuance response according to the user's dialogue interaction state.
[0014] The system may further include a recontrol means for controlling the response language expression output device to perform classification by the first classification means and modification by the modification means on the new language expression input by the user from the language expression input means after the first response determination means has determined that there is no response, and to output a response language expression, which is a response to the new language expression, from the output means.
[0015] The system may further include a second response determination means for determining whether to respond with silence or not based on the dialogue interaction state classified by the first classification means, and a second device control means for controlling the response language expression output device to output information representing a silent state in response to the determination by the second response determination means to respond with silence, and to output the response language expression modified by the modification means in response to the determination by the second response determination means not to respond with silence.
[0016] The second device control means, for example, outputs information indicating a silent state when the second response determination means determines that a silent response is to be given, or when the second response determination means determines that a silent response is not to be given, after a longer period of time has elapsed than the time required from the determination result of the second response determination means to output the response language expression modified by the modification means.
[0017] Furthermore, the system may include a third response determination means that determines whether to make a response that does not contain meaningful content words, based on the dialogue interaction state classified by the first classification means. In this case, it is preferable that the modification means modifies the response language expression, which is a response to the language expression input from the language expression input means, so as not to contain meaningful content words, in response to the determination by the third response determination means that the response does not contain meaningful content words.
[0018] Furthermore, the system may include a second classification means that, based on the user's dialogue interaction state classified by the first classification means, classifies which of several groups, each classified according to similar dialogue interaction states, the user's dialogue interaction state falls into. In this case, the modification means, for example, modifies the response language expression to the language expression input from the language expression input means so that it matches the dialogue interaction state represented by the group classified by the second classification means.
[0019] Furthermore, the system may include a storage control means that controls the storage device to store, for each user, the relationship between the linguistic expression portion, excluding meaningful content words, from the linguistic expression input means and the user's dialogue interaction state, and an update means that updates the relationship stored in the storage device for each user. In this case, the classification means may classify the user's dialogue interaction state based on the relationship updated by the update means, using the linguistic expression portion, excluding meaningful content words, from the linguistic expression input means.
[0020] Furthermore, the empathetic dialogue support system according to this invention is characterized by classifying the user's dialogue interaction state into index bands based on non-semantic linguistic features such as word endings, intonation, inter-word spacing, vocabulary selection, and utterance frequency contained in the user's natural language utterances, and controlling the presence or absence of natural language responses, word endings, nuances, intonation, and response timing according to the index bands.
[0021] Also, when the classified index band corresponds to a non-recommended response range, it may be constitutively selected not to generate a response and skip the dialogue output process. A learning module may be provided that can dynamically adjust the parameters of state index classification based on the end-word and intonation tendencies of each user.
[0022] A response generation module using non-semantic language elements such as end-words, language senses, and between-words may be provided to construct a dialogue experience without a semantic response.
Advantages of the Invention
[0023] According to this invention, as a response to the user's language expression, a response language expression that matches the user's dialogue interaction state can be output. Since a response can be made according to the user's dialogue interaction state, the user can empathize with the responded language expression. Also, although what is described in Patent Document 1 is to share the joy of movement with a partner as described in Patent Document 1, only the recognition of the user's emotion is described in paragraph
[0032] of Patent Document 1, and there is no specific description on how to recognize it. In contrast, the present invention classifies the user's dialogue interaction state based on the language expression part excluding the content words with meaning in the language expression from the user, so the user dialogue interaction state can be classified relatively easily.
[0024] [[ID=I4]] For example, even if the user's utterance has no meaning, a reaction policy can be derived from the state index band. The AI can autonomously select the judgment of not making a response. Responses that are only based on resonance due to language sense such as "I see" and "Yes" are possible. For each user, the end-word, interval, and way of replying can be individually optimized. A structure can be provided that classifies the "relationship temperature" from elements such as end-words, language senses, between-words, and speaking timing without depending on the linguistic meaning, and performs response generation or response omission according to the index band. In particular, the present invention has a fundamental difference from conventional dialogue-type AIs in that "not replying" itself is constitutively implemented as a technical judgment.
Brief Description of the Drawings
[0025] [Figure 1] This is an overview of the response language representation output system. [Figure 2] This is a block diagram showing the electrical configuration of a response language expression output device. [Figure 3] This shows the pre-trained models for dialogue interaction state classification and the pre-trained models for dialogue interaction state-specific responses stored on the SSD. [Figure 4] This shows the pre-trained models included in the group of pre-trained models for dialogue interaction state classification. [Figure 5] This shows the trained models included in the group of trained models for state-specific responses in dialogue interaction. [Figure 6] This is an example of a dialogue interaction state classification table. [Figure 7] This shows the relationship between the metrics representing the user's dialogue interaction state and the elements that influence it. [Figure 8] This flowchart shows the processing procedure of the response language representation output system. [Figure 9] This flowchart shows the processing procedure of the response language representation output system. [Figure 10] This is a flowchart showing the procedure for classifying the state of dialogue interaction. [Figure 11] This shows the normal response region R1. [Figure 12] This shows a low response region R21. [Figure 13] This shows the low response region R22. [Figure 14] This shows the silent response region R3. [Figure 15] This shows regions R1, R21, R22, and R3. [Figure 16] This flowchart shows the procedure for modifying the response language expression. [Figure 17] This is an example of a user's linguistic expression and the response's linguistic expression. [Figure 18] This is an example of a user's linguistic expression and the response's linguistic expression. [Figure 19]This is an example of a user's linguistic expression and the response's linguistic expression. [Figure 20] This is an example of a history information table. [Figure 21] This is a flowchart illustrating the dialogue interaction state classification process. [Figure 22] This is a dialogue interaction state classification table for individual users. [Figure 23] This is a flowchart illustrating the dialogue interaction state classification process. [Figure 24] This flowchart shows part of the processing procedure for the response language representation output system. [Figure 25] This flowchart shows part of the processing procedure for the response language representation output system. [Figure 26] This flowchart shows part of the processing procedure for the response language representation output system. [Figure 27] This is an overall system diagram (speech acquisition → indicator classification → response control → output). [Figure 28] This is an indicator zone classification and response strategy matrix. [Figure 29] This is a flowchart for classifying and adjusting data based on each user's history. [Figure 30A] This shows an example of generating and processing responses for word endings and nuances. [Figure 30B] This shows an example of generating and processing responses for word endings and nuances. [Figure 31] This demonstrates the characteristics of a dialogue-driven, state-adaptive personality switching structure. [Modes for carrying out the invention]
[0026] Figure 1 shows an embodiment of this invention and is an overview of the response language expression output system.
[0027] The response language expression output system comprises a response language expression output device 1 and an AI (Artificial Intelligence) server 20, which can communicate with each other via the internet.
[0028] In this embodiment, the response language expression output device 1 and the AI server 20 communicate to output the response language expression from the response language expression output device 1. However, by performing the classification to obtain the language expression of the normal response in the AI server 20 in the response language expression output device 1, the AI server 20 is not necessarily required.
[0029] Figure 2 is a block diagram showing the electrical configuration of the response language expression output device 1.
[0030] The overall operation of the response language expression output device 1 is controlled by the CPU 2.
[0031] The response language expression output device 1 includes memory 5 for temporarily storing data. The response language expression output device 1 also includes a GPU (Graphics Processing Unit) 3 (an example of a first classification means and modification means), and the GPU 3 performs AI learning, generation of response language expressions using the trained model, and classification of the dialogue interaction state. A display device 4 (an example of an output means) that displays the generated response language expressions, etc., is connected to the GPU 3 via an interface (not shown).
[0032] The response language expression output device 1 includes a PCH (Platform Controller Hub) 6, which controls data communication between the CPU 2 and the communication device 7 that communicates with the AI server 20, a microphone 8 (an example of a language expression input means), a speaker 9 (an example of an output means), a keyboard 10 (an example of a language expression input means), an SSD (Solid State Drive) 11, a CD (Compact Disc) drive 12, etc. The SSD stores various trained models and other data. A CD 13 containing a program that controls the operation of the response language expression output device 1 is inserted into the CD drive 12, and the program is read from the CD 13 and installed in the response language expression output device 1. Alternatively, the program may be received via the internet and installed in the response language expression output device 1.
[0033] Furthermore, the system may be equipped with a scanner (an example of a language expression input means) for reading language expressions written on paper or other materials, so that the language expression is read by the scanner and the response language expression output device 1 recognizes the language expression by OCR (Optical Character Recognition).
[0034] Figure 3 shows the group of trained models 30 for dialogue interaction state classification and the group of trained models 40 for dialogue interaction state-specific responses stored in SSD11.
[0035] The group of trained models 30 for classifying dialogue interaction states are used to classify the user's dialogue interaction state from the language expressions input by the user to the response language expression output device 1 (language expressions made from voice input from the microphone 8, language expressions represented by data input from the keyboard 10, etc.).
[0036] The group of trained models 40 for dialogue interaction state-specific responses modifies response language expressions to language expressions from the user (language expressions that represent responses to language expressions from the user as if in a dialogue with the user) to match the classified dialogue interaction state of the user. For example, at least one of the following elements of the response language expression—ending, length, nuance, interval, and frequency—is modified to match the user's dialogue interaction state.
[0037] Figure 4 shows the group of 30 pre-trained models for classifying dialogue interaction states.
[0038] The group of pre-trained models 30 for classifying dialogue interaction states includes a pre-trained model for classifying dialogue interaction states (for word endings) 31, a pre-trained model for classifying emotional states (for word length) 32, a pre-trained model for classifying dialogue interaction states (for word nuance) 33, a pre-trained model for classifying emotional states (for utterance interval) 34, and a pre-trained model for classifying dialogue interaction states (for utterance frequency) 35. The pre-trained model for classifying dialogue interaction states (for word endings) 31 was obtained by inputting a large number of word endings (or sentences with word endings, etc.) as training data (teaching data), and pre-training which of the following dialogue interaction states—active, joyful, extroverted, calm, normal, casual conversation, silence, fatigue, and introspection—corresponds to those word endings. As shown in Figure 6 later, the correspondence between certain word endings and dialogue interaction states is generally fixed. Therefore, by using the pre-trained model for dialogue interaction state classification (for word endings) 31, it is possible to classify the dialogue interaction state of the user who provided the input language expression based on the word endings of that language expression. In this embodiment, the correspondence between non-semantic features such as word endings and intonation and the user's dialogue interaction state is obtained by collecting a large amount of dialogue data in advance and statistically analyzing it together with labels assigned based on the subject's self-reporting and expert annotations. This allows for a quantitative understanding of, for example, which dialogue interaction states are associated with specific word endings or utterance intervals. By training a machine learning model with this relationship as training data, it becomes possible to classify the user's dialogue interaction state even for unknown inputs. Similarly, for the other models 32-35, by training each judgment element in advance using training data, it is possible to classify the dialogue interaction state of the user who provided the input language expression based on the judgment elements of that language expression. The pre-trained model for general state classification (for word length) 32 classifies the state based on the user's external characteristics, such as the number of characters, syllables, and sounds in the linguistic expression input by the user. The pre-trained model for dialogue interaction state classification (for linguistic nuance) 33 classifies the user's dialogue interaction state using the linguistic nuance of the linguistic expression input by the user.The pre-trained model for dialogue interaction state classification (for utterance intervals) 34 classifies the user's emotional state using the utterance intervals of the linguistic expressions input by the user. The pre-trained model for dialogue interaction state classification (for utterance frequency) 35 classifies the user's dialogue interaction state using the utterance frequency of the linguistic expressions input by the user (these pre-trained models 31-35 are used to classify the user's dialogue interaction state based on the linguistic portion of the linguistic expression, excluding meaningful content words).
[0039] The user's dialogue interaction state is classified from all the classification results of the pre-trained models for dialogue interaction state classification (for word endings) 31, dialogue interaction state classification (for word length) 32, dialogue interaction state classification (for word nuance) 33, dialogue interaction state classification (for utterance interval) 34, and dialogue interaction state classification (for utterance frequency) 35. From the classification results of all five pre-trained models, such as the pre-trained model for dialogue interaction state classification (for word endings) 31, the pre-trained model for dialogue interaction state classification (for word length) 32, the pre-trained model for dialogue interaction state classification (for word nuance) 33, the pre-trained model for dialogue interaction state classification (for utterance interval) 34, and the pre-trained model for dialogue interaction state classification (for utterance frequency) 35, the user's dialogue interaction state is not classified, and the pre-trained model for dialogue interaction state classification (for word endings) 31 Alternatively, a single pre-trained model for classifying dialogue interaction states may be created by combining all or more of the following pre-trained models: 32 (for word length), 33 (for nuance), 34 (for utterance interval), and 35 (for utterance frequency). The classification result of this single pre-trained model may then be used as the user's dialogue interaction state. In this case, a large number of linguistic expressions are input as training data, and these numerous linguistic expressions are pre-trained for each judgment element or the linguistic expressions themselves, thereby generating a pre-trained model that shows which linguistic expressions lead to which dialogue interaction states. Furthermore, each user may have their own set of pre-trained models for classifying dialogue interaction states, 30 of which are unique to that user.
[0040] Furthermore, the group of trained models 30 for dialogue interaction state classification may also include a history-trained model 36, as will be described later.
[0041] Figure 5 shows the group of 40 trained models for responding to different dialogue interaction states.
[0042] The group of pre-trained models 40 for responding to different dialogue interaction states includes a pre-trained model 41 for active states, a pre-trained model 42 for joyful states, a pre-trained model 43 for extroverted states, a pre-trained model 44 for calm states, a pre-trained model 45 for normal states, a pre-trained model 46 for casual conversation, a pre-trained model 47 for silent states, a pre-trained model 48 for fatigued states (a "fatigue pattern" application model based on external features), and a pre-trained model 49 for introspective states. The pre-trained model 41 for active states is a pre-trained model used when the user's dialogue interaction state is active, and its response language expression is modified to produce an active response language expression. The pre-trained model 41 for active states has also been pre-trained using a large amount of training data to determine what kind of expression will produce an active response language expression. As shown in Figure 7 later, the expressions used to output active response language expressions are generally determined. By classifying using the pre-trained model 41 for active states, the response language expression can be changed (for example, by using fine-tuning) after being transmitted from the AI server 20 and input to the response language expression output device 1. This allows for outputting a response language expression suitable for when the user's dialogue interaction state is active. For example, it has been revealed through statistical analysis of a large amount of dialogue data, including self-reported data from subjects and expert annotations, that specific endings and tones correspond to specific dialogue interaction states. Such relationships are collected as training data, and the pre-trained model can perform similar classifications even for unknown inputs. Similarly, for the other models 42-49, by training them in advance using training data for each dialogue interaction state, the input response language expression can be classified and output as a response language expression suitable for the dialogue interaction state. The pre-trained model 42 for joy is a pre-trained model used when the user's dialogue interaction state is joyful, and the response language expression is changed to express joy. The outward-facing pre-trained model 43 is a pre-trained model used when the user's dialogue interaction state is outward-facing, and the response language expression is modified to become an outward-facing response language expression.The pre-trained model 44 for calm states is a pre-trained model used when the user's dialogue interaction state is calm, and the response language expression is modified to reflect a calm response language expression. The pre-trained model 45 for normal states is a pre-trained model used when the user's dialogue interaction state is normal, and the response language expression is modified to reflect a normal response language expression. The pre-trained model 46 for casual conversation is a pre-trained model used when the user's dialogue interaction state is casual conversation, and the response language expression is modified to reflect a casual conversation response language expression. The pre-trained model 47 for silence states is a pre-trained model used when the user's dialogue interaction state is silent, and the response language expression is modified to reflect a silent response language expression. The pre-trained model 48 for fatigue states is a pre-trained model used when the user's dialogue interaction state is fatigued, and the response language expression is modified to reflect a fatigued response language expression. The pre-trained model 49 for introspection is a trained model used when the user's dialogue interaction state is in an introspective state, and the response language expression is modified to reflect the response language expression of the introspective state.
[0043] Depending on the user's dialogue interaction state, one of the following trained models is used to modify the response language expression: trained model 41 for active state, trained model 42 for joyful state, trained model 43 for extroverted state, trained model 44 for calm state, trained model 45 for normal state, trained model 46 for casual conversation, trained model 47 for silent state, trained model 48 for fatigued state (a "fatigue pattern" application model based on external features), or trained model 49 for introspective state. In Figure 5, the group of trained models 40 for dialogue interaction state-specific responses is composed of separate trained models 41 to 49. However, a single trained model for dialogue interaction state-specific responses may be generated from all or more of these trained models 41 to 49, and the response language expression may be modified to match the user's dialogue interaction state. In this case, a large number of language expressions representing various types of dialogue interaction states are pre-trained as training data, and a pre-trained model is generated that can be modified in a way that changes the language expression to a response language expression corresponding to the dialogue interaction state that can be determined from the input language expression.
[0044] Figure 6 is an example of a table showing the relationship between the decision elements contained in the linguistic expressions provided by the user and the user's dialogue interaction state.
[0045] In this embodiment, the user's dialogue interaction state is classified based on the linguistic expression portion, which includes function words (conjunctions, articles, auxiliary verbs, prepositions, pronouns, etc., which mainly perform grammatical functions) excluding content words (nouns, verbs, adjectives, adverbs, etc., which convey substantial meaning in a sentence and play a major role in information transmission), non-verbal elements such as tone, speed, intonation, rhythm, pauses, volume, and quality of voice, as well as paralinguistic information such as emojis and symbols in written communication, and word length, nuance, interval between utterances, and frequency of utterances.
[0046] In this embodiment, judgment elements such as "ending," "length," "sound," "interval," and "frequency" are used as examples of linguistic expression parts excluding content words. However, other judgment elements are used to classify the user's dialogue interaction state. In addition to these, other judgment elements may be used to classify the user's dialogue interaction state, or all of these judgment elements may be used, or some of these judgment elements may be used to classify the user's dialogue interaction state. Furthermore, in this embodiment, the types of user dialogue interaction states to be classified are "active," "joyful," "extroverted," "calm," "normal," "small talk," "silent," "tired," and "introspective." However, other dialogue interaction states may be classified, or some of these dialogue interaction states may be classified. Moreover, in this embodiment, similar types of dialogue interaction states are grouped into the same group. For example, the dialogue interaction states "active," "joyful," and "extroverted" are considered similar and are classified as a high-index dialogue interaction state group; the dialogue interaction states "calm," "normal," and "casual conversation" are considered similar and are classified as a medium-index dialogue interaction state group; and the dialogue interaction states "silent," "tired," and "introspective" are considered similar and are classified as a low-index dialogue interaction state group.
[0047] For example, if a user's linguistic expressions end with "~da ze," have short syllable lengths, have a rhythmic feel, have an interval of about 0.1 to 0.2 seconds between utterances, and have an utterance frequency of 20 to 30 turns per minute, the user's dialogue interaction state is more likely to be classified as "active." The same applies to other dialogue interaction states.
[0048] The table shown in Figure 6 represents the case where the user provides linguistic expressions to the response linguistic expression output device 1 via voice from the microphone 8. However, even when the user provides linguistic expressions to the response linguistic expression output device 1 as text data using the keyboard 10, the user's external dialogue interaction state can be similarly classified according to the feel and speed of typing on the keyboard 10's keypad, the symbols entered (such as smiling faces and crying faces), etc.
[0049] In this embodiment, a table as shown in Figure 6 is stored in the SSD11. When a linguistic expression from the user contains a suffix that is classified as one of the following dialogue interaction states stored in this table: "active," "joyful," "extroverted," "calm," "normal," "small talk," "silent," "tired," or "introspective," the user's dialogue interaction state is pre-trained using a large amount of training data to classify it as "active," "joyful," "extroverted," "calm," "normal," "small talk," "silent," "tired," or "introspective," and this pre-trained model for dialogue interaction state classification (for suffixes) 31 is stored in the SSD11. Similarly, other judgment elements such as word length, nuance, interval between utterances, or frequency of utterances are also stored in the SSD11 as pre-trained dialogue interaction state classification models (for word length) 32, dialogue interaction state classification model (for nuance) 33, dialogue interaction state classification model (for utterance interval) 34, and dialogue interaction state classification model (for utterance frequency) 35, respectively, which are pre-trained to classify the user's state as "active," "joyful," "extroverted," "calm," "normal," "small talk," "silent," "tired," or "introspective," depending on the judgment element.
[0050] Even if the table itself shown in Figure 6 is not stored in SSD11, if the pre-trained models for dialogue interaction state classification (for word endings) 31, dialogue interaction state classification (for word length) 32, dialogue interaction state classification (for nuance) 33, dialogue interaction state classification (for utterance interval) 34, and dialogue interaction state classification (for utterance frequency) 35 are stored in SSD11 or elsewhere, the state can be classified based on the user's external characteristics based on the linguistic expressions from the user. Furthermore, the aforementioned pre-trained models for dialogue interaction state classification 31-35 can also be generated by pre-training using word endings, word lengths, nuances, and the sentences themselves from dictionaries such as emotion expression dictionaries.
[0051] Furthermore, once the user's dialogue interaction state is classified, the response language expression—the response to the user's language expression—can be changed to an expression that matches the classified user dialogue interaction state, using a table like the one shown in Figure 6. For example, if the user's dialogue interaction state is classified as "calm," the response language expression can be changed to a polite expression such as, "Yes, that's right."
[0052] When a user's dialogue interaction state is classified as "active," "joyful," "extroverted," "calm," "normal," "chatty," "silent," "tired," or "introspective," the response language expression is modified to match the user's dialogue interaction state, or the part of the response language expression excluding content words is modified, and then provided to the user, using the pre-trained models 41 (active), 42 (joyful), 43 (extroverted), 44 (calm), 45 (normal), 46 (chatty), 47 (silent), 48 (tired), or 49 (introspective). Because the response matches the user's dialogue interaction state, it is possible to achieve dialogue that is attentive to the user's dialogue interaction state. Instead of the table shown in Figure 6, response language expressions corresponding to the user's dialogue interaction state can be generated and output using these pre-trained models 41-49. In this case as well, by using sentences found in dictionaries such as emotion expression dictionaries, the pre-trained models 41-45 described above can be generated by pre-training the model to produce response language expressions that match the user's dialogue interaction state.
[0053] Figure 7 is an example of a table showing the relationship between the index range and the endings, length, nuance, interval, and frequency of responses in the language expression, when the user's dialogue interaction state is classified as being in the high index range, medium index range, or low index range.
[0054] In particular, in this embodiment, when the user's dialogue interaction state is "active," "joyful," or "extroverted," the user's dialogue interaction state is classified as a high-index state, and a response language expression matching the high-index state can be output. Similarly, when the user's dialogue interaction state is "calm," "normal," or "chatty," the user's dialogue interaction state is classified as a medium-index state, and when the user's dialogue interaction state is "silent," "tired," or "introspective," the user's dialogue interaction state is classified as a low-index state, and a response language expression corresponding to each index pair is output.
[0055] For example, when a user's dialogue interaction state is in the high-indicator range, the ending of sentences should be made brighter (e.g., "That's great!", "Yes!"), the word length should be shorter, the tone should be clear and slightly rising, the interval between utterances should be faster, and the utterance frequency should be higher. This allows for a bright response that matches the user's dialogue interaction state when dealing with a user in the high-indicator range, and the user in the high-indicator range can empathize with the response language expression. Similarly, when a user's dialogue interaction state is in the medium-indicator range, the ending of sentences should be empathetic to the user's language expression (e.g., "That's right."), the intonation should be less, the word length should match the length of the user's language expression, the tone should be flat, the interval between utterances should be normal (e.g., 0.2 to 0.3 seconds), and the utterance frequency should be normal (e.g., 20 utterances per minute). This allows for normal responses that match the user's dialogue interaction state for users in the medium index range, making the user feel comfortable with the response language. Furthermore, if the user's dialogue interaction state is in the low index range, the system may not respond (for example, if the dialogue interaction state is silent), or it may use receptive endings (for example, "I see..."), make the word length very short, give a slow and slightly falling tone, slow the interval between utterances, and reduce the frequency of utterances. This allows for responses that are empathetic to the user's dialogue interaction state for users in the low index range, making the user feel comfortable with the response language.
[0056] It is also possible to pre-generate pre-trained models 41A for high index bandwidth, 44A for medium index bandwidth, and 45A for low index bandwidth, so that the response language representation is modified as shown in Figure 7.
[0057] The high-index trained model 41A processes the input response language expression to make it sound brighter by adding "ne," "yo," and "kana?" to the end of sentences, exclamation marks to the end of sentences, setting the word length to 20 to 30 characters per minute to improve the tempo, using positive words such as "fun," "happy," and "wonderful" to create a positive tone, setting the interval between utterances to be relatively fast, and increasing the frequency of utterances. The response language expression input to the active state trained model 41, the joyful state trained model 42, or the outward state trained model 43 may also be used as the output of the high-index trained model 41A, which is the response language expression output by the input model.
[0058] The pre-trained model 44A for the intermediate index range is designed to evoke empathy in the input response language expression. This is achieved by using questioning or speculative endings such as "Isn't it?", setting word lengths to approximately 30 to 40 characters, using soft and modest phrasing, and adjusting the utterance interval to allow the user time to breathe. The speech frequency is processed to mix questions with restrained reactions. The response language expression input to either the pre-trained model 44 for special occasions, the pre-trained model 45 for normal times, or the pre-trained model 46 for casual conversation may be used as the output of the pre-trained model 44A for the mid-range.
[0059] The low-index trained model 47A deliberately refrains from responding (silent response). When it does respond, it gently replies to the input response language expression, uses sentence-ending particles that reflect the other person's feelings without making a definitive statement, keeps the word length short, uses nuances that directly reflect the other person's feelings such as "That's tough," has short pauses between utterances, avoids repetition, and prioritizes positive sympathy. The response language expression input to the silent trained model 47, the fatigue trained model 48, or the introspective trained model 49 may be used as the output of the low-index trained model 47A, which is the response language expression output by the model that received the input.
[0060] Figures 8 and 9 are flowcharts illustrating the processing procedure of the CPU 2 of the response language expression output device 1. It is.
[0061] The user provides words (linguistic expressions) to the response linguistic expression output device 1 via voice using the microphone 8, or provides words (linguistic expressions) as text data using the keyboard 10 (or other input device). When the linguistic expressions provided by the user are input to the response linguistic expression output device 1 (step 51), the CPU 2 controls the GPU 3 to classify the user's dialogue interaction state from the parts of the input linguistic expressions excluding content words, using trained models 31-35 included in the trained model group 30 for dialogue interaction state classification (step 52). The GPU 3 performs the dialogue interaction state classification process. The dialogue interaction state classification process will be described in more detail later (see Figure 10). Once the user's dialogue interaction state is classified, the CPU 2 causes the GPU 3 to classify the index pair of the user's dialogue interaction state (high index, medium index, or low index) (step 53).
[0062] Next, it is determined whether or not to respond to the linguistic expression provided by the user (step 54) (this is an example of the processing by the CPU 2's first response determination means). For example, if the user's dialogue interaction state is classified as "introspection," no response is given (YES in step 54). The CPU 2 temporarily stops the output of response linguistic expressions from the display device 4, speaker 9, etc. (this is an example of the processing by the CPU 2's first output control means). When the user's dialogue interaction state is "introspection," the user is not expecting a response from the other party in the dialogue, but is in a state of deep reflection on their own thoughts, words, and actions, and since the user themselves is aware of their own actions, it is sometimes best to leave them alone. For this reason, the processing of deliberately not responding is performed. For example, the response is stopped until the next linguistic expression is input from the user.
[0063] If a response is to be given (NO in step 54), the language expression entered by the user is transmitted from the response language expression output device 1 to the AI server 20 under the control of the CPU 2 via API (Application Programming Interface) communication. The AI server 20 generates a response language expression in response to the language expression from the user and sends it to the response language expression output device 1. The transmitted response language expression is input to the response language expression output device 1 (step 55).
[0064] Next, CPU2 determines whether to respond without content words (Figure 8, step 56) (this is an example of processing by the third response determination means; it may also be performed by GPU3). If the user's dialogue interaction state is classified as, for example, "silent" or "tired," it is determined to respond without content words (YES in step 56), and the response language expression is modified so that it does not contain content words (step 57). This is because, when the user's dialogue interaction state is "silent" or "tired," it is thought that the user does not want a response to their language expression, but rather wants empathy that is attuned to the user's dialogue interaction state. If the user's dialogue interaction state is "active," "joyful," "extroverted," "calm," "normal," or "chatty," it is determined to respond with content words (NO in step 56). Even if the user's dialogue interaction state is "active," "joyful," "extroverted," "calm," "normal," or "casual conversation," if the response language expression sent from the AI server 20 does not contain content words, the response will be made without content words.
[0065] Furthermore, the CPU 2 controls the GPU 3 so that the response language expression sent from the AI server 20 is modified to match the user's dialogue interaction state (step 58). The response language expression is modified by the GPU 3. This process will be described in more detail later (see Figure 11). For example, if the user's dialogue interaction state is classified into a high index, medium index, or low index range, and the response policy for the response language expression is determined according to one of the index ranges, the response language expression is modified using one of the models: the high index trained model 41A, the medium index trained model 44A, or the low index trained model 47A, as described above, and the modified response language expression is output from the response language expression output device 1. For example, the modified response language expression is displayed on the display screen of the display device 4 or output from the speaker 9. Furthermore, even if the user's dialogue interaction state is not classified into high, medium, or low index ranges, the response language expression may be modified according to the type of user's dialogue interaction state ("active," "happy," "extroverted," "calm," "normal," "small talk," "silent," "tired," etc.) using the trained model 41 for active states, trained model 42 for happy states, trained model 43 for extroverted states, trained model 44 for calm states, trained model 45 for normal states, trained model 46 for small talk, trained model 47 for silent states, and trained model 48 for tired states (a "tired pattern" application model based on external features), as described above, and the modified response language expression may be output from the response language expression output device 1. In such cases, the response language expression output from the response language expression output device 1 may be displayed on the display screen of the display device 4 or output from the speaker 9.
[0066] Figure 10 is a flowchart showing the procedure for classifying the dialogue interaction state performed by GPU3 (the processing procedure for step 52 in Figure 8).
[0067] The language expression input to the response language expression output device 1 is classified according to the dialogue interaction state indicated by the ending of the language expression using the trained model for dialogue interaction state classification (for word endings) 31 (step 61). The dialogue interaction state indicated by the length of the language expression is also classified using the trained model for dialogue interaction state classification (for word length) 32 (step 62). Furthermore, the dialogue interaction state indicated by the nuance of the language expression is classified using the trained model for dialogue interaction state classification (for word connotation) 33 (step 63). The dialogue interaction state indicated by the utterance of the language expression is classified using the trained model for dialogue interaction state classification (for utterance interval) 34 (step 64). Furthermore, the dialogue interaction state indicated by the utterance frequency of the language expression is classified using the trained model for dialogue interaction state classification (for utterance frequency) 35 (step 65). Next, the type of user interaction state is classified comprehensively from all or some of the classification results obtained (step 66).
[0068] The following structurally explains an example of the conditions under which the dialogue interaction state is judged to be introspection and the silent response mode is activated (the conditions under which it is judged that there is no response in step 54 of Figure 8).
[0069] For example, a user relationship temperature score derived from non-semantic features (calculated from non-semantic features) The obtained temperature value. The lower the value, the lower the external reaction and the lower the activity state of the interaction state. For example, nine types may be represented by scores from 0 to 1 according to the interaction state shown in FIG. 6) as T, the end pitch drop rate (the tendency of the pitch to drop at the end of the sentence. The larger the value, the less intonation and the more serene. The rate at which the intonation at the end of the user's linguistic expression drops) as Pe, the amplitude decay rate (the rate of change of sound pressure at the end of the utterance. The larger the value, the weaker the volume. The decay rate of the amplitude of the sound obtained from the user's linguistic expression.) as Ae, and the utterance interval as ΔT (all of which can be digitized), response control is performed based on the mutual relationship (conditional linkage) of these four judgment elements, namely, the relationship temperature score T, the end pitch drop rate Pe, and the amplitude decay rate Ae. That is, when at least two or more of these four elements simultaneously satisfy a specific condition, the system suppresses response generation and transitions to the silent response mode. Specifically, when any one of the following conditional expressions (Expression 1) to (Expression 3), or a combination thereof, is satisfied, an output signal representing a command to generate a response linguistic expression to the response generation unit (AI server 20 or GPU 3) is stopped.
[0070] T < Tt and Pe < Pt ··· Expression 1 T < Tt and Ae < At ··· Expression 2 T < Tt and ΔT > ΔTt ··· Expression 3 Here, Tt is the dynamic threshold for the relationship temperature, Pt is the dynamic threshold for the end pitch drop rate, At is the dynamic threshold for the amplitude decay rate, and ΔTt is the dynamic threshold for the utterance interval, which are sequentially updated based on the user's utterance history (history of linguistic expressions) or average value (average value obtained by digitizing the judgment elements), and further based on environmental conditions.
[0071] With the above condition linkage, the system can detect the correlated changes of multiple feature quantities (judgment elements) and perform response control without depending on a single threshold value. Thus, for example, when the empathy temperature score T is low and the pitch drop rate Pe at the end of a sentence is high, it can be highly accurately classified that the user is in an introspective state (a dialogue interaction state where the external reaction is in a low-activity state, a state of mental block), and can shift to a silent response mode in which the output of the response language expression is stopped. Also, even if the relevant temperature score T is above a certain value, when the amplitude attenuation rate Ae is high and the utterance interval ΔT is long, it can be detected as a temporary hesitation or a gap between thoughts, and short-term output suspension can be performed. At this time, a response generation suppression signal S0 that controls the response control unit is issued as a logical value "1", and a silent state flag Fs is turned on. The silent state flag Fs is released when a language expression from the user is input to the response language expression output device 1 again.
[0072] With such a configuration, the determination of response suppression is defined not as a single parameter but as a relational expression of multiple features, and the non-verbal timing and atmosphere of a human can be technically reproduced. Therefore, the response language expression output device 1 in this embodiment does not depend on a fixed threshold value and can dynamically optimize the response according to the history and state of each user. Note that these conditions and threshold values are examples, and can be appropriately changed or updated by learning according to the usage environment, age, language, or type of dialogue application.
[0073] Also, in addition to Expressions 1 to 3, the response language expression output device 1 may be controlled so as to be a normal response R1, a light response R21 or a light response R22, and a silent response R3 as shown in Expressions 4 to 7 as follows.
[0074] R1 = T > 0.5 and Pe > -20 and Ae > -5 ··· Expression 4 R21 = 0.3 < T < 0.5 and Pe < -25 ··· Expression 5 R22 = 0.3 < T < 0.5 and Ae < -8 ··· Expression 6 R3 = T < 0.3 and Pe < -30 and Ae < -10 ··· Expression 7
[0075] Figures 11 to 15 show the normal response region R1, the mild response regions R21 and R22, and the silent response region R3, respectively, with the relational temperature score T, the sentence-end pitch drop rate Pe, and the amplitude attenuation rate Ae as axes. In other words, they show a response suppression determination structure based on the correlation of three judgment elements: relational temperature score T, sentence-end pitch drop rate Pe, and amplitude attenuation rate Ae. The T axis represents the relational temperature score (0.0 to 1.0), the Pe axis represents the sentence-end pitch drop rate (Hz / s), and the Ae axis represents the amplitude attenuation rate at the end of the utterance (dB / s).
[0076] Figures 11 to 15 show the regions represented by equations 4 to 7. For clarity, the Pe axis is defined as a space ranging from -40 to +40, the Ae axis from -30 to +30, and the T axis from -1 to +.
[0077] Figure 11 shows the normal response region R1 represented by Equation 4, which is a region where the relative temperature score T is high and the pitch decrease rate Pe and amplitude attenuation rate Ae are small. When the system enters this normal response region R1 (i.e., when Equation 4 is satisfied, and the user is judged to be in a normal response dialogue interaction state), a normal response is performed.
[0078] Figure 12 shows the mild response region R21, represented by Equation 5, which is an intermediate region between normal response and silent response. When entering this mild response region R21 (when Equation 5 is satisfied, and the user is judged to be in a mild response dialogue interaction state), a sentence-ending response or empathetic response in a short phrase is made.
[0079] Figure 13 shows the mild response region R22, represented by Equation 6, which is an intermediate region between normal response and silent response. When entering this mild response region R22 (when Equation 6 is satisfied, and the user is judged to be in a mild response dialogue interaction state), a sentence-ending response or empathetic response in a short phrase is made.
[0080] Figure 14 shows the silent response region R3 represented by Equation 7, which is a region where the relational temperature score T is low and the pitch decrease rate Pe and amplitude attenuation rate Ae are below a threshold. When entering this silent response region R3 (when Equation 7 is satisfied, and the user is judged to be in a silent response dialogue interaction state), a word-ending response or empathetic response in a short phrase is made. When entering the silent response region R3, the response control unit issues a response suppression signal S0, and if its logical value is "1", the output generation unit stops operating and the silent state flag F_silent is set to "ON". During this time, speech output is suppressed, reanalysis of non-semantic features continues, and when the relational temperature T or other features fall outside a predetermined range, the silent flag is released and normal response processing resumes. With this configuration, the silent response is realized not as a simple stop of a fixed time, but as conditional response control based on the correlation of multiple features. As described above, the generation of response language expressions stops when the output suppression signal is issued.
[0081] Figure 15 combines the regions R1, R21, R22, and R3 shown in Figures 11 to 14 into a single figure. The relationships between regions R1, R21, R22, and R3 can be understood relatively easily.
[0082] Therefore, the present invention can control responses using correlations between non-semantic features such as speech rate, sound pressure changes, and intonation changes, without relying on judgment based on a single threshold. This provides a technical means to detect non-verbal nuances such as pauses and hesitations in human dialogue in the signal space and select a response or silence according to that nuance. This configuration is applicable to dialogue AI, voice interfaces, educational support systems, and the medical and welfare fields, and can realize non-verbal empathetic expressions in the relationship between humans and AI through engineering.
[0083] Figure 16 is a flowchart showing the procedure for modifying the response language representation performed by GPU3 (the procedure for step 58 in Figure 9).
[0084] In the process of modifying the response language expression, first, the ending of the response language expression sent from the AI server 20 is modified or generated so that it corresponds to the indicator range of the user's dialogue interaction state (or the user's dialogue interaction state itself) (step 71). Also, the ending of the response language expression sent from the AI server 20 is modified so that it corresponds to the length of the indicator range of the user's dialogue interaction state (or the user's dialogue interaction state itself) (step 72). Furthermore, the nuance of the response language expression sent from the AI server 20 is modified so that it corresponds to the indicator range of the user's dialogue interaction state (or the user's dialogue interaction state itself) (step 73).
[0085] Furthermore, the utterance interval of the response language expression sent from the AI server 20 is generated so that the ending corresponds to an indicator range of the user's dialogue interaction state (or the user's dialogue interaction state itself) (step 74), and the ending of the response language expression sent from the AI server 20 is modified so that the utterance frequency corresponds to an indicator range of the user's dialogue interaction state (or the user's dialogue interaction state itself) (step 75).
[0086] The processes from steps 71 to 75 in Figure 16 may all be performed, or only some of them may be performed.
[0087] In this embodiment, the process of changing the response language expression may be performed using a word sense response generation unit (generation process in GPU3). The word sense response generation unit comprises a word ending dictionary (stored in SSD11, the data about word endings shown in Figures 6 and 7 may be used as the word ending dictionary) that stores candidate word endings or word senses, a word sense adjustment unit (which can be implemented in GPU3) that modifies the selected word ending by changing vowel extension and intonation to correspond to the index band and dialogue interaction state, and a timing control unit (which controls the time between responses and the time until a response is made) (which can be implemented in CPU2) that adjusts the timing of the response output. As a result, when the user's dialogue interaction state is in a low index band state, it is possible to output word sense responses such as "..." or "I see", and it is also possible to adjust the timing of the response and word ending expression according to the medium index band and high index band state. With this word sense response generation unit, in this embodiment, instead of simply returning semantic content, it is possible to present a sense of atmosphere such as "word sense" and "timing" as a response based on non-semantic features. Furthermore, the response generation unit may perform processes such as adjusting the nuance of the word, adjusting the waiting time, and adding reasoning to candidate responses obtained from LLMs (Large Language Models), for example, to determine the final output response. This series of processes is positioned as a verification phase in the implementation.
[0088] Figures 17 to 19 show the language expression provided by the user, the normal response language expression (the response language expression sent from the AI server 20) (referred to as the normal response), and the response language expression output device. The response language expression output from position 1 (which is assumed to be an empathetic response that the user empathizes with) This is an example.
[0089] Figure 17 shows an example of a case where the user's dialogue interaction state is classified as active.
[0090] Let's assume that the user inputs the linguistic expression "I had dinner with my friends today!" to the response language expression output device 1. When this linguistic expression from the user is input to the trained models 31-35 included in the trained model group 30 for classifying dialogue interaction states shown in Figure 4, a classification process is performed, and the user's dialogue interaction state is classified as active (see also Figure 6). When such a linguistic expression is input to the AI server 20, the AI server 20 outputs a normal response language expression. For example, the AI server 20 outputs a response language expression with the normal response "What did you eat?" and inputs it to the response language expression output device 1. Since the user's dialogue interaction state is active and in the high index range, the response language expression is modified to match that dialogue interaction state. In this case, for example, the response language expression is modified to become an empathetic response corresponding to an active dialogue interaction state, such as "Wow, that's great! What did you eat? Where?" and output from the response language expression output device 1. The user empathizes with the response language expression output from the response language expression output device 1, which allows for a smoother subsequent conversation.
[0091] Since the user's dialogue interaction state is active and in the high-indicator range, the response language expression may be changed to one that corresponds to active interaction as described above, or it may be changed to one that corresponds to a non-active dialogue interaction state of pleasure or outward-facing interaction.
[0092] Figure 18 shows an example of a user interaction state that is classified as casual conversation.
[0093] Let's assume that the user inputs the linguistic expression "I had dinner with a friend today" to the response language expression output device 1. When this linguistic expression from the user is input to the trained models 31-35 included in the trained model group 30 for classifying dialogue interaction states shown in Figure 4, a classification process is performed, and the user's dialogue interaction state is classified as being in a casual conversation state (see also Figure 6). When such a linguistic expression is input to the AI server 20, the AI server 20 outputs a normal response language expression. For example, the response language expression "What did you eat?" is output and input to the response language expression output device 1. Since the user's dialogue interaction state is casual conversation, the response language expression is modified to match that dialogue interaction state. In this case, for example, the response language expression is modified and output as "What did you eat?". Since the user's dialogue interaction state is casual conversation, it is in the medium index range, so for users in the medium index range, the normal response language expression may be output as is. For example, the AI server 20 could output "What did you eat?". In this case as well, the AI will empathize with the response language expression output from the response language expression output device 1, allowing for a smoother subsequent conversation while maintaining empathy.
[0094] Since the user's conversational interaction state is casual conversation and falls within the medium indicator range, the response language expression may be changed to one that corresponds to casual conversation, as described above, or it may be changed to one that corresponds to a calm or normal conversational interaction state other than casual conversation.
[0095] Figure 19 shows an example of a case where a user's conversational interaction state is classified as fatigue.
[0096] Let's assume the user inputs the following linguistic expression to the response language expression output device 1: "I had dinner with a friend today, and I'm exhausted." When this linguistic expression from the user is input to the trained models 31-35 included in the trained model group 30 for classifying dialogue interaction states shown in Figure 4, a classification process is performed, and the user's dialogue interaction state is classified as fatigued (see also Figure 6). When such a linguistic expression is input to the AI server 20, the AI server 20 outputs a normal response language expression. For example, the response language expression "What did you eat?" is output and input to the response language expression output device 1. Since the user's dialogue interaction state is fatigued and in the low index range, the response language expression is modified to match that dialogue interaction state. In this case, for example, the response language expression is modified to be more empathetic to the user, such as "I see," and output. Because the user empathizes with the response language expression output from the response language expression output device 1, they feel understood, and subsequent conversations become smoother.
[0097] Since the user's dialogue interaction state is fatigue and in the low index range, the response language expression may be changed to one that corresponds to fatigue, as described above, or it may be changed to one that corresponds to a dialogue interaction state other than fatigue, such as silence or introspection.
[0098] A history information table may be created that accurately classifies the dialogue interaction states for a specific user by recording the endings of words used, the nuances of words selected, and the intervals between utterances in the user's past language expressions in chronological order, and by associating these with the corresponding dialogue interaction states.
[0099] Figure 20 shows an example of a history information table for a specific user. The history information table is managed by a user ID unique to each user. The user ID is requested when logging in using the response language expression output device 1, and the history information table corresponding to the user ID is updated. The history information table is stored in SSD 11. The history information table stores the relationship between the language expression portion, excluding content words, and the user's dialogue interaction state for each user.
[0100] The history information table includes columns for features, judgment elements, and dialogue interaction state. The judgment elements column stores the features of the language expression provided by the user, the judgment elements for the dialogue interaction state determined from those features, and the dialogue interaction state column stores the user's dialogue interaction state when that feature is included in the language expression, under the control of CPU2 (this is an example of an update method). When the suffix "yes" is included in the language expression, it is classified as a general user emotional state (see Figure 6), but in the case of the user corresponding to the history information table in Figure 20, the dialogue interaction state is classified as "joyful." The reason why a particular user's dialogue interaction state can be determined to be "joyful" when the suffix "yes" is included in the language expression is explained below.
[0101] Even with the same linguistic expression, different users may experience different conversational interaction states. By considering the tendencies specific to a particular user and classifying that user's conversational interaction state accordingly, it is possible to classify a user's conversational interaction state with relatively high accuracy.
[0102] Figure 21 is an example flowchart showing the dialogue interaction state classification and correction process performed by GPU3. The process shown in Figure 21 is performed under the control of CPU2 at some point between the input of language expression from the user and the output of the response language expression to the user, as shown in Figures 8 and 9.
[0103] In step 81, judgment elements such as word endings and linguistic expressions in the classification table for a specific user are input into the history-trained model 36. In step 82, the history-trained model 36 outputs the dialogue interaction state of that specific user according to the input judgment elements and linguistic expressions. For example, if the user uses the linguistic expression "yes" at the end of a word, it would normally be inferred to be an outward dialogue interaction state, but for a specific user, it is classified and output as a casual conversation dialogue interaction state. The user's dialogue interaction state can be determined by the change in the linguistic expressions given by the user in relation to the response language expression output from the response language output device. For example, if a user uses a language expression ending in "yes," it is usually classified as an outward dialogue interaction state. Therefore, the response language expression is modified and output to match the outward dialogue interaction state. However, if the user expresses dissatisfaction with the response language expression that matches the outward dialogue interaction state, for example, "Hmm, something's not right, I'm not in the mood," then the dialogue interaction state for that particular user's "yes" expression is classified as not outward. If the user responds to a response language expression that matches the casual conversation dialogue interaction state with a positive expression, for example, "That's nice," then the dialogue interaction state for that particular user's "yes" expression becomes casual conversation, and the history information table is updated by CPU2. In this way, the "yes" of a particular user is output from the casual conversation dialogue interaction state and the history-trained model 36. This applies not only to cases where the decision element is the ending of a word, but also to other decision elements.
[0104] The output of the history-trained model 36 among the trained dialogue interaction state classification models 31-35 is given to the corresponding trained dialogue interaction state classification model 31-35 so that a dialogue interaction state corresponding to the judgment elements of a specific user is output, and the classification parameters are adjusted by the GPU3 (step 83).
[0105] Figure 22 corresponds to Figure 6 and is an example of an individual user determination table (relationship) for determining the dialogue interaction state for a specific user.
[0106] For example, generally, if a user's given language expression ends in "~dayo," the user's dialogue interaction state is classified as casual conversation. However, for a particular user, if the language expression ends in "~dayo," the user's dialogue interaction state is considered active. In that case, the judgment table is updated so that if the user's language expression ends in "~dayo," the user's dialogue interaction state is classified as active (in Figure 16, the judgment element is the ending, and "~dayo" is added to the column for active dialogue interaction state). This improves the accuracy of classifying the user's dialogue interaction state. Individual user judgment tables are also managed for each user ID and stored in SSD11.
[0107] Figure 23 is a flowchart illustrating the dialogue interaction state classification process, which is performed by GPU3.
[0108] As described above, when a linguistic expression from the user is input to the response linguistic expression output device 1, and a response linguistic expression that matches the user's dialogue interaction state is output to the user (or when the dialogue concludes), a question inquiring about the user's dialogue interaction state is displayed on the display screen 4 of the response linguistic expression output device 1 (step 84). The user answers the question with their dialogue interaction state and inputs it to the response linguistic expression output device 1 (step 85). The user's emotional state in the process shown in Figure 16 may be handled in the same way as in steps 84 and 85.
[0109] In the response language expression output device 1, the individual user judgment table is updated so that the judgment elements such as word endings used to determine the language expression input by the user correspond to the dialogue interaction state input by the user (step 86). Furthermore, once the individual user judgment table is updated, the GPU 3 adjusts the trained dialogue interaction state classification models 31-35 so that the judgment elements such as word endings used to determine the language expression input by the user correspond to the dialogue interaction state input by the user (step 87).
[0110] As shown in Figures 21 to 23, when dialogue interaction state classification processing is performed to match each user, the trained model group 30 for dialogue interaction state classification shown in Figure 4, etc., is also managed for each user by a user ID, and the model group 30 corresponding to each individual user is used. In addition, a trained model group 40 for dialogue interaction state-specific responses may also be provided for each user, so that a response language expression that matches each user can be obtained.
[0111] Figures 24 to 26 are flowcharts showing the processing steps corresponding to a part of the process shown in Figure 8.
[0112] Referring to Figure 24, if it is determined that there is no response (YES in step 54), there is a period of time during which there is no response (NO in step 91), and once that period has elapsed (YES in step 91), the process proceeds to step 55 and the response language expression is output.
[0113] Since the system will remain silent for a certain period before responding, the user can reflect on themselves during that time, and the silence can give the impression that the system is considerate of the user's well-being. Furthermore, because the system remains silent for a certain period before responding, it can also obtain a response language that corresponds to the language expression provided by the user.
[0114] Alternatively, the process in step 54 of Figure 8 may be skipped, and the process in step 54 shown in Figure 24 may be inserted after step 58 of Figure 9. If a response is given, the response language expression may be output (step 59 of Figure 9). If no response is given, the process in step 91, which checks whether a certain amount of time has elapsed, may be performed, and if a certain amount of time has elapsed, the process may proceed to output the response language expression (step 59 of Figure 9).
[0115] Referring to Figure 25, if it is determined that there is no response (YES in step 54), there is a period of time during which there is no response, similar to the process in Figure 19 (NO in step 91). After the specified period has elapsed (YES in step 91), the process proceeds to step 55. For example, the period during which a response language expression is output from the response language expression output device 1 without determining a silent response is considered a specified period. If the user inputs a new language expression before the specified period has elapsed (NO in step 91) (YES in step 92), the process proceeds to step 55. However, it is not necessary to determine whether the specified period has elapsed.
[0116] If there is no response, the user may become anxious and enter new language expressions such as "What's wrong?", and this can be addressed in such cases. If a new language expression is entered (YES in step 92), a response language expression for that new language expression will be obtained from the AI server 20 (step 55).
[0117] In this case as well, the process in step 54 of Figure 8 may be skipped, and the process in step 54 shown in Figure 25 may be inserted after step 58 of Figure 9. If a response is given, the response language expression may be output (step 59 of Figure 9). If no response is given, the process in step 91 may be performed to check whether a certain amount of time has elapsed. If a certain amount of time has elapsed, the process may proceed to outputting the response language expression (step 59 of Figure 9). If a new language expression is input before a certain amount of time has elapsed, the process may proceed to step 55.
[0118] In the example described above, the process in step 91, which determines whether a certain amount of time has elapsed, may be omitted, and the process may proceed to step 92 if it is determined in step 54 that there is no response.
[0119] Referring to Figure 26, following the processing in step 53 of Figure 7, it is determined whether to give a silent response (step 93) (processing of the second response determination means). For example, if the user's dialogue interaction state is determined to be silent, it is determined to give a silent response. If it is determined to give a silent response (YES in step 93), it is determined whether a certain amount of time has elapsed (step 94). For example, the time when a response language expression is output from the response language expression output device 1 without determining whether to give a silent response is considered a certain amount of time. For example, even if it is determined to give a silent response, the processing in steps 55 to 58 is performed (no response language expression is output), and this time is considered necessary for those processes. After a certain amount of time has elapsed, information indicating the silent state is displayed, for example, the symbol "..." is displayed on the display screen of the display device 4, or "I am deliberately remaining silent." is displayed, or the voice "I am deliberately remaining silent." is output from the speaker 9 (step 95). The user can see that the response language expression output device 1 is deliberately remaining silent. If it is not determined that the response is silent (NO in step 93), the response language expression for the input language expression is obtained from the AI server 20 (step 55) (the CPU 2 that controls it in this way is an example of a second output control means).
[0120] In the above embodiment, the process in step 94, which determines whether a certain amount of time has elapsed, may be omitted, and information indicating a silent state may be output in response to the determination that a silent response has been made (YES in step 93) (step 95).
[0121] In the above-described embodiment, state classification is performed based on the user's external characteristics and linguistic expressions from the user using trained models 31-35, etc. However, instead of performing such classification, state classification may be performed based on the user's external characteristics and linguistic expressions from the user's linguistic expressions using a table like the one in Figure 6. Also, while the response linguistic expression is modified by classifying it to correspond to the user's dialogue interaction state using trained models 41-49, etc., instead of performing such classification, the response linguistic expression may be modified to correspond to the user's dialogue interaction state using a table like the one in Figure 6.
[0122] Furthermore, in the above embodiment, processing is performed using GPU3, but part or all of the processing using GPU3 may be performed by CPU2, or part of the processing performed by CPU2 may be performed by GPU3. Alternatively, instead of GPU3, or in addition to GPU3, an NPU (Neural Processing Unit) may be provided in the response language representation output device 1, and part or all of the processing of GPU3 may be performed by the NPU.
[0123] The empathetic dialogue support system according to the present invention is a conversational AI technology capable of generating emotional responses to a user's natural language utterances. When incorporated as software, it can be widely applied to smartphones, chatbots, digital assistants, car navigation systems, robots, medical and nursing care communication devices, educational support systems, and the like. Therefore, the present invention has industrial applicability.
[0124] This invention relates to an empathetic dialogue support system that classifies the user's dialogue interaction state (e.g., low-activity / active system) as a "relationship temperature zone" based on non-semantic elements such as word endings, intonation, utterance intervals, vocabulary, and nuances contained in the user's natural language utterances, and controls the presence, content, tone, and word endings of natural language responses according to that temperature.
[0125] Conventional conversational AIs primarily rely on semantic analysis of user utterances to generate informative responses. However, the technology to construct responses such as "not speaking," "not replying," or "gently offering support" in response to silences, hesitations, and the emotional temperature conveyed by the choice of sentence endings, which do not necessarily carry linguistic meaning, had not yet been established.
[0126] According to the present invention, non-semantic control becomes possible, such as responding only with resonant endings depending on the relationship temperature range, omitting the response, or softening the ending, thereby providing an AI response experience that can naturally resonate with the user's emotions.
[0127] In this specification, "non-semantic elements" refer to linguistic features based on non-semantic or phonological information structures, such as intonation at the end of words, softness of tone, spacing between words, number of syllables, and tempo of speech, without relying on the semantic or logical meaning of the utterance.
[0128] Figure 27 is an overall block diagram of the main components in one embodiment of the relational temperature classification type dialogue generation system according to the present invention.
[0129] This system acquires the user's natural language utterances, classifies the user's dialogue interaction state as a "relationship temperature zone" based on the non-semantic features contained therein (sentence endings, intonation, nuance, interval between utterances, frequency of utterances, etc.), and generates and outputs natural language responses (or no responses) corresponding to the index zone. Specifically, it starts with a user utterance acquisition unit (Figure 22, far left) that receives the user's utterances, and the utterance data is input to a relation temperature classification module (Figure 22, top center), where index zones (high index zone, medium index zone, low index zone, etc.) are classified.
[0130] Next, the response control decision module (Figure 27, center) determines the response policy, such as "to respond / not to respond" and "how to adjust the ending," based on the classification results.
[0131] If it is determined that a response should be given, the resonance word sense generation module (Figure 27, bottom center) generates non-semantic word sense expressions such as "I see" and "Yeah," and these are presented in the dialogue output unit (Figure 27, far right) in the form of audio, text, or screen display.
[0132] Furthermore, within each process, the user history learning module (not shown in the diagram) accumulates information such as the individual user's word ending tendencies and response tolerance, and this is reflected in the tuning of the relational temperature classification module and the word sense generation module.
[0133] Figure 28 is a response matrix diagram showing the correspondence between the user's dialogue interaction state, the corresponding response control policy, and specific response examples in the relational temperature classification type dialogue generation system according to the present invention. In this figure, the relational temperature zones classified in the present invention (e.g., high index zone, medium temperature, low index zone) are arranged in the left column, and to the right, the corresponding user state examples, response control policies, and actually generated response examples are shown as a correspondence. For example, if the user is in an active state and is classified as a high index zone state where emotions are outward and include joy, the system is configured to adjust the endings and tone of the sentences to be brighter and to provide an immediate response, and a resonant expression with intonation and emotion, such as "That's great!", is generated as an example of such a response.
[0134] On the other hand, in the case of the medium index range, which represents a calm and normal state, soft and restrained responses such as "That's right." are selected. Furthermore, when a low index range state, including silence, fatigue, and introspection, is classified, non-invasive and receptive responses are selected, such as omitting the response or slowly replying only with the end of a sentence, like "I see..."
[0135] Thus, the response matrix shown in this figure visually demonstrates how the response generation strategy of the present invention adaptively corresponds to the user's dialogue interaction state (relevant temperature range, such as low activity / active system).
[0136] Figure 29 is a block diagram showing a learning configuration in the relational temperature classification type dialogue generation system according to the present invention, which dynamically adjusts the accuracy of index band classification based on the utterance history of each user. As shown in this figure, the present invention is provided with a user utterance history database that continuously records and learns the user's utterance content and response tendencies, where the usage tendencies of word endings, selection tendencies of word nuances, utterance intervals, relational temperature response patterns, etc., in each user's past utterances are accumulated in chronological order.
[0137] This historical data is analyzed by the history learning module, and parameters such as the following are learned: When a user responds with "yes," they are in a relatively high index state. When a suffix like "I see..." is selected, they tend to be introspective. When silence lasts for 30 seconds or more, they tend not to want a response. The history learning module feeds these analysis results back to the relational temperature estimation module, which dynamically adjusts the index classification thresholds and response control parameters according to the user's unique response tendencies. This ensures that different contexts are considered for each user, even for the same utterance, resulting in individually optimized relational temperature classification and appropriate response control based on it.
[0138] As shown in Figure 24, the present invention does not rely on a fixed index classification model, but rather has an evolutionary dialogue support structure based on user-specific history learning.
[0139] Figure 30A is a block diagram showing the generation process configuration of a "resonance response based on word endings and nuances" that is executed when a response is deemed necessary in the relational temperature classification type dialogue generation system according to the present invention.
[0140] The resonance word sense generation module shown in this figure is a module that sequentially designs and constructs non-semantic responses (word endings) based on the user's state index classification results and history information.
[0141] This module consists of the following three processing units. 1. Ending generation section Based on an index, context, and nuance database, appropriate suffixes (e.g., "Oh, I see," "Yeah," "That's right," etc.) are selected. Candidates are extracted based on conditions such as "tone tags," "words without meaning," and "intensity of impression." 2. Speech feel adjustment section For expressions selected in the word-ending generation section, adjustments are made to the number of syllables (length of elongated sounds), the selection of voiced / nasal consonants, and the softness / intensity of the word's feel. For example, it controls the difference in feel between the same word or phrase, such as "I see." and "I see..." 3. Inter-word control unit By controlling the intervals between responses and the pauses within phrases, immediate responses, delayed responses, and even "silence that feels like a response" can be expressed. This pause control is particularly important when the index range is low, and the response is configured to feel as if it has been "gently placed." Through these processes, the responses in this invention achieve resonant responses that do not rely on semantic content, but are composed solely of word endings and their emotional design. As shown in Figure 30A, the resonant word sense generation module achieves flexibility and empathy in response design, including "not saying anything," by performing a structurally stepwise generation process.
[0142] Figure 30B is a word-sense response mapping table showing specific output examples of word-ending expressions corresponding to the user's state index range in the relational temperature classification type dialogue generation system according to the present invention, and the corresponding relationships between the word-sense features (tone, number of syllables, pauses, etc.) included in each.
[0143] This figure visualizes the processing results in the resonance word sense generation module shown in the previous figure (Figure 30A) as an output format, and shows the interrelationships of the following elements in a table format. (1) Indicator zone classification: Interaction state classified by the relational temperature classification module (high indicator zone / medium indicator zone / low indicator zone) (2) Examples of sentence endings: Resonant sentence endings selected in the relevant index range (e.g., "That's right," "I see," "Yeah!" etc.) (3) Tonal characteristics: The intonation and tone (rising / flat / falling) conveyed in the final expression of a word. (4) Moral length and word spacing: length of syllable endings, pauses between words, and the immediacy or delay of responses. (5) Intent and nuance of response: The impression and atmosphere created by the response (empathy, affirmation, quiet reasoning) (Solutions, etc.)
[0144] For example, in a low-index state, a longer, falling tone ending like "I see..." is selected, and the longer pause in response gives the impression of "gently empathizing." On the other hand, in a high-index state, a short, immediate response ending like "Yeah!" is selected, resulting in a high level of activity and emphasizing emotional sharing, thus achieving resonance.
[0145] Thus, this figure shows that the linguistic response generated according to the relational temperature is designed as a trinity of word ending, intonation, and space between words, visualizing as a concrete example the structure of dialogue control that "returns as air" rather than merely conveying meaning.
[0146] An example of an embodiment of the present invention is shown below. The empathetic dialogue support system of the present invention is mainly composed of the following components.
[0147] 1. User utterance acquisition unit It receives natural language input via text or voice. In the case of voice input, it works in conjunction with a mechanism to extract acoustic features (intonation, pitch at the end of words, tempo, etc.). 2. Related Temperature Classification Module Non-semantic features such as the type of word endings included in utterances, the softness of the tone, the interval between utterances, word repetition, and lexical index tendencies are analyzed and classified into "relationship temperature zones" (e.g., low index zone = silence / fatigue, medium index zone = calm, high index zone = activity / joy) using rule-based or machine learning models. 3. Response control decision module Based on the classified indicator range, the system determines whether to respond or not, selects appropriate word endings if a response is given, adjusts the nuance of the wording, and delays the response timing. 4. Resonant word sense generation module It generates sentence endings that convey empathy even without having any inherent meaning (e.g., "Oh, I see," "Yeah," "That's right"). The phonemes, syllable count, duration, and tempo of the sentence endings can be controlled. 5. User History Learning Module The system learns each user's speech history, including their tendency towards certain word endings, tolerance for responses in each index range, and speech rhythm, and uses this information to optimize index range classification and word ending generation. 6. Dialogue Output Section Based on the output control results, the system responds in text or voice format. Even if no response is given, it can implement "presence acknowledgments" such as hiding or flashing a display to maintain the connection status.
[0148] Furthermore, in this embodiment, it is possible to enhance the "naturalness" and "resonance of atmosphere" for the user by adjusting the threshold for state index classification, adding tone tags, inter-word length, and randomness to the output interval for generating response endings.
[0149] Furthermore, even if the user provides incomplete input via voice or typing, the system can classify the state indicators from the intermediate steps and automatically select an appropriate response (no response, wait patiently, or gently resonate).
[0150] The present invention has the following configuration. 1. User utterance acquisition unit Receive natural language text or voice input. 2. Related Temperature Classification Module State index ranges (high index range / medium index range / low index range, etc.) are classified based on the type of word ending, nuance, interval between utterances, vocabulary selection, and speaking speed. 3. Response control decision module Based on the classification results, output control such as "return / don't return" and "select tone" is performed. 4. Resonant word sense generation module It generates meaningless suffixes (e.g., "Oh, I see," "Yeah," "That's right," etc.) and responds empathetically. 5. User History Learning Module The system learns each user's use of sentence endings, pauses, and emotional nuances, and reflects this in classification and word selection. 6. Dialogue Output Section The response will be either output or omitted, and presented via text, audio, screen, etc.
[0151] An example of the relationship between user metrics status and AI. Example 1: Users in a low performance index User: "...(Silence)" AI: "(No response)" or "I see." Example 2: Medium indicator range (calm) User: "Nothing particularly noteworthy happened today." AI: "That's right." Example 3: High indicator (activity) state User: "Something wonderful happened today!" AI: "That's great!" (replies with a slightly bouncy tone at the end) Example 4: Remaining silent and connected to the system. User: "(Several tens of seconds pass without a word)" AI: Based on user history and the length of utterance intervals, it classifies the state as a low-indicator zone and generates a soft response such as "(no response)" or "I see...". Furthermore, even if no utterance is present, it is possible to configure the system to interpret silence as a form of "non-semantic input information" based on connection status, time elapsed, past history, etc., and perform relational temperature classification and response control.
[0152] This invention has a different structure and purpose from conventional semantic and emotional response generation, and is a unique technology established by the following technical differences.
[0153] 1. Differences in input targets: While conventional technologies focused on "semantic content" or "biometric information (facial expressions, heart rate, etc.)" in user input, this invention analyzes only non-semantic linguistic features such as word endings, intonation, intervals between words, syllable count, and speech rhythm as input information. 2. Differences in classification axes: Conventional technologies primarily focused on classifying emotions such as joy, anger, sadness, and pleasure, or on semantic evaluation using a two-axis system of positive / negative. However, this invention uses non-emotional states of the atmosphere, such as "the activity level of dialogue interaction" and "relationship temperature range (low index range, medium index range, high index range)," as classification axes. 3. Differences in response control structures: While conventional technologies are information-responsive, returning appropriate information or meaning according to the input content, the present invention controls whether or not to return information, the selection of word endings when returning information, the intonation of the tone, the pauses, and the delay timing based on intuitive judgment. 4. Nature of the response: While conventional responses primarily consisted of meaningful responses (e.g., "Good luck"), the novel aspect of this invention is that it allows responses to be formed using only non-meaningful, nuanced expressions such as "That's right," "Yeah," and "...". 5. Whether or not the system is designed to "omit" responses: While many conventional technologies are based on the premise of "always returning a response," the present invention has a structure in which the AI itself decides and executes not to respond to silence or low-indicator states, and "not responding" is incorporated into the technical specifications. Based on the above, the present invention is a technological domain that is fundamentally different from conventional conversational AI in terms of input structure, classification concept, control design, and response mode.
[0154] Figure 31 illustrates the characteristics of the conversational state-adaptive personality switching structure. This embodiment is a technology for realizing "human-centered responses" in conversational AI, and in particular, it is an "air-reactive" technology that classifies the user's relationship temperature range (from low index range to high index range) based on non-semantic linguistic features (word endings, nuances, utterance intervals, etc.) and enables response selection such as "no response" or "respond based only on nuance." [Explanation of symbols]
[0155] 1: Response language expression output device, 2: CPU, 3: GPU, 4: Display device, 5: Memory, 7: Communication device, 8: Microphone, 9: Speaker, 10: Keyboard, 12: CD drive, 13: CD, 20: AI server, 30: Pre-trained models for dialogue interaction state classification, 31-35: Pre-trained models for dialogue interaction state classification, 36: History-trained model, 40: Pre-trained models for dialogue interaction state-specific responses, 41: Pre-trained model for active state, 4 1A: Trained model for high index frequency, 42: Trained model for joyful situations, 43: Trained model for outward-facing situations, 44: Trained model for calm situations, 44A: Trained model for medium index frequency, 45: Trained model for normal situations, 45A: Trained model for low index frequency, 46: Trained model for casual conversation, 47: Trained model for silence, 47A: Trained model for low index frequency, 48: Trained model for fatigued situations (model applying "fatigue patterns" based on external features), 49: Trained model for introspective situations
Claims
1. A language expression input means that inputs language expressions provided by the user. A first classification means that classifies the user's dialogue interaction state based on the portion of the language expression input from the above-mentioned language expression input means, excluding the content words that have meaning. A modification means for modifying a response language expression, which is a response to a language expression input from the above language expression input means, so as to match the user's dialogue interaction state classified by the first classification means. Output means for outputting a response language expression modified by the above modification means, A first response determination means that determines whether or not to respond based on the dialogue interaction state classified by the first classification means described above, and In response to the determination by the first response determination means that no response is expected, a first output control means controls the output means to stop outputting the response language expression. A response language expression output device equipped with the following features.
2. The above modification means modifies the response language expression portion of the response language expression, which is a response to the language expression input from the above language expression input means, excluding the meaningful content words, so as to match the user's dialogue interaction state classified by the first classification means. The response language expression output device according to claim 1.
3. The first output control means described above is: After a predetermined time has elapsed since the first response determination means determined that there is no response, the modified response language expression is output by the modification means. The response language expression output device according to claim 1.
4. The above modification means is, A word ending dictionary that stores candidate word endings or word nuances, A word feel adjustment unit that modifies the selected word ending by vowel extension or intonation change, It includes an inter-control unit that adjusts the timing of the response output, A word sense response generation unit that combines the above word ending dictionary, the above word sense adjustment unit, and the above inter-control unit to generate a word sense response according to the user's dialogue interaction state. A response language expression output device according to any one of claims 1 to 3, characterized by including the following:
5. After the first response determination means determines that there is no response, a recontrol means controls the response language expression output device to cause the first classification means to classify the new language expression and the modification means to modify it, in response to the user inputting a new language expression from the language expression input means, and to output a response language expression which is a response to the new language expression from the output means. The response language expression output device according to claim 1, further comprising:
6. A second response determination means that determines whether to respond with silence or not based on the dialogue interaction state classified by the first classification means described above, and A second output control means controls the output means to output information representing a silent state in response to a determination by the second response determination means that a silent response is made, and to output a response language expression modified by the modification means in response to a determination by the second response determination means that a silent response is made. An output device for a response language expression according to claim 1, comprising:
7. The second output control means described above is: If the second response determination means determines that a silent response is to be given, or if the second response determination means determines that a silent response is not to be given, information indicating a silent state is output after a period of time longer than the time required from the determination result of the second response determination means until the modified response language expression is output by the modification means. The response language expression output device according to claim 6.
8. The system includes a third response determination means that determines whether to make a response that does not contain meaningful content words, based on the dialogue interaction state classified by the first classification means described above. The above modification means is, In response to the determination by the third response determination means described above that the response does not contain meaningful content words, the response language expression, which is a response to the language expression input from the language expression input means described above, is modified so that it does not contain meaningful content words. The response language expression output device according to claim 1.
9. From the user's dialogue interaction state classified by the first classification means described above, the user's dialogue interaction state is: It includes a second classification means for classifying which of several groups, each classified according to a similar dialogue interaction state, it belongs to. The above modification means is, The response language expression to the language expression input from the above language expression input means is modified to match the dialogue interaction state represented by the group classified by the second classification means. The response language expression output device according to claim 1.
10. A memory control means that controls a memory device to store, for each user, the relationship between the language expression portion of the language expression input from the above language expression input means, excluding meaningful content words, and the user's dialogue interaction state, and The above storage device includes an update means for updating the relationships stored for each user, The first classification means described above is Based on the relationships updated by the above update means, the user's dialogue interaction state is classified based on the linguistic expression portion of the linguistic expression input from the above linguistic expression input means, excluding the meaningful content words. The response language expression output device according to claim 1.
11. The language expression input means inputs language expressions provided by the user, The first classification means classifies the user's dialogue interaction state based on the linguistic expression portion of the linguistic expression input from the linguistic expression input means, excluding the content words that have meaning. The modification means modifies the response language expression, which is a response to the language expression input from the language expression input means, so that it matches the user's dialogue interaction state classified by the first classification means. The output means outputs the response language expression modified by the modification means, The first response determination means determines whether or not to respond based on the dialogue interaction state classified by the first classification means. The first output control means controls the output means to stop outputting the response language expression in response to the first response determination means determining that no response is given. A method for outputting a response language representation.
12. A computer-readable program that controls a response language expression output device and takes language expressions provided by the user as input. The system classifies the user's dialogue interaction state based on the linguistic expression portion of the input language expression, excluding the meaningful content words. The response language expression, which is a response to the input language expression, is modified to match the classified user's dialogue interaction state. Output the modified response language representation. Based on the above-classified dialogue interaction state, determine whether or not to respond. A program that controls the computer of the response language expression output device to stop outputting the response language expression in response to the determination that it is not responding as described above.
13. A recording medium storing the program described in claim 12.
Citation Information
Patent Citations
Recognition device, learning device, method for same, and program
WO2021166207A1
system
JP2025047448A