Response language expression output device and method

The co-creation AI system addresses the lack of dynamic personality adjustment in dialogue AI by classifying user states and adjusting AI personalities based on non-semantic features, resulting in empathetic and trustworthy interactions.

JP7843094B1Active Publication Date: 2026-04-09ITO CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional dialogue AI systems lack the ability to dynamically adjust their personality and expression style to match the user's conversational state, leading to a lack of deep empathy and trust in interactions.

Method used

A co-creation AI system that classifies user dialogue states based on non-semantic features such as word endings, intonation, and utterance intervals, and adjusts AI personalities to generate responses that match the user's conversational style, using pre-trained models to determine appropriate dialogue states and personalities.

Benefits of technology

The system enables AI responses that align with the user's conversational state, fostering deep empathy and trust by dynamically adapting its personality, enhancing the conversational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843094000001_ABST
    Figure 0007843094000001_ABST
Patent Text Reader

Abstract

Conventional conversational AIs have a fixed response style and are unable to flexibly change their tone or demeanor according to the user's state during the conversation, making it difficult to foster deep empathy and trust. [Solution] The present invention provides a dialogue AI structure that achieves empathetic and adaptive responses by classifying the state of the conversation (e.g., calm, serene, compassionate) from the tone, endings, and pauses of the user's utterances, and dynamically selecting and switching a corresponding personality profile (e.g., endings, perspective, tempo, tone).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to outputting a response language expression that responds to a language expression given by a user (including language expressions such as those spoken or written by the user and input through a microphone as voice, language expressions input by the user using an input device such as a keyboard, and language expressions such as words and sentences obtained by reading the character strings described on paper).

[0002] In particular, this invention relates to a co - creation AI system that dynamically switches the response personality of AI (Artificial Intelligence) according to the user's dialogue state, and more particularly to a dialogue control technology that adjusts the personality through state changes such as concentration, tranquility, and compassion. The "dialogue interaction state" and "dialogue state" in the present invention refer to a classification of the dialogue interaction state based on external (non - semantic) feature quantities such as utterance length, utterance interval, word endings, intonation, tempo, etc. (a state that can be seen from the outside when the user interacts with the response language expression output device, for example, the user's appearance, attitude, behavior, activity state, etc.), and do not classify or judge an individual's emotional, psychological, emotional, or mental state.

Background Art

[0003] In conventional dialogue AI, a fixed response style is used. There are also some attempts to give machines emotions through emotion AI.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] It has been difficult to change personality and expression style according to the user's conversational state. For this reason, while there are changes in response tone based on emotion classification, deep structural changes including personality switching have not been implemented. Furthermore, the invention described in Patent Document 1, as described in claim 1 of Patent Document 1, involves sharing a cheerful movement pattern with a partner. This invention aims to realize conversations that foster deep empathy and trust by having the AI ​​respond with a relatively optimal personality in accordance with the user's conversational interaction style. [Means for solving the problem]

[0006] The response language output device according to this invention is a language table that inputs language expressions provided by the user. The system is characterized by comprising: current input means; dialogue state classification means for classifying the user's dialogue state based on non-semantic features such as word endings, intonation, interval between utterances, length of utterances, and tempo (for example, semantic features that can be deduced from external features recognized from the language expression given by the user) among the language expressions input from the language expression input means; AI personality determination means for determining an AI personality suitable for responding to a partner (user) having the dialogue state classified by the dialogue state classification means; acquisition means for acquiring a response language expression which is a response to the language expression input from the language expression input means, and which is a response to the language expression input from the language expression input means (including generating a response language expression with a standard AI personality and modifying the generated response language expression with the AI ​​personality determined by the AI ​​personality determination means).

[0007] This invention also provides a method for outputting response language. Specifically, this method involves a language expression input means inputting a language expression provided by a user, a dialogue state classification means classifying the user's dialogue state from the language expression input means based on non-semantic features such as word endings, intonation, utterance intervals, utterance length, and tempo, an AI personality determination means determining an AI personality suitable for responding to a partner having the dialogue state classified by the dialogue state classification means, an acquisition means acquiring a response language expression which is a response to the language expression input from the language expression input means and is generated by the AI ​​personality determined by the AI ​​personality determination means, an acquisition means acquiring a response language expression which is a response to the language expression input from the language expression input means by the AI ​​personality determined by the AI ​​personality determination means, and an output means outputting the response language expression acquired by the acquisition means.

[0008] Furthermore, this invention relates to a program for controlling the computer of a response language expression output device and We also provide a storage medium containing that program.

[0009] The acquisition means described above can be obtained, for example, by generating a response language expression, which is a response to a language expression input from the language expression input means, using the AI ​​personality determined in the AI ​​personality determination means, or by having a device other than the response language expression output device generate a response language expression, which is a response to a language expression input from the language expression input means, using the AI ​​personality determined in the AI ​​personality determination means.

[0010] The system may also include: dialogue state specification means for specifying a dialogue state; determination means for determining whether the dialogue state classified by the dialogue state classification means does not match the user's dialogue state; storage control means for controlling the storage device to store, in accordance with the determination means's finding that the dialogue state does not match, the characteristics of the language expression used for classification in the dialogue state classification means and the dialogue state specified by the dialogue state specification means for each user; and adjustment means for adjusting the dialogue state classification means to classify the dialogue state stored in the storage device in accordance with the characteristics of the language expression stored in the storage device when a language expression having the same characteristics as those stored in the storage device is input to the language expression input means.

[0011] The system may also include: a means for specifying a dialogue state; a means for selecting whether to determine an AI personality suitable for responding to an opponent having a dialogue state classified by the dialogue state classification means, or to determine an AI personality suitable for responding to an opponent having a dialogue state specified by the dialogue state specification means; and a control means for controlling the AI ​​personality determination means to determine an AI personality suitable for responding to an opponent having a classified dialogue state in response to the selection means selecting the determination of an AI personality suitable for responding to an opponent having a classified dialogue state, and to determine an AI personality suitable for responding to an opponent having a specified dialogue state in response to the selection means selecting the determination of an AI personality suitable for responding to an opponent having a specified dialogue state.

[0012] If an AI personality profile defining the AI ​​personality is defined for each AI personality, and the AI ​​personality profile stores at least one of the following: tone of voice, endings of words, nuance, response speed, speaking tempo, speech style, viewpoint, emotional intonation pattern, dialogue rhythm, and position in the dialogue, then the acquisition means may, for example, acquire a response language expression generated based on the AI ​​personality profile corresponding to the AI ​​personality determined by the AI ​​personality determination means. The dialogue state classification means classifies the user's dialogue state based, for example, on the language expression portion of the language expression given by the user, excluding meaningful content words (the language expression portion excluding meaningful content words includes not only the words themselves, such as the endings of the language expression portion, but also the length of the language expression, the nuance derived from the language expression, such as rhythmicity and the presence or absence of long vowels, as well as the interval between utterances, the frequency of utterances, the intonation of the voice, and, if the language expression is input from a keyboard, the speed of the language input from the keyboard and the interval of input time).

[0013] The dialogue-adaptive personality switching control system according to this invention is a co-creation AI system that controls AI responses generated through dialogue with a user, and is characterized by including: a dialogue interaction state determination unit that classifies or designates multiple dialogue states (hereinafter also referred to as "dialogue interaction states") such as calm mind, tranquil mind, compassionate mind, and emptiness of mind based on context, tone, endings, and pauses included in the user's natural language utterances, or selection operations by the user; a personality selection unit that selects a corresponding personality from a plurality of AI personality profiles in which dialogue styles such as tone, vocabulary, endings, nuances, response tempo, and perspective are predetermined, according to the corresponding dialogue interaction state; a response generation unit that generates a natural language response based on the selected personality profile; and an output unit that outputs the response to the user.

[0014] Furthermore, the system may include a history management unit that stores and analyzes utterance patterns, transition trends in dialogue interaction states, and personality selection history included in past user interaction history. Based on this information, a dialogue interaction state determination unit dynamically classifies the next dialogue interaction state to be applied and reflects this in the personality switching control by the personality selection unit.

[0015] Furthermore, the dialogue interaction state determination unit may have switching means for classifying or determining a mode based on an explicit mode selection operation by the user or on unstructured information (such as tone, keywords, and pauses) inherent in natural language utterances.

[0016] Furthermore, the personality profile includes tone, sentence endings, nuance, response speed, speaking tempo, speech style (polite, informal, etc.), point of view (first-person, second-person), emotional intonation patterns, dialogue rhythm, and dialogueal stance (advisory, empathetic, attentive, etc.), and each personality may be selected or switched depending on the state of the dialogue interaction. [Effects of the Invention]

[0017] According to this invention, the user's conversational state is classified based on the input language expression, an AI personality suitable for responding to the user with the classified conversational state is determined, and a response language expression generated by the determined AI personality is output. Therefore, the output response language expression is expressed in a style (writing style, expression method, communication method) that is suitable for the user's conversational state. Since the user receives a response language expression in a style that matches their conversational state, a conversation that fosters deep empathy and trust can be realized. Furthermore, while the invention described in Patent Document 1 shares a cheerful movement pattern with a partner, as described in Patent Document 1, paragraph

[0032] of Patent Document 1 only states that it recognizes the user's emotions, without specifying how it does so. In contrast, the present invention classifies the user's conversational state and is different from recognizing the user's emotions.

[0018] Not only the tone of the response, but also the entire personality "changes as if getting closer to the heart", enabling the user to feel empathy and a sense of being guided. Depending on the combination of the dialogue interaction state and the personality, various forms can be constituted, from an "AI that quietly watches" to an "AI that gives a push". With the flexibility to accept both operational switching / non-explicit dialogue classification, a co-creative experience is established.

Brief Description of the Drawings

[0019] [Figure 1] It is an overview of the response language expression output system. [Figure 2] It is a block diagram showing the electrical configuration of the response language expression output device. [Figure 3] It shows the pre-trained model groups for dialogue state classification and the pre-trained model groups for AI personality-specific responses. [Figure 4] It shows the pre-trained model group for dialogue state classification. [Figure 5] It shows the determination structure of the mind-calming mode. [Figure 6] It shows the determination structure of the pure-heart mode. [Figure 7] It shows the determination structure of the kind-heart mode. [Figure 8] It shows the determination structure of the heartless mode. [Figure 9] It is a combination of regions M1 to M4 into one. [Figure 10] It shows the pre-trained model group for AI personality-specific responses. [Figure 11] It is an example of a table showing the correspondence between the judgment elements of language expressions and the dialogue state. [Figure 12] It is an example of a table showing the characteristics of language expressions corresponding to the user's dialogue state. [Figure 13] It is a flowchart showing the processing procedure of the response language expression output device. [Figure 14] It is a flowchart showing the classification processing procedure of the dialogue state. [Figure 15] It is a flowchart showing the response language expression generation processing procedure. [Figure 16] This shows the input language expression and the response language expression. [Figure 17] This shows the input language expression and the response language expression. [Figure 18] This shows the input language expression and the response language expression. [Figure 19] This flowchart shows part of the processing procedure for the response language expression output device. [Figure 20] This is an example of a history information table. [Figure 21] Overall Structure Diagram [Figure 22] Dialogue interaction state determination ~ personality switching processing flow [Figure 23A] Personality profile diagram (examples of personalities corresponding to each dialogue interaction state) [Figure 23B] Dialogue interaction state × Personality selection flow [Figure 24] Components of personality (detailing of sentence endings, point of view, tempo, etc.) [Figure 25] This example illustrates the flow of this embodiment. [Figure 26] This describes the characteristics of the state during the dialogue. [Modes for carrying out the invention]

[0020] Figure 1 shows an embodiment of this invention and is an overview of the response language expression output system.

[0021] The response language expression output system comprises a response language expression output device 1 and an AI (Artificial Intelligence) server 20, which can communicate with each other via the internet.

[0022] In this embodiment, the AI ​​server 20 generates a response language expression, which is a response to the language expression input by the user. However, the response language expression may also be generated by the response language expression output device 1. For this reason, the AI ​​server 20 includes a group of trained AI personality-specific response models 40 (see Figures 3 and 5).

[0023] Figure 2 is a block diagram showing the electrical configuration of the response language expression output device 1.

[0024] The overall operation of the response language expression output device 1 is controlled by the CPU 2.

[0025] The response language expression output device 1 includes memory 5 for temporarily storing data. The response language expression output device 1 also includes a GPU (Graphics Processing Unit) 3 (an example of a means for classifying the state of dialogue, determining the AI ​​personality, and adjusting), and the GPU 3 is used to classify the state of the dialogue. The CPU 2 (an example of an acquisition means) instructs the AI ​​server 20 to generate a response language expression, but the GPU 3 may perform AI learning and generate the response language expression using a trained model. A display device 4 (an example of an output means) that displays the generated response language expression, etc., is connected to the GPU 3 via an interface (not shown).

[0026] The response language expression output device 1 includes a PCH (Platform Controller Hub) 6, which controls data communication between the CPU 2 and the communication device 7 that communicates with the AI ​​server 20, a microphone 8 (an example of a language expression input means), a speaker 9 (an example of an output means), a keyboard 10 (an example of a language expression input means), an SSD (Solid State Drive) 11 (an example of a storage device), a CD (Compact Disc) drive 12, etc. The SSD 11 stores various trained models and other data. A CD 13 containing a program that controls the operation of the response language expression output device 1 is inserted into the CD drive 12, and the program is read from the CD 13 and installed in the response language expression output device 1. Alternatively, the program may be received via the internet and installed in the response language expression output device 1.

[0027] Furthermore, the system may be equipped with a scanner (an example of a language expression input means) for reading language expressions written on paper or other materials, so that the language expression is read by the scanner and the response language expression output device 1 recognizes the language expression by OCR (Optical Character Recognition).

[0028] Figure 3 shows the group of trained models 30 for dialogue state classification stored in SSD11.

[0029] The pre-trained models 30 for classifying dialogue states are used to classify the user's dialogue state from the language expressions input by the user to the response language expression output device 1 (language expressions represented by voice input from the microphone 8, language expressions represented by data input from the keyboard 10, etc.).

[0030] A group of trained AI models 40 for responding to different personalities may be stored in the SSD 11.

[0031] The AI ​​personality-specific response trained model group 40 generates response language expressions (language expressions that represent responses to language expressions from the user as if in a conversation with the user) using AI personalities suitable for responding to users with classified (or specified) conversational states. The AI ​​personality-specific response trained model group 40 is for the response language expression output device 1.

[0032] Figure 4 shows the group of trained models 30 used for classifying states in dialogue.

[0033] The group of pre-trained models 30 for classifying dialogue states includes a pre-trained model for classifying dialogue states (context) 31, a pre-trained model for classifying dialogue states (intonation) 32, a pre-trained model for classifying dialogue states (endings) 33, and a pre-trained model for classifying dialogue states (intervals) 34. The pre-trained model for classifying dialogue interaction states (context) 31 was obtained by inputting the contexts of a large number of sentences as training data (teaching data) and learning in advance which dialogue state from among the dialogue states of balanced mind (e.g., a state of mind and body that is balanced and can concentrate), pure mind (e.g., a clear dialogue state that is free from distracting thoughts and worldly desires) or compassionate mind (e.g., a loving and affectionate dialogue state). As shown in Figure 6 later, the correspondence between different contexts and dialogue states is generally fixed. Therefore, by training the model with many sentences representing the contexts used for "straightforward mind," "pure mind," and "compassionate mind," a pre-trained dialogue interaction state classification model (for context) 31 can be generated. By using this pre-trained dialogue interaction state classification model (for context) 31 to perform classification, the user's dialogue state can be classified from the context of the linguistic expressions provided by the user. In this embodiment, the correspondence between judgment elements such as word endings and intonation and the user's dialogue state is obtained by collecting a large amount of dialogue data in advance and statistically analyzing it together with labels assigned based on the subject's self-reporting and expert annotations. This allows for a quantitative understanding of, for example, which dialogue states are associated with specific word endings or utterance intervals. By training the machine learning model with this relationship as training data, it becomes possible to classify the user's dialogue interaction state even for unknown inputs.

[0034] Similarly, for each of the other pre-trained models 32-34 for classifying dialogue states, by pre-training them using training data for each judgment element (feature) of language expression, it is possible to classify the dialogue state of the user who provided the input language expression based on the judgment elements of that language expression. The pre-trained model for classifying dialogue states (for intonation) 32 classifies the user's dialogue state based on the intonation, rhythm, etc., of the language expression from the input user. The pre-trained model for classifying dialogue states (for word endings) 33 classifies the user's dialogue state based on the word endings of the language expression from the input user. For example, if the word ending is "desu" or "masu," the user's dialogue state is considered to be calm; if the word ending is "masu nee" or "deyoude yo," the user's dialogue state is considered to be pure; and if the word ending is "shite kudasai ne," "shimasen ka?", "desu yo," or "desu yo ne," the user's dialogue state is considered to be compassionate, and so they are classified accordingly. The pre-trained model for dialogue interaction state classification (for utterance intervals) 34 classifies the user's dialogue state using the utterance intervals of the linguistic expressions input by the user. A slightly longer pause suggests the user's dialogue state is considered calm, a longer pause suggests a pure state, and a shorter pause suggests a compassionate state; and so the models are classified accordingly.

[0035] In this embodiment, the dialogue states are balanced mind, pure mind, and compassionate mind, but other dialogue states may be classified, or response language expressions may be generated by an AI personality suitable for users with other dialogue states. The elements for determining the dialogue state are not limited to context, but other elements may also be used.

[0036] The user's dialogue state is classified from all the classification results of the pre-trained dialogue state classification models (context) 31, (tone) 32, (ending) 33, and (interval) 34. Alternatively, instead of classifying the user's dialogue state from the classification results of four pre-trained models such as 31 (contextual), 32 (intonation), 33 (endings), and 34 (interval), a single pre-trained model for dialogue state classification can be created by combining all or more of these pre-trained models, and the classification result of this single pre-trained model for dialogue state classification can be used as the user's dialogue state. In this case, a large number of linguistic expressions are input as training data, and these numerous linguistic expressions are pre-trained for each judgment element or the linguistic expressions themselves, thereby generating a pre-trained model that shows which linguistic expressions correspond to which dialogue states. Furthermore, as will be described later, each user may have a set of pre-trained models 30 for classifying the interactive state specific to that user.

[0037] In this embodiment, the user's dialogue state is classified not only based on semantic information of the content of utterances (an example of linguistic expression from the user), but also based on the correlation of multiple non-semantic features of the utterances, namely, the utterance interval (a feature indicating the length of thought or silence) Δt (seconds), the rate of change in pitch at the end of sentences (voice tone and intonation tendency at the end of words) Pe (Hz / s), and the rate of change in amplitude (utterance intensity) Ae (dB / s). Here, the relational temperature score T is a dialogue interaction intensity index calculated from a combination of external features such as the user's utterance length, utterance interval, word endings, and intonation, which are calculated based on these non-semantic features, and is used as auxiliary information for determining the dialogue interaction state. The user's dialogue interaction state can be classified based on non-semantic information such as context, tone, word endings, and utterance intervals contained in the linguistic information from the user. For example, a user's dialogue interaction state can be classified into one of the following: "high activity pattern," "lively pattern," "outward-facing speech pattern," "calm speech pattern," "normal tempo," "casual tempo," "mostly silent pattern," "low activity pattern," or "static low-response pattern." If a user's linguistic expressions end with "~da ze," have short syllable lengths, have a rhythmic feel, have an utterance interval of approximately 0.1 to 0.2 seconds, and have an utterance frequency of 20 to 30 turns per minute, the user's dialogue interaction state is likely to be classified as a "high activity pattern." The same applies to the other dialogue interaction states. For example, for each of these nine dialogue interaction states, a score from 0 to 1 corresponding to the state from static low-response to high activity is defined as the relational temperature score T. When these features simultaneously satisfy specific conditions, the system determines the corresponding dialogue interaction state (e.g., calm, peaceful, compassionate, detached). In this way, it is also possible to classify the user's conversational state using the non-semantic features (the parts of the linguistic expression that exclude meaningful content words) from the user's linguistic expression.

[0038] Specifically, if the relational temperature score T decreases, the utterance interval Δt is longer than the average value Δta, and the sentence-end pitch drop rate Pe falls below a predetermined threshold Pt, the user's speech state is judged to be calm and their thinking is judged to be a static, low-response pattern, and they are classified as the calm mind mode. On the other hand, if the relational temperature score T is high, the utterance speed is faster than the average value, and the sound pressure change Ae is on an upward trend, it is judged that a highly active, organized and structured thinking pattern is taking place, and they are classified as the organized mind mode. Furthermore, if the relational temperature score T is at an intermediate value, the utterance speed and pitch change are gradual, and the sound pressure change rate is low, they are classified as a gentle speech pattern and judged as the compassionate mind mode.

[0039] Thus, the classification of dialogue interaction states is not based on a single threshold, but rather on a conditional structure of multiple non-semantic features (Equations 1 to 3). T <TtかつPe<PtかつΔt> Δtt...Equation 1 T>Tt and Ae>At... Formula 2 If T ≈ Tt and Pe and Ae are in the stable region... Equation 3 For example, if equation 1 is satisfied, the user is in the "Pure Heart" mode; if equation 2 is satisfied, the user is in the "Energized Heart" mode; and if equation 3 is satisfied, the user is in the "Compassionate Heart" mode. Here, Tt, Pt, At, and Δtt are dynamic thresholds that are updated sequentially based on the individual user's history or environmental conditions. This allows for the learning-based reflection of the user's unique speaking style and emotional tendencies, continuously improving the accuracy of dialogue interaction state classification. At this time, the mode determination signal S_mode is output as a logical value "1", and the corresponding personality switching trigger F_switch is turned on.

[0040] Figures 5 to 9 show a three-dimensional space with relational temperature score T, utterance interval Δt, and sentence-end pitch drop rate Pe as axes, respectively, illustrating a dialogue interaction state determination structure based on the correlation between relational temperature score T, utterance interval Δt, and sentence-end pitch drop rate Pe. Each mode is determined based on equations 1 to 3.

[0041] Figure 5 shows the determination structure for the harmonious mode; entering region M1 results in the harmonious mode. For example, the relational temperature score T is high, the utterance interval Δt is short, and the sentence-end pitch drop rate Pe is on an upward trend. Figure 6 shows the determination structure for the peaceful mode; entering region M2 results in the peaceful mode M2. For example, the relational temperature score T is low, the utterance interval Δt is long, and the sentence-end pitch drop rate Pe is on a downward trend. Figure 7 shows the determination structure for the compassionate mode; entering region M3 results in the compassionate mode. For example, the relational temperature score T is moderate, and the sentence-end pitch drop rate Pe and sound pressure change Ae are stable. Figure 8 shows the determination structure for the unconscious mode; entering region M4 results in the unconscious mode. For example, the relational temperature score T is low, the utterance interval Δt is extremely long, and the sentence-end pitch drop rate Pe is extremely short. Figure 9 combines these regions M1 to M4 into one.

[0042] Thus, for example, if the relational temperature score T is low, the utterance interval Δt is long, and the sentence-end pitch drop rate Pe is small, it is determined to be M2 (calm state). On the other hand, if T is high, Δt is short, and Pe is large, it is determined to be M1 (balanced state). With this configuration, the system can define the dialogue interaction state as a region in the correlation space of multiple features, rather than relying on a single threshold, and it becomes possible to dynamically classify the user's psychological state from the structure of non-semantic signals.

[0043] Figure 10 shows the group of trained AI personality-specific response models 40 stored in the AI ​​server 20.

[0044] The AI ​​personality-specific pre-trained model group 40 includes the organized AI personality response pre-trained model 41, the listening AI personality response pre-trained model 42, and the kind AI personality response pre-trained model 43. Other AI personality response pre-trained models may also be generated. The pre-trained model 41 is a pre-trained model used when the user's conversational state is calm and composed. The listening-type AI personality response pre-trained model 42 is a pre-trained model used when the user's conversational state is peaceful and composed. The kind-type AI personality response pre-trained model 43 is a pre-trained model used when the user's conversational state is compassionate and caring. In the organized AI personality response pre-trained model 41 as well, the appropriate response language expressions to output when the user's conversational state is calm and composed have been pre-trained using a large amount of training data. As shown in Figure 7 described later, the appropriate response language expressions to output when the user's conversational state is calm and composed are generally determined, so by classifying them using the organized AI personality response pre-trained model 41, it is possible to output appropriate response language expressions when the user's conversational state is calm and composed. For example, it has been revealed that specific endings and tones correspond to specific conversational states by statistically analyzing a large amount of conversational data with self-reported responses from subjects and expert annotations. Such relationships are collected as training data, and the trained model can perform similar classifications even for unknown inputs. Similarly, for the other models 42-43, by training them in advance using training data for each dialogue state, the input language expression can be classified and output as a response language expression appropriate to the dialogue state. The listening-type AI personality response trained model 42 is a trained model used when the user's dialogue state is pure, and outputs a response language expression that listens attentively to the user's language expression. The kindness-type AI personality response trained model 43 is a trained model used when the user's dialogue state is compassionate, and it classifies and outputs a response language expression that resonates with the user's language expression.

[0045] Depending on the user's dialogue state, one of the pre-trained models—the organized AI personality response pre-trained model 41, the listening AI personality response pre-trained model 42, or the kind AI personality response pre-trained model 43—is used to modify the response language expression. In Figure 5, the AI ​​personality-specific response pre-trained model group 40 is composed of separate pre-trained models 41 to 43. However, a single dialogue state-specific response pre-trained model may be generated from all or more of these pre-trained models 41 to 43 to output a response language expression corresponding to the user's dialogue state. In this case, a large number of language expressions representing various types of dialogue states are pre-trained as training data, and a pre-trained model is generated in advance that outputs a response language expression corresponding to the dialogue state that can be determined from the input language expression, depending on what kind of response language expression is output.

[0046] Figure 11 is an example of a table showing the relationship between the decision elements contained in the linguistic expressions provided by the user and the user's state in the dialogue.

[0047] In this embodiment, the linguistic expression provided by the user is analyzed based on context, tone, ending, and utterance. The user's conversational state is classified using judgment factors (features) such as intervals, but other judgment factors may also be used to classify the user's conversational state. For example, the user's conversational state may be classified based on the linguistic expression provided by the user, excluding content words (nouns, verbs, adjectives, adverbs, etc., which convey substantial meaning in a sentence and play a major role in information transmission), function words (conjunctions, articles, auxiliary verbs, prepositions, pronouns, etc., which mainly perform grammatical functions), non-verbal elements such as tone, speed, intonation, rhythm, pauses, volume, and quality of voice, paralinguistic information such as emojis and symbols in written communication, and linguistic expression elements including word length, nuance, utterance intervals, and utterance frequency.

[0048] For example, a user's linguistic expression is classified as "well-structured" if it clearly states the purpose and main point at the beginning, develops logically without unnecessary details, has a stable intonation in the mid-to-high range with little variation in stress, emphasizes conclusions with "desu" and "masu" at the end of sentences, and has slightly longer pauses between utterances, ranging from 0.8 to 1.2 seconds. Other dialogue states can be classified in a similar manner.

[0049] The table shown in Figure 11 represents the case where the user provides linguistic expressions to the response linguistic expression output device 1 via voice from the microphone 8. However, even when the user provides linguistic expressions to the response linguistic expression output device 1 as text data using the keyboard 10, the user's conversational state can be similarly classified according to the context of the sentences typed on the keyboard 10's keypad, the rhythm (tone) of input, the endings of words, the intervals between inputs, etc.

[0050] In this embodiment, a table as shown in Figure 11 (only a portion is shown in Figure 11, and in reality a very large amount of data is stored) is stored in SSD 11. When the contexts that classify the dialogue states of "settled mind," "pure mind," or "compassionate mind" stored in this table are included in the language expression from the user, the system is pre-trained using a large amount of training data to classify the user's dialogue state as "settled mind," "pure mind," or "compassionate mind," respectively, and is stored in SSD 11 as the pre-trained dialogue state classification model (for context) 31 described above. Similarly, for other judgment elements such as tone, word endings, or utterance intervals, when they are included in the language expression from the user, the system is pre-trained to classify the user's dialogue state as "settled mind," "pure mind," or "compassionate mind," respectively, and is stored in SSD 11 as the pre-trained dialogue state classification model (for tone) 32, pre-trained dialogue state classification model (for word endings) 33, and pre-trained dialogue state classification model (for utterance intervals), respectively, and is stored in SSD 11 as described above.

[0051] Even if the table itself, as shown in Figure 11, is not stored in SSD 11, if the pre-trained models for dialogue state classification (context) 31, dialogue state classification (intonation) 32, dialogue state classification (endings) 33, and dialogue state classification (interval) are stored in SSD 11 or elsewhere, the user's dialogue state can be classified based on the linguistic expressions from the user. It is not necessary to use the table, and mental states can be determined using a table like the one shown in Figure 116 without using the pre-trained models 31, etc.

[0052] When the user's conversational state is classified as "calm mind," "pure mind," or "compassionate mind," a response language expression suitable for the user's conversational state is generated using the pre-trained model 41 for organized AI personality responses, the pre-trained model 42 for attentive AI personality responses, or the pre-trained model 43 for kind AI personality responses. Since the response matches the user's conversational state, it is possible to realize a dialogue that corresponds to the user's conversational interaction style. Instead of the table shown in Figure 6, response language expressions corresponding to the user's conversational state can be generated and output using these pre-trained models 41-43.

[0053] Figure 12 is an example of a table (an example of an AI personality profile) that shows the characteristics of linguistic expressions corresponding to the user's conversational state.

[0054] For example, when the user's conversational state is "calm," the response language output will have a clear, positive tone, emphasize the conclusion by ending sentences with "let's do it," keep the phrasing short and clear, have a slightly faster tempo, and adopt a perspective that considers things together with the user. When the user's conversational state is "pure-hearted," the response language output will have a softer tone, listen attentively to the user's language by ending sentences with "right," have a relaxed phrasing, a slow tempo, and adopt a perspective that observes the user's viewpoint. When the user's conversational state is "compassionate," the response language output will have a gentle, resonant tone, gently resonate by ending sentences with "that's right," have drawn-out syllables, have a paused tempo, and adopt a perspective that is empathetic to the user. Figure 12 shows only a part of the data, and a very large amount of data is stored. In addition, response speed, speech emotion intonation patterns, and dialogue rhythm may also be included as features. The table shown in Figure 12 defines an AI personality profile for each AI personality. However, it is not necessary to use the table; any method that can determine the AI ​​personality is sufficient.

[0055] As described above, pre-trained models 41 for organized AI personality responses, 42 for attentive AI personality responses, and 43 for kind AI personality responses are generated in advance to output response language expressions as shown in Figure 12. Alternatively, instead of using the pre-trained models 41, etc., a table like the one shown in Figure 12 may be used to modify the response language expressions to correspond to the determined AI personality, thereby generating response language expressions corresponding to the AI ​​personality.

[0056] Figure 13 is a flowchart showing the processing procedure of the CPU 2 of the response language expression output device 1.

[0057] When the user speaks into the microphone 8 (or enters text using the keyboard 10, etc.), the user's linguistic expression is input to the response language expression output device 1 (step 51). Next, it is determined whether to generate a response language expression in the conversational state specified by the user (step 52). For example, the display screen of the display device 4 may display the string "Do you want the AI ​​to respond in the conversational state you have specified?" and ask the user to respond. A question area with the string "Do you want the AI ​​to respond in the conversational state you have specified?", as well as answer areas with the string "Yes" and the string "No", may be displayed. If the user wants the AI ​​to respond in the conversational state they have specified (generate a response language expression), they should input "Yes". They may also be asked to select the answer area labeled "Yes" (selection using the keypad on the keyboard 10, clicking with a mouse (not shown): this is an example of a selection method). If the user wants the AI ​​to classify the user's conversational state and respond to it, they should input "No". They may also be asked to select the answer area labeled "No". Selection using microphone 8, selection by entering text, etc., are also acceptable methods.

[0058] If the user specifies a conversational state and the AI ​​is to generate a response language expression (YES in step 52), the display screen of the display device 4 will display a question asking "Which is it: 'Rectified mind,' 'Pure mind,' or 'Compassionate heart'?" and the user will be asked to answer (step 53). The display screen of the display device 4 will display conversational state areas containing the strings "Rectified mind," "Pure mind," and "Compassionate heart," and the user may be asked to select their conversational state using the keyboard 10 (an example of a means for specifying a conversational state). Other conversational states may also be displayed, or the user may be asked to select a conversational state using voice or other means. If the user is asked to classify the user's conversational state and respond to it (NO in step 52), the CPU 2 will have the GPU 3 classify the user's conversational state from the language expression input by the user (step 54). This classification process will be described in more detail later (see Figure 9).

[0059] Next, the CPU2 (an example of an AI personality determination means, control means) determines an AI personality that generates response language expressions from a specified dialogue state or a classified dialogue state (step 55). If the dialogue state is "calm mind," an "organized" AI personality is determined; if the dialogue state is "pure mind," an "attentive" AI personality is determined; and if the dialogue state is "compassionate," a "kind" AI personality is determined (see Figure 12).

[0060] Once an AI personality is determined, the AI ​​server 20 is instructed to generate a response language expression for the input language expression using the determined AI personality (step 56). For example, if the dialogue state is "calm mind," the AI ​​personality is determined to be "organized," so the trained model 41 for organized AI personality responses stored in the AI ​​server 20 is used; if the dialogue state is "pure mind," the AI ​​personality is determined to be "listening," so the trained model 42 for listening AI personality responses stored in the AI ​​server 20 is used; and if the dialogue state is "compassionate," the AI ​​personality is determined to be "kind," so the trained model 43 for kind AI personality responses stored in the AI ​​server 20 is used. The CPU 2 sends a command to the AI ​​server 20 in this manner. The response language expression generated in the AI ​​server 20 is sent to the response language expression output device 1.

[0061] The response language expression sent from the AI ​​server 20 is output from the response language expression output device 1. Step 57).

[0062] Figure 14 is a flowchart (Step 54 in Figure 13) showing the processing procedure for classifying the user's conversational state from the input language expression. This is performed by GPU3.

[0063] The language expression input to the response language expression output device 1 is classified using a pre-trained model for dialogue state classification (for context) 31 to classify the dialogue state indicated by the ending of the language expression (step 71). The dialogue state indicated by the tone of the language expression is then classified using a pre-trained model for dialogue state classification (for tone) 32 (step 72). Furthermore, the dialogue state indicated by the ending of the language expression is classified using a pre-trained model for dialogue state classification (for endings) 33 (step 73). The dialogue state indicated by the utterance interval of the language expression is then classified using a pre-trained model for dialogue state classification (for utterance interval) 34 (step 74). Subsequently, the user's dialogue state is comprehensively classified from all or some of the obtained classification results (step 66).

[0064] Figure 15 is a flowchart showing how to generate response language expressions with a determined AI personality. be.

[0065] Data representing the language expression provided by the user and the determined AI personality are transmitted from the response language expression output device 1 to the AI ​​server 20 (step 81).

[0066] When the AI ​​server 20 receives the data representing the language expression and AI personality transmitted from the response language expression output device 1, the AI ​​server 20 uses the trained models 41-43 included in the group of trained AI personality response models 40 stored in the AI ​​server 20 to generate a response language expression for the language expression represented by the data representing the transmitted language expression (step 84). The trained model corresponds to the AI ​​personality represented by the transmitted AI personality data. The data to be represented is transmitted from the AI ​​server 20 to the response language expression output device 1 (step 85).

[0067] When data representing a response language expression transmitted from the AI ​​server 20 is received by the response language expression output device 1, the GPU 3 classifies whether the response language expression represented by the transmitted data is compatible with the determined AI personality, and the CPU 2 makes a determination. Similar to the dialogue state classification process for language expressions input by the user as shown in Figure 9, the dialogue state represented by the received response language expression is classified, and the AI ​​personality can be determined from the classified dialogue state. The CPU 2 then determines whether the determined AI personality is compatible with the determined AI personality (the AI ​​personality determined in step 55 of Figure 13). If it is compatible (YES in step 82), the response language expression generated by that AI personality is considered appropriate for the user's dialogue state and is output to the user from the response language expression output device 1. If it is not compatible (NO in step 82), a correction request is sent to the AI ​​server 20 to generate a response language expression that is compatible with the determined AI personality (step 83).

[0068] If a correction request from the response language expression output device 1 is received by the AI ​​server 20 within a certain time after the transmission of the response language expression (YES in step 86), the CPU 2 of the AI ​​server 20 outputs an instruction to regenerate the response language expression to the classification unit of the AI ​​server 20, which performs classification using the trained response model (step 87). If the response language expression corresponding to the determined AI personality is generated using some of the features such as nuance and tone used in the trained response model corresponding to the determined AI personality, the trained response model corresponding to the determined AI personality should be modified by increasing the types of features used to generate the response language expression (for example, a trained response model generated using only nuance and tone should be modified to generate a trained response model corresponding to the AI ​​personality that also uses features other than nuance and tone). Alternatively, multiple pre-trained response models for each AI personality may be generated with varying degrees of AI personality strength. Initially, a pre-trained model that represents a weaker AI personality (a pre-trained model with fewer features representing that AI personality) may be used to generate the response language expression. When a correction request is received, a pre-trained model that represents a stronger AI personality (a pre-trained model with many features representing that AI personality) may be used to generate the response language expression. The corrected response language expression is generated (step 84), and the response language expression is again sent from the AI ​​server 20 to the response language expression output device 1 (step 85). The response language expression output device 1 then determines whether the received response language expression is suitable for the AI ​​personality (step 82). As the response language expression is corrected, a more appropriate response language expression can be obtained for the determined AI personality. If a correction request is not received within a certain time after the transmission of the response language expression (NO in step 86), the response language expression will not be corrected.

[0069] In the above embodiment, if the response language expression generated by the AI ​​server 20 does not match the determined AI personality, a correction request is sent to the AI ​​server 20 to generate a response language expression that matches it. However, the response language expression output device 1 may also modify the language expression portion of the response language expression generated by the AI ​​server, excluding content words such as word endings, to adjust it to match the determined AI personality. Since there is no need to send a correction request to the AI ​​server 20 and have the AI ​​server 20 make corrections, it becomes possible to output a response language expression that is suitable for responding to the user's conversational state relatively quickly.

[0070] Figures 16 to 18 show examples of linguistic expressions provided by the user and response linguistic expressions generated by an AI personality that matches the user's conversational state.

[0071] Figure 16 shows an example of a case where the user's conversational state is classified as "calm and focused."

[0072] Let's assume the user inputs the phrase "My head's all jumbled up..." into the response language expression output device 1. When this phrase from the user is input into the trained models 31-34, which are included in the trained model group 30 for classifying dialogue states shown in Figure 4, a classification process is performed using the characteristics of the language expression, and the user's dialogue state is classified as "organized" (see also Figure 11). Since the user's dialogue state is "organized," the AI ​​personality is determined to be "organized." The AI ​​server 20 performs a classification of the phrase "My head's all jumbled up..." using the trained model 41 for organized AI personality responses, and a response language expression that thinks along with the user, "Let's organize it little by little together," is generated. The generated response language expression is output from the response language expression output device 1. The user can resonate with the response language expression and organize their thoughts while being guided by the response language expression output device 1.

[0073] For example, if the response language expression output device 1 determines that the response language expression "Shall we organize it little by little together?" generated by the AI ​​server 20 does not conform to the state of "organization" in the dialogue, the AI ​​server 20 will modify it as described above, and the modified response language expression will be sent to the response language expression output device 1. Also, if the response language expression output device 1 determines that changing the ending of this response language expression (the language expression part excluding the content word), "Shall we do it?", to "Shall we?" would conform to the state of "organization" in the dialogue, the response language expression output device 1 may modify and output it in that way.

[0074] Figure 17 shows an example of a case where the user's conversational state is classified as "calm and composed."

[0075] Let's assume the user inputs the following linguistic expression to the response language expression output device 1: "I've been tired lately... but I can't put it into words." When this linguistic expression from the user is input to the trained models 31-34 included in the trained model group 30 for classifying dialogue states shown in Figure 4, a classification process is performed, and the user's dialogue state is classified as calm (see also Figure 6). Since the user's dialogue state is calm, the AI ​​personality is determined to be listening type. The AI ​​server 20 performs classification on the linguistic expression "I've been tired lately... but I can't put it into words" using the trained model 42 for listening type AI personality responses, and a response language expression that looks out for the user, "...I see. That's right," is generated. The generated response language expression is transmitted to the response language expression output device 1 and output. In this case as well, the user resonates with the response language expression and is guided to the response language expression output device 1. In this case, the response language expression does not try to force a conversation but listens to the user's words in order to elicit the user's words, so the user becomes able to solve the problem themselves.

[0076] Figure 18 shows an example of a case where the user's conversational state is classified as compassionate.

[0077] Let's assume the user inputs the following linguistic expression to the response language expression output device 1: "I failed again, and I'm disappointed in myself..." When this linguistic expression from the user is input to the trained models 31-34 included in the trained model group 30 for classifying dialogue states shown in Figure 4, a classification process is performed, and the user's dialogue state is classified as compassionate (see also Figure 11). Since the user's dialogue state is compassionate, the AI ​​personality is determined to be kind. The AI ​​server 20 performs classification on the linguistic expression "I failed again, and I'm disappointed in myself..." using the trained model 43 for kind AI personality responses, and a response language expression that empathizes with the user, "But you still faced it head-on. That's great," is generated. In this case as well, the user resonates with the response language expression and is guided to the response language expression output device 1. Through such response language expressions, the AI ​​empathizes with the user's dialogue state, and the user gains confidence.

[0078] A history information table may be created to improve the accuracy of classifying the state of dialogue for each user by remembering the context, tone, endings, and utterance intervals used in past language expressions.

[0079] Figure 19 is a flowchart showing part of the processing procedure of the response language expression output device 1, and is intended to improve the accuracy of mental state classification. Step 57 of the flowchart shown in Figure 13 This shows an example of a process that may be followed as needed after the initial processing.

[0080] When the response language expression generated in step 57 of Figure 13 is output, it is checked whether the AI ​​personality was determined using the dialogue state specified by the user and whether the response language expression was generated (step 58). If the dialogue state specified by the user was not used (NO in step 58), it means that the AI ​​personality was determined using the dialogue state classified from the language expression from the user and the response language expression was generated. For this reason, the classified dialogue state may not be accurate. The correspondence between the characteristics (judgment elements) of the input language expression and the dialogue state is stored in the user's history information table (step 59).

[0081] Figure 20 shows an example of a history information table specific to a particular user. The history information table is managed by a user ID unique to that user. The user ID is requested when logging in using the response language expression output device 1, and the history information table corresponding to the user ID is updated. The history information table is stored in SSD 11 (an example of a storage device).

[0082] The information stored in the history information table is identified by an identification number (No.) and stored in chronological order. For example, in the case of identification number No. 1, the endings of the judgment elements included in the language information were "desu" and "masu," so the user's conversational state is classified (or designated) as "calm and composed." In the case of identification number No. 2, the endings of the judgment elements included in the language information were "desu ne" and "masu ne," so the user's conversational state is classified (or designated) as "soothing and composed," and both are stored in the history information table.

[0083] Next, the display screen of the display device 4 displays the strings for the interactive states "Setting the mind," "Pure mind," and "Compassionate mind," and the user specifies their own interactive state (step 60). The CPU 2 (an example of a determination means) checks whether the user's interactive state specified by the user matches the interactive states classified in step 54 of Figure 13 (step 61). If they do not match, the history information table shown in Figure 15 is updated under the control of the CPU 2 (an example of a memory control means), and the relationship between the characteristics of the input language expression and the interactive state is stored in the history information table, and the history information table is updated (step 62). If they match, the process in step 62 is skipped, but even if they match, the information may still be stored in the history information table.

[0084] For example, as shown in identification number No. 3, if the suffix of the judgment element included in the language information is "desu ne" or "masu ne," and the classified user's dialogue state is "Seishin," but the dialogue state specified by the user is "Seishin," or as shown in identification number No. 4, if the suffix of the judgment element included in the language information is "desu" or "masu," and the classified user's dialogue state is "Seishin," but the dialogue state specified by the user is "Seishin," then the classified dialogue state and the specified dialogue state do not match. Therefore, the relationship between the characteristics of the input language expression and the dialogue state is stored and the history information table is updated. When the information corresponding to identification numbers No. 3 and 4 is stored in the history information table, this user is in a dialogue state of "stressed" if the ending of the judgment element included in the language information is "desu ne" or "masu ne," and in a dialogue interaction state of "clear mind" if the ending of the judgment element included in the language information is "desu" or "masu." Thereafter, the parameters of the trained model 33 for dialogue state classification are adjusted by the GPU 3 (step 63) so that the dialogue state is classified as "stressed" if the ending of the judgment element included in the language information of this particular user is "desu ne" or "masu ne." Similarly, the parameters of the trained model 33 for dialogue state classification (an example of a state classification means that classifies the user's dialogue state based on non-semantic features such as word endings, intonation, interval between utterances, length of utterances, and tempo) are adjusted by the GPU 3 (an example of an adjustment means) so that the dialogue state is classified as "clear mind" if the ending of the judgment element included in the language information of this particular user is "desu" or "masu." This enables the classification of dialogue states tailored to specific users. A set of pre-trained models 30 for dialogue state classification will also be provided for each user. This will improve the accuracy of classifying the user's dialogue state. In this case, a set of pre-trained models 30 for dialogue state classification (for example, pre-trained models 31-34 for state classification that classify the user's dialogue state based on non-semantic features such as word endings, intonation, utterance intervals, utterance length, and tempo) will be prepared specifically for each user, and the user's dialogue state will be classified using the set of pre-trained models 30 for dialogue state classification that is tailored to that user.

[0085] In the above embodiment, classification processing is performed by GPU3, but NPU (Neural Processing Unit) may be used to perform some or all of the processing that GPU3 does. In addition, the AI ​​server 20 may generate response language expressions for input language expressions using a standard AI personality, and then modify the generated response language expressions using an AI personality that corresponds to the user's conversational state. Furthermore, some or all of the trained model group 30 for conversational state classification and the trained model group 40 for AI personality-specific responses may be stored in the memory of the AI ​​server 20, and the AI ​​server 20 may perform mental state classification, AI personality determination, and generation of response language expressions using the determined AI personality.

[0086] (Summary of the invention) The present invention provides a dialogue support system that structurally classifies the dialogue interaction state based on the tone of voice, intervals between utterances, response tendencies, and historical information that appear during the user's dialogue, and automatically switches the personality profile according to the mode.

[0087] This invention is applicable as an AI that empathizes with people's feelings in areas such as conversational AI, educational support, medical and nursing care support, mental care, home AI assistants, and customer support, and has broad industrial applicability.

[0088] The personality profile is composed of non-semantic elements such as sentence endings, nuances, speaking tempo, the speaker's perspective, and their stance (affirmative, supportive, suggestive, etc.). By switching between these elements, it becomes possible to respond in ways that suit the user's state, such as "gentle speaking," "listening silently," or "organizing." According to this invention, even without the user explicitly stating "this is what I want," it becomes possible to capture the state of the conversation from within the dialogue and change the personality accordingly, realizing an empathetic and dynamic dialogue experience.

[0089] Furthermore, by learning and accumulating the user's dialogue history, interaction state transition trends, and personality switching history, and by realizing individually optimized response styles and personality configuration switching, it becomes possible to build a continuous and trustworthy dialogue relationship while adapting to changes in the user's speech patterns.

[0090] (Modes for carrying out the invention) An example of an embodiment of the present invention will be described below with reference to the drawings. (Overall structure) As shown in Figure 21, the interactive state-adaptive personality switching control system of the present invention mainly consists of the following modules Composed of: 1. Input section: It receives natural language input (text or voice) or operational input (mode selection, etc.) from the user. 2. Dialogue interaction state determination unit: Based on the tone, endings, context, utterance intervals, or selected actions in the input, the system classifies or specifies the user's conversational state (e.g., calm, peaceful, compassionate, detached). 3. History Management Department: Past conversation history with the user, interaction state transition trends, and personality response history are accumulated and analyzed, and used as feedback information for the conversation interaction state determination unit and the personality selection unit. 4. Personality Selection Section: A personality profile corresponding to the determined dialogue interaction state is selected from a pre-configured set of profiles. The personality profile includes components such as tone, sentence endings, nuance, perspective, speech style, speaking tempo, and dialogueal position (listening type, empathetic type, advisory type, etc.). 5. Response generation unit: Based on the selected personality profile, natural language responses are generated. The responses are composed of content and tone that are in line with the user's conversational state, and include patterns such as frequent silences, questioning, and positive responses. 6. Response output section: The generated response is output to the user in the form of text, audio, or facial expressions / animations.

[0091] (Operation Flow) As shown in Figure 22, the control method of this system consists of the following steps: Step S1: Input Acquisition The input unit receives the user's speech or actions. Step S2: Determining the state of the dialogue interaction The state of dialogue interaction is classified based on tone, keywords, word order, history, etc. Step S3: Personality Selection The personality selection unit selects a personality profile that corresponds to the conversational interaction state. Step S4: Response generation Constructs natural language responses with tone, perspective, and rhythm that correspond to the personality. Step S5: Output The response output unit presents a response.

[0092] (Examples) For example, if a user says, "I want to relax today," the dialogue interaction state determination unit classifies it as a calm state, and the personality selection unit selects a "listening personality with a slow tempo and frequent use of positive sentence endings." Based on this, the response generation unit generates expressions such as "I see. That's right," and the output unit presents these as audio or text, resulting in a response that naturally harmonizes with the user's conversational state.

[0093] This invention includes the following configuration: Dialogue interaction state determination unit → Classify and specify dialogueal states (dialogue interaction states) such as calmness, serenity, compassion, and emptiness, based on speech and actions. Personality Selection Department → Select a personality profile (tone, sentence endings, perspective, tempo, etc.) that corresponds to the state of the dialogue interaction. Response generation unit → Generate natural language responses based on personality profiles History Management Department → Learns the user's conversation history and mode transition tendencies, and reflects this in classification accuracy and personality selection. Response output section → Provide responses via text, audio, etc.

[0094] (Examples) Mindfulness Mode × Organized Personality User: "My head is all jumbled up." AI: "Let's look at them together, little by little, in order." Calm Mind Mode × Attentive Listening Personality User: "I just want to be quiet." AI: "...I see. That's right." Compassionate Mode × Kind Personality User: "I've been making so many mistakes today..." AI: "But even so, you faced it head-on, didn't you? That's admirable."

[0095] This invention classifies abstract dialogueal states (dialogue interaction states) based on non-semantic features such as the user's utterance structure, sentence endings, intonation, and pauses, and structurally switches the personality of the response accordingly.

[0096] Conventional technologies have employed the following approaches: • Emotion-based response control (e.g., changes in response based on emotion classification such as anger / cheerfulness patterns) • Explicit mode-switching chatbot (e.g., select "gentle mode" or "strict mode") • Emotional classification and corresponding response conversion based on non-verbal signals such as acoustic characteristics and facial expressions. In contrast to this, the present invention: • It does not presuppose emotional classification, but instead adopts a structure that classifies dialogue interaction styles based on the rhythm of speech and changes in word endings, enabling responses to dialogueal states that cannot be expressed by classification-based processing. Rather than explicit mode selection, the system dynamically switches personality profiles based on user utterances, providing a more natural personality shift and conversational experience. The switchable personalities are not merely based on tone or manner of speaking, but are systematized using components such as "sentence endings," "sound," "tempo," "speaker's perspective," and "position," realizing for the first time a dialogue support interface in which personalities can be structurally designed and switched.

[0097] As a result, this invention goes beyond existing emotion-responding AI and style-switching chatbots, providing a co-creative dialogue platform that autonomously changes its personality while responding to the user's "emotional changes."

[0098] Figure 21 is an overall configuration diagram of the interactive state-adaptive personality switching control system according to the present invention.

[0099] This system receives input from the user (natural language or user interaction) and determines the dialogue interaction state (e.g., calm, peaceful, compassionate) based on its content. Furthermore, it selects an appropriate personality from several predefined personality profiles according to the determined dialogue interaction state and generates and outputs a response based on that. It also includes a history management unit that records and analyzes past dialogue history and interaction state transition trends, contributing to improved accuracy in dialogue interaction state determination and personality selection.

[0100] Figure 22 is a processing flow diagram of the interactive state-adaptive personality switching control method of the present invention.

[0101] This flow consists of the following steps: acquiring user input (step S1), classifying or confirming the dialogue interaction state (step S2), selecting a personality profile (step S3), generating a response according to the personality (step S4), and outputting the response to the user (step S5). This process enables dynamic switching to a dialogue style according to the user's state in the dialogue.

[0102] Figure 23A is a structural diagram showing the relationships between the components of the AI ​​personality profile selected in response to the user's dialogue interaction state.

[0103] This diagram shows, in matrix format, how personality components such as tone, sentence endings, nuance, tempo, perspective, and stance correspond to multiple dialogue interaction states, such as calmness, serenity, and compassion. This diagram clearly demonstrates that a personality profile is not a single expressive characteristic, but rather a complex structure composed of multiple expressive dimensions.

[0104] Figure 23B is a flowchart showing the processing procedure for determining the dialogue interaction state based on user input and reflecting the corresponding personality profile in the selection response generation.

[0105] The process, starting with input acquisition, followed by classification of the dialogue interaction state, selection of a personality profile, response generation, and response output, is shown sequentially, illustrating a structure in which dynamic personality switching control is performed systematically.

[0106] Figure 24 is a block diagram showing the components of the AI ​​personality profile according to the present invention.

[0107] Each personality profile is composed of elements such as tone, sentence endings, nuance, tempo, perspective, speech style, and stance, and combinations of these elements are switched according to the user's dialogue interaction state. This diagram defines personality as a "collection of impressions" and visually represents the decomposition of its elements to make it structurally controllable.

[0108] Figure 25 shows a series of examples in which the system classifies the dialogue interaction state in response to the user's natural language utterances and selects and responds with an AI personality accordingly.

[0109] This diagram shows that even if the user does not explicitly state their wishes in their utterance, the system will pick up on tone, sentence endings, pauses, and historical information. This process classifies the dialogue interaction state (calm, composed, compassionate, etc.) based on external features and dynamically selects and reflects a personality profile. For example, in response to a vague complaint such as "I've been tired lately...but I can't put it into words," the calm mode is classified, and an attentive personality (gentle tone, affirmative endings, paused tempo, watchful perspective, etc.) is automatically selected, generating an expression like "...I see. That's right." from the AI.

[0110] Furthermore, this diagram also includes examples of responses to the "calming mode" and "compassionate mode," clearly illustrating, in a comparative format, how the differences in dialogue interaction states are specifically expressed in the personality profile and generative responses. This structure demonstrates that a dialogue structure is realized that goes beyond mere emotional responses, instead empathizing with the user's state during the dialogue and building trust while changing personality.

[0111] The invention according to this embodiment is a technology for realizing "human-centered responses" in conversational AI, and in particular, it is a "personality-responsive" technology that changes the conversational style as a way of being, such as by switching the AI's personality itself (tone, perspective, tempo, speech style, position, etc.) according to the user's conversational state (dialogue interaction state), such as "adjusting," "quietly accepting," or "gently guiding."

[0112] Figure 26 shows the characteristics of the dialogue state. [Explanation of Symbols]

[0113] 1: Response language expression output device, 2: CPU, 3: GPU, 4: Display device, 5: Memory, 7: Communication device, 8: Microphone, 9: Speaker, 10: Keyboard, 12: CD drive, 13: CD, 20: AI server, 30: Pre-trained models for dialogue state classification, 31-34: Pre-trained models for dialogue state classification, 40: Pre-trained models for AI personality-specific responses, 41-43: Pre-trained models for AI personality responses

Claims

1. A language expression input means that inputs language expressions provided by the user. A dialogue state classification means that classifies the user's dialogue state based on non-semantic features such as word endings, intonation, utterance intervals, utterance length, and tempo, from among the language expressions input from the above language expression input means. AI personality determination means for determining an AI personality suitable for responding to an opponent having a dialogue state classified in the above dialogue state classification means, An acquisition means for acquiring a response language expression which is a response to a language expression input from the language expression input means, and which is a response language expression generated by the AI ​​personality determined in the above AI personality determination means, and Output means for outputting the response language expression obtained in the above acquisition means, A response language expression output device equipped with the following features.

2. The acquisition means is obtained by generating a response language expression, which is a response to a language expression input from the language expression input means, using the AI ​​personality determined by the AI ​​personality determination means, or by having a device other than the response language expression output device generate a response language expression, which is a response to a language expression input from the language expression input means, using the AI ​​personality determined by the AI ​​personality determination means. The response language expression output device according to claim 1.

3. Means for specifying the state of the dialogue, A determination means for determining whether the dialogue state classified by the above dialogue state classification means does not match the user's dialogue state. In response to the determination means that determines that they do not match, a storage control means controls the storage device to store, for each user, the characteristics of the language expression used for classification in the dialogue state classification means and the dialogue state specified by the dialogue state specification means, and An adjustment means adjusts the dialogue state classification means to classify the dialogue state stored in correspondence with the characteristics of the language expression stored in the memory device, in response to inputting a language expression having the same characteristics as the language expression stored in the memory device into the language expression input means. An output device for a response language expression according to claim 1, comprising:

4. Means for specifying the state of the dialogue, A selection means for selecting whether to determine an AI personality suitable for responding to an opponent having a dialogue state classified by the above-mentioned dialogue state classification means, or to determine an AI personality suitable for responding to an opponent having a dialogue state specified by the above-mentioned dialogue state specification means, and Control means for controlling the AI ​​personality determination means to determine an AI personality suitable for responding to An output device for a response language expression according to claim 1, comprising:

5. An AI personality profile is defined for each AI personality, and this AI personality profile stores at least one of the following: tone of voice, sentence endings, nuance of speech, response speed, speaking tempo, speech style, perspective, emotional intonation patterns, dialogue rhythm, and dialogue position. The acquisition method described above is: The system obtains a response language expression generated based on the AI ​​personality profile corresponding to the AI ​​personality determined by the AI ​​personality determination means described above. The response language expression output device according to claim 1.

6. The above means for classifying the state in the dialogue is, The system classifies the user's conversational state based on the linguistic expression portion of the linguistic expression provided by the user, excluding the content words that have meaning. The response language expression output device according to claim 1.

7. The language expression input means inputs language expressions provided by the user, The dialogue state classification means classifies the user's dialogue state based on non-semantic features such as word endings, intonation, utterance intervals, utterance length, and tempo, from the language expressions input from the language expression input means. The AI ​​personality determination means determines an AI personality suitable for responding to an opponent having a dialogue state classified by the dialogue state classification means, The acquisition means is a response language expression generated by the AI ​​personality determined in the AI ​​personality determination means, and the response language expression is acquired from the language expression input means, which is a response to the language expression input. The output means outputs the response language expression generated by the acquisition means. A response language expression output method equipped with a response language representation.

8. A computer-readable program for controlling the computer of a response language expression output device, The user provides a linguistic expression as input. Based on non-semantic features such as word endings, intonation, utterance intervals, utterance length, and tempo, the system classifies the user's conversational state from the input language expressions. Determine the AI ​​personality best suited to responding to an opponent with a classified conversational state. This is a response language expression generated by the determined AI personality, a response language expression that is a response to the input language expression, and the determined AI personality obtains the response language expression that is a response to the input language expression. A program that controls the computer of a response language expression output device to output the acquired response language expression.

9. A recording medium storing the program described in claim 8.

Citation Information

Patent Citations

  • Recognition device, learning device, method for same, and program

    WO2021166207A1

  • system

    JP2025047448A