Dialog system, dialog method, and dialog program

WO2026163524A1PCT designated stage Publication Date: 2026-08-06ARAKAWA YOHEI +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ARAKAWA YOHEI
Filing Date
2025-10-16
Publication Date
2026-08-06

Smart Images

  • Figure JP2025036536_06082026_PF_FP_ABST
    Figure JP2025036536_06082026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a dialog system that provides a dialog with a user, the dialog system enabling a dialog corresponding to the user's emotion. In order to solve the problem, the present invention provides a dialog system that analyzes the psychology of a user from a dialog with the user. The dialog system comprises an acquisition unit and a state generation unit. The acquisition unit acquires a plurality of types of modality signals obtained from the dialog with the user and user signal information including an acquisition timing of each of the modality signals. The state generation unit performs analytical processing on the acquired modality signals, performs, for each of the modality signals, generation processing on a psychological element probability vector indicating a subconscious structure having a plurality of different psychologies of the user and observation probabilities of the respective psychologies, performs integrated computing on the plurality of different psychologies of the user and the observation probabilities of the respective psychologies in the plurality of types of modality signals on the basis of the plurality of psychological element probability vectors having the same acquisition timing, and performs generation processing on a psychological probability integration vector indicating the subconscious structure based on the plurality of types of modality signals.
Need to check novelty before this filing date? Find Prior Art

Description

Dialogue systems, dialogue methods, and dialogue programs

[0001] The present invention relates to a dialogue system, a dialogue method, and a dialogue program using a terminal that provides dialogue with a user, and more particularly to a dialogue system that enables dialogue in response to the user's emotions.

[0002] Dialogue systems (dialogue agents) that take speech or natural language input, analyze this linguistic information, and provide responses are well known. Furthermore, with the spread of mobile devices and the advancement of AI processing technology, many dialogue systems utilizing AI mobile devices have also been proposed.

[0003] For example, Patent Document 1 discloses a dialogue system that enables a user (speaker) to engage in dialogue without feeling any discomfort. This system estimates the user's emotions based on expression information related to the user's expressions and determines the response characteristics of the AI ​​dialogue agent based on this estimation. Here, the expression information used is either audio information spoken by the speaker or image or video information of the speaker. The system creates a comment based on the response characteristics and integrates this comment into the dialogue text that will be the response.

[0004] Japanese Patent Publication No. 2024-164644

[0005] Incidentally, in fields such as mental healthcare, there is a trend that mechanical dialogue systems are more psychologically accessible than human interaction, but it is important that they are more empathetic and responsive to the user's emotions. However, these user emotions are often even more complex.

[0006] The present invention has been made in view of the above circumstances, and its purpose is to provide a dialogue system that enables dialogue in response to the user's emotions in a dialogue system that provides dialogue with a user.

[0007] The system according to the present invention is a dialogue system using an AI dialogue agent that utilizes non-verbal information such as images (facial expressions) and voice (tone of voice) in a multimodal manner.

[0008] The dialogue system according to the present invention is a dialogue system using a mobile terminal that provides a dialogue with a user. The mobile terminal receives language data that the user has input in characters or converted from input speech uttered, and uses a dialogue engine to generate response data and read it out as response speech by the reading means of the mobile terminal as conversation data. The system includes a process of capturing facial expression image data of the user who provides the language data to the mobile terminal and voice tone data during speech, extracting a single feature amount for the voice tone data and the facial expression image data, and displaying an image determined in advance corresponding to the feature amount on the mobile terminal according to an emotion table while reading out the response speech.

[0009] According to such a feature, it is possible to provide a dialogue according to the user's emotion, and moreover, it is possible to perform processing at a speed that does not cause a sense of discomfort in the dialogue.

[0010] In the above invention, the emotion table may be characterized in that an image is determined corresponding to each index of a plurality of emotion elements, and includes an emotion extraction engine that extracts the feature amount and extracts the corresponding index. According to such a feature, it is possible to provide a dialogue according to the user's emotion, and moreover, it is possible to perform processing at a speed that does not cause a sense of discomfort in the dialogue.

[0011] In the above invention, the image may be characterized as a moving image. According to such a feature, it is possible to provide a dialogue according to the user's emotion.

[0012] In the above invention, it may be characterized in that it includes voice correction means for correcting the response data generated by the dialogue engine based on the index and using it as the conversation data. According to such a feature, it is possible to provide a dialogue according to the user's emotion.

[0013] This is a block diagram showing one embodiment of the dialogue system according to the present invention. This is a diagram showing the flow of information by the dialogue system in one embodiment of the present invention. This is a diagram listing the specific processing flow in one embodiment of the present invention. This is an example of a block diagram showing the hardware configuration of the server of the dialogue system in another embodiment of the present invention. This is an example of a block diagram showing the hardware configuration of the terminal device of the dialogue system in another embodiment of the present invention. This is an example of a block diagram showing the functional configuration of the dialogue system in another embodiment of the present invention. This is an example of a flowchart of a dialogue method using the dialogue system in another embodiment of the present invention.

[0014] (Embodiment 1) Hereinafter, an interactive system as one embodiment of the present invention will be described with reference to Figure 1. However, it can be implemented in many different forms and is not limited to the embodiments described herein.

[0015] For example, in this embodiment, the configuration and operation of the dialogue system will be described, but similar configurations, devices, computer programs, etc., can achieve similar effects. Furthermore, the program may be stored on a recording medium. Using this recording medium, the program can be installed on a computer, for example, thereby configuring a dialogue device or dialogue system. Here, the recording medium on which the program is stored may be a non-transient recording medium such as a CD-ROM.

[0016] The present invention relates to an invention for analyzing the potential psychology of a user who conducts a dialogue. A person's mind always contains contradictions, fluctuations, and conflicts. For example, a state where multiple emotions, thoughts, and psychologies such as "want to do it but afraid" and "want to talk to someone, but may be disliked if I talk" coexist simultaneously is a natural phenomenon for the human mind. That is, many people do not directly convey their true feelings at once, but express themselves through speech, expressions, etc. in multiple contexts such as behind words or in silence. In contrast, conventional conversational AI only returns a superficial response according to the modality signal obtained from the dialogue with the user, and there is a problem that it cannot consider the fluctuations, confusion, and coexistence of psychology obtained from such a user. Also, in the field of psychological clinical practice, it is possible to observe, hypothesize, and verify the meaning of a conversation from multiple perspectives such as "the real problem behind the symptoms" and take appropriate measures. However, conventional AI cannot support such a psychological clinical scene, and there has been a demand for a technology to support the resolution of the troubles of users with psychosomatic symptoms.

[0017] In view of such a background, the present invention provides a technology that comprehensively considers the modality signals obtained from a user who conducts a dialogue and describes a plurality of different potential psychologies held by the user in parallel and probabilistically.

[0018] In addition to fields such as education, medical care, and counseling, the present invention is also applicable to applications that realize dialogue and interaction via an AI avatar in a virtual space such as a metaverse.

[0019] As shown in FIG. 1, the dialogue system 1 includes a mobile terminal 10 such as a smartphone used by a user, and a server 20 that can be connected to the mobile terminal 10 via a network 2 from a base station 3 or the like. The dialogue system 1 is a dialogue agent that conducts a dialogue in response to a user's consultation or the like.

[0020] The mobile terminal 10 includes a voice acquisition unit 11, a text acquisition unit 12, and an image acquisition unit 13 that receive input from the user. The voice acquisition unit 11 includes a microphone and can obtain voice data from the voice spoken by the user as input voice. The text acquisition unit 12 can obtain text data from the user's input of characters through the operation of the mobile terminal 10. The image acquisition unit 13 includes a camera and can capture images of the user's facial expressions and gestures while they are inputting voice or text and acquire them as facial image data.

[0021] The mobile terminal 10 further includes a language data creation unit 14 that creates language data by converting the user's text input or spoken input. The language data creation unit 14 can convert text input directly into text-formatted language data, or it can process the input speech using a known voice assistant to convert it into text data and then use that as language data. The language data creation unit 14 can also transmit the input speech to the server 20 along with the language data as tone data that reflects the user's emotions, etc.

[0022] The mobile terminal 10 also includes a reading means 15 that reads out conversation data received from the server 20 as voice, and an image display means 16 that displays image data received from the server 20. This conversation data and image data can be received from the server 20 as response data.

[0023] On the other hand, the server 20 includes a response data generation unit 21 and an image selection unit 22, as well as an emotion table 23 (database) that associates images to be sent back to the mobile terminal 10 with feature quantities described later.

[0024] The response data generation unit 21 receives language data from the mobile terminal 10 and uses a dialogue engine to generate response data that includes conversational data as a response to the user. The response data generation unit 21 also combines the image selected by the image selection unit 22 (described later) with the conversational data to form the response data. The response data generation unit 21 can send the generated response data back to the mobile terminal 10.

[0025] The image selection unit 22 receives tone data and facial expression image data from the mobile terminal 10 and can extract feature quantities from them. The feature quantities are a single, integrated entity derived from the tone data and facial expression image data, respectively. The image selection unit 22 can further select a predetermined image from the emotion table 23 corresponding to the obtained feature quantities and pass it to the response data generation unit 21.

[0026] Next, the operation of the dialogue system 1 will be explained, referring to Figure 2 in conjunction with Figure 1.

[0027] Referring to Figure 2 in conjunction with Figure 1, first, the user inputs into the mobile terminal 10 (S1: input). At this time, the user can choose between voice input and / or text input. The user's main action is to speak into the mobile terminal 10, with text input being used only as a supplement. The dialogue system 1 is envisioned to provide consultation and advice to the user (consultant) through a dialogue that feels natural to the user using the mobile terminal 10.

[0028] Meanwhile, in the mobile terminal 10, each unit converts user input into data and converts it into transmission data to be sent to the server 20 (S2: creation of transmission data).

[0029] In detail, the voice acquisition unit 11 acquires the voice spoken by the user as input voice, and as described above, converts the input voice into text data using a known voice assistant to obtain language data. The text acquisition unit 12 also converts the user's text input directly into text-format language data. Furthermore, the voice acquisition unit 11 uses the input voice as tone data to infer the user's emotions. In other words, the tone data is data obtained by converting the user's spoken input voice into audio signals. It is preferable for the voice acquisition unit 11 to also acquire voices that the user did not specifically intend to input, such as coughs, interjections, and voices that cannot be converted into text, as input voice. Therefore, it is preferable to acquire input voice even when the user selects text input as the input method.

[0030] Furthermore, the image acquisition unit 13 acquires facial expression image data, capturing the facial expressions of the user who is inputting into the mobile terminal 10, when the user is speaking in conjunction with voice input, when they are inputting text, and even between inputs. Preferably, the facial expression image data can capture not only the user's facial expressions but also the user's gestures and other actions. Also, preferably, the facial expression image data is a video consisting of multiple consecutive images captured at a predetermined frame rate.

[0031] The transmission data, including the language data, tone data, and facial expression image data created in the manner described above, is sent from the mobile terminal 10 to the server 20.

[0032] Upon receiving data from the mobile terminal 10, the server 20 generates response data such as conversation data (S3). Specifically, upon receiving language data, a known dialogue engine, such as artificial intelligence that performs natural language processing, generates conversation data that responds to the language data from the user. It is also preferable to include a voice correction means for modifying the conversation data generated by the emotion score described later. For example, the endings of the conversation data are corrected according to the emotion score index, so that in the response output (S5) described later, conversation data that can give the user a sense of security or calmness is read aloud.

[0033] Server 20, on the one hand, analyzes the tone data and facial expression image data separately to infer the user's emotions, and extracts a single feature that integrates both. For example, the feature can be a state vector consisting of multiple emotional elements that infer the user's emotions. Examples of emotional elements include anxiety, hope, and anger, and each emotional element can be assigned an emotional score, for example, 60%, 30%, and 10%. In addition, language data may also be processed into emotion inference language data from which words for emotion inference have been extracted, analyzed, and then integrated into the feature described above.

[0034] For feature extraction, for example, a known sentiment extraction engine can be used. Furthermore, in order to perform processing by the server 20 in a short time, it is preferable to apply a quantum-inspired algorithm. For example, by using the concept of quantum superposition, the time required for analysis can be shortened, and by using the concept of quantum entanglement, the correlation and dependence between multiple emotions of the user can be evaluated. Parallel and distributed processing, such as cloud computing via network 2, may be used for these analysis processes.

[0035] Furthermore, an image is selected from the emotion table 23 in accordance with the feature. When an emotion score is defined as a feature, the emotion table 23 includes images that, for example, provide a sense of security or lead the user to a calm emotion, corresponding to the emotion score. As such images, for example, a predetermined character can be set and multiple expressions can be defined. The predetermined character may be a character that the user likes to select, or it may be possible to create a user avatar in anticipation of a dialogue with oneself. Alternatively, such images may be videos that change the character's expression or make the character perform actions.

[0036] Next, the server 20 returns the generated conversation data and selected image data as response data, which is received by the mobile terminal 10 (S4: Response data received).

[0037] Furthermore, the mobile terminal 10 that receives the response data reads the conversation data aloud as a response voice using the reading means 15, and displays the image selected according to the emotion table using the image display means 16 (S5: response output). As a result, the user can hear the outputted response voice, see the image, and feel as if they are interacting with the character in the image.

[0038] In particular, by performing multimodal processing using language data, tone data, and facial expression image data, it is possible to provide dialogue that responds to the user's emotions, while also enabling processing at a speed that does not impair the naturalness of the dialogue. Furthermore, by extracting an emotion index using an emotion extraction engine, the selected image becomes more responsive to the user's emotions, and as a result, the dialogue can also become more responsive to the user's emotions. The user can then feel as if they are conversing with an image such as a character or avatar through the response audio and image that responds to their questions or inquiries. This can, for example, allow the user to gain emotional stability, such as alleviating feelings of loneliness.

[0039] Next, referring to Figure 3, the specific processing flow is outlined in bullet points. Note that (1) to (8) indicate the order of processing, and the letters correspond to the content of each processing step (1) to (8). The face symbol on the left represents the user (consultant). On the other hand, the face symbol on the right represents the conversational AI avatar (an image of a character, etc., displayed on a mobile device), which represents the conversation partner from the user's subjective perspective. In other words, when considering the actual processing, the conversational AI avatar corresponds to server 20.

[0040] (1) A. Input and spoken data of the person seeking advice (voice) - Facial expressions and gestures (video data) - Text data entered directly as needed

[0041] (2) B. Receive multimodal data from the client and send an analysis quest to the quantum-inspired algorithm.

[0042] (3) C - Integration of features (voice, facial expression, text data) into a single state vector - Simultaneous data analysis by applying the concept of "quantum superposition" - Evaluation of correlations and dependencies between emotions by applying the concept of "quantum entanglement"

[0043] (4) D. Cloud computing receives input data in real time and quickly extracts data features through parallel computing.

[0044] (5) E. Send to quantum-inspired algorithm

[0045] (6) F - Emotion score: "Anxiety 60%, Hope 30%, Anger 10%" - Response instructions: "Generate a reassuring response," "Speak with calm emotions" - Emotion score and response instructions are sent to the conversational AI avatar

[0046] (7) G - Receives "emotion score" and "response instructions" - Generates optimal responses, gestures, and facial expressions using Fine Tuning, RAG (GPT-4 if necessary) - Speaks in a natural voice using speech synthesis (TTS). Facial expressions and gestures are controlled in real time.

[0047] (8) H. The conversational AI avatar responds and processes through voice, facial expressions, and actions. The entire process from (1) to (8) is completed in approximately 1 second.

[0048] The processing flow described above allows the user (consultant) to engage in dialogue without feeling any discomfort, and the processing speed is sufficient to maintain the naturalness of the conversation. Furthermore, the images corresponding to the user's emotions allow the user to recognize that they are interacting with a conversational AI avatar.

[0049] Although representative embodiments of the present invention and their variations have been described above, the present invention is not necessarily limited thereto and can be modified as appropriate by those skilled in the art. That is, those skilled in the art will be able to find various alternative embodiments and modifications without departing from the scope of the attached claims.

[0050] The above-mentioned functional configuration in the present invention is realized by the following hardware configuration. Figure 4 is a block diagram showing the hardware configuration of the dialogue device 20 (server).

[0051] As shown in Figure 4, the interactive device 20 comprises a processing unit 2001, a storage unit 2002, and a communication unit 2003. The processing unit 2001 has a processor such as a CPU capable of executing instruction sets and controls the overall operation of the interactive device 20 by executing the interactive program, OS, and other applications according to the present invention. The storage unit 2002 has a volatile memory such as RAM capable of storing instruction sets, and a non-volatile recording medium such as an HDD or SSD capable of recording the OS and interactive programs. The communication unit 2003 has a communication interface device for connecting to the network 2 and performs communication control with the network 2 to input and output information.

[0052] Figure 5 is a block diagram showing the hardware configuration of a user terminal device 10 (corresponding to the mobile terminal 10 in Embodiment 1) according to one embodiment. As shown in Figure 5, the terminal device (user terminal device 10) has a processing unit 101, a storage unit 102, a communication unit 103, an input unit 104, and an output unit 105.

[0053] The processing unit 101 has a processor such as a CPU capable of executing instruction sets and controls the entire operation of the terminal device 9 by executing the OS and other applications. The storage unit 102 has volatile memory such as RAM capable of storing instruction sets, and non-volatile recording media such as HDDs and SSDs capable of recording the OS, etc. The communication unit 103 has a communication interface device for connecting to the network 2 and performs communication control with the network 2 to input and output information. The input unit 104 has input devices such as a keyboard, touch panel, microphone, and camera capable of processing input for receiving modality signals from the user. The output unit 105 has output devices such as a display and speaker capable of output processing.

[0054] (Embodiment 2) Next, in Embodiment 2, we will describe in detail a process that predicts the user's potential psychology based on multiple types of modality signals obtained from interaction with the user through multimodal processing, by introducing a quantum mechanical description.

[0055] In terms of a quantum mechanical description, specifically, from the fundamental principles of quantum mechanics, the observable E is considered a self-adjoint operator in Hilbert space, and the result that can be obtained from observing the observable E is one of the eigenvalues ​​of the operator E. Therefore, the eigenvalue equation for the operator E is Therefore, |Φ i > represents the orthonormal basis of the eigenvectors (basis vectors) of operator E (when there is no degeneracy in the eigenvalues). In this embodiment, the eigenvectors in equation (1) are eigenvectors related to psychology. Also, the eigenvalue e i Φ represents an observed value of psychology, and hereafter, the observed value itself will be referred to as psychology. Furthermore, the subscript e in Φ represents the number of dimensions of psychology, and the number of dimensions of psychology is the number of distinct psychological elements that can be distinguished from emotions such as joy, anger, sadness, pleasure, and surprise, as well as the pleasure and displeasure of emotions, the intensity of emotions, and the sense of control over emotions (whether or not one can control their emotions).

[0056] In this embodiment, acquisition timing t n By analyzing each modality signal obtained from the interaction with the user, the following psychological element probability vectors are constructed, probabilistically representing multiple different latent psychological states of the user using the eigenvectors of equation (1): Here, |ψ mod (t n )> is the acquisition timing t n This indicates that the psychological element probability vector was obtained in [location].

[0057] Furthermore, the acquisition timing in this embodiment is a certain point in the dialogue, starting from the beginning of the series of dialogues, but it may also be the time when the modality signal is acquired, or a predetermined time interval from the start point to the end point of the series of dialogues. Specifically, the predetermined time interval from the start point to the end point of the series of dialogues is the time from when the dialogue text data, described later, is output to the user terminal device to the time when the modality signal of text data, audio data, or image data is received as a response to the dialogue text data. Note that the start and end points of the time interval can be adjusted arbitrarily.

[0058] Furthermore, from the fundamental principles of quantum mechanics, the time t of the observable En in which the probability that the mind is e i is observed is shown as follows:

[0059] In the present embodiment, a mental probability integration vector (described later) obtained by integrating the mental element probability vectors (Equation (2)) obtained from each modality signal under predetermined conditions is generated, and the latent mind of the user is predicted in real time from a series of interactions with the user.

[0060] Also, in the present embodiment, for the mental probability integration vector generated at the acquisition timing t n the mental probability integration vector generated at the acquisition timing t n―1 is updated by considering correction values based on a plurality of different latent minds at the immediately preceding acquisition timing t n in the series of interactions.

[0061] Also, in the present embodiment, based on the generated mental probability integration vector, the mind of the user at each acquisition timing is observed by the generative AI, and the observed mind of the user (hereinafter referred to as the observed mind) is presented to the user. Further, based on the fitness score indicating the reliability of the observation result, the generative AI re-observes the mind of the user.

[0062] Here, the generative AI is a neural network model pre-trained using a large amount of text, voice, and numerical data, and is an inference engine that generates related output data (such as a mental element probability vector) based on the input natural language text or voice.

[0063] Note that the generative AI in the present embodiment is a generative neural network pre-trained for natural language tasks, and a representative example is a GPT (Generative Pre-trained Transformer) - based model, but its parameter scale and network configuration are not limited. For example, a hybrid model that combines a convolutional neural network (CNN) and a recurrent neural network (LSTM, GRU, etc.) in an auxiliary manner while centering on a Transformer may also be used.

[0064] Furthermore, the present invention includes a form in which the speech, facial expressions, and actions of an AI avatar, as well as economic activities within the virtual space, are observed in a virtual space such as the metaverse, and these are processed based on a state update formula described later, thereby presenting the visualization of multiple different latent psychological states (subconscious structures) within the virtual space.

[0065] In this embodiment, psychology is described using discrete eigenvectors, but psychology may also be described using continuous eigenvectors.

[0066] Figure 4 is a block diagram showing the functional blocks of the dialogue device 20 in this embodiment. As shown in Figure 4, the dialogue device 20 includes an acquisition unit 201, a state generation unit 202, a corrected state generation unit 203, an update unit 204, an observation unit 205, a fitness calculation unit 206, an output processing unit 207, and a database 23. This represents the concrete implementation of information processing by software (stored in the storage unit 102) by hardware (the processing unit 101). The acquisition unit 201 corresponds to the voice acquisition unit 11, the text acquisition unit 12, and the image acquisition unit 13 in Embodiment 1. The output processing unit 207 corresponds to the reading means 15 and the image display means 16 in Embodiment 1. The acquisition unit 201 may acquire voice data and execute the processing of the language data creation unit 14.

[0067] In this embodiment, the user terminal device 10 (client) is a so-called web application in which the client receives the processing results performed by the dialogue device 20 (server) in response to the client's request. Alternatively, it may be a so-called standalone type in which the dialogue program is launched on the client terminal. In this case, the user terminal device 10 may include some or all of the functional components (parts) of the dialogue device 20. For example, the user terminal device 10 may include an acquisition unit 201, a state generation unit 202, a corrected state generation unit 203, an update unit 204, an observation unit 205, a fitness calculation unit 206, and an output processing unit 207, and the dialogue device 20 may be a cloud storage that records and manages the observation information described later.

[0068] The acquisition unit 201 acquires user signal information from the interaction with the user. The acquisition unit 201 acquires multiple types of modality signals obtained from the interaction with the user, and the acquisition timing (t) for each modality signal. n ) is obtained to acquire user signal information including this.

[0069] The state generation unit 202 processes a vector representing a subconscious structure that simultaneously possesses multiple different potential psychological states. The state generation unit 202 analyzes the acquired modality signals and processes a psychological element probability vector for each modality signal, representing a subconscious structure that possesses multiple different psychological states of the user and the observation probability of each psychological state.

[0070] Furthermore, the state generation unit 202 generates a psychological probability integration vector by integrating the psychological element probability vectors. Based on multiple psychological element probability vectors acquired at the same time, the state generation unit 202 integrates and calculates multiple different psychological states of the user included in multiple types of modality signals and the observation probabilities of each psychological state to generate a psychological probability integration vector that shows the subconscious structure based on multiple types of modality signals.

[0071] The correction state generation unit 203 generates a psychological probability correction vector for correcting the user's latent psychology at a given moment based on past subconscious structures. The correction state generation unit 203 generates a psychological probability correction vector for correcting a new subconscious structure based on past subconscious structures, based on the changes in the observed probabilities of each psychology included in the psychological probability integration vector generated at consecutive acquisition timings in a series of conversations.

[0072] Furthermore, the correction state generation unit 203 determines whether or not to generate a psychological probability correction vector based on the changes in the observed probabilities of each psychology included in the psychological probability integration vector generated at consecutive acquisition timings in the time series of a series of conversations, and registers the generated psychological probability correction vector in the database 23 in association with the conversation.

[0073] The update unit 204 updates the integrated psychological probability vector. Based on the generated corrected psychological probability vector, the update unit 204 updates the observed probability of each psychology in the integrated psychological probability vector, thereby updating the integrated psychological probability vector. In this embodiment, the update unit 204 inputs the corrected psychological probability vector and the integrated psychological probability vector into a mathematical model and updates the observed probability of each psychology in the integrated psychological probability vector. The specific mathematical model will be described later.

[0074] Furthermore, the update unit 204 determines whether or not to update the psychoprobability integration vector based on the goodness-of-fit score calculated by the goodness-of-fit calculation unit 206, and updates the psychoprobability integration vector.

[0075] Furthermore, the update unit 204 updates the psychological probability integration vector based on the convergence of observations. If the convergence of observations is poor, the update unit 204 updates the psychological probability integration vector by generating new eigenvectors for each psychology.

[0076] The observation unit 205 causes the generative AI to observe the user's psychology. The observation unit 205 transmits a generated psychology probability integration vector and an observation instruction to the generative AI to observe the user's psychology based on the observation probability of each psychology in the psychology probability integration vector, causing the AI ​​to observe one of several different potential psychology at a certain acquisition timing.

[0077] The observation unit 205 also registers observation information in the database 23 in chronological order, which is linked to a set of psychological probability values, which is a combination of multiple different psychological states and their respective observation probabilities, and the observed psychological states observed based on the integrated psychological probability vector having said psychological probability value.

[0078] The goodness-of-fit calculation unit 206 calculates a goodness-of-fit score indicating the observation confidence of the newly observed psychology. Based on the observation information, the goodness-of-fit calculation unit 206 calculates a transition rate indicating the confidence of each observed psychology when transitioning from one set of psychology probability to another set of psychology probability. Then, based on the number of psychology probability correction vectors associated with the dialogue prior to the observation time of the newly observed psychology, and / or the transition rate indicating the confidence of the newly observed psychology, the goodness-of-fit calculation unit 206 calculates a goodness-of-fit score indicating the observation confidence of the newly observed psychology.

[0079] The output processing unit 207 processes the dialogue generated by the conversational AI and the observed psychological state observed during the dialogue for output. The output processing results are then output to the user terminal device 10. In this embodiment, the output processing includes processing to output the dialogue and observed psychological state using text data, and processing to output using audio data.

[0080] The following describes a dialogue method using the dialogue system of the present invention with reference to Figure 7. Figure 7 is a processing flowchart showing the process from acquiring user signal information, generating and updating a psychological probability integration vector, observing the psychology of the user corresponding to the dialogue AI based on the psychological probability integration vector, and presenting the observed psychology to the user. In the following processes, the acquisition of user signal information, generation of the psychological probability integration vector, and observation of psychology can be performed synchronously.

[0081] First, in step S101 (hereinafter, "step SXXXX" will be simply referred to as "SXXXX"), response data is output. In this embodiment, the response data generation unit 21 receives an instruction to start a dialogue with the user and generates dialogue text data for counseling the user. Then, the output processing unit 207 processes the generated dialogue text data and outputs it to the user terminal device 10.

[0082] In S102, the acquisition unit 201 acquires user signal information. In this embodiment, the user terminal device 10 receives from the user multiple types of modality signals (text data corresponding to the dialogue text data, voice data, facial images (eye gaze, facial stiffness, etc.), and biometric information, etc.) for the dialogue text data output in S101. The acquisition unit 201 then receives and acquires these multiple types of modality signals at each acquisition timing. The acquisition unit 201 may also transmit the acquired multiple types of modality signals to the generation system AI and acquire the noise-removed signals as modality signals.

[0083] In S103, the state generation unit 202 processes the generation of psychological element probability vectors. In this embodiment, as a generation process, the state generation unit 202 transmits the multiple types of modality signals acquired in S102 and an instruction to analyze the modality signals to the generation system AI, causing the AI ​​to analyze each modality signal as an analysis process, and generate psychological element probability vectors for each of the multiple types of modality signals that simultaneously possess multiple different potential psychological states. The state generation unit 202 then acquires the psychological element probability vectors for each generated modality signal.

[0084] Specifically, the state generation unit 202 transmits the acquired modality signals and instructions to analyze the modality signals to the generation system AI, causing it to generate eigenvectors corresponding to each psychology and psychological element probability vectors that represent the subconscious structure having the observation probability of each psychology (see equation (3) for specific equations). More specifically, the state generation unit 202 generates eigenvectors (so-called eigenfunctions) that have position variables as eigenvectors. Here, any vector can be used as the eigenvector, but for example, each basis in Hermitian polynomials, a basis in Fourier series expansion, etc., can be used.

[0085] Preferably in the present invention, the generative AI is trained to predict an accurate psychological element probability vector from the modality signal, using modality signal information and the psychological elements and their respective observed probabilities included in the psychological element probability vector as training data.

[0086] In S104, the state generation unit 202 processes a psychological probability integration vector. In this embodiment, the state generation unit 202 transmits modality signal information and multiple types of psychological element probability vectors to the generation system AI, which generates a psychological probability integration vector by integrating the multiple types of psychological element probability vectors corresponding to the same acquisition timing based on predetermined weights. The state generation unit 202 then acquires the generated psychological probability integration vector and registers it in the database 23 in chronological order.

[0087] Specifically, the state generation unit 202 calculates predetermined weights based on the confidence level of the acquired modality signal information to generate a psychological probability integration vector. For example, if the modality signal is the user's "facial expression" acquired by the camera, the face recognition score based on the acquired face image is used as the confidence level, and the predetermined weight is calculated to be smaller for those with a low face recognition score (for example, those in which only a portion of the user's face is included in the face image), thereby generating a psychological probability integration vector. Alternatively, if the modality signal is the user's "voice" acquired by the microphone, the conversion accuracy score, which indicates the accuracy of the conversion of the acquired voice data to text data, is used as the confidence level, and the predetermined weight is calculated to be smaller for those with a low conversion accuracy score (for example, those in which the amplitude of the voice data is small and there are parts that cannot be converted to text data that are above a predetermined threshold), thereby generating a psychological probability integration vector.

[0088] In S105, the correction state generation unit 203 determines whether correction based on past subconscious structure is necessary in the psychology of the user engaging in the dialogue. In this embodiment, the correction state generation unit 203 refers to the observed probability of each psychology among the psychological element probability vectors that are registered in a series of dialogues and have a continuous time series, and determines the acquisition timing t n―1 The probability of observing psychology in the acquisition timing t n The acquisition timing t is determined according to the degree of change, which indicates how much the data has changed (increased or decreased). n Regarding the user's psychology, the timing of acquisition t n―1 It is determined that correction based on the subconscious structure is necessary.

[0089] Specifically, the correction state generation unit 203 acquires the timing t n―1 The observed probability of each psychological element in the probability vector of each psychological element, and the acquisition timing t n In the probability vector of each psychological element, the observed probability of each psychological state and the acquisition timing t n―1 While each psychological element probability vector has a large observation probability, the acquisition timing t n The observed probability of a psychological trait in each psychological element probability vector has decreased significantly (at acquisition timing t). n―1 Observation probability and acquisition timing t n The observed probability in and the psychological state in which there is a decrease of more than a predetermined threshold in the acquisition timing t n For the probability vector of psychological elements, acquisition timing t n―1 It is determined that a correction based on the subconscious structure is necessary. For example, if the observed probabilities of the first, second, and third psychological states change from 10%, 50%, and 40% respectively to 35%, 25%, and 40% respectively, it is determined that the observed probability of the second psychological state has decreased significantly, and a correction based on the past subconscious structure is necessary.

[0090] In this embodiment, the determination of whether or not a correction based on past subconscious structure is necessary is made based on the psychological element probability vector, but it is also possible to determine whether or not a correction based on past subconscious structure is necessary based on the psychological probability integration vector.

[0091] Then, if it is determined in S105 that no correction based on past subconscious structure is necessary (NO in S105), the process proceeds to S108. On the other hand, if it is determined in S105 that a correction based on past subconscious structure is necessary (YES in S105), the correction state generation unit 203 generates a psychological probability correction vector in S106.

[0092] In this embodiment, the correction state generation unit 203 transmits to the generation system AI the observed probability of each psychology in the psychological element probability vector for each acquisition timing, and an instruction to generate a psychological probability correction vector having each psychology included in the psychological element probability vector and a correction value for each observed probability of psychology based on the transition of the observed probability, causing the AI ​​to generate an eigenvector corresponding to the same psychology as the psychological element probability vector generated in S103, and a psychological probability correction vector having a correction value for the observed probability of each psychology. Specifically, the correction state generation unit 203 transmits to the generation system AI the observed probability of each psychology in the psychological element probability vector for each acquisition timing t n―1 The observation probability of each psychological state and the acquisition timing t n Based on the observed probability of each psychology and the degree of change indicating the increase or decrease in the observed probability of each psychology, a psychology probability correction vector is generated in which the absolute value of the increase or decrease in the observed probability of each psychology is used as the coefficient of the eigenvector corresponding to each psychology. For example, if the observed probabilities of the first psychology, second psychology, and third psychology change from 10%, 50%, and 40% respectively to 35%, 25%, and 40% respectively, the observed probability of the second psychology has decreased significantly, and the correction state generation unit 203 generates a psychology probability correction vector having a psychology correction value that increases the observed probability of the second psychology. This makes it possible to consider the state of the psychology immediately preceding the current psychology, which should have an effect on the current psychology.

[0093] In addition, the correction state generation unit 203 calculates the degree of change in the observed probability of each psychology at two consecutive acquisition timings in the time series of a single dialogue, and sends the degree of change at multiple acquisition timings and the statistical values ​​of the multiple degrees of change (e.g., sum, average, etc.) to the generation system AI as an instruction to generate a psychology probability correction vector having the observed probability of each psychology as the coefficient of the eigenvector, thereby generating the psychology probability correction vector. Here, the correction state generation unit 203 may generate eigenfunctions instead of eigenvectors, similar to the psychology element probability vector.

[0094] The correction state generation unit 203 then acquires the psychological probability correction vector and registers it in the database 23, linked to the psychological element probability vector used to generate the psychological probability correction vector. In this embodiment, a psychological probability correction vector is generated for each psychological element probability vector, but the psychological probability correction vector may be generated based on a psychological probability integration vector instead of the psychological element probability vectors.

[0095] In S107 of Figure 7, the update unit 204 updates the psychological probability integration vector. In this embodiment, the update unit 204 updates the psychological probability integration vector based on the following calculation formula, which is based on the psychological probability integration vector registered in S104, the psychological probability correction vector registered in association with the psychological element probability vector, and a predetermined weighting for each modality signal corresponding to the psychological element probability vector: Furthermore, by expressing it based on the description of quantum mechanics as follows, the psychological stochastic integration vector can be represented as having position dependence based on eigenfunctions: Here, |ψ int (t n )> represents the integrated psychological probability vector. Also, |G j (t n )> indicates the psychological probability correction vector (the subscript j is the number of modality signals). Also, b j |ψ represents a predetermined weight for each predetermined modality signal. int '(t n +Δt)> is |ψ int (t n )> Newly generated time t n The integrated psychological probability vector at +Δt is shown.

[0096] Furthermore, the updated psychological stochastic integration vector has the normalized observed probability of each psychology so as to satisfy the condition of equation (4). The predetermined weights for each modality signal can be determined in the same way as when the psychological stochastic integration vector was generated. Note that the time of the newly generated psychological stochastic integration vector may be the same as the time before the update, i.e., |ψ int '(t n ) > is also acceptable.

[0097] In S108, the observation unit 205 observes psychology based on the psychological probability integration vector. In this embodiment, the observation unit 205 transmits the psychological probability integration vector from S104 or S107, an observation instruction to observe one of the psychology in the psychological probability integration vector according to the observation probability of the psychological probability integration vector, to the generating AI, causing it to observe one of several different potential psychology at the acquisition timing corresponding to the generated psychological probability integration vector.

[0098] Specifically, the observation unit 205 uses an observation operator to observe the psychological probability integration vector. Here, a general observation is a psychological observation operator M that satisfies the following completeness relation. i It can be expressed using the set: Here, I represents the identity matrix.

[0099] And the observation operator M i Psychology by i The observed probability of is expressed as follows: Here, the first term on the right-hand side of equation (8) is the probability in equation (4) (the projection operator P below). i This represents a term that matches the probability (based on ). Here, I represents the identity matrix. On the other hand, the second term on the right-hand side of equation (8) is a term that represents the error based on the observation accuracy and affects the observation probability obtained in S104. The observation unit 205 then uses the observation operator M i This is applied to a psychological probability integration vector, causing the generative AI to observe psychology.

[0100] The observation unit 205 may also transmit a psychological probability integration vector and an instruction to the generating AI to select an eigenvector of 1 according to the observed probability coefficient of each psychological eigenvector, thereby causing the AI ​​to observe the psychological state corresponding to the selected eigenvector as the observed psychological state.

[0101] In S109, the observation unit 205 determines whether the convergence is poor (the accuracy of the observation is poor). In this embodiment, the observation unit 205 determines the observation operator M iWhen multiple psychological states are observed through observation, it is determined that the convergence of observations is poor. For example, if the eigenfunctions generated in S103 are not orthogonal to each other, or if the observation operator M i It is not diagonalized (for example, M 0 =|Φ 1 ><Φ 1 | + cosθ | Φ 2 ><Φ 2 |, M 1 = sinθ|Φ 2 ><Φ 2 | (M 0 and M 1 If it is obvious that equation (7) is satisfied, etc., then multiple psychological states will be observed, and therefore the convergence of observations will be judged as poor.

[0102] If it is not determined in S109 that the convergence of observations is poor (NO in S109), the process proceeds to S111. On the other hand, if it is determined that the convergence of observations is poor (YES in S109), the update unit 204 updates the psychological probability integration vector.

[0103] In this embodiment, if the update unit 204 determines that the convergence of observations is poor, it determines that the eigenfunctions generated in S203 are not orthogonal to each other, and updates the psychological probability integration vector by generating a new eigenfunction different from the generated eigenfunction for each psychological state.

[0104] In addition, if the update unit 204 determines that the convergence of observations is poor, it determines that it is necessary to correct the observation probability of each psychology based on the error term of the observation operator used for the observation. It obtains the observation probability based on the error term and calculates the observation probability by subtracting the portion of the error term's observation probability from the observation probability of each psychology. The update unit 204 then updates the psychology probability integration vector with the new observation probability as the observation probability of each psychology. The update unit 204 continues to perform the processes S108 to S110 until the convergence of observations improves. This makes it possible to generate a new psychology probability integration vector that has the new observation probability as the observation probability of each psychology.

[0105] In S111, the fitness calculation unit 206 calculates the fitness score of the observed psychology. In this embodiment, in S108, the fitness calculation unit 206 calculates the time (t) during which the psychology was observed. n The generation system AI receives observation information linked to the psychological probability integration vectors for each acquisition timing registered prior to the first acquisition timing, and the observed psychology observed based on each psychological probability integration vector, along with an instruction to calculate the transition rate of the observed psychology when transitioning from one psychological probability integration vector to another, causing the AI ​​to calculate a transition rate that indicates the confidence (probability) that the second psychological probability set will be observed after the first psychological probability set has been observed. The goodness-of-fit calculation unit 206 then obtains this transition rate.

[0106] For example, if, at a certain acquisition timing in a series of dialogues, a transition occurs from the first psychological probability set to the second psychological probability set, and the first observed psychology and the second observed psychology are observed and registered based on each psychological probability set, then, at another acquisition timing, if a transition occurs from the first psychological probability set to the second psychological probability set, and the first observed psychology is observed in the first psychological probability set, the reliability of observing the second observed psychology based on the second psychological probability set is considered high, and the transition rate of the second observed psychology is calculated to be relatively higher than the transition rates of other psychology.

[0107] In a preferred embodiment of this model, the dialogue system of the present invention includes a learning unit (not shown), which inputs a first and second set of consecutive psychological probability sets, each containing a combination of observed probabilities for each psychology in a psychological probability integration vector for each acquisition timing, into a generative AI, causing the generative AI to learn a transition rate that indicates the probability that the second psychological probability set will be observed after the first psychological probability set has been observed.

[0108] Then, the goodness-of-fit calculation unit 206 calculates the psychological probability correction vectors registered in association with the user during the series of conversations, and in S108, the time (t) at which the psychological state was observed. nThe goodness-of-fit score is calculated based on the number of psychological probability correction vectors registered within a predetermined time prior to the specified date and the calculated transition rate. Specifically, the goodness-of-fit calculation unit 206 calculates the goodness-of-fit score as the sum of, for example, the value of the transition rate and a negative value corresponding to the number of psychological probability correction vectors (for example, the product of a predetermined negative value and the number of psychological probability correction vectors, or a negative value corresponding to a threshold set for the number of psychological probability correction vectors). This makes it possible to calculate a goodness-of-fit score that can evaluate the reliability of the observed psychology by considering not only the transition rate but also the instability of psychology generation due to how much the user does not express their own psychology as modality signals.

[0109] In S112, the update unit 204 determines whether the goodness-of-fit score exceeds a predetermined threshold. In this embodiment, if the goodness-of-fit score exceeds the threshold (YES in S112), the process proceeds to S115. On the other hand, if the goodness-of-fit score does not exceed the threshold (NO in S112), in S113, the update unit 204 determines that the newly observed observed psychology is abnormal (unstable) and updates the psychology probability integration vector.

[0110] In this embodiment, the update unit 204 compares the transition rate for each observed psychology between the psychological probability set corresponding to the newly observed psychology and the psychological probability set corresponding to the psychology observed immediately before the observation of the said psychology with the observed probability of the psychological probability set corresponding to the newly observed psychology, identifies an abnormal psychology among the psychology in the psychological probability integration vector used to observe the newly observed psychology, and generates a new psychological probability integration vector with updated observed probabilities corresponding to that psychology.

[0111] Specifically, the update unit 204 transmits the transition rate for each observed psychology and the observed probability of the psychology probability integration vector used to observe the newly observed psychology to the generation system AI, causing the AI ​​to identify psychology with a relatively high observed probability and a relatively low transition rate, and / or psychology with a relatively low observed probability and a relatively high transition rate, as abnormal psychology, and to generate a new psychology probability integration vector with updated observed probabilities corresponding to the abnormal psychology. Then, the update unit 204 acquires the new psychology probability integration vector and updates the psychology probability integration vector.

[0112] In S114, the observation unit 205 causes the user's psychology to be re-observed based on the updated psychological probability integration vector. In this embodiment, the observation unit 205 transmits the psychological probability integration vector updated in S113 and observation instructions based on the observation probability of each psychology in the psychological probability integration vector to the generation system AI, causing it to re-observe the psychological probability integration vector. More preferably, the observation unit 205 adjusts the observation probability based on the error term acquired in S110 to observe psychology 1.

[0113] In S115, the output processing unit 207 presents the observed psychology and dialogue text data corresponding to that psychology to the user. In this embodiment, the output processing unit 207 outputs the psychology observed in S114 and dialogue data (text data, audio data, and image data, etc.) with a writing style and tone appropriate to the user's psychology, which has been generated by the response data generation unit 21 based on that psychology. The output processing results are then output to the user terminal device 10.

[0114] As described above, by executing processes S101 to S115, the AI ​​can accurately and precisely infer the user's underlying psychology and output corresponding dialogue data, thereby enabling appropriate user counseling. Furthermore, by actively allowing the AI ​​to acquire the user's psychology, it is possible to objectively present the user's underlying psychology without the arbitrariness or bias of a human such as a psychotherapist. As a result, the user can respond to the answers received without bias and stress.

[0115] In this embodiment, the update unit 204 updates a new psychological probability integration vector based on the transition rate in S113. However, the update unit 204 may also update the psychological probability integration vector by referring to a plurality of psychological probability integration vectors generated at multiple acquisition timings within a predetermined time after the observation time of the newly observed observed psychology. Based on the psychology and observed probability of each of the plurality of psychological probability integration vectors, the update unit 204 generates a new psychological probability integration vector in which the observed probability of the psychological probability integration vector used to observe the newly observed psychology has been updated.

[0116] Specifically, the update unit 204 updates the psychological probability integration vectors by referring to multiple psychological probability integration vectors generated at multiple acquisition timings within a predetermined time after the observation time of the newly observed psychological state, and generating a new psychological probability integration vector in which the average of the psychological state and the observed probability of the psychological state for each of the multiple psychological probability integration vectors is used as the observed probability of each psychological state.

[0117] In a preferred embodiment of the present invention, an acquisition timing t having a time width n When multiple modality signals acquired in satisfy predetermined conditions, the acquisition timing t n-1 The psychological probability integration vector is updated using a psychological probability correction vector based on each psychological observation probability in the psychological probability integration vector. Specifically, the correction state generation unit 203 uses an acquisition timing t which has a time width. n The modality signal acquired in this process is analyzed using well-known speech analysis and image analysis techniques to determine whether it contains predetermined user instability elements (for example, a state in which there are silent parts for a predetermined time or longer in a series of audio data, a state in which speech is intermittent and unclear, and a state in which the gaze is not directed forward for a predetermined time or longer in a series of facial image data). The correction state generation unit 203 then determines whether the acquisition timing t n If a user instability element is included in the modality signal, the modality signal is determined to satisfy predetermined conditions, and acquisition timing t n―1Based on the observed probability of each psychology in the integrated psychology probability vector, a psychology probability correction vector is generated. The update unit 204 then updates the integrated psychology probability vector by inputting the psychology probability correction vector into equation (5).

[0118] Furthermore, although the calculation process, generation process, and update process in this embodiment involve sending input data to an external generation AI and obtaining the generation result, these processes may also be performed using an internal generation AI within this system.

[0119] Furthermore, although the dialogue system 1 in this embodiment is composed of a dialogue device 20 and a user terminal device 10, it may be composed of only the dialogue device 20, or only the terminal device of the user terminal device 10. Also, the dialogue system 1 may appropriately share the processing described above between the dialogue device 20 and the terminal device of the user terminal device 10, and is not limited to the processing sharing described above.

[0120] [Effects, etc.] The following describes examples of inventions that can be obtained from the disclosures of this specification, and explains the effects, etc. that can be obtained from the examples of inventions.

[0121] Invention 1 is a dialogue system for analyzing a user's psychology from a conversation with the user, wherein the dialogue system comprises an acquisition unit and a state generation unit, the acquisition unit acquires user signal information including a plurality of modality signals obtained from a conversation with the user and the acquisition timing of each of the modality signals, the state generation unit analyzes the acquired modality signals and generates a psychological element probability vector for each modality signal that shows a subconscious structure having multiple different psychology of the user and the observation probability of each psychology, and based on the plurality of psychological element probability vectors with the same acquisition timing, integrates the multiple different psychology of the user and the observation probability of each psychology in the plurality of modality signals and generates a psychological probability integration vector that shows the subconscious structure based on the plurality of modality signals.

[0122] This configuration allows for the simultaneous representation of multiple latent psychological states of the user based on various modality signals obtained from user interaction. This makes it possible to bring to the surface user's latent psychological states that were previously impossible to achieve.

[0123] Invention 2 further comprises a correction state generation unit and an update unit, wherein the correction state generation unit generates a psychological probability correction vector for correcting a new subconscious structure based on past subconscious structures, based on the changes in the observed probabilities of each psychology between the consecutive psychological element probability vectors in a time series of dialogues, and the update unit inputs the psychological probability integration vector and the psychological probability correction vector into a mathematical model to update the observed probabilities of each psychology in the psychological probability integration vector.

[0124] From a neuroscience perspective, emotions manifest as patterns of neural activity, and the brain networks activated by the most recent emotion (particularly the amygdala and prefrontal cortex) do not immediately return to their original state but persist like inertia. On the other hand, humans sometimes engage in dialogue in a way that avoids outwardly expressing such recent emotions. Against this backdrop, the present invention, by adopting the above-described configuration, makes it possible to accurately represent the user's latent psychology by considering the influence of multiple coexisting past psychological states when describing the user's current psychology in a series of conversations.

[0125] Invention 3 is that the updating unit generates the psychological probability correction vector |G| based on the transition of the observed probability between the psychological probability integration vector |ψ> at a certain acquisition timing, the psychological element probability vector at the acquisition timing, and the psychological element probability vector at the acquisition timing immediately preceding the acquisition timing. j > and a predetermined weight b for each of the multiple modality signals corresponding to the psychological element probability vector j Mathematical models including: Using this, we obtain a psychological probability integration vector |ψ'> having the updated observed probabilities of each of the aforementioned psychological states.

[0126] This configuration allows for updating the user's subconscious structure by considering the weights of each of the multiple modality signals corresponding to the psychological element probability vectors.

[0127] Invention 4 further comprises an observation unit and an output processing unit, wherein the dialogue system transmits the psychological probability integration vector and an observation instruction to observe the user's psychology based on the observed probability of each psychology in the psychological probability integration vector to a generation system AI, causing it to observe one of the multiple different potential psychology at the acquisition timing, and the output processing unit outputs dialogue data based on the observed psychology, which is the observed psychology of the user, to the user.

[0128] By adopting this configuration, the AI ​​can proactively acquire multiple coexisting psychological states within the user's mind, enabling it to engage in dialogue that is tailored to the user's psychology without human bias.

[0129] Invention 5 further comprises a goodness-of-fit calculation unit, wherein the psychological probability integration vector has psychological probability sets which are combinations of multiple different psychological states and their respective observed probabilities, the observation unit registers observation information in a database in chronological order, which is linked to the psychological probability sets and the observed psychological states observed based on the psychological probability integration vector having the psychological probability sets, and the goodness-of-fit calculation unit calculates a transition rate based on the observation information, which indicates the confidence level for each observed psychological state observed based on the other psychological probability set when transitioning from one psychological probability set to another psychological probability set.

[0130] This configuration allows for the calculation of a score that can evaluate the validity of newly observed psychological states, based on past observational information from a series of conversations.

[0131] Invention 6 states that the observation unit transmits the integrated psychological probability vector used to observe the newly observed psychological state and the observation instruction to the generative AI, in accordance with the newly observed psychological state, the psychological probability set corresponding to the newly observed psychological state, and the transition rate between the psychological probability set corresponding to the observed psychological state observed immediately before the observation of the said psychological state, causing the AI ​​to re-observe the user's potential psychological state.

[0132] This configuration increases the reliability of observed psychological states and allows for more accurate predictions of the user's underlying psychology.

[0133] Invention 7 further comprises a correction state generation unit, which determines whether or not to generate a psychological probability correction vector for correcting a new subconscious structure based on past subconscious structures, based on the transition of the observed probability of each psychological state between the consecutive psychological element probability vectors in a time series of dialogues, registers the generated psychological probability correction vector in association with the dialogue in the database, and calculates a goodness-of-fit score indicating the observation reliability of the newly observed psychological state based on the number of psychological probability correction vectors associated with the dialogue prior to the observation time of the newly observed psychological state, and / or the transition rate indicating the reliability of the newly observed psychological state.

[0134] This configuration allows for the calculation of a score that can evaluate the validity of newly observed psychological states, using the user's past subconscious thoughts and the level of trust based on a series of past conversations.

[0135] Invention 8 further comprises an update unit, which determines, based on the fitness score, that the newly observed observed psychology is abnormal, and compares the transition rate for each observed psychology between the psychology probability set corresponding to the newly observed psychology and the psychology probability set corresponding to the psychology observed immediately before the observation of the said psychology with the observed probability of the psychology probability set corresponding to the newly observed psychology to identify an abnormal observed probability among the observed probabilities of the psychology probability integration vector used to observe the newly observed psychology, and updates the observed probability, and the observation unit transmits the updated psychology probability integration vector and the observation instruction to the generation system AI to re-observe the user's potential psychology.

[0136] By adopting this configuration, it becomes possible to efficiently and accurately observe psychology by identifying which observed probabilities are abnormal in the psychological probability integration vector where abnormal psychology is observed.

[0137] Invention 9 further comprises an update unit, which determines, based on the goodness-of-fit score, that the newly observed observed psychology is abnormal, and if it determines that it is abnormal, it refers to a plurality of psychology probability integration vectors generated at a plurality of acquisition timings within a predetermined time after the observation time of the newly observed psychology, and updates the observed probability of the psychology probability integration vector used to observe the newly observed psychology based on the psychology and observed probability of the psychology for each of the plurality of psychology probability integration vectors, and the observation unit transmits the updated psychology probability integration vector and the observation instruction to the generation system AI to re-observe the user's potential psychology.

[0138] This configuration allows for updating the integrated psychological stochastic vector based on multiple new modality signals.

[0139] Invention 10 further comprises an update unit, wherein the state generation unit processes the generation of basis vectors corresponding to the plurality of different potential psychology, the observation unit transmits the psychology probability integration vector composed of the basis vectors and the observation instruction to the generation system AI to observe the potential psychology, the update unit generates a new basis vector if one of the potential psychology is not observed and updates the psychology probability integration vector based on the new basis vector, and the observation unit transmits the psychology probability integration vector composed of the new basis vector and the observation instruction to the generation system AI to re-observe the potential psychology.

[0140] This configuration improves the poor convergence that occurs when the basis vectors generated by the generative AI are abnormal.

[0141] Invention 11 is a dialogue method using a dialogue system that analyzes the user's psychology from a dialogue with the user, wherein the dialogue system performs the following processes: a process of acquiring user signal information including multiple types of modality signals obtained from a dialogue with the user and the acquisition timing of each of the modality signals; a process of analyzing the acquired modality signals and generating a psychological element probability vector for each modality signal that shows a subconscious structure having multiple different psychological states of the user and the observation probability of each psychological state; and a process of integrating the multiple psychological element probability vectors that have the same acquisition timing with the multiple types of modality signals to generate a psychological probability integration vector that shows the subconscious structure based on the multiple types of modality signals.

[0142] Invention 12 is a dialogue program that analyzes a user's psychology from a dialogue with the user, wherein the dialogue program causes a computer to function as an acquisition unit and a state generation unit, the acquisition unit acquires user signal information including a plurality of modality signals obtained from the dialogue with the user and the acquisition timing of each of the modality signals, the state generation unit analyzes the acquired modality signals and generates a psychological element probability vector for each modality signal that shows a subconscious structure having multiple different psychology of the user and the observation probability of each psychology, and based on the plurality of psychological element probability vectors with the same acquisition timing, integrates the multiple different psychology of the user and the observation probability of each psychology in the plurality of modality signals and generates a psychological probability integration vector that shows the subconscious structure based on the plurality of modality signals.

[0143] Furthermore, according to the present invention, it is possible to visualize the interaction of subconscious structures among multiple users not only in the real world but also in virtual spaces such as the metaverse, thereby obtaining new effects such as learning support and psychological care.

[0144] 1 Dialogue System 10 Mobile Terminal (User Terminal Device) 11 Voice Acquisition Unit 12 Text Acquisition Unit 13 Image Acquisition Unit 14 Language Data Creation Unit 15 Text-to-Speech Means 16 Image Display Means 101 Processing Unit 102 Storage Unit 103 Communication Unit 104 Input Unit 105 Output Unit 20 Server 21 Response Data Generation Unit 22 Image Generation Unit 201 Acquisition Unit 202 State Generation Unit 203 Correction State Generation Unit 204 Update Unit 205 Observation Unit 206 Fit Calculation Unit 207 Output Processing Unit 2001 Processing Unit 2002 Storage Unit 2003 Communication Unit S2 Transmission Data Creation S3 Response Data Creation

Claims

1. A dialogue system for analyzing a user's psychology from a conversation with the user, wherein the dialogue system comprises an acquisition unit and a state generation unit, the acquisition unit acquires user signal information including multiple types of modality signals obtained from a conversation with the user and the acquisition timing of each of the modality signals, the state generation unit analyzes the acquired modality signals and generates a psychological element probability vector for each modality signal that shows a subconscious structure having multiple different psychological states and the observation probability of each psychological state, and the dialogue system integrates the multiple different psychological states and the observation probability of each psychological state in the multiple types of modality signals based on the multiple psychological element probability vectors having the same acquisition timing, and generates a psychological probability integration vector that shows the subconscious structure based on the multiple types of modality signals.

2. The dialogue system further comprises a correction state generation unit and an update unit, wherein the correction state generation unit generates a psychological probability correction vector for correcting a new subconscious structure based on past subconscious structures, based on the changes in the observed probabilities of each psychology between the consecutive psychological element probability vectors in a time series of dialogues, and the update unit inputs the psychological probability integration vector and the psychological probability correction vector into a mathematical model to update the observed probabilities of each psychology in the psychological probability integration vector, according to claim 1.

3. The update unit generates the psychological probability correction vector |G| based on the transition of the observed probability between the psychological probability integration vector |ψ> at a certain acquisition timing, the psychological element probability vector at that acquisition timing, and the psychological element probability vector at the acquisition timing immediately preceding that acquisition timing. j > and a predetermined weight b for each of the multiple modality signals corresponding to the psychological element probability vector j The mathematical model including: The dialogue system according to claim 2, which uses to obtain the psychological probability integration vector |ψ'> having the observed probability of each updated psychological state.

4. The dialogue system further comprises an observation unit and an output processing unit, wherein the observation unit transmits the psychological probability integration vector and an observation instruction to observe the user's psychology based on the observed probability of each psychology in the psychological probability integration vector to a generating AI, causing it to observe one of the plurality of different potential psychology at the acquisition timing, and the output processing unit outputs dialogue data based on the observed psychology, which is the observed psychology of the user, to the user, as described in claim 1 or 2.

5. The dialogue system according to claim 4, relating to claim 1, further comprising an update unit, wherein the state generation unit processes the generation of basis vectors corresponding to the plurality of different potential psychology, the observation unit transmits the psychology probability integration vector composed of the basis vectors and the observation instruction to the generation system AI to observe the potential psychology, the update unit generates a new basis vector if one of the potential psychology is not observed and updates the psychology probability integration vector based on the new basis vector, and the observation unit transmits the psychology probability integration vector composed of the new basis vector and the observation instruction to the generation system AI to re-observe the potential psychology.

6. The dialogue system further comprises a goodness-of-fit calculation unit, wherein the psychological probability integration vector has psychological probability sets which are combinations of a plurality of different psychological states and the observed probabilities of each psychological state, the observation unit registers observation information in a database in chronological order, which is linked to the psychological probability sets and the observed psychological states observed based on the psychological probability integration vector having the psychological probability sets, and the goodness-of-fit calculation unit calculates a transition rate that indicates the confidence level for each observed psychological state observed based on the other psychological probability set when transitioning from one psychological probability set to another psychological probability set, based on the observation information.

7. The dialogue system according to claim 6, wherein the observation unit transmits the integrated psychological probability vector used to observe the newly observed psychological state and the observation instruction to the generative AI, in accordance with the newly observed psychological state, the psychological probability set corresponding to the newly observed psychological state, and the transition rate between the psychological probability set corresponding to the observed psychological state observed immediately before the observation of the said psychological state, causing the AI ​​to re-observe the user's latent psychological state.

8. The dialogue system further comprises a correction state generation unit, the correction state generation unit determines whether or not to generate a psychological probability correction vector for correcting a new subconscious structure based on past subconscious structures, based on the transition of the observed probability of each psychological state between the consecutive psychological element probability vectors in a time series of dialogues, registers the generated psychological probability correction vector in association with the dialogue in the database, and the goodness-of-fit calculation unit calculates a goodness-of-fit score indicating the observation reliability of the newly observed psychological state based on the number of psychological probability correction vectors associated with the dialogue prior to the observation time of the newly observed psychological state, and / or the transition rate indicating the reliability of the newly observed psychological state. The dialogue system according to claim 6, which is dependent on claim 4, which references claim 1.

9. The dialogue system further comprises an update unit, which determines, based on the goodness-of-fit score, that the newly observed observed psychology is abnormal, and compares the transition rate for each observed psychology between the psychology probability set corresponding to the newly observed psychology and the psychology probability set corresponding to the psychology observed immediately before the observation of the said psychology with the observed probability of the psychology probability set corresponding to the newly observed psychology to identify an abnormal observed probability among the observed probabilities of the psychology probability integration vector used to observe the newly observed psychology, and updates the observed probability, and the observation unit transmits the updated psychology probability integration vector and the observation instruction to the generation AI to re-observe the user's potential psychology, as described in claim 8.

10. The dialogue system further comprises an update unit, which determines, based on the goodness-of-fit score, that the newly observed observed psychology is abnormal, and if it determines that it is abnormal, it refers to a plurality of psychology probability integration vectors generated at a plurality of acquisition timings within a predetermined time after the observation time of the newly observed psychology, and updates the observed probability of the psychology probability integration vector used to observe the newly observed psychology based on the psychology and the observed probability of the psychology for each of the plurality of psychology probability integration vectors, and the observation unit transmits the updated psychology probability integration vector and the observation instruction to the generation system AI to re-observe the user's potential psychology, the dialogue system according to claim 8.

11. A dialogue method using a dialogue system that analyzes the user's psychology from a dialogue with the user, wherein the dialogue system performs the following steps: a process of acquiring user signal information including multiple types of modality signals obtained from a dialogue with the user and the acquisition timing of each of the modality signals; a process of analyzing the acquired modality signals and generating a psychological element probability vector for each of the modality signals that shows a subconscious structure having multiple different psychological states and the observation probability of each psychological state; and a process of integrating the multiple psychological element probability vectors that have the same acquisition timing with respect to the multiple types of modality signals to generate a psychological probability integration vector that shows the subconscious structure based on the multiple types of modality signals.

12. A dialogue program for analyzing a user's psychology from a conversation with the user, wherein the dialogue program causes a computer to function as an acquisition unit and a state generation unit, the acquisition unit acquires user signal information including multiple types of modality signals obtained from a conversation with the user and the acquisition timing of each of the modality signals, the state generation unit analyzes the acquired modality signals and generates a psychological element probability vector for each modality signal that shows a subconscious structure having multiple different psychological states and the observation probability of each psychological state, and based on the multiple psychological element probability vectors having the same acquisition timing, the dialogue program integrates the multiple different psychological states of the user and the observation probability of each psychological state in the multiple types of modality signals and generates a psychological probability integration vector that shows the subconscious structure based on the multiple types of modality signals.