system

The system addresses the challenge of creating tailored conversational scenarios by generating natural-sounding text using user input and natural language processing, enhancing communication skills and interpersonal relationships.

JP2026068312APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Users who are not good at communication find it difficult to create a natural conversation flow, and conventional conversation example materials are limited, failing to provide variations tailored to individual needs, making it hard to improve communication skills and maintain interpersonal relationships effectively.

Method used

A system that receives user input, selects a generation algorithm based on specified information, and generates conversational text using natural language processing technology to provide tailored conversation examples.

Benefits of technology

Enables users to practice and improve their communication skills in specific situations by providing natural-sounding conversational scenarios that meet individual needs, facilitating smoother interpersonal relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068312000001_ABST
    Figure 2026068312000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving information specified by the user, means for selecting a generation algorithm based on the aforementioned information, Means for generating conversational text using the aforementioned generation algorithm, A means for sending the generated conversation text to the user's terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Users who are not good at communication find it difficult to create a natural conversation flow, and as a result, they have obstacles in building and maintaining interpersonal relationships. Conventional conversation example materials are limited to general scenarios and have the problem that they cannot provide variations of conversations tailored to the individual needs of users. Also, due to the cost and time constraints in conducting conversation practice, an effective solution that anyone can easily use is required.

Means for Solving the Problems

[0005] This invention provides a system that receives information specified by a user, selects a generation algorithm based on that information, and generates conversational text. This system includes means for transmitting the generated conversational text to the user's terminal, and by applying natural language processing technology, it can provide natural conversational examples that meet the needs of individual users. As a result, users can effectively practice conversation even in specific situations, improving their communication skills and enabling the smooth construction and maintenance of interpersonal relationships.

[0006] A "user" is an individual or group that intends to generate conversational text using this system.

[0007] "Information" refers to data specified by the user, such as keywords and conversational tone necessary for generating dialogue.

[0008] A "generation algorithm" is an algorithm used to generate natural-sounding conversational sentences based on specified information.

[0009] "Dialogue text" refers to conversational text that is generated by a generative algorithm and has a natural flow.

[0010] A "user terminal" is an electronic device, such as a personal computer or mobile device, that has the function of receiving and displaying generated conversation text.

[0011] "Natural language processing technology" is a technology that enables computers to understand, process, and generate natural human language. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPUs (Central Processing Units), GPUs (Graphics Processing Units), GPGPUs (General-Purpose computing on Graphics Processing Units), APUs (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memories (SSDs (Solid State Drives)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention is a system for users to generate natural conversational example sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm. The specific roles and processing details of each element are described below.

[0034] First, the user uses their device to input information such as conversation topics, keywords, and desired conversation tone. This clarifies the direction and style of conversation the user desires. The device then collects this input data and prepares to send it to the server.

[0035] Next, the server receives this information. Based on the received information, the server selects the optimal generation algorithm. This algorithm is designed based on natural language processing technology and works to generate the most appropriate conversational sentences from the given information.

[0036] The server inputs information into a selected generation algorithm and automatically generates natural-sounding conversational text based on the user's conditions. The generated conversational text is then formatted in a way that is intuitively understandable to the user.

[0037] Next, the server sends the generated conversation text to the user's terminal. The terminal displays the received conversation text on its screen, making it easy for the user to use.

[0038] For example, if a user wants a casual conversation about "planning a trip," the generation algorithm might create a casual conversational text that includes travel-related phrases and questions. For instance, the conversation might unfold as follows:

[0039] When a user asks, "Which country will you visit on your next vacation?", the system may generate responses such as, "How about Italy? It has amazing food and plenty of tourist attractions."

[0040] Thus, by utilizing the system of the present invention, users can practice a variety of conversation scenarios and improve their practical communication skills tailored to specific situations.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user uses the device to input conversation topics, keywords, and desired conversation tone. The device receives this information through the on-screen input form and organizes it as a data package.

[0044] Step 2:

[0045] The terminal sends the organized data package to the server. The sent data package contains all the conversation conditions specified by the user.

[0046] Step 3:

[0047] The server receives the data package and analyzes the information within it. Based on the analyzed information, it selects the most suitable generation algorithm and configures the necessary settings.

[0048] Step 4:

[0049] The server invokes a selected generation algorithm to generate natural-sounding conversational text that meets the user's specified conditions. The generation algorithm uses natural language processing techniques to create conversational patterns appropriate to the specified topic and tone.

[0050] Step 5:

[0051] The server-generated conversation text is formatted to make it easier for the user to understand. Proofreading and grammar checks are also performed as needed.

[0052] Step 6:

[0053] The server sends the formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it easy for the user to view.

[0054] Step 7:

[0055] Users can review the conversation text displayed on their device and use it as a reference or practice tool for actual conversations. By repeating this process, users can try out various conversation scenarios and improve their communication skills.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] In today's world, many people need to communicate effectively in diverse situations. However, generating natural conversation requires specialized knowledge, and creating conversational content tailored to each individual is not easy. As a result, there is a lack of effective means to practice specific conversational scenarios and improve practical communication skills. To address this challenge, there is a need for a new system that automatically generates natural conversational sentences tailored to individual conversational conditions.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for receiving information specified by the user and storing that information in a database, means for selecting a generative model using natural language processing based on the received information, and means for inputting prompt data into the selected generative model to generate conversational data. This makes it possible to automatically generate and provide natural conversational scenarios that meet the user's needs.

[0061] A "user" refers to a person who wants to operate the system and generate conversations.

[0062] "Information" refers to data that includes the conversation topic, keywords, and tone of conversation specified by the user.

[0063] A "database" refers to an internal storage device within a system used to temporarily or permanently store received information.

[0064] A "generative model" refers to an algorithm that uses natural language processing techniques to generate conversational text based on specified prompts.

[0065] "Prompt data" refers to data containing instructions and prerequisite information to be input into a generative model.

[0066] "Conversation data" refers to natural-sounding conversational text generated by generative models.

[0067] "Information equipment" refers to electronic devices owned or used by users, and is the subject to which conversation data is transmitted.

[0068] This invention is a system that allows users to easily generate natural conversational sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm.

[0069] First, the user uses their device to input the topic, keywords, and tone of the conversation they want to generate. For example, a user might want a casual conversation on the topic of "planning a trip." This information is collected by the user's device and sent to the server.

[0070] The server stores the received information in a database and selects the most suitable generative model. Publicly available natural language processing frameworks are used for this selection. This generative model is based on libraries such as Transformers, and is capable of generating natural-sounding conversational text from large amounts of training data.

[0071] Next, the server inputs prompts into the selected generative model to generate conversation data. The prompts are based on the topic and tone specified by the user. For example, a prompt might say, "Generate a casual conversation about travel planning. The tone should be friendly, and Italy should be included as a travel destination."

[0072] The generated conversation data is formatted by the server and sent to the user's information device. The information device can then display this conversation data, which the user can view and use to practice real-life conversations.

[0073] In this way, the system of the present invention provides conversation scenarios tailored to user needs quickly and efficiently, and helps users practically improve their communication skills.

[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0075] Step 1:

[0076] The user inputs the topic, keywords, and tone of the conversation they want to generate using their device. Upon receiving this input, the device temporarily stores it in its internal memory and prepares to send it to the server. The input consists of text data specified by the user, and the output is a formatted data packet.

[0077] Step 2:

[0078] The terminal sends formatted data packets to the server. The server receives these packets and records the information in its internal database system. Specifically, the server analyzes the received data and stores it in the appropriate database table based on attributes such as topic, keywords, and tone. The input is the data packets sent from the terminal, and the output is the writing of information to the database.

[0079] Step 3:

[0080] The server selects the most appropriate generative model based on the information stored in the database. The server uses an algorithm based on a natural language processing framework to identify a suitable model. Based on the selected model, it forms a prompt sentence and prepares it for input to the model. The input is information read from the database, and the output is the prompt sentence.

[0081] Step 4:

[0082] The server inputs a prompt sentence into the selected generative model and performs the generation of conversational data. Here, the AI ​​model generates natural conversational sentences according to the given prompt. The input is a prompt sentence, and the output is the generated conversational data. Specifically, the generative AI engine is instantiated, and the prompt is fed into the AI ​​model.

[0083] Step 5:

[0084] The generated conversation data is formatted by the server and sent to the terminal. The server converts the conversation data into a user-friendly format, i.e., a highly readable text format, and sends it to the terminal as a data packet. The input is conversation data obtained from an AI model, and the output is formatted data packets sent to the terminal.

[0085] Step 6:

[0086] The terminal receives conversation data packets sent from the server and displays them on the screen. At this time, the terminal interprets the data and visualizes it in a way that is easily understandable to the user. The input is data packets received from the server, and the output is a screen display that the user can see. Specifically, text is rendered on the display and presented through the user interface.

[0087] (Application Example 1)

[0088] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0089] In today's world, improving effective and natural communication skills is required in many situations. However, conventional learning methods make it difficult to acquire conversational skills appropriate to individual contexts, and suggesting appropriate conversations in real time is a major challenge. To solve this problem, there is a need to develop a system that generates and presents conversations in real time, tailored to the user's utterances.

[0090] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0091] In this invention, the server includes means for an information processing device that includes a function for receiving information specified by a user, means for a processing device that includes a function for selecting a generation method based on the information, and means for converting voice data into text data. This enables the generation and presentation of conversations that are natural and contextual to the user in real time.

[0092] "User-specified information" refers to topics, keywords, and conversational tone that users input into the system, and serves as a guideline for conversation generation.

[0093] "Means for converting audio data to text data" refers to a device or program that uses speech recognition technology to perform the process of converting audio input into text format.

[0094] A "processing device that includes a function for selecting a generation method" is a device or program that can select the optimal natural language processing method based on the received information.

[0095] "Means for generating conversation data" refers to an algorithm or program that automatically generates appropriate conversation content based on specified information or converted text data.

[0096] A "display device" is a device that provides generated conversation data to the user visually, and includes screens, smart glasses, and the like.

[0097] This invention presents a specific form of a system that supports the improvement of a user's natural conversational abilities. First, the user inputs information such as topics, keywords, and conversational tone using the terminal. This information is intended to clarify the user's intentions and desires. The terminal processes this information and prepares it for transmission to the server.

[0098] The server converts speech input into text data using speech recognition technologies such as Google® Cloud Speech-to-Text API. This converted text data is analyzed along with the user-specified information mentioned above, and the optimal generation method is selected. Specifically, conversation data is generated based on the obtained information using the Python language and natural language processing libraries (e.g., NLTK and SpaCy). The generated conversation data is then transmitted in real time to the user's display device, such as smart glasses, using WebSocket technology. This display allows the user to receive appropriate conversation suggestions in real time, improving their practical communication skills.

[0099] As a concrete example, in an international trade show setting, it is envisioned that conversation examples related to the user's utterances would be displayed to facilitate smooth conversations with new customers. An example of a prompt would be, "The topic of conversation is music technology. Please generate appropriate conversation examples in a natural tone and with a light touch."

[0100] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0101] Step 1:

[0102] The user inputs the conversation topic, keywords, and desired tone into the device. This clarifies the user's intent and prepares the necessary input data for subsequent processing. The input data is temporarily stored on the device and prepared to be sent to the server later.

[0103] Step 2:

[0104] The terminal sends the prepared input data to the server. This transmission utilizes an internet connection to ensure the data reaches the server securely and quickly. The transmitted data is structured in text format.

[0105] Step 3:

[0106] The server receives the transmitted data and, if it contains audio data, converts it to text data using the Google Cloud Speech-to-Text API. The converted text data is recorded in a database and used to further select parsing and generation algorithms.

[0107] Step 4:

[0108] The server uses Python and related natural language processing libraries to determine which generative AI model is most appropriate based on the analyzed text data. This determination is based on the content and patterns of the input data, and the selected model is then used for subsequent conversation generation.

[0109] Step 5:

[0110] The selected generative AI model is given prompt sentences to generate conversational data. These prompt sentences may include phrases such as, "The topic of conversation is music technology. Please generate appropriate conversational examples with a natural tone and lightheartedness." The generated conversational data is then formatted according to a specific format.

[0111] Step 6:

[0112] The server sends the generated conversation data to the terminal using a dedicated communication protocol (e.g., WebSocket). This communication is conducted in real time and is designed to minimize latency.

[0113] Step 7:

[0114] The device visually presents the received conversation data to the user, for example, through the display of smart glasses. This allows the user to engage in appropriate conversations in real time with visual support.

[0115] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0116] This invention provides a system that offers a richer communication experience by recognizing the user's emotions and incorporating them into the conversation generation process. This system is implemented using a user terminal, a server, a generation algorithm, and an emotion engine. The role and processing of each element are described below.

[0117] First, when the user uses the device to input information such as conversation topics, keywords, and tone of voice, the emotion engine simultaneously activates to analyze the user's emotions. This analysis is performed using the content of the input text and, if necessary, the user's voice data. The emotion engine uses a specific algorithm to recognize the user's tension and emotional state in real time.

[0118] Next, the device combines this information with the recognized emotion data and sends it to the server. The server analyzes the received data and selects the optimal generation algorithm. This selection also takes into account the emotions recognized by the emotion engine. For example, if the user is excited, the server can prioritize selecting an algorithm that generates more lively and energetic conversational text.

[0119] The server generates conversational text that matches the user's conditions according to a selected generation algorithm. This process utilizes natural language processing techniques to create conversational text that reflects the specified topic, tone, and emotion.

[0120] Furthermore, the server formats the generated conversation text into a format that is easily understandable to the user before sending it to the user's terminal. The conversation text received on the terminal is displayed in the user interface, and the user can use it as a reference for actual conversations.

[0121] For example, if a user is having a conversation about "planning a trip," and the emotion engine detects the user's excitement, the generation algorithm can generate lively conversational sentences that match that state, such as "Are you looking forward to new adventures on your next trip? What are your plans?"

[0122] This format allows users to obtain conversation examples that better suit their own emotions, and the system provides situation-appropriate feedback, enabling a mutually enjoyable and enriching communication experience.

[0123] The following describes the processing flow.

[0124] Step 1:

[0125] The user uses their device to input conversation topics, keywords, and desired conversational tone. In addition, the user's voice and text data are provided to the sentiment engine.

[0126] Step 2:

[0127] The device sends the input information and the user's voice or text data to the emotion engine. The emotion engine then analyzes the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0128] Step 3:

[0129] The emotion engine returns the analysis results to the device. The device then packages this emotion data along with topic and tone information into a data package.

[0130] Step 4:

[0131] The terminal sends a data package to the server. The package contains conversational topics, keywords, tone, and sentiment data.

[0132] Step 5:

[0133] The server analyzes the received data package and selects the optimal generation algorithm. This selection includes adjustments that reflect the user's emotional state.

[0134] Step 6:

[0135] The server uses a selected generation algorithm to generate natural-sounding conversational text that matches the user's criteria. The conversational text reflects the specified topic, tone, and emotion.

[0136] Step 7:

[0137] The server formats the generated dialogue into a format that is easily understandable to the user. The formatted dialogue is then provided in a context that is appropriate to the emotion and tone of the conversation.

[0138] Step 8:

[0139] The server sends formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it intuitive for the user to use.

[0140] Step 9:

[0141] The system allows users to review conversational text presented on their device and incorporate content that matches their emotions into their actual conversations. This enables users to practice smoother conversational scenarios and improve their communication skills.

[0142] (Example 2)

[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0144] Conventional conversation generation systems have faced challenges in generating conversational text that adequately considers the user's emotional state, making it difficult to provide a natural and appropriate communication experience. In particular, it has been difficult to achieve flexible responses that adapt to changes in the user's emotions.

[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0146] In this invention, the server includes means for analyzing information and emotions specified by the user, means for selecting a generation algorithm based on the information and analyzed emotions, and means for generating conversational text that reflects the user's emotional state using the generation algorithm. This makes it possible to generate conversational text adapted to the user's emotional state in real time, providing natural and rich communication.

[0147] "User-specified information" refers to data that users input for conversation generation, including conversation topics, keywords, and tone of conversation.

[0148] "Means of analyzing emotions" refers to processes and algorithms that recognize a user's emotional state in real time based on their input or voice data.

[0149] A "generation algorithm" is a computational method for generating optimal conversational text based on information obtained from users and analyzed sentiment data.

[0150] "Natural language processing technology" refers to a set of algorithms and techniques that enable computers to understand, generate, or interpret human language.

[0151] A "prompt sentence" is a sentence that serves as an instruction or hint, used by a generative AI model when generating conversational text.

[0152] "User's terminal" refers to a device operated by the user that enables the input of conversations and the display of generated conversational text.

[0153] This invention provides a more personalized communication experience by analyzing the user's emotions and generating conversations accordingly. This system is implemented using a user-operated terminal, a data processing server, a generation algorithm, and an emotion analysis engine.

[0154] The user enters conversation topics and keywords using the device, and the sentiment engine activates at that time. The sentiment analysis engine on the device analyzes the entered text and audio data. The specific analysis uses natural language processing and speech signal processing techniques, and a particular algorithm is applied to recognize the user's emotional state in real time.

[0155] Next, the device sends the collected information and analyzed sentiment data to the server. Based on the received data, the server selects a generation algorithm and creates a prompt that is appropriate for the user's sentiment. For example, if the user selects the topic "Planning a Trip" and the sentiment engine detects exhilaration, the prompt will include a phrase such as "What are you looking forward to on your next trip?"

[0156] The server uses a selected generative AI model to generate conversational text based on the prompt. The generated conversational text is refined using natural language processing techniques, reflecting the user's emotions and the theme of the conversation.

[0157] The completed conversation is formatted by the server, converted into an easily understandable format, and then sent to the terminal. The terminal displays this conversation on its user interface, allowing the user to utilize it to create richer and more meaningful conversations.

[0158] This configuration allows the system to have the flexibility to provide appropriate feedback in response to the user's emotional state and communication needs, enabling a high-quality user experience.

[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0160] Step 1:

[0161] The user operates the terminal to input the conversation topic, keywords, and desired tone. The terminal receives this information and compiles it as prompt data. The input data is stored in text format for later use in sentiment analysis.

[0162] Step 2:

[0163] The device activates an emotion analysis engine and analyzes the user's input text. If voice input is available, it performs voice signal processing to estimate the emotional state from the tone, speed, and intonation of the voice. The analysis engine processes the data using natural language processing techniques to identify the user's emotional state. The output is metadata indicating the user's emotional state.

[0164] Step 3:

[0165] The terminal sends the user's topics and keywords, along with the analyzed sentiment data, to the server. The server reviews the received data and selects the optimal generation algorithm. In doing so, it considers the user's emotional state and sets conditions for generating appropriate prompt sentences.

[0166] Step 4:

[0167] The server uses a generation algorithm to generate conversational text based on prompts. A generative AI model is employed here to output responses that match the user's input and emotions. The generated data is conversational text in text format.

[0168] Step 5:

[0169] The server formats the generated dialogue, changing it into a format that is easily understandable to the user. This includes optimizing word choice and adjusting grammar. The formatted dialogue is the final data that aids user comprehension.

[0170] Step 6:

[0171] The server sends the formatted conversation text back to the terminal. The terminal displays the received conversation text in its user interface, allowing the user to review it. The user then uses this displayed information as a reference for their actual conversation.

[0172] (Application Example 2)

[0173] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0174] Traditional communication systems generate conversations without considering the user's emotions, making it difficult to respond flexibly to the individual user's situation and feelings. Furthermore, enriching the customer experience in online shopping requires customer service tailored to the user's emotions, but conventional technologies lack the means to achieve this. This results in the challenge of not being able to sufficiently increase user satisfaction and purchasing intent.

[0175] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0176] In this invention, the server includes means for selecting a generation algorithm based on information specified by the user and the user's emotional state, means for recognizing the user's emotions using emotion analysis technology, and means for transmitting the generated conversation text to the user's terminal. This makes it possible to generate conversation text that responds to the user's emotions in real time and provide a communication experience tailored to each individual.

[0177] "User-specified information" refers to data such as arbitrary topics, keywords, and tones that the user provides to the system.

[0178] A "generative algorithm" is a series of processes that create the most suitable dialogue based on specified conditions.

[0179] "Emotional state" refers to data that indicates the emotional state exhibited by the user, and includes emotions such as joy, sadness, and excitement.

[0180] "Emotion analysis technology" is a technology that recognizes, classifies, and analyzes a user's emotions based on their statements and voice data.

[0181] "Natural language processing technology" refers to technologies that enable computers to understand and process human language, and includes sentence generation and sentiment analysis.

[0182] "User terminal" refers to a computer, smartphone, or other communication device used by the user.

[0183] This invention relates to a system for generating conversations that respond to a user's emotions, and its implementation is achieved using a server, a terminal, and various software. This system receives information specified by the user using a terminal and recognizes the user's emotional state using emotion analysis technology. Based on this, the server selects the optimal generation algorithm and generates conversational text using natural language processing technology.

[0184] The server uses emotion analysis technology, specifically an emotion engine (e.g., Watson® Tone Analyzer), to analyze the user's emotions in real time based on their text and voice data. The analyzed emotion data is then input as prompts into a generative AI model (e.g., the GPT model).

[0185] This system allows a generative AI model to generate conversational text that matches the user's emotions, which is then sent from the server to the terminal and displayed on the user interface. Ultimately, the user can use this generated conversational text to have a smooth conversation on the terminal.

[0186] As a concrete example, consider a scenario where a device receives information from a user saying, "I want to buy new clothes," and sentiment analysis technology identifies the user's level of excitement. As a result, the server can generate an enthusiastic conversational message such as, "I think these clothes would look great on you! You'll be the center of attention at the party!"

[0187] A concrete example of a prompt message is an instruction given to the generation AI model, such as, "The user is highly excited, so please generate an energetic and positive product description."

[0188] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0189] Step 1:

[0190] The user uses the device to input their desired topic, keywords, and conversational tone. The device uses this information as input data, activates an emotion engine, and analyzes the user's emotional state. This process yields emotion analysis results from the input text data.

[0191] Step 2:

[0192] The terminal sends analyzed emotion data and user-input data to the server. The server receives both sets of data and selects an appropriate generation algorithm. In this selection process, the emotion data is used as a criterion for choosing a generation algorithm that is appropriate for the user's emotions.

[0193] Step 3:

[0194] The server uses a selected generation algorithm to provide prompt sentences to the generative AI model, generating conversational text. The input includes the user's emotional state, topic, and keywords, and the output is conversational text that matches these elements. In this process, the generative AI model constructs appropriate sentences based on natural language processing techniques.

[0195] Step 4:

[0196] The server formats the generated conversation text into a format that is easily understandable to the user. The formatted conversation text is then sent to the terminal as the final output. In this way, the generated text obtained in step 3 is fed back into the user's conversation in a practical way.

[0197] Step 5:

[0198] The terminal displays received conversation text on the user interface, allowing the user to communicate smoothly by referring to the text. This output is based on user interaction and is used to enrich the overall communication experience of the system.

[0199] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0200] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0201] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0202] [Second Embodiment]

[0203] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0204] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0205] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0206] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0207] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0208] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0209] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0210] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0211] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0212] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0213] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0214] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0215] This invention is a system for users to generate natural conversational example sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm. The specific roles and processing details of each element are described below.

[0216] First, the user uses their device to input information such as conversation topics, keywords, and desired conversation tone. This clarifies the direction and style of conversation the user desires. The device then collects this input data and prepares to send it to the server.

[0217] Next, the server receives this information. Based on the received information, the server selects the optimal generation algorithm. This algorithm is designed based on natural language processing technology and works to generate the most appropriate conversational sentences from the given information.

[0218] The server inputs information into a selected generation algorithm and automatically generates natural-sounding conversational text based on the user's conditions. The generated conversational text is then formatted in a way that is intuitively understandable to the user.

[0219] Next, the server sends the generated conversation text to the user's terminal. The terminal displays the received conversation text on its screen, making it easy for the user to use.

[0220] For example, if a user wants a casual conversation about "planning a trip," the generation algorithm might create a casual conversational text that includes travel-related phrases and questions. For instance, the conversation might unfold as follows:

[0221] When a user asks, "Which country will you visit on your next vacation?", the system may generate responses such as, "How about Italy? It has amazing food and plenty of tourist attractions."

[0222] Thus, by utilizing the system of the present invention, users can practice a variety of conversation scenarios and improve their practical communication skills tailored to specific situations.

[0223] The following describes the processing flow.

[0224] Step 1:

[0225] The user uses the device to input conversation topics, keywords, and desired conversation tone. The device receives this information through the on-screen input form and organizes it as a data package.

[0226] Step 2:

[0227] The terminal sends the organized data package to the server. The sent data package contains all the conversation conditions specified by the user.

[0228] Step 3:

[0229] The server receives the data package and analyzes the information within it. Based on the analyzed information, it selects the most suitable generation algorithm and configures the necessary settings.

[0230] Step 4:

[0231] The server invokes a selected generation algorithm to generate natural-sounding conversational text that meets the user's specified conditions. The generation algorithm uses natural language processing techniques to create conversational patterns appropriate to the specified topic and tone.

[0232] Step 5:

[0233] The server-generated conversation text is formatted to make it easier for the user to understand. Proofreading and grammar checks are also performed as needed.

[0234] Step 6:

[0235] The server sends the formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it easy for the user to view.

[0236] Step 7:

[0237] Users can review the conversation text displayed on their device and use it as a reference or practice tool for actual conversations. By repeating this process, users can try out various conversation scenarios and improve their communication skills.

[0238] (Example 1)

[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0240] In today's world, many people need to communicate effectively in diverse situations. However, generating natural conversation requires specialized knowledge, and creating conversational content tailored to each individual is not easy. As a result, there is a lack of effective means to practice specific conversational scenarios and improve practical communication skills. To address this challenge, there is a need for a new system that automatically generates natural conversational sentences tailored to individual conversational conditions.

[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0242] In this invention, the server includes means for receiving information specified by the user and storing that information in a database, means for selecting a generative model using natural language processing based on the received information, and means for inputting prompt data into the selected generative model to generate conversational data. This makes it possible to automatically generate and provide natural conversational scenarios that meet the user's needs.

[0243] A "user" refers to a person who wants to operate the system and generate conversations.

[0244] "Information" refers to data that includes the conversation topic, keywords, and tone of conversation specified by the user.

[0245] A "database" refers to an internal storage device within a system used to temporarily or permanently store received information.

[0246] A "generative model" refers to an algorithm that uses natural language processing techniques to generate conversational text based on specified prompts.

[0247] "Prompt data" refers to data containing instructions and prerequisite information to be input into a generative model.

[0248] "Conversation data" refers to natural-sounding conversational text generated by generative models.

[0249] "Information equipment" refers to electronic devices owned or used by users, and is the subject to which conversation data is transmitted.

[0250] This invention is a system that allows users to easily generate natural conversational sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm.

[0251] First, the user uses their device to input the topic, keywords, and tone of the conversation they want to generate. For example, a user might want a casual conversation on the topic of "planning a trip." This information is collected by the user's device and sent to the server.

[0252] The server stores the received information in a database and selects the most suitable generative model. Publicly available natural language processing frameworks are used for this selection. This generative model is based on libraries such as Transformers, and is capable of generating natural-sounding conversational text from large amounts of training data.

[0253] Next, the server inputs prompts into the selected generative model to generate conversation data. The prompts are based on the topic and tone specified by the user. For example, a prompt might say, "Generate a casual conversation about travel planning. The tone should be friendly, and Italy should be included as a travel destination."

[0254] The generated conversation data is formatted by the server and sent to the user's information device. The information device can then display this conversation data, which the user can view and use to practice real-life conversations.

[0255] In this way, the system of the present invention provides conversation scenarios tailored to user needs quickly and efficiently, and helps users practically improve their communication skills.

[0256] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0257] Step 1:

[0258] The user inputs the topic, keywords, and tone of the conversation they want to generate using their device. Upon receiving this input, the device temporarily stores it in its internal memory and prepares to send it to the server. The input consists of text data specified by the user, and the output is a formatted data packet.

[0259] Step 2:

[0260] The terminal sends formatted data packets to the server. The server receives these packets and records the information in its internal database system. Specifically, the server analyzes the received data and stores it in the appropriate database table based on attributes such as topic, keywords, and tone. The input is the data packets sent from the terminal, and the output is the writing of information to the database.

[0261] Step 3:

[0262] The server selects the most appropriate generative model based on the information stored in the database. The server uses an algorithm based on a natural language processing framework to identify a suitable model. Based on the selected model, it forms a prompt sentence and prepares it for input to the model. The input is information read from the database, and the output is the prompt sentence.

[0263] Step 4:

[0264] The server inputs a prompt sentence into the selected generative model and performs the generation of conversational data. Here, the AI ​​model generates natural conversational sentences according to the given prompt. The input is a prompt sentence, and the output is the generated conversational data. Specifically, the generative AI engine is instantiated, and the prompt is fed into the AI ​​model.

[0265] Step 5:

[0266] The generated conversation data is formatted by the server and sent to the terminal. The server converts the conversation data into a user-friendly format, i.e., a highly readable text format, and sends it to the terminal as a data packet. The input is conversation data obtained from an AI model, and the output is formatted data packets sent to the terminal.

[0267] Step 6:

[0268] The terminal receives conversation data packets sent from the server and displays them on the screen. At this time, the terminal interprets the data and visualizes it in a way that is easily understandable to the user. The input is data packets received from the server, and the output is a screen display that the user can see. Specifically, text is rendered on the display and presented through the user interface.

[0269] (Application Example 1)

[0270] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0271] In today's world, improving effective and natural communication skills is required in many situations. However, conventional learning methods make it difficult to acquire conversational skills appropriate to individual contexts, and suggesting appropriate conversations in real time is a major challenge. To solve this problem, there is a need to develop a system that generates and presents conversations in real time, tailored to the user's utterances.

[0272] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0273] In this invention, the server includes means for an information processing device that includes a function for receiving information specified by a user, means for a processing device that includes a function for selecting a generation method based on the information, and means for converting voice data into text data. This enables the generation and presentation of conversations that are natural and contextual to the user in real time.

[0274] "User-specified information" refers to topics, keywords, and conversational tone that users input into the system, and serves as a guideline for conversation generation.

[0275] "Means for converting audio data to text data" refers to a device or program that uses speech recognition technology to perform the process of converting audio input into text format.

[0276] A "processing device that includes a function for selecting a generation method" is a device or program that can select the optimal natural language processing method based on the received information.

[0277] "Means for generating conversation data" refers to an algorithm or program that automatically generates appropriate conversation content based on specified information or converted text data.

[0278] A "display device" is a device that provides generated conversation data to the user visually, and includes screens, smart glasses, and the like.

[0279] This invention presents a specific form of a system that supports the improvement of a user's natural conversational abilities. First, the user inputs information such as topics, keywords, and conversational tone using the terminal. This information is intended to clarify the user's intentions and desires. The terminal processes this information and prepares it for transmission to the server.

[0280] The server converts speech input into text data using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is analyzed along with the user-specified information mentioned earlier, and the optimal generation method is selected. Specifically, conversation data is generated based on the obtained information using the Python language and natural language processing libraries (e.g., NLTK and SpaCy). The generated conversation data is then transmitted in real time to the user's display device, such as smart glasses, using WebSocket technology. This display allows the user to receive appropriate conversation suggestions in real time, improving their practical communication skills.

[0281] As a concrete example, in an international trade show setting, it is envisioned that conversation examples related to the user's utterances would be displayed to facilitate smooth conversations with new customers. An example of a prompt would be, "The topic of conversation is music technology. Please generate appropriate conversation examples in a natural tone and with a light touch."

[0282] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0283] Step 1:

[0284] The user inputs the conversation topic, keywords, and desired tone into the terminal. This concretizes the user's intention and prepares the input data necessary for subsequent processing. The input data is temporarily stored in the terminal and is then readied for transmission to the server later.

[0285] Step 2:

[0286] The terminal transmits the prepared input data to the server. This transmission utilizes the Internet connection to ensure that the data reaches the server safely and quickly. The data to be transmitted is structured in text format.

[0287] Step 3:

[0288] The server receives the transmitted data and, if the data includes audio data, uses the Google Cloud Speech-to-Text API to convert it into text data. The converted text data is recorded in the database and is further utilized for the selection of analysis and generation algorithms.

[0289] Step 4:

[0290] The server uses Python and related natural language processing libraries to determine which generation AI model is most appropriate based on the analyzed text data. This determination is made based on the content and pattern of the input data, and the selected model is then utilized for subsequent conversation generation.

[0291] Step 5:

[0292] An input prompt sentence is entered into the selected generation AI model to generate conversation data. This prompt sentence includes, for example, sentences such as "The conversation theme is music technology. Please generate a natural tone and lively appropriate conversation example." The generated conversation data is formatted in a certain format.

[0293] Step 6:

[0294] The server sends the generated conversation data to the terminal using a dedicated communication protocol (e.g., WebSocket). This communication is conducted in real time and is designed to minimize latency.

[0295] Step 7:

[0296] The device visually presents the received conversation data to the user, for example, through the display of smart glasses. This allows the user to engage in appropriate conversations in real time with visual support.

[0297] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0298] This invention provides a system that offers a richer communication experience by recognizing the user's emotions and incorporating them into the conversation generation process. This system is implemented using a user terminal, a server, a generation algorithm, and an emotion engine. The role and processing of each element are described below.

[0299] First, when the user uses the device to input information such as conversation topics, keywords, and tone of voice, the emotion engine simultaneously activates to analyze the user's emotions. This analysis is performed using the content of the input text and, if necessary, the user's voice data. The emotion engine uses a specific algorithm to recognize the user's tension and emotional state in real time.

[0300] Next, the terminal summarizes this information and the recognized emotion data and sends it to the server. The server analyzes the received data and selects an optimal generation algorithm. This selection also takes into account the emotions recognized by the emotion engine. For example, if the user is in an excited state, an algorithm that generates more lively and energetic conversation sentences can be preferentially selected.

[0301] The server generates a conversation sentence that meets the user's conditions according to the selected generation algorithm. In this process, natural language processing technology is used to create a conversation sentence that reflects the specified topic, tone, and emotion.

[0302] Furthermore, the server formats the generated conversation sentence into a form that the user can easily understand and then sends it to the user's terminal. The conversation sentence received on the terminal is displayed on the user interface, and the user can use it as a reference for actual conversations.

[0303] As a specific example, if the user is having a conversation about "travel plans" and the emotion engine detects the user's excitement, the generation algorithm can generate lively conversation sentences such as "Are you looking forward to new adventures in this trip? What plans do you have?" that match that state.

[0304] In this way, the user can obtain conversation examples adapted to their own emotions, and the system can provide feedback according to the situation, making it possible to realize a mutually enjoyable and rich communication experience.

[0305] The following explains the processing flow.

[0306] Step 1:

[0307] The user uses the terminal to input the topic, keywords, and desired tone of the conversation. In addition, the user's voice and text data are provided to the emotion engine.

[0308] Step 2:

[0309] The device sends the input information and the user's voice or text data to the emotion engine. The emotion engine then analyzes the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0310] Step 3:

[0311] The emotion engine returns the analysis results to the device. The device then packages this emotion data along with topic and tone information into a data package.

[0312] Step 4:

[0313] The terminal sends a data package to the server. The package contains conversational topics, keywords, tone, and sentiment data.

[0314] Step 5:

[0315] The server analyzes the received data package and selects the optimal generation algorithm. This selection includes adjustments that reflect the user's emotional state.

[0316] Step 6:

[0317] The server uses a selected generation algorithm to generate natural-sounding conversational text that matches the user's criteria. The conversational text reflects the specified topic, tone, and emotion.

[0318] Step 7:

[0319] The server formats the generated dialogue into a format that is easily understandable to the user. The formatted dialogue is then provided in a context that is appropriate to the emotion and tone of the conversation.

[0320] Step 8:

[0321] The server sends formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it intuitive for the user to use.

[0322] Step 9:

[0323] The system allows users to review conversational text presented on their device and incorporate content that matches their emotions into their actual conversations. This enables users to practice smoother conversational scenarios and improve their communication skills.

[0324] (Example 2)

[0325] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0326] Conventional conversation generation systems have faced challenges in generating conversational text that adequately considers the user's emotional state, making it difficult to provide a natural and appropriate communication experience. In particular, it has been difficult to achieve flexible responses that adapt to changes in the user's emotions.

[0327] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0328] In this invention, the server includes means for analyzing information and emotions specified by the user, means for selecting a generation algorithm based on the information and analyzed emotions, and means for generating conversational text that reflects the user's emotional state using the generation algorithm. This makes it possible to generate conversational text adapted to the user's emotional state in real time, providing natural and rich communication.

[0329] "User-specified information" refers to data that users input for conversation generation, including conversation topics, keywords, and tone of conversation.

[0330] "Means of analyzing emotions" refers to processes and algorithms that recognize a user's emotional state in real time based on their input or voice data.

[0331] A "generation algorithm" is a computational method for generating optimal conversational text based on information obtained from users and analyzed sentiment data.

[0332] "Natural language processing technology" refers to a set of algorithms and techniques that enable computers to understand, generate, or interpret human language.

[0333] A "prompt sentence" is a sentence that serves as an instruction or hint, used by a generative AI model when generating conversational text.

[0334] "User's terminal" refers to a device operated by the user that enables the input of conversations and the display of generated conversational text.

[0335] This invention provides a more personalized communication experience by analyzing the user's emotions and generating conversations accordingly. This system is implemented using a user-operated terminal, a data processing server, a generation algorithm, and an emotion analysis engine.

[0336] The user enters conversation topics and keywords using the device, and the sentiment engine activates at that time. The sentiment analysis engine on the device analyzes the entered text and audio data. The specific analysis uses natural language processing and speech signal processing techniques, and a particular algorithm is applied to recognize the user's emotional state in real time.

[0337] Next, the device sends the collected information and analyzed sentiment data to the server. Based on the received data, the server selects a generation algorithm and creates a prompt that is appropriate for the user's sentiment. For example, if the user selects the topic "Planning a Trip" and the sentiment engine detects exhilaration, the prompt will include a phrase such as "What are you looking forward to on your next trip?"

[0338] The server uses a selected generative AI model to generate conversational text based on the prompt. The generated conversational text is refined using natural language processing techniques, reflecting the user's emotions and the theme of the conversation.

[0339] The completed conversation is formatted by the server, converted into an easily understandable format, and then sent to the terminal. The terminal displays this conversation on its user interface, allowing the user to utilize it to create richer and more meaningful conversations.

[0340] This configuration allows the system to have the flexibility to provide appropriate feedback in response to the user's emotional state and communication needs, enabling a high-quality user experience.

[0341] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0342] Step 1:

[0343] The user operates the terminal to input the conversation topic, keywords, and desired tone. The terminal receives this information and compiles it as prompt data. The input data is stored in text format for later use in sentiment analysis.

[0344] Step 2:

[0345] The device activates an emotion analysis engine and analyzes the user's input text. If voice input is available, it performs voice signal processing to estimate the emotional state from the tone, speed, and intonation of the voice. The analysis engine processes the data using natural language processing techniques to identify the user's emotional state. The output is metadata indicating the user's emotional state.

[0346] Step 3:

[0347] The terminal sends the user's topics and keywords, along with the analyzed sentiment data, to the server. The server reviews the received data and selects the optimal generation algorithm. In doing so, it considers the user's emotional state and sets conditions for generating appropriate prompt sentences.

[0348] Step 4:

[0349] The server uses a generation algorithm to generate conversational text based on prompts. A generative AI model is employed here to output responses that match the user's input and emotions. The generated data is conversational text in text format.

[0350] Step 5:

[0351] The server formats the generated dialogue, changing it into a format that is easily understandable to the user. This includes optimizing word choice and adjusting grammar. The formatted dialogue is the final data that aids user comprehension.

[0352] Step 6:

[0353] The server sends the formatted conversation text back to the terminal. The terminal displays the received conversation text in its user interface, allowing the user to review it. The user then uses this displayed information as a reference for their actual conversation.

[0354] (Application Example 2)

[0355] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0356] Traditional communication systems generate conversations without considering the user's emotions, making it difficult to respond flexibly to the individual user's situation and feelings. Furthermore, enriching the customer experience in online shopping requires customer service tailored to the user's emotions, but conventional technologies lack the means to achieve this. This results in the challenge of not being able to sufficiently increase user satisfaction and purchasing intent.

[0357] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0358] In this invention, the server includes means for selecting a generation algorithm based on information specified by the user and the user's emotional state, means for recognizing the user's emotions using emotion analysis technology, and means for transmitting the generated conversation text to the user's terminal. This makes it possible to generate conversation text that responds to the user's emotions in real time and provide a communication experience tailored to each individual.

[0359] "User-specified information" refers to data such as arbitrary topics, keywords, and tones that the user provides to the system.

[0360] A "generative algorithm" is a series of processes that create the most suitable dialogue based on specified conditions.

[0361] "Emotional state" refers to data that indicates the emotional state exhibited by the user, and includes emotions such as joy, sadness, and excitement.

[0362] "Emotion analysis technology" is a technology that recognizes, classifies, and analyzes a user's emotions based on their statements and voice data.

[0363] "Natural language processing technology" refers to technologies that enable computers to understand and process human language, and includes sentence generation and sentiment analysis.

[0364] "User terminal" refers to a computer, smartphone, or other communication device used by the user.

[0365] This invention relates to a system for generating conversations that respond to a user's emotions, and its implementation is achieved using a server, a terminal, and various software. This system receives information specified by the user using a terminal and recognizes the user's emotional state using emotion analysis technology. Based on this, the server selects the optimal generation algorithm and generates conversational text using natural language processing technology.

[0366] The server uses emotion analysis technology, specifically an emotion engine (e.g., Watson Tone Analyzer), to analyze the user's emotions in real time based on their text and voice data. The analyzed emotion data is then input as prompts into a generative AI model (e.g., the GPT model).

[0367] This system allows a generative AI model to generate conversational text that matches the user's emotions, which is then sent from the server to the terminal and displayed on the user interface. Ultimately, the user can use this generated conversational text to have a smooth conversation on the terminal.

[0368] As a concrete example, consider a scenario where a device receives information from a user saying, "I want to buy new clothes," and sentiment analysis technology identifies the user's level of excitement. As a result, the server can generate an enthusiastic conversational message such as, "I think these clothes would look great on you! You'll be the center of attention at the party!"

[0369] A concrete example of a prompt message is an instruction given to the generation AI model, such as, "The user is highly excited, so please generate an energetic and positive product description."

[0370] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0371] Step 1:

[0372] The user uses the device to input their desired topic, keywords, and conversational tone. The device uses this information as input data, activates an emotion engine, and analyzes the user's emotional state. This process yields emotion analysis results from the input text data.

[0373] Step 2:

[0374] The terminal sends analyzed emotion data and user-input data to the server. The server receives both sets of data and selects an appropriate generation algorithm. In this selection process, the emotion data is used as a criterion for choosing a generation algorithm that is appropriate for the user's emotions.

[0375] Step 3:

[0376] The server uses a selected generation algorithm to provide prompt sentences to the generative AI model, generating conversational text. The input includes the user's emotional state, topic, and keywords, and the output is conversational text that matches these elements. In this process, the generative AI model constructs appropriate sentences based on natural language processing techniques.

[0377] Step 4:

[0378] The server formats the generated conversation text into a format that is easily understandable to the user. The formatted conversation text is then sent to the terminal as the final output. In this way, the generated text obtained in step 3 is fed back into the user's conversation in a practical way.

[0379] Step 5:

[0380] The terminal displays received conversation text on the user interface, allowing the user to communicate smoothly by referring to the text. This output is based on user interaction and is used to enrich the overall communication experience of the system.

[0381] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0382] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0383] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0384] [Third Embodiment]

[0385] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0386] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0387] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0388] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0389] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0390] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0391] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0392] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0393] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0394] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0395] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0396] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0397] This invention is a system for users to generate natural conversational example sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm. The specific roles and processing details of each element are described below.

[0398] First, the user uses their device to input information such as conversation topics, keywords, and desired conversation tone. This clarifies the direction and style of conversation the user desires. The device then collects this input data and prepares to send it to the server.

[0399] Next, the server receives this information. Based on the received information, the server selects the optimal generation algorithm. This algorithm is designed based on natural language processing technology and works to generate the most appropriate conversational sentences from the given information.

[0400] The server inputs information into a selected generation algorithm and automatically generates natural-sounding conversational text based on the user's conditions. The generated conversational text is then formatted in a way that is intuitively understandable to the user.

[0401] Next, the server sends the generated conversation text to the user's terminal. The terminal displays the received conversation text on its screen, making it easy for the user to use.

[0402] For example, if a user wants a casual conversation about "planning a trip," the generation algorithm might create a casual conversational text that includes travel-related phrases and questions. For instance, the conversation might unfold as follows:

[0403] When a user asks, "Which country will you visit on your next vacation?", the system may generate responses such as, "How about Italy? It has amazing food and plenty of tourist attractions."

[0404] Thus, by utilizing the system of the present invention, users can practice a variety of conversation scenarios and improve their practical communication skills tailored to specific situations.

[0405] The following describes the processing flow.

[0406] Step 1:

[0407] The user uses the device to input conversation topics, keywords, and desired conversation tone. The device receives this information through the on-screen input form and organizes it as a data package.

[0408] Step 2:

[0409] The terminal sends the organized data package to the server. The sent data package contains all the conversation conditions specified by the user.

[0410] Step 3:

[0411] The server receives the data package and analyzes the information within it. Based on the analyzed information, it selects the most suitable generation algorithm and configures the necessary settings.

[0412] Step 4:

[0413] The server invokes a selected generation algorithm to generate natural-sounding conversational text that meets the user's specified conditions. The generation algorithm uses natural language processing techniques to create conversational patterns appropriate to the specified topic and tone.

[0414] Step 5:

[0415] The server-generated conversation text is formatted to make it easier for the user to understand. Proofreading and grammar checks are also performed as needed.

[0416] Step 6:

[0417] The server sends the formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it easy for the user to view.

[0418] Step 7:

[0419] Users can review the conversation text displayed on their device and use it as a reference or practice tool for actual conversations. By repeating this process, users can try out various conversation scenarios and improve their communication skills.

[0420] (Example 1)

[0421] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0422] In today's world, many people need to communicate effectively in diverse situations. However, generating natural conversation requires specialized knowledge, and creating conversational content tailored to each individual is not easy. As a result, there is a lack of effective means to practice specific conversational scenarios and improve practical communication skills. To address this challenge, there is a need for a new system that automatically generates natural conversational sentences tailored to individual conversational conditions.

[0423] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0424] In this invention, the server includes means for receiving information specified by the user and storing that information in a database, means for selecting a generative model using natural language processing based on the received information, and means for inputting prompt data into the selected generative model to generate conversational data. This makes it possible to automatically generate and provide natural conversational scenarios that meet the user's needs.

[0425] A "user" refers to a person who wants to operate the system and generate conversations.

[0426] "Information" refers to data that includes the conversation topic, keywords, and tone of conversation specified by the user.

[0427] A "database" refers to an internal storage device within a system used to temporarily or permanently store received information.

[0428] A "generative model" refers to an algorithm that uses natural language processing techniques to generate conversational text based on specified prompts.

[0429] "Prompt data" refers to data containing instructions and prerequisite information to be input into a generative model.

[0430] "Conversation data" refers to natural-sounding conversational text generated by generative models.

[0431] "Information equipment" refers to electronic devices owned or used by users, and is the subject to which conversation data is transmitted.

[0432] This invention is a system that allows users to easily generate natural conversational sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm.

[0433] First, the user uses their device to input the topic, keywords, and tone of the conversation they want to generate. For example, a user might want a casual conversation on the topic of "planning a trip." This information is collected by the user's device and sent to the server.

[0434] The server stores the received information in a database and selects the most suitable generative model. Publicly available natural language processing frameworks are used for this selection. This generative model is based on libraries such as Transformers, and is capable of generating natural-sounding conversational text from large amounts of training data.

[0435] Next, the server inputs prompts into the selected generative model to generate conversation data. The prompts are based on the topic and tone specified by the user. For example, a prompt might say, "Generate a casual conversation about travel planning. The tone should be friendly, and Italy should be included as a travel destination."

[0436] The generated conversation data is formatted by the server and sent to the user's information device. The information device can then display this conversation data, which the user can view and use to practice real-life conversations.

[0437] In this way, the system of the present invention provides conversation scenarios tailored to user needs quickly and efficiently, and helps users practically improve their communication skills.

[0438] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0439] Step 1:

[0440] The user inputs the topic, keywords, and tone of the conversation they want to generate using their device. Upon receiving this input, the device temporarily stores it in its internal memory and prepares to send it to the server. The input consists of text data specified by the user, and the output is a formatted data packet.

[0441] Step 2:

[0442] The terminal sends formatted data packets to the server. The server receives these packets and records the information in its internal database system. Specifically, the server analyzes the received data and stores it in the appropriate database table based on attributes such as topic, keywords, and tone. The input is the data packets sent from the terminal, and the output is the writing of information to the database.

[0443] Step 3:

[0444] The server selects the most appropriate generative model based on the information stored in the database. The server uses an algorithm based on a natural language processing framework to identify a suitable model. Based on the selected model, it forms a prompt sentence and prepares it for input to the model. The input is information read from the database, and the output is the prompt sentence.

[0445] Step 4:

[0446] The server inputs a prompt sentence into the selected generative model and performs the generation of conversational data. Here, the AI ​​model generates natural conversational sentences according to the given prompt. The input is a prompt sentence, and the output is the generated conversational data. Specifically, the generative AI engine is instantiated, and the prompt is fed into the AI ​​model.

[0447] Step 5:

[0448] The generated conversation data is formatted by the server and sent to the terminal. The server converts the conversation data into a user-friendly format, i.e., a highly readable text format, and sends it to the terminal as a data packet. The input is conversation data obtained from an AI model, and the output is formatted data packets sent to the terminal.

[0449] Step 6:

[0450] The terminal receives conversation data packets sent from the server and displays them on the screen. At this time, the terminal interprets the data and visualizes it in a way that is easily understandable to the user. The input is data packets received from the server, and the output is a screen display that the user can see. Specifically, text is rendered on the display and presented through the user interface.

[0451] (Application Example 1)

[0452] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0453] In today's world, improving effective and natural communication skills is required in many situations. However, conventional learning methods make it difficult to acquire conversational skills appropriate to individual contexts, and suggesting appropriate conversations in real time is a major challenge. To solve this problem, there is a need to develop a system that generates and presents conversations in real time, tailored to the user's utterances.

[0454] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0455] In this invention, the server includes means for an information processing device that includes a function for receiving information specified by a user, means for a processing device that includes a function for selecting a generation method based on the information, and means for converting voice data into text data. This enables the generation and presentation of conversations that are natural and contextual to the user in real time.

[0456] "User-specified information" refers to topics, keywords, and conversational tone that users input into the system, and serves as a guideline for conversation generation.

[0457] "Means for converting audio data to text data" refers to a device or program that uses speech recognition technology to perform the process of converting audio input into text format.

[0458] A "processing device that includes a function for selecting a generation method" is a device or program that can select the optimal natural language processing method based on the received information.

[0459] "Means for generating conversation data" refers to an algorithm or program that automatically generates appropriate conversation content based on specified information or converted text data.

[0460] A "display device" is a device that provides generated conversation data to the user visually, and includes screens, smart glasses, and the like.

[0461] This invention presents a specific form of a system that supports the improvement of a user's natural conversational abilities. First, the user inputs information such as topics, keywords, and conversational tone using the terminal. This information is intended to clarify the user's intentions and desires. The terminal processes this information and prepares it for transmission to the server.

[0462] The server converts speech input into text data using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is analyzed along with the user-specified information mentioned earlier, and the optimal generation method is selected. Specifically, conversation data is generated based on the obtained information using the Python language and natural language processing libraries (e.g., NLTK and SpaCy). The generated conversation data is then transmitted in real time to the user's display device, such as smart glasses, using WebSocket technology. This display allows the user to receive appropriate conversation suggestions in real time, improving their practical communication skills.

[0463] As a concrete example, in an international trade show setting, it is envisioned that conversation examples related to the user's utterances would be displayed to facilitate smooth conversations with new customers. An example of a prompt would be, "The topic of conversation is music technology. Please generate appropriate conversation examples in a natural tone and with a light touch."

[0464] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0465] Step 1:

[0466] The user inputs the conversation topic, keywords, and desired tone into the device. This clarifies the user's intent and prepares the necessary input data for subsequent processing. The input data is temporarily stored on the device and prepared to be sent to the server later.

[0467] Step 2:

[0468] The terminal sends the prepared input data to the server. This transmission utilizes an internet connection to ensure the data reaches the server securely and quickly. The transmitted data is structured in text format.

[0469] Step 3:

[0470] The server receives the transmitted data and, if it contains audio data, converts it to text data using the Google Cloud Speech-to-Text API. The converted text data is recorded in a database and used to further select parsing and generation algorithms.

[0471] Step 4:

[0472] The server uses Python and related natural language processing libraries to determine which generative AI model is most appropriate based on the analyzed text data. This determination is based on the content and patterns of the input data, and the selected model is then used for subsequent conversation generation.

[0473] Step 5:

[0474] The selected generative AI model is given prompt sentences to generate conversational data. These prompt sentences may include phrases such as, "The topic of conversation is music technology. Please generate appropriate conversational examples with a natural tone and lightheartedness." The generated conversational data is then formatted according to a specific format.

[0475] Step 6:

[0476] The server sends the generated conversation data to the terminal using a dedicated communication protocol (e.g., WebSocket). This communication is conducted in real time and is designed to minimize latency.

[0477] Step 7:

[0478] The device visually presents the received conversation data to the user, for example, through the display of smart glasses. This allows the user to engage in appropriate conversations in real time with visual support.

[0479] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0480] This invention provides a system that offers a richer communication experience by recognizing the user's emotions and incorporating them into the conversation generation process. This system is implemented using a user terminal, a server, a generation algorithm, and an emotion engine. The role and processing of each element are described below.

[0481] First, when the user uses the device to input information such as conversation topics, keywords, and tone of voice, the emotion engine simultaneously activates to analyze the user's emotions. This analysis is performed using the content of the input text and, if necessary, the user's voice data. The emotion engine uses a specific algorithm to recognize the user's tension and emotional state in real time.

[0482] Next, the device combines this information with the recognized emotion data and sends it to the server. The server analyzes the received data and selects the optimal generation algorithm. This selection also takes into account the emotions recognized by the emotion engine. For example, if the user is excited, the server can prioritize selecting an algorithm that generates more lively and energetic conversational text.

[0483] The server generates conversational text that matches the user's conditions according to a selected generation algorithm. This process utilizes natural language processing techniques to create conversational text that reflects the specified topic, tone, and emotion.

[0484] Furthermore, the server formats the generated conversation text into a format that is easily understandable to the user before sending it to the user's terminal. The conversation text received on the terminal is displayed in the user interface, and the user can use it as a reference for actual conversations.

[0485] For example, if a user is having a conversation about "planning a trip," and the emotion engine detects the user's excitement, the generation algorithm can generate lively conversational sentences that match that state, such as "Are you looking forward to new adventures on your next trip? What are your plans?"

[0486] This format allows users to obtain conversation examples that better suit their own emotions, and the system provides situation-appropriate feedback, enabling a mutually enjoyable and enriching communication experience.

[0487] The following describes the processing flow.

[0488] Step 1:

[0489] The user uses their device to input conversation topics, keywords, and desired conversational tone. In addition, the user's voice and text data are provided to the sentiment engine.

[0490] Step 2:

[0491] The device sends the input information and the user's voice or text data to the emotion engine. The emotion engine then analyzes the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0492] Step 3:

[0493] The emotion engine returns the analysis results to the device. The device then packages this emotion data along with topic and tone information into a data package.

[0494] Step 4:

[0495] The terminal sends a data package to the server. The package contains conversational topics, keywords, tone, and sentiment data.

[0496] Step 5:

[0497] The server analyzes the received data package and selects the optimal generation algorithm. This selection includes adjustments that reflect the user's emotional state.

[0498] Step 6:

[0499] The server uses a selected generation algorithm to generate natural-sounding conversational text that matches the user's criteria. The conversational text reflects the specified topic, tone, and emotion.

[0500] Step 7:

[0501] The server formats the generated dialogue into a format that is easily understandable to the user. The formatted dialogue is then provided in a context that is appropriate to the emotion and tone of the conversation.

[0502] Step 8:

[0503] The server sends formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it intuitive for the user to use.

[0504] Step 9:

[0505] The system allows users to review conversational text presented on their device and incorporate content that matches their emotions into their actual conversations. This enables users to practice smoother conversational scenarios and improve their communication skills.

[0506] (Example 2)

[0507] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0508] Conventional conversation generation systems have faced challenges in generating conversational text that adequately considers the user's emotional state, making it difficult to provide a natural and appropriate communication experience. In particular, it has been difficult to achieve flexible responses that adapt to changes in the user's emotions.

[0509] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0510] In this invention, the server includes means for analyzing information and emotions specified by the user, means for selecting a generation algorithm based on the information and analyzed emotions, and means for generating conversational text that reflects the user's emotional state using the generation algorithm. This makes it possible to generate conversational text adapted to the user's emotional state in real time, providing natural and rich communication.

[0511] "User-specified information" refers to data that users input for conversation generation, including conversation topics, keywords, and tone of conversation.

[0512] "Means of analyzing emotions" refers to processes and algorithms that recognize a user's emotional state in real time based on their input or voice data.

[0513] A "generation algorithm" is a computational method for generating optimal conversational text based on information obtained from users and analyzed sentiment data.

[0514] "Natural language processing technology" refers to a set of algorithms and techniques that enable computers to understand, generate, or interpret human language.

[0515] A "prompt sentence" is a sentence that serves as an instruction or hint, used by a generative AI model when generating conversational text.

[0516] "User's terminal" refers to a device operated by the user that enables the input of conversations and the display of generated conversational text.

[0517] This invention provides a more personalized communication experience by analyzing the user's emotions and generating conversations accordingly. This system is implemented using a user-operated terminal, a data processing server, a generation algorithm, and an emotion analysis engine.

[0518] The user enters conversation topics and keywords using the device, and the sentiment engine activates at that time. The sentiment analysis engine on the device analyzes the entered text and audio data. The specific analysis uses natural language processing and speech signal processing techniques, and a particular algorithm is applied to recognize the user's emotional state in real time.

[0519] Next, the device sends the collected information and analyzed sentiment data to the server. Based on the received data, the server selects a generation algorithm and creates a prompt that is appropriate for the user's sentiment. For example, if the user selects the topic "Planning a Trip" and the sentiment engine detects exhilaration, the prompt will include a phrase such as "What are you looking forward to on your next trip?"

[0520] The server uses a selected generative AI model to generate conversational text based on the prompt. The generated conversational text is refined using natural language processing techniques, reflecting the user's emotions and the theme of the conversation.

[0521] The completed conversation is formatted by the server, converted into an easily understandable format, and then sent to the terminal. The terminal displays this conversation on its user interface, allowing the user to utilize it to create richer and more meaningful conversations.

[0522] This configuration allows the system to have the flexibility to provide appropriate feedback in response to the user's emotional state and communication needs, enabling a high-quality user experience.

[0523] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0524] Step 1:

[0525] The user operates the terminal to input the conversation topic, keywords, and desired tone. The terminal receives this information and compiles it as prompt data. The input data is stored in text format for later use in sentiment analysis.

[0526] Step 2:

[0527] The device activates an emotion analysis engine and analyzes the user's input text. If voice input is available, it performs voice signal processing to estimate the emotional state from the tone, speed, and intonation of the voice. The analysis engine processes the data using natural language processing techniques to identify the user's emotional state. The output is metadata indicating the user's emotional state.

[0528] Step 3:

[0529] The terminal sends the user's topics and keywords, along with the analyzed sentiment data, to the server. The server reviews the received data and selects the optimal generation algorithm. In doing so, it considers the user's emotional state and sets conditions for generating appropriate prompt sentences.

[0530] Step 4:

[0531] The server uses a generation algorithm to generate conversational text based on prompts. A generative AI model is employed here to output responses that match the user's input and emotions. The generated data is conversational text in text format.

[0532] Step 5:

[0533] The server formats the generated dialogue, changing it into a format that is easily understandable to the user. This includes optimizing word choice and adjusting grammar. The formatted dialogue is the final data that aids user comprehension.

[0534] Step 6:

[0535] The server sends the formatted conversation text back to the terminal. The terminal displays the received conversation text in its user interface, allowing the user to review it. The user then uses this displayed information as a reference for their actual conversation.

[0536] (Application Example 2)

[0537] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0538] Traditional communication systems generate conversations without considering the user's emotions, making it difficult to respond flexibly to the individual user's situation and feelings. Furthermore, enriching the customer experience in online shopping requires customer service tailored to the user's emotions, but conventional technologies lack the means to achieve this. This results in the challenge of not being able to sufficiently increase user satisfaction and purchasing intent.

[0539] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0540] In this invention, the server includes means for selecting a generation algorithm based on information specified by the user and the user's emotional state, means for recognizing the user's emotions using emotion analysis technology, and means for transmitting the generated conversation text to the user's terminal. This makes it possible to generate conversation text that responds to the user's emotions in real time and provide a communication experience tailored to each individual.

[0541] "User-specified information" refers to data such as arbitrary topics, keywords, and tones that the user provides to the system.

[0542] A "generative algorithm" is a series of processes that create the most suitable dialogue based on specified conditions.

[0543] "Emotional state" refers to data that indicates the emotional state exhibited by the user, and includes emotions such as joy, sadness, and excitement.

[0544] "Emotion analysis technology" is a technology that recognizes, classifies, and analyzes a user's emotions based on their statements and voice data.

[0545] "Natural language processing technology" refers to technologies that enable computers to understand and process human language, and includes sentence generation and sentiment analysis.

[0546] "User terminal" refers to a computer, smartphone, or other communication device used by the user.

[0547] This invention relates to a system for generating conversations that respond to a user's emotions, and its implementation is achieved using a server, a terminal, and various software. This system receives information specified by the user using a terminal and recognizes the user's emotional state using emotion analysis technology. Based on this, the server selects the optimal generation algorithm and generates conversational text using natural language processing technology.

[0548] The server uses emotion analysis technology, specifically an emotion engine (e.g., Watson Tone Analyzer), to analyze the user's emotions in real time based on their text and voice data. The analyzed emotion data is then input as prompts into a generative AI model (e.g., the GPT model).

[0549] This system allows a generative AI model to generate conversational text that matches the user's emotions, which is then sent from the server to the terminal and displayed on the user interface. Ultimately, the user can use this generated conversational text to have a smooth conversation on the terminal.

[0550] As a concrete example, consider a scenario where a device receives information from a user saying, "I want to buy new clothes," and sentiment analysis technology identifies the user's level of excitement. As a result, the server can generate an enthusiastic conversational message such as, "I think these clothes would look great on you! You'll be the center of attention at the party!"

[0551] A concrete example of a prompt message is an instruction given to the generation AI model, such as, "The user is highly excited, so please generate an energetic and positive product description."

[0552] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0553] Step 1:

[0554] The user uses the device to input their desired topic, keywords, and conversational tone. The device uses this information as input data, activates an emotion engine, and analyzes the user's emotional state. This process yields emotion analysis results from the input text data.

[0555] Step 2:

[0556] The terminal sends analyzed emotion data and user-input data to the server. The server receives both sets of data and selects an appropriate generation algorithm. In this selection process, the emotion data is used as a criterion for choosing a generation algorithm that is appropriate for the user's emotions.

[0557] Step 3:

[0558] The server uses a selected generation algorithm to provide prompt sentences to the generative AI model, generating conversational text. The input includes the user's emotional state, topic, and keywords, and the output is conversational text that matches these elements. In this process, the generative AI model constructs appropriate sentences based on natural language processing techniques.

[0559] Step 4:

[0560] The server formats the generated conversation text into a format that is easily understandable to the user. The formatted conversation text is then sent to the terminal as the final output. In this way, the generated text obtained in step 3 is fed back into the user's conversation in a practical way.

[0561] Step 5:

[0562] The terminal displays received conversation text on the user interface, allowing the user to communicate smoothly by referring to the text. This output is based on user interaction and is used to enrich the overall communication experience of the system.

[0563] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0564] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0565] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0566] [Fourth Embodiment]

[0567] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0568] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0569] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0570] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0571] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0572] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0573] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0574] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0575] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0576] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0577] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0578] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0579] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0580] This invention is a system for users to generate natural conversational example sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm. The specific roles and processing details of each element are described below.

[0581] First, the user uses their device to input information such as conversation topics, keywords, and desired conversation tone. This clarifies the direction and style of conversation the user desires. The device then collects this input data and prepares to send it to the server.

[0582] Next, the server receives this information. Based on the received information, the server selects the optimal generation algorithm. This algorithm is designed based on natural language processing technology and works to generate the most appropriate conversational sentences from the given information.

[0583] The server inputs information into a selected generation algorithm and automatically generates natural-sounding conversational text based on the user's conditions. The generated conversational text is then formatted in a way that is intuitively understandable to the user.

[0584] Next, the server sends the generated conversation text to the user's terminal. The terminal displays the received conversation text on its screen, making it easy for the user to use.

[0585] For example, if a user wants a casual conversation about "planning a trip," the generation algorithm might create a casual conversational text that includes travel-related phrases and questions. For instance, the conversation might unfold as follows:

[0586] When a user asks, "Which country will you visit on your next vacation?", the system may generate responses such as, "How about Italy? It has amazing food and plenty of tourist attractions."

[0587] Thus, by utilizing the system of the present invention, users can practice a variety of conversation scenarios and improve their practical communication skills tailored to specific situations.

[0588] The following describes the processing flow.

[0589] Step 1:

[0590] The user uses the device to input conversation topics, keywords, and desired conversation tone. The device receives this information through the on-screen input form and organizes it as a data package.

[0591] Step 2:

[0592] The terminal sends the organized data package to the server. The sent data package contains all the conversation conditions specified by the user.

[0593] Step 3:

[0594] The server receives the data package and analyzes the information within it. Based on the analyzed information, it selects the most suitable generation algorithm and configures the necessary settings.

[0595] Step 4:

[0596] The server invokes a selected generation algorithm to generate natural-sounding conversational text that meets the user's specified conditions. The generation algorithm uses natural language processing techniques to create conversational patterns appropriate to the specified topic and tone.

[0597] Step 5:

[0598] The server-generated conversation text is formatted to make it easier for the user to understand. Proofreading and grammar checks are also performed as needed.

[0599] Step 6:

[0600] The server sends the formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it easy for the user to view.

[0601] Step 7:

[0602] Users can review the conversation text displayed on their device and use it as a reference or practice tool for actual conversations. By repeating this process, users can try out various conversation scenarios and improve their communication skills.

[0603] (Example 1)

[0604] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0605] In today's world, many people need to communicate effectively in diverse situations. However, generating natural conversation requires specialized knowledge, and creating conversational content tailored to each individual is not easy. As a result, there is a lack of effective means to practice specific conversational scenarios and improve practical communication skills. To address this challenge, there is a need for a new system that automatically generates natural conversational sentences tailored to individual conversational conditions.

[0606] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0607] In this invention, the server includes means for receiving information specified by the user and storing that information in a database, means for selecting a generative model using natural language processing based on the received information, and means for inputting prompt data into the selected generative model to generate conversational data. This makes it possible to automatically generate and provide natural conversational scenarios that meet the user's needs.

[0608] A "user" refers to a person who wants to operate the system and generate conversations.

[0609] "Information" refers to data that includes the conversation topic, keywords, and tone of conversation specified by the user.

[0610] A "database" refers to an internal storage device within a system used to temporarily or permanently store received information.

[0611] A "generative model" refers to an algorithm that uses natural language processing techniques to generate conversational text based on specified prompts.

[0612] "Prompt data" refers to data containing instructions and prerequisite information to be input into a generative model.

[0613] "Conversation data" refers to natural-sounding conversational text generated by generative models.

[0614] "Information equipment" refers to electronic devices owned or used by users, and is the subject to which conversation data is transmitted.

[0615] This invention is a system that allows users to easily generate natural conversational sentences and thereby improve their communication skills. This system is implemented using a user terminal, a server, and a generation algorithm.

[0616] First, the user uses their device to input the topic, keywords, and tone of the conversation they want to generate. For example, a user might want a casual conversation on the topic of "planning a trip." This information is collected by the user's device and sent to the server.

[0617] The server stores the received information in a database and selects the most suitable generative model. Publicly available natural language processing frameworks are used for this selection. This generative model is based on libraries such as Transformers, and is capable of generating natural-sounding conversational text from large amounts of training data.

[0618] Next, the server inputs prompts into the selected generative model to generate conversation data. The prompts are based on the topic and tone specified by the user. For example, a prompt might say, "Generate a casual conversation about travel planning. The tone should be friendly, and Italy should be included as a travel destination."

[0619] The generated conversation data is formatted by the server and sent to the user's information device. The information device can then display this conversation data, which the user can view and use to practice real-life conversations.

[0620] In this way, the system of the present invention provides conversation scenarios tailored to user needs quickly and efficiently, and helps users practically improve their communication skills.

[0621] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0622] Step 1:

[0623] The user inputs the topic, keywords, and tone of the conversation they want to generate using their device. Upon receiving this input, the device temporarily stores it in its internal memory and prepares to send it to the server. The input consists of text data specified by the user, and the output is a formatted data packet.

[0624] Step 2:

[0625] The terminal sends formatted data packets to the server. The server receives these packets and records the information in its internal database system. Specifically, the server analyzes the received data and stores it in the appropriate database table based on attributes such as topic, keywords, and tone. The input is the data packets sent from the terminal, and the output is the writing of information to the database.

[0626] Step 3:

[0627] The server selects the most appropriate generative model based on the information stored in the database. The server uses an algorithm based on a natural language processing framework to identify a suitable model. Based on the selected model, it forms a prompt sentence and prepares it for input to the model. The input is information read from the database, and the output is the prompt sentence.

[0628] Step 4:

[0629] The server inputs a prompt sentence into the selected generative model and performs the generation of conversational data. Here, the AI ​​model generates natural conversational sentences according to the given prompt. The input is a prompt sentence, and the output is the generated conversational data. Specifically, the generative AI engine is instantiated, and the prompt is fed into the AI ​​model.

[0630] Step 5:

[0631] The generated conversation data is formatted by the server and sent to the terminal. The server converts the conversation data into a user-friendly format, i.e., a highly readable text format, and sends it to the terminal as a data packet. The input is conversation data obtained from an AI model, and the output is formatted data packets sent to the terminal.

[0632] Step 6:

[0633] The terminal receives conversation data packets sent from the server and displays them on the screen. At this time, the terminal interprets the data and visualizes it in a way that is easily understandable to the user. The input is data packets received from the server, and the output is a screen display that the user can see. Specifically, text is rendered on the display and presented through the user interface.

[0634] (Application Example 1)

[0635] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0636] In today's world, improving effective and natural communication skills is required in many situations. However, conventional learning methods make it difficult to acquire conversational skills appropriate to individual contexts, and suggesting appropriate conversations in real time is a major challenge. To solve this problem, there is a need to develop a system that generates and presents conversations in real time, tailored to the user's utterances.

[0637] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0638] In this invention, the server includes means for an information processing device that includes a function for receiving information specified by a user, means for a processing device that includes a function for selecting a generation method based on the information, and means for converting voice data into text data. This enables the generation and presentation of conversations that are natural and contextual to the user in real time.

[0639] "User-specified information" refers to topics, keywords, and conversational tone that users input into the system, and serves as a guideline for conversation generation.

[0640] "Means for converting audio data to text data" refers to a device or program that uses speech recognition technology to perform the process of converting audio input into text format.

[0641] A "processing device that includes a function for selecting a generation method" is a device or program that can select the optimal natural language processing method based on the received information.

[0642] "Means for generating conversation data" refers to an algorithm or program that automatically generates appropriate conversation content based on specified information or converted text data.

[0643] A "display device" is a device that provides generated conversation data to the user visually, and includes screens, smart glasses, and the like.

[0644] This invention presents a specific form of a system that supports the improvement of a user's natural conversational abilities. First, the user inputs information such as topics, keywords, and conversational tone using the terminal. This information is intended to clarify the user's intentions and desires. The terminal processes this information and prepares it for transmission to the server.

[0645] The server converts speech input into text data using speech recognition technologies such as the Google Cloud Speech-to-Text API. This converted text data is analyzed along with the user-specified information mentioned earlier, and the optimal generation method is selected. Specifically, conversation data is generated based on the obtained information using the Python language and natural language processing libraries (e.g., NLTK and SpaCy). The generated conversation data is then transmitted in real time to the user's display device, such as smart glasses, using WebSocket technology. This display allows the user to receive appropriate conversation suggestions in real time, improving their practical communication skills.

[0646] As a concrete example, in an international trade show setting, it is envisioned that conversation examples related to the user's utterances would be displayed to facilitate smooth conversations with new customers. An example of a prompt would be, "The topic of conversation is music technology. Please generate appropriate conversation examples in a natural tone and with a light touch."

[0647] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0648] Step 1:

[0649] The user inputs the conversation topic, keywords, and desired tone into the device. This clarifies the user's intent and prepares the necessary input data for subsequent processing. The input data is temporarily stored on the device and prepared to be sent to the server later.

[0650] Step 2:

[0651] The terminal sends the prepared input data to the server. This transmission utilizes an internet connection to ensure the data reaches the server securely and quickly. The transmitted data is structured in text format.

[0652] Step 3:

[0653] The server receives the transmitted data and, if it contains audio data, converts it to text data using the Google Cloud Speech-to-Text API. The converted text data is recorded in a database and used to further select parsing and generation algorithms.

[0654] Step 4:

[0655] The server uses Python and related natural language processing libraries to determine which generative AI model is most appropriate based on the analyzed text data. This determination is based on the content and patterns of the input data, and the selected model is then used for subsequent conversation generation.

[0656] Step 5:

[0657] The selected generative AI model is given prompt sentences to generate conversational data. These prompt sentences may include phrases such as, "The topic of conversation is music technology. Please generate appropriate conversational examples with a natural tone and lightheartedness." The generated conversational data is then formatted according to a specific format.

[0658] Step 6:

[0659] The server sends the generated conversation data to the terminal using a dedicated communication protocol (e.g., WebSocket). This communication is conducted in real time and is designed to minimize latency.

[0660] Step 7:

[0661] The device visually presents the received conversation data to the user, for example, through the display of smart glasses. This allows the user to engage in appropriate conversations in real time with visual support.

[0662] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0663] This invention provides a system that offers a richer communication experience by recognizing the user's emotions and incorporating them into the conversation generation process. This system is implemented using a user terminal, a server, a generation algorithm, and an emotion engine. The role and processing of each element are described below.

[0664] First, when the user uses the device to input information such as conversation topics, keywords, and tone of voice, the emotion engine simultaneously activates to analyze the user's emotions. This analysis is performed using the content of the input text and, if necessary, the user's voice data. The emotion engine uses a specific algorithm to recognize the user's tension and emotional state in real time.

[0665] Next, the device combines this information with the recognized emotion data and sends it to the server. The server analyzes the received data and selects the optimal generation algorithm. This selection also takes into account the emotions recognized by the emotion engine. For example, if the user is excited, the server can prioritize selecting an algorithm that generates more lively and energetic conversational text.

[0666] The server generates conversational text that matches the user's conditions according to a selected generation algorithm. This process utilizes natural language processing techniques to create conversational text that reflects the specified topic, tone, and emotion.

[0667] Furthermore, the server formats the generated conversation text into a format that is easily understandable to the user before sending it to the user's terminal. The conversation text received on the terminal is displayed in the user interface, and the user can use it as a reference for actual conversations.

[0668] For example, if a user is having a conversation about "planning a trip," and the emotion engine detects the user's excitement, the generation algorithm can generate lively conversational sentences that match that state, such as "Are you looking forward to new adventures on your next trip? What are your plans?"

[0669] This format allows users to obtain conversation examples that better suit their own emotions, and the system provides situation-appropriate feedback, enabling a mutually enjoyable and enriching communication experience.

[0670] The following describes the processing flow.

[0671] Step 1:

[0672] The user uses their device to input conversation topics, keywords, and desired conversational tone. In addition, the user's voice and text data are provided to the sentiment engine.

[0673] Step 2:

[0674] The device sends the input information and the user's voice or text data to the emotion engine. The emotion engine then analyzes the user's emotional state (e.g., joy, sadness, surprise, etc.).

[0675] Step 3:

[0676] The emotion engine returns the analysis results to the device. The device then packages this emotion data along with topic and tone information into a data package.

[0677] Step 4:

[0678] The terminal sends a data package to the server. The package contains conversational topics, keywords, tone, and sentiment data.

[0679] Step 5:

[0680] The server analyzes the received data package and selects the optimal generation algorithm. This selection includes adjustments that reflect the user's emotional state.

[0681] Step 6:

[0682] The server uses a selected generation algorithm to generate natural-sounding conversational text that matches the user's criteria. The conversational text reflects the specified topic, tone, and emotion.

[0683] Step 7:

[0684] The server formats the generated dialogue into a format that is easily understandable to the user. The formatted dialogue is then provided in a context that is appropriate to the emotion and tone of the conversation.

[0685] Step 8:

[0686] The server sends formatted conversation text to the user's terminal. The terminal displays the received conversation text in its user interface, making it intuitive for the user to use.

[0687] Step 9:

[0688] The system allows users to review conversational text presented on their device and incorporate content that matches their emotions into their actual conversations. This enables users to practice smoother conversational scenarios and improve their communication skills.

[0689] (Example 2)

[0690] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0691] Conventional conversation generation systems have faced challenges in generating conversational text that adequately considers the user's emotional state, making it difficult to provide a natural and appropriate communication experience. In particular, it has been difficult to achieve flexible responses that adapt to changes in the user's emotions.

[0692] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0693] In this invention, the server includes means for analyzing information and emotions specified by the user, means for selecting a generation algorithm based on the information and analyzed emotions, and means for generating conversational text that reflects the user's emotional state using the generation algorithm. This makes it possible to generate conversational text adapted to the user's emotional state in real time, providing natural and rich communication.

[0694] "User-specified information" refers to data that users input for conversation generation, including conversation topics, keywords, and tone of conversation.

[0695] "Means of analyzing emotions" refers to processes and algorithms that recognize a user's emotional state in real time based on their input or voice data.

[0696] A "generation algorithm" is a computational method for generating optimal conversational text based on information obtained from users and analyzed sentiment data.

[0697] "Natural language processing technology" refers to a set of algorithms and techniques that enable computers to understand, generate, or interpret human language.

[0698] A "prompt sentence" is a sentence that serves as an instruction or hint, used by a generative AI model when generating conversational text.

[0699] "User's terminal" refers to a device operated by the user that enables the input of conversations and the display of generated conversational text.

[0700] This invention provides a more personalized communication experience by analyzing the user's emotions and generating conversations accordingly. This system is implemented using a user-operated terminal, a data processing server, a generation algorithm, and an emotion analysis engine.

[0701] The user enters conversation topics and keywords using the device, and the sentiment engine activates at that time. The sentiment analysis engine on the device analyzes the entered text and audio data. The specific analysis uses natural language processing and speech signal processing techniques, and a particular algorithm is applied to recognize the user's emotional state in real time.

[0702] Next, the device sends the collected information and analyzed sentiment data to the server. Based on the received data, the server selects a generation algorithm and creates a prompt that is appropriate for the user's sentiment. For example, if the user selects the topic "Planning a Trip" and the sentiment engine detects exhilaration, the prompt will include a phrase such as "What are you looking forward to on your next trip?"

[0703] The server uses a selected generative AI model to generate conversational text based on the prompt. The generated conversational text is refined using natural language processing techniques, reflecting the user's emotions and the theme of the conversation.

[0704] The completed conversation is formatted by the server, converted into an easily understandable format, and then sent to the terminal. The terminal displays this conversation on its user interface, allowing the user to utilize it to create richer and more meaningful conversations.

[0705] This configuration allows the system to have the flexibility to provide appropriate feedback in response to the user's emotional state and communication needs, enabling a high-quality user experience.

[0706] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0707] Step 1:

[0708] The user operates the terminal to input the conversation topic, keywords, and desired tone. The terminal receives this information and compiles it as prompt data. The input data is stored in text format for later use in sentiment analysis.

[0709] Step 2:

[0710] The device activates an emotion analysis engine and analyzes the user's input text. If voice input is available, it performs voice signal processing to estimate the emotional state from the tone, speed, and intonation of the voice. The analysis engine processes the data using natural language processing techniques to identify the user's emotional state. The output is metadata indicating the user's emotional state.

[0711] Step 3:

[0712] The terminal sends the user's topics and keywords, along with the analyzed sentiment data, to the server. The server reviews the received data and selects the optimal generation algorithm. In doing so, it considers the user's emotional state and sets conditions for generating appropriate prompt sentences.

[0713] Step 4:

[0714] The server uses a generation algorithm to generate conversational text based on prompts. A generative AI model is employed here to output responses that match the user's input and emotions. The generated data is conversational text in text format.

[0715] Step 5:

[0716] The server formats the generated dialogue, changing it into a format that is easily understandable to the user. This includes optimizing word choice and adjusting grammar. The formatted dialogue is the final data that aids user comprehension.

[0717] Step 6:

[0718] The server sends the formatted conversation text back to the terminal. The terminal displays the received conversation text in its user interface, allowing the user to review it. The user then uses this displayed information as a reference for their actual conversation.

[0719] (Application Example 2)

[0720] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0721] Traditional communication systems generate conversations without considering the user's emotions, making it difficult to respond flexibly to the individual user's situation and feelings. Furthermore, enriching the customer experience in online shopping requires customer service tailored to the user's emotions, but conventional technologies lack the means to achieve this. This results in the challenge of not being able to sufficiently increase user satisfaction and purchasing intent.

[0722] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0723] In this invention, the server includes means for selecting a generation algorithm based on information specified by the user and the user's emotional state, means for recognizing the user's emotions using emotion analysis technology, and means for transmitting the generated conversation text to the user's terminal. This makes it possible to generate conversation text that responds to the user's emotions in real time and provide a communication experience tailored to each individual.

[0724] "User-specified information" refers to data such as arbitrary topics, keywords, and tones that the user provides to the system.

[0725] A "generative algorithm" is a series of processes that create the most suitable dialogue based on specified conditions.

[0726] "Emotional state" refers to data that indicates the emotional state exhibited by the user, and includes emotions such as joy, sadness, and excitement.

[0727] "Emotion analysis technology" is a technology that recognizes, classifies, and analyzes a user's emotions based on their statements and voice data.

[0728] "Natural language processing technology" refers to technologies that enable computers to understand and process human language, and includes sentence generation and sentiment analysis.

[0729] "User terminal" refers to a computer, smartphone, or other communication device used by the user.

[0730] This invention relates to a system for generating conversations that respond to a user's emotions, and its implementation is achieved using a server, a terminal, and various software. This system receives information specified by the user using a terminal and recognizes the user's emotional state using emotion analysis technology. Based on this, the server selects the optimal generation algorithm and generates conversational text using natural language processing technology.

[0731] The server uses emotion analysis technology, specifically an emotion engine (e.g., Watson Tone Analyzer), to analyze the user's emotions in real time based on their text and voice data. The analyzed emotion data is then input as prompts into a generative AI model (e.g., the GPT model).

[0732] This system allows a generative AI model to generate conversational text that matches the user's emotions, which is then sent from the server to the terminal and displayed on the user interface. Ultimately, the user can use this generated conversational text to have a smooth conversation on the terminal.

[0733] As a concrete example, consider a scenario where a device receives information from a user saying, "I want to buy new clothes," and sentiment analysis technology identifies the user's level of excitement. As a result, the server can generate an enthusiastic conversational message such as, "I think these clothes would look great on you! You'll be the center of attention at the party!"

[0734] A concrete example of a prompt message is an instruction given to the generation AI model, such as, "The user is highly excited, so please generate an energetic and positive product description."

[0735] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0736] Step 1:

[0737] The user uses the device to input their desired topic, keywords, and conversational tone. The device uses this information as input data, activates an emotion engine, and analyzes the user's emotional state. This process yields emotion analysis results from the input text data.

[0738] Step 2:

[0739] The terminal sends analyzed emotion data and user-input data to the server. The server receives both sets of data and selects an appropriate generation algorithm. In this selection process, the emotion data is used as a criterion for choosing a generation algorithm that is appropriate for the user's emotions.

[0740] Step 3:

[0741] The server uses a selected generation algorithm to provide prompt sentences to the generative AI model, generating conversational text. The input includes the user's emotional state, topic, and keywords, and the output is conversational text that matches these elements. In this process, the generative AI model constructs appropriate sentences based on natural language processing techniques.

[0742] Step 4:

[0743] The server formats the generated conversation text into a format that is easily understandable to the user. The formatted conversation text is then sent to the terminal as the final output. In this way, the generated text obtained in step 3 is fed back into the user's conversation in a practical way.

[0744] Step 5:

[0745] The terminal displays received conversation text on the user interface, allowing the user to communicate smoothly by referring to the text. This output is based on user interaction and is used to enrich the overall communication experience of the system.

[0746] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0747] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0748] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0749] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0750] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0751] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0752] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0753] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0754] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0755] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0756] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0757] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0758] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0759] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0760] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0761] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0762] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0763] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0764] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0765] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0766] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0767] The following is further disclosed regarding the embodiments described above.

[0768] (Claim 1)

[0769] A means of receiving information specified by the user,

[0770] means for selecting a generation algorithm based on the aforementioned information,

[0771] Means for generating conversational text using the aforementioned generation algorithm,

[0772] A means for sending the generated conversation text to the user's terminal,

[0773] A system that includes this.

[0774] (Claim 2)

[0775] The system according to claim 1, wherein the generation algorithm uses natural language processing technology.

[0776] (Claim 3)

[0777] The system according to claim 1, wherein the aforementioned information includes keywords and conversational tone.

[0778] "Example 1"

[0779] (Claim 1)

[0780] A means of receiving information specified by the user and storing that information in a database,

[0781] A means for selecting a generative model using natural language processing based on the received information,

[0782] A means of inputting prompt data into a selected generative model to generate conversational data,

[0783] A means for formatting the generated conversation data and transmitting it to the user's information device,

[0784] A system that includes this.

[0785] (Claim 2)

[0786] The system according to claim 1, wherein the generative model uses a publicly available natural language processing framework.

[0787] (Claim 3)

[0788] The system according to claim 1, wherein the information includes keywords and conversational tone.

[0789] "Application Example 1"

[0790] (Claim 1)

[0791] A means comprising an information processing device that includes a function for receiving information specified by the user,

[0792] A means comprising a processing device that includes a function for selecting a generation method based on the aforementioned information,

[0793] A means of converting audio data into text data,

[0794] A means for generating conversation data using a generation method based on the aforementioned text data,

[0795] Means for transmitting the generated conversation data to a display device,

[0796] A system that includes this.

[0797] (Claim 2)

[0798] The system according to claim 1, wherein the generation method uses natural language processing technology.

[0799] (Claim 3)

[0800] The system according to claim 1, in which the aforementioned information includes feature information and the type of conversation.

[0801] "Example 2 of combining an emotion engine"

[0802] (Claim 1)

[0803] A means of analyzing information and emotions specified by the user,

[0804] A means for selecting a generation algorithm based on the aforementioned information and analyzed emotions,

[0805] A means for generating conversational text that reflects the user's emotional state using the aforementioned generation algorithm,

[0806] A means for sending the generated conversation text to the user's terminal and displaying it,

[0807] A system that includes this.

[0808] (Claim 2)

[0809] The system according to claim 1, wherein the generation algorithm uses natural language processing technology to create prompt sentences that correspond to the user's emotional state.

[0810] (Claim 3)

[0811] The system according to claim 1, wherein the information includes a topic, keywords, and tone of conversation.

[0812] "Application example 2 when combining with an emotional engine"

[0813] (Claim 1)

[0814] A means of receiving information specified by the user,

[0815] Means for selecting a generation algorithm based on the aforementioned information and the user's emotional state,

[0816] A means of recognizing a user's emotions using emotion analysis technology,

[0817] Means for generating conversational text using the aforementioned generation algorithm,

[0818] A means for sending the generated conversation text to the user's terminal,

[0819] A system that includes this.

[0820] (Claim 2)

[0821] The system according to claim 1, wherein the generation algorithm uses natural language processing technology.

[0822] (Claim 3)

[0823] The system according to claim 1, wherein the information includes keywords and conversational tone, and further includes emotional data recognized in real time. [Explanation of Symbols]

[0824] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving information specified by the user, means for selecting a generation algorithm based on the aforementioned information, Means for generating conversational text using the aforementioned generation algorithm, A means for sending the generated conversation text to the user's terminal, A system that includes this.

2. The system according to claim 1, wherein the generation algorithm uses natural language processing technology.

3. The system according to claim 1, wherein the aforementioned information includes keywords and the tone of conversation.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A