system

The system addresses the challenge of providing personalized interactions and learning content for children by registering user information, creating dialogue profiles, and adapting dialogues based on user responses, ensuring effective engagement and learning support.

JP2026064609APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Modern parenting generations and the senior population face challenges in interacting with children, particularly when they cannot leave them unattended, and there is a lack of means to provide appropriate learning content and continuous conversations tailored to individual children's needs, especially during home working or growth stages.

Method used

A system that registers user information, creates individual dialogue profiles, selects dialogue scenarios, initiates and conducts dialogues, analyzes user responses, provides learning-enhancing content, and stores data for feedback to adapt subsequent dialogues, using devices like pet robots and smart devices.

Benefits of technology

The system provides personalized conversations and learning content tailored to each user's age and interests, supporting children's learning and play when parents are unavailable, and adapts dialogue content based on real-time user responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064609000001_ABST
    Figure 2026064609000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for registering user information, A means of creating individual dialogue profiles based on registered information, A means of selecting a dialogue scenario based on a dialogue profile, A means of initiating and conducting a dialogue based on a selected dialogue scenario, A means of obtaining and analyzing user responses during a conversation, Means for selecting and providing content to promote learning, A means of saving data collected during a conversation and applying the feedback to the next conversation, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including: receiving a user utterance; adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For modern parenting generations and the senior population, there is a problem that it is difficult to interact with children when parents cannot leave them unattended. In particular, there are problems in providing appropriate learning content when a child is absent due to illness during home working, or when the child is in a growth stage. Furthermore, in these situations, there is a problem of a lack of means to continuously provide conversations suitable for individual children.

Means for Solving the Problems

[0005] The present invention first provides means for registering user information. Based on this information, it includes means for creating individual dialogue profiles. It also includes means for selecting dialogue scenarios based on the dialogue profiles. It provides means for initiating and conducting dialogues based on the dialogue scenarios, and means for obtaining and analyzing user responses during the dialogue. Furthermore, it includes means for selecting and providing content to facilitate learning, and means for storing data collected during the dialogue and applying the feedback to the next dialogue, thereby solving these problems.

[0006] "User information" refers to personal information such as the name, age, gender, interests, and learning progress of the system user.

[0007] "Means of registration" refers to devices or software functions used to input and store user information in a system.

[0008] A "dialogue profile" refers to a dataset containing individual dialogue scenarios and response patterns generated based on registered user information.

[0009] A "dialogue scenario" refers to a plan that defines the flow and content of a series of conversations that take place between the user and the system.

[0010] "Means for initiating and conducting a dialogue" refers to devices and software functions that enable actual dialogue with the user based on a selected dialogue scenario.

[0011] The means of acquiring and analyzing "user responses" refers to devices and software functions that record verbal or nonverbal responses exhibited by users during a conversation and analyze that data.

[0012] "Learning-enhancing content" refers to educational information and activities provided based on the user's age and interests.

[0013] "Means of saving data and applying feedback to the next dialogue" refers to devices or software functions that store information collected during a dialogue in a database and reflect that information in the next dialogue session. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention relates to an interactive pet robot system that registers user information, creates a dialogue profile based on that information, and provides individually adapted dialogue scenarios. This system is designed to support children's learning and play, especially for parents and seniors, when parents are unable to attend to their children.

[0036] User Information Registration

[0037] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0038] Terminal: Send this information to the server.

[0039] Server: Stores received information in a database and generates individual interaction profiles.

[0040] Creating and adapting dialogue profiles

[0041] Server: The conversation profile created based on user information reflects age, interests, and learning progress, and includes individual conversation scenarios.

[0042] Terminal: Synchronize conversation profiles with pet robots, allowing the system to provide appropriate interactions for each user.

[0043] Initiating and facilitating the dialogue

[0044] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[0045] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. It says, "Hi Taro. What would you like to do today?"

[0046] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[0047] Acquiring and analyzing user responses

[0048] User: The child responds.

[0049] Device: The device analyzes the child's response using a speech recognition engine to understand its meaning. For example, if the child says "I like cars," it retrieves that information.

[0050] Server: Based on the analyzed information, the server generates the following dialogue, and the pet robot responds, "Do you know what colors cars come in?"

[0051] Providing content to promote learning

[0052] Server: Selects learning-enhancing content (e.g., English vocabulary, place names, etc.) based on the child's age and interests.

[0053] Terminal: The pet robot teaches, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0054] User: The child keeps repeating "red."

[0055] Device: Recognizes repeated content and records accuracy and learning effectiveness.

[0056] Data storage and feedback

[0057] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0058] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[0059] Terminal: Update and prepare a new interaction profile for the next session.

[0060] Specific example

[0061] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then says "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[0062] In this way, the system can utilize user information to provide personalized conversations and learning content tailored to each user.

[0063] The following describes the processing flow.

[0064] Step 1:

[0065] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0066] Step 2:

[0067] Terminal: Sends the entered user information to the system server.

[0068] Step 3:

[0069] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[0070] Step 4:

[0071] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[0072] Step 5:

[0073] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[0074] Step 6:

[0075] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0076] Step 7:

[0077] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[0078] Step 8:

[0079] User: Taro replies, "Car."

[0080] Step 9:

[0081] Terminal: The speech recognition engine analyzes Taro's response to understand its content.

[0082] Step 10:

[0083] Server: Analyzes the response and generates the next appropriate dialogue.

[0084] Step 11:

[0085] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[0086] Step 12:

[0087] User: Taro replies "Red".

[0088] Step 13:

[0089] Terminal: The speech recognition engine analyzes Taro's response again to verify its accuracy.

[0090] Step 14:

[0091] Server: Based on this response, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[0092] Step 15:

[0093] Device: Provides selected learning content and says, "The English word for red is 'Red'. Let's say it together, red."

[0094] Step 16:

[0095] User: Taro repeats "Red".

[0096] Step 17:

[0097] Terminal: Recognizes Taro's pronunciation and records its accuracy and learning effectiveness.

[0098] Step 18:

[0099] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[0100] Step 19:

[0101] Server: Stores data collected during the interaction (Taro's responses, learning progress, etc.) in a database.

[0102] Step 20:

[0103] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[0104] (Example 1)

[0105] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0106] One of the problems faced by modern families, particularly by parents and seniors, is the difficulty in adequately supporting children's learning and play when parents are unable to attend to them directly. Traditional educational tools and interactive learning toys have limited effectiveness because they cannot provide individualized dialogue and learning content tailored to the child's age and interests. Furthermore, the technology to monitor learning progress in real time during dialogue and incorporate that information into subsequent dialogues has been insufficient.

[0107] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0108] In this invention, the server includes means for registering user information, means for storing user information in a database and generating individual dialogue profiles, and means for synchronizing the generated dialogue profiles with a dialogue robot. This makes it possible to provide individual dialogue scenarios tailored to the user's age and interests in real time, and to effectively support learning progress.

[0109] "Means for registering user information" refers to devices or software that allow users to input and register basic information about their child and parent using a smartphone app or a dedicated web interface.

[0110] "Means for sending registered information to the server" refers to a communication device or software that has the role of sending information entered by the user to the server.

[0111] "Means for storing received information in a database and generating individual dialogue profiles" refers to a device or software that stores user information received by a server in a database and generates individual dialogue profiles based on that information.

[0112] "Means for synchronizing the generated dialogue profile with the dialogue robot" refers to a communication device or software that transmits the generated dialogue profile to the dialogue robot's memory, causing the robot to operate based on that profile.

[0113] "Means of initiating dialogue based on parental instructions" refers to a device or software that initiates a dialogue session when a parent gives instructions to a dialogue robot via voice or other interface.

[0114] "Means for selecting and implementing an appropriate dialogue scenario based on a dialogue profile" refers to a device or software that selects the optimal dialogue scenario based on the content of the dialogue profile and proceeds with the dialogue according to that scenario.

[0115] "Means for analyzing and understanding the content of a user's response using a speech recognition engine" refers to a device or software that uses a speech recognition engine to convert a user's response into text and understand its content.

[0116] "Means for generating the next dialogue based on analyzed information" refers to a device or software that uses a generative AI model to generate the content of the next dialogue based on information analyzed by a speech recognition engine.

[0117] "Means for selecting and providing learning-promoting content based on a child's age and interests" refers to a device or software that selects appropriate learning content according to a child's specific age and interests and provides it to the user.

[0118] "Means of sending data collected during a dialogue to a server and feeding it back into the next dialogue" refers to a device or software that sends detailed data collected during a dialogue to a server and processes it to reflect in the next dialogue profile.

[0119] This invention relates to an interactive pet robot system that performs a series of processes including user information registration, dialogue profile generation, initiation and progression of dialogue, acquisition and analysis of user responses, provision of content to promote learning, and data storage and feedback. This system is designed particularly for parents and seniors to support children's learning and play when parents are unable to attend to them.

[0120] User Information Registration

[0121] User: Use a smartphone app or a dedicated web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter information such as "Taro, 3 years old, boy, interest: cars."

[0122] Terminal: Sends user-entered information to the server. This process uses communication methods such as HTTP POST requests.

[0123] Server: Stores received information in a database (e.g., MySQL® or PostgreSQL) and generates individual interaction profiles. A generative AI model (e.g., OpenAI® GPT-4®) is used for profile generation.

[0124] Creating and adapting dialogue profiles

[0125] Server: Generates a conversation profile based on user information. The profile reflects age, interests, and learning progress, and includes individual conversation scenarios.

[0126] Terminal: Synchronizes dialogue profiles with the conversational robot, enabling the system to provide appropriate dialogues for each user. Specifically, it downloads dialogue profiles from the server to the terminal and saves them to the robot's memory.

[0127] Initiating and facilitating the dialogue

[0128] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[0129] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[0130] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content. A NoSQL database (e.g., MongoDB) is used to store the logs.

[0131] Acquiring and analyzing user responses

[0132] User: The child responds. For example, the child says, "I like cars."

[0133] Device: The child's responses are analyzed using a speech recognition engine (e.g., Google® Cloud Speech-to-Text API) to understand the content.

[0134] Server: Based on the analyzed information, it generates the following dialogue. For example, it generates the following question: "Do you know what colors cars come in?"

[0135] Providing content to promote learning

[0136] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools (e.g., NLTK or SpaCy) are used for content selection.

[0137] Device: The pet robot provides children with selected learning content. For example, it might teach them, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0138] User: The child keeps repeating "red."

[0139] Device: Analyzes repeated content using speech recognition and records accuracy and learning effectiveness.

[0140] Data storage and feedback

[0141] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0142] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[0143] Terminal: Update and prepare a new interaction profile for the next session. Specifically, download the updated profile from the server and synchronize it with the robot.

[0144] Specific example

[0145] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then mentions the color "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[0146] Example of a prompt

[0147] The following is an example of a prompt statement for a generative AI model:

[0148] User information:

[0149] Name: Taro

[0150] Age: 3 years old

[0151] Gender: Boy

[0152] Interests: Cars

[0153] Start conversation:

[0154] User: Playing with my pet robot, Taro

[0155] Robot: Hi Taro. What do you want to do today?

[0156] User: I like cars

[0157] Robot: Do you know what colors cars come in?

[0158] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0159] Step 1: Enter user information

[0160] User: Use a smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter "Taro, 3 years old, boy, interest: cars" into the app's form.

[0161] Input: Basic information entered by the user

[0162] Output: Input basic information data

[0163] Step 2: Submit User Information

[0164] Terminal: Sends the entered user information to the server using an HTTP POST request.

[0165] Input: Basic information data entered by the user

[0166] Output: Request data sent to the server

[0167] Step 3: Saving user information and generating a profile

[0168] Server: The server stores the received information in a database (e.g., MySQL or PostgreSQL) and uses a generative AI model (e.g., OpenAI GPT-4) to generate individual interaction profiles. This information includes prompts for the generative AI model.

[0169] Input: Request data sent to the server

[0170] Output: Saved database entries and generated interaction profiles

[0171] Specific operation: Use an "INSERT" SQL query to save information to the database, create a prompt statement for profile generation, and send it to the generating AI model.

[0172] Step 4: Synchronize conversation profiles

[0173] Terminal: Downloads the dialogue profile sent from the server and synchronizes it with the dialogue robot.

[0174] Input: Interaction profile sent from the server

[0175] Output: Dialogue profile synchronized with the conversational robot

[0176] Specific operation: Download the profile using an HTTP GET request and save it to the conversational robot's memory.

[0177] Step 5: Start the conversation

[0178] User: The parent gives voice commands to the conversational robot, saying, "Pet robot, play with Taro."

[0179] Input: Parental voice commands

[0180] Output: Trigger signal to start interaction

[0181] Specific operation: A microphone system for receiving voice commands recognizes the command and sends a trigger to the robot system to start the dialogue.

[0182] Step 6: Select and implement a dialogue scenario

[0183] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and starts the conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[0184] Input: Interaction Profile

[0185] Output: Selected dialogue scenario and its implementation

[0186] Specific operation: The robot selects the optimal scenario from the dialogue profiles and plays the scenario aloud.

[0187] Step 7: Obtain and analyze user responses

[0188] Device: The system analyzes the child's response using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, if the child says "I like cars," the audio is converted into text.

[0189] Input: Child's voice response

[0190] Output: Analyzed data in text format

[0191] Specific operation: Audio data is sent to the API and returned as text data.

[0192] Step 8: Generating the next dialogue

[0193] Server: Based on the analyzed text data, it generates the next dialogue using an AI model. For example, it generates the following question: "Do you know what colors cars come in?"

[0194] Input: Text-formatted analysis data

[0195] Output: The following dialogue

[0196] Specific operation: A prompt message is sent to the generative AI model, and the generated dialogue content is retrieved.

[0197] Step 9: Provide content to facilitate learning

[0198] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools are used to select content. For example, it might teach, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0199] Input: Interaction profile and user response data

[0200] Output: Selected learning content

[0201] Specific operation: The robot selects content from a database that matches the child's interests and teaches them using voice.

[0202] Step 10: Data storage and feedback

[0203] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0204] Input: Collected data

[0205] Output: Feedback data sent to the server

[0206] Specific operation: The collected data is sent to the server via an HTTP POST request.

[0207] Step 11: Save feedback data and update the next interaction profile.

[0208] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[0209] Input: Feedback data sent to the server

[0210] Output: Updated interaction profile

[0211] Specific actions: Record feedback data in a database and use a generative AI model to update the next interaction profile.

[0212] (Application Example 1)

[0213] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0214] Traditional interactive systems were designed to support user learning and play, but they were difficult to apply to new employee training in factories and other similar settings. These systems struggled with providing complex operational instructions and training content tailored to skill levels, resulting in limited practicality. Furthermore, they lacked the ability to adapt training content in real time or analyze user instructions using speech recognition technology. This highlighted the challenge of improving the efficiency of training in new environments.

[0215] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0216] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, and means for providing training programs based on the user information. This makes it possible to provide personalized training content according to the user's skill level and learning progress.

[0217] "Means of registering user information" refers to interfaces or mechanisms for inputting basic information such as a user's name, age, job title, and skill level, and saving it in a database.

[0218] "Means for creating dialogue profiles" refers to a mechanism that generates individual dialogue scenarios based on registered user information and sets up dialogue content tailored to the user.

[0219] "Means for selecting dialogue scenarios" refers to a mechanism that selects a scenario that fits the created dialogue profile and facilitates effective dialogue with the user.

[0220] "Means for initiating and conducting a dialogue" refers to a mechanism for initiating a dialogue with a user based on a selected dialogue scenario and continuing the dialogue while providing timely responses.

[0221] "Means for acquiring and analyzing user responses" refers to a mechanism that analyzes the voice and actions obtained from the user during a conversation and dynamically adjusts the content of the next conversation based on that analysis.

[0222] "Means for selecting and providing content to promote learning" refers to a mechanism for selecting and providing educational content (e.g., vocabulary and operating procedures) that is appropriate for the user's age, interests, and skill level.

[0223] "Means of saving data collected during a conversation and applying feedback to the next conversation" refers to a mechanism that records the user's responses and learning progress obtained during a conversation and reflects them in subsequent conversations.

[0224] "Means of providing training programs" refers to a mechanism for generating and providing suitable training content based on user information.

[0225] "A means of analyzing user instructions using a speech recognition engine and generating the next dialogue content" refers to a mechanism that uses speech recognition technology to convert user statements into text and generates the next response based on that text.

[0226] "Means for providing instructions and training content using voice output means" refers to a mechanism for providing text-based instructions and training content to the user as audio.

[0227] To implement this invention, a server, terminal, and user must work together to build a dialogue system. The system of this invention consists of the following steps: user information registration, dialogue profile creation, training program provision, dialogue initiation and progress, user response acquisition and analysis, learning promotion content provision, data storage and feedback. The details are described below.

[0228] User Information Registration

[0229] Terminal: Users enter their information (name, job title, skill level, etc.) using a smartphone app or head-mounted display (HMD). This information is sent to the server and stored in the database.

[0230] Server: Based on the received information, it generates individual interaction profiles. These interaction profiles are customized according to the user's skill level and role.

[0231] Training program provision

[0232] Server: Based on user information, generates an appropriate training scenario and sends it to the terminal.

[0233] Terminal: Displays and provides the received training scenario to the user. For example, if a new employee says, "Teach me how to operate this machine," the terminal will display instructions from the server in both audio and display format.

[0234] Initiating and facilitating the dialogue

[0235] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[0236] Terminal: Uses a speech recognition engine to analyze the user's words and sends them to the server.

[0237] Server: Based on the analysis results, it generates the next dialogue content and sends it to the terminal. The dialogue content is dynamically adjusted in real time.

[0238] Acquiring and analyzing user responses

[0239] Terminal: Analyzes the user's voice input in real time and sends the results to the server.

[0240] Server: Based on the received information, it generates the next dialogue content and sends it to the terminal. This provides personalized training content.

[0241] Providing learning-enhancing content

[0242] Server: Select appropriate learning content (e.g., machine operation procedures, troubleshooting methods, etc.) according to the user's skill level and role.

[0243] Terminal: Provides selected content using an audio output device. For example, it might instruct, "First, press the power button. Then, press the start button to operate the machine."

[0244] Data storage and feedback

[0245] Terminal: Sends data collected during the interaction to the server.

[0246] Server: Receives data and saves it to the database, applying the feedback to the next interaction. This makes subsequent interactions more tailored to the user.

[0247] Examples of specific cases and prompt statements

[0248] As a concrete example, let's consider a scenario where a new employee, "Yamada," learns the basic operation of a new machine. In this case, the following prompt messages are used.

[0249] Example of a prompt

[0250] User: Yamada

[0251] Position: Operator

[0252] Skill level: beginner

[0253] User: Please tell me how to operate this machine.

[0254] Robot: "I'll teach you the basic operation of the machine. First, press the power button. Then, press the start button to start the machine."

[0255] In this way, the system of the present invention can utilize user information to provide personalized conversations and training content to each user.

[0256] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0257] Step 1: Register User Information

[0258] Terminal: Users (new recruits and employees) use a smartphone app or head-mounted display (HMD) to enter basic information such as name, job title, and skill level. This information is sent from the terminal to the server.

[0259] Input: User's basic information (name, job title, skill level)

[0260] Output: User information sent to the server

[0261] Step 2: Create an interaction profile

[0262] Server: Based on the received user information, it generates an individual dialogue profile. The dialogue profile includes dialogue scenarios based on the user's skill level and job title.

[0263] Input: User information

[0264] Output: Individual interaction profiles

[0265] Step 3: Providing the training program

[0266] Server: Based on the interaction profile, generates an appropriate training scenario and sends it to the terminal.

[0267] Input: Interaction Profile

[0268] Output: Training scenario

[0269] Step 4: Start the conversation

[0270] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[0271] Input: User's voice instructions

[0272] Output: Startup of speech recognition engine

[0273] Step 5: Voice Analysis

[0274] Terminal: Using a voice recognition engine, convert the user's voice instructions into text data and send it to the server.

[0275] Input: User's voice instructions

[0276] Output: Text data

[0277] Step 6: Generation of dialogue content

[0278] Server: Based on the received text data, generate the next dialogue content and send it to the terminal.

[0279] Input: Text data, dialogue profile

[0280] Output: Dialogue content (training content)

[0281] Step 7: Instructions by voice and display

[0282] Terminal: Provide the user with the dialogue content received from the server through voice output and display. For example, give instructions such as "First, press the power button. Then, press the start button to operate the machine."

[0283] Input: Dialogue content

[0284] Output: Voice and display

[0285] Step 8: Acquisition and analysis of user responses

[0286] Terminal: Analyze the user's voice input in real time and send the result to the server. For example, when the user asks "Where is the power button?", analyze the voice and convert it into text data.

[0287] Input: User's voice input

[0288] Output: Analyzed text data

[0289] Step 9: Generating Feedback

[0290] Server: Based on the analyzed text data, it generates the next dialogue and sends it to the terminal. This provides personalized training content in real time.

[0291] Input: Analyzed text data, dialogue profile

[0292] Output: The following dialogue

[0293] Step 10: Data storage and feedback

[0294] Terminal: Sends data collected during the interaction to the server.

[0295] Server: Receives data and saves it to a database, which is then used in subsequent interactions. This data includes user responses, learning progress, and other information.

[0296] Input: Data collected during the conversation

[0297] Output: Saved to the database, feedback that will be reflected in the next interaction.

[0298] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0299] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles based on that information, and provides dialogue scenarios. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized dialogue and learning content.

[0300] Registration of User Information

[0301] User: Using a smartphone app or other interface, enter basic information (name, age, gender, interests, learning progress, etc.) of the child and the parent.

[0302] Terminal: Send this information to the server.

[0303] Server: Save the received information in the database and generate an interaction profile based on it.

[0304] Integration of the Emotion Engine

[0305] Terminal: Install an emotion engine in the pet robot and endow it with the function of recognizing emotions from the user's voice tone, expression, actions, etc.

[0306] Server: Implement an algorithm for analyzing emotion data to identify the user's current emotional state.

[0307] Creation and Adaptation of the Interaction Profile

[0308] Server: Generate an interaction profile based on user information and emotion data. The interaction profile includes individualized interaction scenarios according to the user's age, interests, learning progress, and emotional state.

[0309] Terminal: Synchronize the interaction profile to the pet robot so that the system can provide an optimal interaction for each user.

[0310] Start and Progress of the Interaction

[0311] User: The parent gives an instruction such as "Pet robot, play with Taro" through a voice instruction or an app.

[0312] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0313] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[0314] Acquiring and analyzing user responses

[0315] User: Taro replies, "Car."

[0316] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions.

[0317] Server: Analyzes response content and sentiment data to generate the next appropriate dialogue.

[0318] Providing content to promote learning

[0319] Server: Selects learning-enhancing content (English vocabulary, place names, etc.) based on the child's age, interests, and emotional state.

[0320] Terminal: The pet robot generates an encouraging response, saying, "The English word for red is 'Red.' Let's say it together, 'Red.' Sounds fun!"

[0321] Data storage and feedback

[0322] Terminal: Sends data collected during the conversation (Taro's responses, learning progress, emotional state, etc.) to the server.

[0323] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[0324] Terminal: Update and prepare a new interaction profile for the next session.

[0325] Specific example

[0326] For example, if a 3-year-old child named Taro is registered, and along with information that he likes cars, the emotion engine analyzes Taro's emotional state as "very interested." Based on this information, the pet robot asks, "Do you know what colors cars come in?" If Taro replies "red," and the emotion engine analyzes Taro's joy and excitement, the server generates a response such as, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for the content of the next conversation.

[0327] In this way, the system utilizes user information and emotional data to provide personalized dialogue and learning content, and can also apply feedback tailored to the user's emotional state.

[0328] The following describes the processing flow.

[0329] Step 1:

[0330] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0331] Step 2:

[0332] Terminal: Sends the entered user information to the system server.

[0333] Step 3:

[0334] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[0335] Step 4:

[0336] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[0337] Step 5:

[0338] Terminal: The pet robot activates its built-in emotion engine to recognize the user's emotional state in real time based on their voice tone, facial expressions, and movements.

[0339] Step 6:

[0340] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[0341] Step 7:

[0342] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0343] Step 8:

[0344] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[0345] Step 9:

[0346] User: Taro replies, "I like cars."

[0347] Step 10:

[0348] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions (excitement, joy, etc.).

[0349] Step 11:

[0350] Server: Based on the response content and sentiment data, it generates the next appropriate dialogue.

[0351] Step 12:

[0352] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[0353] Step 13:

[0354] User: Taro replies "Red".

[0355] Step 14:

[0356] Terminal: Taro's response is analyzed by a speech recognition engine, and along with the content of the response, an emotion engine is used to analyze his emotion of joy.

[0357] Step 15:

[0358] Server: Based on this response and sentiment data, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[0359] Step 16:

[0360] Device: Provides selected learning content and says, "The English word for red is 'Red.' Let's say it together, 'Red.' Isn't that great!"

[0361] Step 17:

[0362] User: Taro repeats "Red".

[0363] Step 18:

[0364] Terminal: Recognizes Taro's pronunciation, records its accuracy and learning effectiveness, and simultaneously analyzes Taro's sense of accomplishment using an emotion engine.

[0365] Step 19:

[0366] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[0367] Step 20:

[0368] Server: Stores data collected during the interaction (Taro's responses, learning progress, emotional state, etc.) in a database.

[0369] Step 21:

[0370] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[0371] (Example 2)

[0372] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0373] Traditional dialogue systems have faced challenges in providing personalized dialogue and learning content due to a lack of individual dialogue profiles based on user information and insufficient emotion recognition. In particular, the lack of functionality to understand the user's emotional state and dynamically adjust dialogue content accordingly makes it difficult to improve the user experience and promote effective learning.

[0374] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0375] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, means for selecting dialogue scenarios based on the dialogue profiles, means for initiating and progressing a dialogue based on the selected dialogue scenario, means for acquiring and analyzing user responses during the dialogue, means for selecting and providing content to promote learning, means for saving data collected during the dialogue and applying the feedback to the next dialogue, means for analyzing emotional data using an emotion engine that recognizes the user's emotions, and means for dynamically adjusting the dialogue content based on the emotional data. This makes it possible to provide personalized dialogues and learning content that correspond to the user's emotional state.

[0376] "Means of registering user information" refers to means that include an interface for users to input their basic information (name, age, gender, interests, learning progress, etc.) and send it to the system.

[0377] "Means for creating individual dialogue profiles based on registered information" refers to a method for storing user-registered information in a database and generating optimal dialogue content for each user based on that information.

[0378] "Means for selecting dialogue scenarios based on dialogue profiles" refers to a means that has the function of automatically selecting an appropriate dialogue scenario from a large number of dialogue scenarios based on the created dialogue profile.

[0379] "Means for initiating and conducting a dialogue based on a selected dialogue scenario" refers to means for initiating a dialogue with the user according to a selected scenario and for smoothly conducting that dialogue.

[0380] "Means for acquiring and analyzing user responses during a conversation" refers to methods for collecting voice and text data obtained from users in real time during a conversation and analyzing that data to understand their responses.

[0381] "Means of selecting and providing content to promote learning" refers to means of selecting appropriate educational content based on the user's age, interests, and learning progress, and providing it to the user.

[0382] "A means of saving data collected during a conversation and applying the feedback to the next conversation" refers to a means of saving data obtained from the user during a conversation and applying that data as feedback for use in the next conversation.

[0383] "Means for analyzing emotional data using an emotion engine that recognizes user emotions" refers to a method for recognizing emotions from a user's voice tone, facial expressions, and actions using an emotion engine, and then analyzing that recognized data.

[0384] "Means for dynamically adjusting dialogue content based on emotional data" refers to means for adjusting and optimizing the content and progress of dialogue in real time based on analyzed emotional data.

[0385] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles, and provides dialogue scenarios based on emotional data. This system is implemented using the following hardware and software.

[0386] Hardware and software

[0387] Devices: Smartphones and tablets (including pet robots)

[0388] Servers: Cloud servers and database systems

[0389] Emotion engine: EmotionAI engine

[0390] Speech recognition engine: VoiceRec

[0391] System Overview

[0392] The system registers user information and creates an individual dialogue profile based on it. Based on the dialogue profile, the optimal dialogue scenario is selected, and the dialogue content is dynamically adjusted using an emotion engine that analyzes the user's emotional state. Appropriate content for learning is provided based on emotion and response data. Furthermore, data collected during the dialogue is used as feedback for subsequent dialogues.

[0393] User information registration and interaction profile creation

[0394] Users (e.g., parents) enter basic information about their children (name, age, gender, interests, learning progress, etc.) using a smartphone app or web interface. This information is sent to the server via the device's "information transmission module." The server receives this information and stores it in its user database. Based on the registered information, it then uses a dialogue profile generation module to create a dialogue profile.

[0395] Emotion engine integration and analysis

[0396] The terminal (pet robot) is equipped with an EmotionAI engine that recognizes emotions from the user's voice tone, facial expressions, and movements. The server analyzes the received emotion data and uses an emotion analysis algorithm to identify the user's emotional state.

[0397] Facilitating and adjusting the content of the dialogue

[0398] The user gives instructions via voice commands or a dedicated app, such as "Play with pet robot Taro." The device recognizes the voice command using the VoiceRec voice recognition engine and selects an appropriate dialogue scenario using the dialogue scenario selection module. The server receives the dialogue log in real time and dynamically adjusts the dialogue content.

[0399] Acquiring and analyzing user responses

[0400] During the conversation, when the user (i.e., the child) responds, the device analyzes the response using its VoiceRec speech recognition engine and analyzes the emotion using its EmotionAI engine. Based on this data, the server generates the next conversation and selects appropriate learning content.

[0401] Learning promotion and content provision

[0402] The server uses a "learning content selection module" to select educational content (such as words or place names in different languages) based on the child's age, interests, and emotional state to facilitate learning. The device uses a "dialogue generation module" and a "speech synthesis engine" to enable the pet robot to generate personalized responses.

[0403] Data storage and feedback

[0404] The device sends data collected during the interaction (responses, learning progress, emotional state, etc.) to the server. The server stores this data in an interaction log database and uses a feedback generation module to generate feedback to be reflected in the next interaction. The device updates its new interaction profile and prepares for the next session.

[0405] Specific example

[0406] For example, if a 3-year-old child named Taro is registered and the system determines that he likes cars and that the emotion engine is "excited," the pet robot might ask, "Do you know what colors cars come in?" If Taro replies "red" and the EmotionAI engine analyzes Taro's excitement, the server will generate a response like, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for future interactions.

[0407] Example of a prompt

[0408] An example of a prompt to input into the generative AI model is: "Based on user information and sentiment data, generate a conversational scenario about cars that a 3-year-old child would be interested in. Also, provide an appropriate response if the child is excited."

[0409] This system is expected to improve the user experience by providing personalized conversational and learning content.

[0410] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0411] Step 1:

[0412] User Information Registration

[0413] Users enter basic information about their child (name, age, gender, interests, learning progress, etc.) through a smartphone app or web interface.

[0414] Input: Basic information about the child

[0415] Data processing: Enter basic information into the input fields and click the submit button.

[0416] Output: Send input information

[0417] The terminal sends the information entered by the user to the server using an "information transmission module".

[0418] Step 2:

[0419] Creating a dialogue profile

[0420] The server saves the received information to the "user database".

[0421] Input: Received user information

[0422] Data processing: Saving operations to the database

[0423] Output: Base data for dialogue profile generation

[0424] The server generates an interaction profile using the "Interaction Profile Generation Module" based on the registration information. The interaction profile includes the user's basic information and interests.

[0425] Step 3:

[0426] Embedding an emotion engine

[0427] The device (pet robot) is equipped with an "EmotionAI engine" that recognizes emotions from the user's voice tone, facial expressions, and movements.

[0428] Input: User's voice tone, facial expressions, and gestures

[0429] Data processing: Audio analysis, image analysis, motion analysis

[0430] Output: Sentiment data

[0431] The server analyzes the received emotional data using an "emotion analysis algorithm" to identify the user's emotional state.

[0432] Step 4:

[0433] Initiating and facilitating the dialogue

[0434] Users can give commands via voice commands or a dedicated app, such as "Play with my pet robot, Taro."

[0435] Input: Parental voice commands

[0436] Data processing: Speech recognition and analysis

[0437] Output: Analysis result (instruction to "play with Taro")

[0438] The device uses the "VoiceRec" voice recognition engine to recognize voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0439] The server receives dialogue logs in real time and dynamically adjusts the dialogue content using a "dialogue adjustment module".

[0440] Step 5:

[0441] Acquiring and analyzing user responses

[0442] The user (Taro) responds to the pet robot's question with "car".

[0443] Input: Taro's voice response

[0444] Data processing: Speech recognition and sentiment analysis

[0445] Output: Analysis results (response content and sentiment data)

[0446] The device analyzes Taro's responses using the "VoiceRec" voice recognition engine, and simultaneously analyzes his emotions using the "EmotionAI engine."

[0447] The server uses a "dialogue content analysis module" to analyze the response content and sentiment data, and then generates the next appropriate dialogue.

[0448] Step 6:

[0449] Providing content to promote learning

[0450] The server selects learning-enhancing content based on the user's age, interests, and emotional state.

[0451] Input: User age, interests, and sentiment data

[0452] Data processing: Execution of content selection algorithm

[0453] Output: Selected learning content

[0454] The device uses a "dialogue generation module" and a "speech synthesis engine" to generate learning content and motivating responses. For example, it might respond, "The English word for red is 'Red.' Let's say it together, red. Sounds fun!"

[0455] Step 7:

[0456] Data storage and feedback

[0457] The device sends the data collected during the conversation to the server.

[0458] Input: Data collected during the interaction (responses, learning progress, emotional state, etc.)

[0459] Data processing: Data transmission processing

[0460] Output: Sent data

[0461] The server saves the data in the "interaction log database" and uses the "feedback generation module" to apply the feedback to the next interaction.

[0462] The device updates its conversation profile with a new one and prepares for the next session.

[0463] (Application Example 2)

[0464] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0465] Traditional food delivery services have struggled to provide personalized menu recommendations and promotions tailored to users' emotions and preferences. This hindered the optimization of the user experience and the effective delivery of services. In addition, the lack of mechanisms to respond immediately to changes in users' emotions led to decreased convenience and satisfaction.

[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0467] In this invention, the server includes means for acquiring and analyzing user responses and emotional data, means for selecting and providing personalized recommended menus and promotions based on the user's emotional state, and means for storing data collected during the interaction and applying the feedback to the next interaction in the food delivery service. This enables personalized interactions and menu provision that are tailored to the user's emotions and preferences, thereby improving the user experience.

[0468] "User information" refers to basic personal information such as the user's name, age, interests, and learning progress.

[0469] A "dialogue profile" refers to a dataset containing personalized dialogue scenarios based on user registration information and sentiment data.

[0470] A "dialogue scenario" refers to a scenario that defines the progression and content of the interaction with the user, selected based on the dialogue profile.

[0471] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice tone, facial expressions, and actions.

[0472] "Recommended menus" refer to personalized menus in food delivery services that are selected based on the user's emotional state and interests.

[0473] "Promotion" refers to sales promotion activities such as special discounts and campaigns offered based on the user's emotional state and interests.

[0474] A "food delivery service" refers to a service that delivers meals ordered by users.

[0475] "Feedback" refers to information collected during a conversation that is used to improve the next conversation, and it is data that contributes to system improvement and enhanced personalization.

[0476] "User experience" refers to the overall satisfaction and convenience that users feel when using a service.

[0477] This invention relates to a food delivery service that provides personalized menu recommendations and promotions based on user information and sentiment data. Embodiments thereof are described below.

[0478] Hardware and software used

[0479] Face recognition camera (e.g., Logitech C920)

[0480] This is a camera used to capture the user's facial expressions.

[0481] Server (e.g., Amazon Web Services EC2 instance)

[0482] This server stores and analyzes user information, sentiment data, and dialogue profiles.

[0483] Emotion recognition libraries (e.g., DeepFace)

[0484] This is a library that analyzes emotional data from a user's facial expressions.

[0485] HTTP request library (e.g., request)

[0486] This is a library for communicating user information, sentiment data, and other data with the server.

[0487] Program processing details

[0488] In this invention, three elements—a server, a terminal, and a user—work together to process information.

[0489] 1. User information registration

[0490] The user uses a smartphone app to enter basic information such as their name, age, interests, and learning progress. The device sends this information to the server via an HTTP POST request, and the server stores the received information in a database. This registers the user's basic information.

[0491] 2. Emotion recognition

[0492] The device uses a facial recognition camera to capture the user's facial expressions and analyzes the emotional data using the DeepFace library. The analyzed emotional data is sent to the server using an HTTP POST request, and the server stores this information in a database. This allows the user's emotional state to be monitored in real time.

[0493] 3. Providing personalized recommended menus

[0494] The server analyzes the user's basic information and sentiment data, and generates personalized recommendations and promotions based on that information. The device retrieves this information from the server using an HTTP GET request and displays it to the user. This provides personalized services tailored to the user's sentiments and interests.

[0495] Specific example

[0496] For example, suppose a user opens a food delivery app and faces the camera. If the emotion recognition engine recognizes the user's emotion as "joy," the server uses that emotion data to generate recommended menus of sushi dishes or special promotions that match the user's interests. The device then displays these recommended menus to the user, who can then order a meal based on them.

[0497] Example of a prompt

[0498] The following are examples of prompts for a generative AI model.

[0499] Design a food delivery application that analyzes user emotions and provides food recommendations based on those emotions. Use DeepFace for emotion recognition and include specific user scenarios and dialogue scenarios.

[0500] The above is a specific description of the embodiment for carrying out the invention. This system enables the provision of highly personalized services that respond to the user's emotions and preferences, thereby improving the user experience.

[0501] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0502] Step 1: Register User Information

[0503] Users enter basic information such as their name, age, interests, and learning progress using a smartphone app. This information is sent from the device to the server via an HTTP POST request. The server stores the received information in a database and generates a conversational profile for each user. This process registers the user's basic information in the database.

[0504] Input: User information such as name, age, interests, and learning progress.

[0505] Data processing: Convert user information to JSON format and construct an HTTP POST request.

[0506] Output: User information stored on the server, and generated interaction profile.

[0507] Step 2: Emotion Recognition

[0508] The device uses a facial recognition camera to capture the user's facial expressions. The captured image data is analyzed using the DeepFace library to identify dominant emotions (e.g., "joy," "sadness," etc.). The analyzed emotion data is sent from the device to the server via an HTTP POST request, and the server stores it in a database.

[0509] Input: Captured user's face image

[0510] Data processing: Sentiment analysis using DeepFace

[0511] Output: Emotional data stored on the server

[0512] Step 3: Generating personalized recommendation menus

[0513] The server analyzes registered user information and sentiment data to generate recommended menus and promotions based on the user's current state. This is done by combining menu and promotion information that has been previously stored in the database.

[0514] Input: User information, sentiment data

[0515] Data processing: Data analysis and menu selection using AI algorithms.

[0516] Output: Personalized menu recommendations and promotional information

[0517] Step 4: Serve the recommended menu

[0518] The device uses an HTTP GET request to retrieve personalized menu recommendations and promotional information from the server. Based on this information, the app displays menus and promotions that are suitable for the user.

[0519] Input: User's request

[0520] Data processing: Retrieving and displaying menu information sent from the server.

[0521] Output: Recommended menus and promotions presented to the user.

[0522] Step 5: Obtain and analyze user feedback

[0523] The user reacts to the presented menu or promotion, and the device captures that reaction again. This data is sent to the server using an HTTP POST request, where it is analyzed and stored.

[0524] Input: User response data

[0525] Data processing: Analysis of emotions and intentions using DeepFace and speech recognition engines.

[0526] Output: User response data stored on the server

[0527] Step 6: Applying Feedback

[0528] The server uses the emotional data and user response data collected during the conversation to apply feedback to the next conversation. This allows the next conversation scenario to be tailored to be more personalized.

[0529] Input: Collected sentiment data and reaction data

[0530] Data processing: Adjusting the next dialogue scenario using past data.

[0531] Output: Updated dialogue profile and next dialogue scenario

[0532] The above outlines the processing flow of the system program that implements the application example, and each step includes specific actions.

[0533] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0534] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0535] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0536] [Second Embodiment]

[0537] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0538] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0539] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0540] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0541] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0542] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0543] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0544] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0545] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0546] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0547] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0548] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0549] This invention relates to an interactive pet robot system that registers user information, creates a dialogue profile based on that information, and provides individually adapted dialogue scenarios. This system is designed to support children's learning and play, especially for parents and seniors, when parents are unable to attend to their children.

[0550] User Information Registration

[0551] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0552] Terminal: Send this information to the server.

[0553] Server: Stores received information in a database and generates individual interaction profiles.

[0554] Creating and adapting dialogue profiles

[0555] Server: The conversation profile created based on user information reflects age, interests, and learning progress, and includes individual conversation scenarios.

[0556] Terminal: Synchronize conversation profiles with pet robots, allowing the system to provide appropriate interactions for each user.

[0557] Initiating and facilitating the dialogue

[0558] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[0559] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. It says, "Hi Taro. What would you like to do today?"

[0560] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[0561] Acquiring and analyzing user responses

[0562] User: The child responds.

[0563] Device: The device analyzes the child's response using a speech recognition engine to understand its meaning. For example, if the child says "I like cars," it retrieves that information.

[0564] Server: Based on the analyzed information, the server generates the following dialogue, and the pet robot responds, "Do you know what colors cars come in?"

[0565] Providing content to promote learning

[0566] Server: Selects learning-enhancing content (e.g., English vocabulary, place names, etc.) based on the child's age and interests.

[0567] Terminal: The pet robot teaches, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0568] User: The child keeps repeating "red."

[0569] Device: Recognizes repeated content and records accuracy and learning effectiveness.

[0570] Data storage and feedback

[0571] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0572] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[0573] Terminal: Update and prepare a new interaction profile for the next session.

[0574] Specific example

[0575] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then says "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[0576] In this way, the system can utilize user information to provide personalized conversations and learning content tailored to each user.

[0577] The following describes the processing flow.

[0578] Step 1:

[0579] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0580] Step 2:

[0581] Terminal: Sends the entered user information to the system server.

[0582] Step 3:

[0583] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[0584] Step 4:

[0585] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[0586] Step 5:

[0587] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[0588] Step 6:

[0589] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0590] Step 7:

[0591] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[0592] Step 8:

[0593] User: Taro replies, "Car."

[0594] Step 9:

[0595] Terminal: The speech recognition engine analyzes Taro's response to understand its content.

[0596] Step 10:

[0597] Server: Analyzes the response and generates the next appropriate dialogue.

[0598] Step 11:

[0599] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[0600] Step 12:

[0601] User: Taro replies "Red".

[0602] Step 13:

[0603] Terminal: The speech recognition engine analyzes Taro's response again to verify its accuracy.

[0604] Step 14:

[0605] Server: Based on this response, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[0606] Step 15:

[0607] Device: Provides selected learning content and says, "The English word for red is 'Red'. Let's say it together, red."

[0608] Step 16:

[0609] User: Taro repeats "Red".

[0610] Step 17:

[0611] Terminal: Recognizes Taro's pronunciation and records its accuracy and learning effectiveness.

[0612] Step 18:

[0613] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[0614] Step 19:

[0615] Server: Stores data collected during the interaction (Taro's responses, learning progress, etc.) in a database.

[0616] Step 20:

[0617] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[0618] (Example 1)

[0619] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0620] One of the problems faced by modern families, particularly by parents and seniors, is the difficulty in adequately supporting children's learning and play when parents are unable to attend to them directly. Traditional educational tools and interactive learning toys have limited effectiveness because they cannot provide individualized dialogue and learning content tailored to the child's age and interests. Furthermore, the technology to monitor learning progress in real time during dialogue and incorporate that information into subsequent dialogues has been insufficient.

[0621] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0622] In this invention, the server includes means for registering user information, means for storing user information in a database and generating individual dialogue profiles, and means for synchronizing the generated dialogue profiles with a dialogue robot. This makes it possible to provide individual dialogue scenarios tailored to the user's age and interests in real time, and to effectively support learning progress.

[0623] "Means for registering user information" refers to devices or software that allow users to input and register basic information about their child and parent using a smartphone app or a dedicated web interface.

[0624] "Means for sending registered information to the server" refers to a communication device or software that has the role of sending information entered by the user to the server.

[0625] "Means for storing received information in a database and generating individual dialogue profiles" refers to a device or software that stores user information received by a server in a database and generates individual dialogue profiles based on that information.

[0626] "Means for synchronizing the generated dialogue profile with the dialogue robot" refers to a communication device or software that transmits the generated dialogue profile to the dialogue robot's memory, causing the robot to operate based on that profile.

[0627] "Means of initiating dialogue based on parental instructions" refers to a device or software that initiates a dialogue session when a parent gives instructions to a dialogue robot via voice or other interface.

[0628] "Means for selecting and implementing an appropriate dialogue scenario based on a dialogue profile" refers to a device or software that selects the optimal dialogue scenario based on the content of the dialogue profile and proceeds with the dialogue according to that scenario.

[0629] "Means for analyzing and understanding the content of a user's response using a speech recognition engine" refers to a device or software that uses a speech recognition engine to convert a user's response into text and understand its content.

[0630] "Means for generating the next dialogue based on analyzed information" refers to a device or software that uses a generative AI model to generate the content of the next dialogue based on information analyzed by a speech recognition engine.

[0631] "Means for selecting and providing learning-promoting content based on a child's age and interests" refers to a device or software that selects appropriate learning content according to a child's specific age and interests and provides it to the user.

[0632] "Means of sending data collected during a dialogue to a server and feeding it back into the next dialogue" refers to a device or software that sends detailed data collected during a dialogue to a server and processes it to reflect in the next dialogue profile.

[0633] This invention relates to an interactive pet robot system that performs a series of processes including user information registration, dialogue profile generation, initiation and progression of dialogue, acquisition and analysis of user responses, provision of content to promote learning, and data storage and feedback. This system is designed particularly for parents and seniors to support children's learning and play when parents are unable to attend to them.

[0634] User Information Registration

[0635] User: Use a smartphone app or a dedicated web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter information such as "Taro, 3 years old, boy, interest: cars."

[0636] Terminal: Sends user-entered information to the server. This process uses communication methods such as HTTP POST requests.

[0637] Server: Stores received information in a database (e.g., MySQL or PostgreSQL) and generates individual interaction profiles. A generative AI model (e.g., OpenAI GPT-4) is used for profile generation.

[0638] Creating and adapting dialogue profiles

[0639] Server: Generates a conversation profile based on user information. The profile reflects age, interests, and learning progress, and includes individual conversation scenarios.

[0640] Terminal: Synchronizes dialogue profiles with the conversational robot, enabling the system to provide appropriate dialogues for each user. Specifically, it downloads dialogue profiles from the server to the terminal and saves them to the robot's memory.

[0641] Initiating and facilitating the dialogue

[0642] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[0643] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[0644] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content. A NoSQL database (e.g., MongoDB) is used to store the logs.

[0645] Acquiring and analyzing user responses

[0646] User: The child responds. For example, the child says, "I like cars."

[0647] Device: The child's responses are analyzed using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to understand the content.

[0648] Server: Based on the analyzed information, it generates the following dialogue. For example, it generates the following question: "Do you know what colors cars come in?"

[0649] Providing content to promote learning

[0650] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools (e.g., NLTK or SpaCy) are used for content selection.

[0651] Device: The pet robot provides children with selected learning content. For example, it might teach them, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0652] User: The child keeps repeating "red."

[0653] Device: Analyzes repeated content using speech recognition and records accuracy and learning effectiveness.

[0654] Data storage and feedback

[0655] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0656] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[0657] Terminal: Update and prepare a new interaction profile for the next session. Specifically, download the updated profile from the server and synchronize it with the robot.

[0658] Specific example

[0659] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then mentions the color "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[0660] Example of a prompt

[0661] The following is an example of a prompt statement for a generative AI model:

[0662] User information:

[0663] Name: Taro

[0664] Age: 3 years old

[0665] Gender: Boy

[0666] Interests: Cars

[0667] Start conversation:

[0668] User: Playing with my pet robot, Taro

[0669] Robot: Hi Taro. What do you want to do today?

[0670] User: I like cars

[0671] Robot: Do you know what colors cars come in?

[0672] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0673] Step 1: Enter user information

[0674] User: Use a smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter "Taro, 3 years old, boy, interest: cars" into the app's form.

[0675] Input: Basic information entered by the user

[0676] Output: Input basic information data

[0677] Step 2: Submit User Information

[0678] Terminal: Sends the entered user information to the server using an HTTP POST request.

[0679] Input: Basic information data entered by the user

[0680] Output: Request data sent to the server

[0681] Step 3: Saving user information and generating a profile

[0682] Server: The server stores the received information in a database (e.g., MySQL or PostgreSQL) and uses a generative AI model (e.g., OpenAI GPT-4) to generate individual interaction profiles. This information includes prompts for the generative AI model.

[0683] Input: Request data sent to the server

[0684] Output: Saved database entries and generated interaction profiles

[0685] Specific operation: Use an "INSERT" SQL query to save information to the database, create a prompt statement for profile generation, and send it to the generating AI model.

[0686] Step 4: Synchronize conversation profiles

[0687] Terminal: Downloads the dialogue profile sent from the server and synchronizes it with the dialogue robot.

[0688] Input: Interaction profile sent from the server

[0689] Output: Dialogue profile synchronized with the conversational robot

[0690] Specific operation: Download the profile using an HTTP GET request and save it to the conversational robot's memory.

[0691] Step 5: Start the conversation

[0692] User: The parent gives voice commands to the conversational robot, saying, "Pet robot, play with Taro."

[0693] Input: Parental voice commands

[0694] Output: Trigger signal to start interaction

[0695] Specific operation: A microphone system for receiving voice commands recognizes the command and sends a trigger to the robot system to start the dialogue.

[0696] Step 6: Select and implement a dialogue scenario

[0697] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and starts the conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[0698] Input: Interaction Profile

[0699] Output: Selected dialogue scenario and its implementation

[0700] Specific operation: The robot selects the optimal scenario from the dialogue profiles and plays the scenario aloud.

[0701] Step 7: Obtain and analyze user responses

[0702] Device: The system analyzes the child's response using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, if the child says "I like cars," the audio is converted into text.

[0703] Input: Child's voice response

[0704] Output: Analyzed data in text format

[0705] Specific operation: Audio data is sent to the API and returned as text data.

[0706] Step 8: Generating the next dialogue

[0707] Server: Based on the analyzed text data, it generates the next dialogue using an AI model. For example, it generates the following question: "Do you know what colors cars come in?"

[0708] Input: Text-formatted analysis data

[0709] Output: The following dialogue

[0710] Specific operation: A prompt message is sent to the generative AI model, and the generated dialogue content is retrieved.

[0711] Step 9: Provide content to facilitate learning

[0712] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools are used to select content. For example, it might teach, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[0713] Input: Interaction profile and user response data

[0714] Output: Selected learning content

[0715] Specific operation: The robot selects content from a database that matches the child's interests and teaches them using voice.

[0716] Step 10: Data storage and feedback

[0717] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[0718] Input: Collected data

[0719] Output: Feedback data sent to the server

[0720] Specific operation: The collected data is sent to the server via an HTTP POST request.

[0721] Step 11: Save feedback data and update the next interaction profile.

[0722] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[0723] Input: Feedback data sent to the server

[0724] Output: Updated interaction profile

[0725] Specific actions: Record feedback data in a database and use a generative AI model to update the next interaction profile.

[0726] (Application Example 1)

[0727] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0728] Traditional interactive systems were designed to support user learning and play, but they were difficult to apply to new employee training in factories and other similar settings. These systems struggled with providing complex operational instructions and training content tailored to skill levels, resulting in limited practicality. Furthermore, they lacked the ability to adapt training content in real time or analyze user instructions using speech recognition technology. This highlighted the challenge of improving the efficiency of training in new environments.

[0729] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0730] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, and means for providing training programs based on the user information. This makes it possible to provide personalized training content according to the user's skill level and learning progress.

[0731] "Means of registering user information" refers to interfaces or mechanisms for inputting basic information such as a user's name, age, job title, and skill level, and saving it in a database.

[0732] "Means for creating dialogue profiles" refers to a mechanism that generates individual dialogue scenarios based on registered user information and sets up dialogue content tailored to the user.

[0733] "Means for selecting dialogue scenarios" refers to a mechanism that selects a scenario that fits the created dialogue profile and facilitates effective dialogue with the user.

[0734] "Means for initiating and conducting a dialogue" refers to a mechanism for initiating a dialogue with a user based on a selected dialogue scenario and continuing the dialogue while providing timely responses.

[0735] "Means for acquiring and analyzing user responses" refers to a mechanism that analyzes the voice and actions obtained from the user during a conversation and dynamically adjusts the content of the next conversation based on that analysis.

[0736] "Means for selecting and providing content to promote learning" refers to a mechanism for selecting and providing educational content (e.g., vocabulary and operating procedures) that is appropriate for the user's age, interests, and skill level.

[0737] "Means of saving data collected during a conversation and applying feedback to the next conversation" refers to a mechanism that records the user's responses and learning progress obtained during a conversation and reflects them in subsequent conversations.

[0738] "Means of providing training programs" refers to a mechanism for generating and providing suitable training content based on user information.

[0739] "A means of analyzing user instructions using a speech recognition engine and generating the next dialogue content" refers to a mechanism that uses speech recognition technology to convert user statements into text and generates the next response based on that text.

[0740] "Means for providing instructions and training content using voice output means" refers to a mechanism for providing text-based instructions and training content to the user as audio.

[0741] To implement this invention, a server, terminal, and user must work together to build a dialogue system. The system of this invention consists of the following steps: user information registration, dialogue profile creation, training program provision, dialogue initiation and progress, user response acquisition and analysis, learning promotion content provision, data storage and feedback. The details are described below.

[0742] User Information Registration

[0743] Terminal: Users enter their information (name, job title, skill level, etc.) using a smartphone app or head-mounted display (HMD). This information is sent to the server and stored in the database.

[0744] Server: Based on the received information, it generates individual interaction profiles. These interaction profiles are customized according to the user's skill level and role.

[0745] Training program provision

[0746] Server: Based on user information, generates an appropriate training scenario and sends it to the terminal.

[0747] Terminal: Displays and provides the received training scenario to the user. For example, if a new employee says, "Teach me how to operate this machine," the terminal will display instructions from the server in both audio and display format.

[0748] Initiating and facilitating the dialogue

[0749] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[0750] Terminal: Uses a speech recognition engine to analyze the user's words and sends them to the server.

[0751] Server: Based on the analysis results, it generates the next dialogue content and sends it to the terminal. The dialogue content is dynamically adjusted in real time.

[0752] Acquiring and analyzing user responses

[0753] Terminal: Analyzes the user's voice input in real time and sends the results to the server.

[0754] Server: Based on the received information, it generates the next dialogue content and sends it to the terminal. This provides personalized training content.

[0755] Providing learning-enhancing content

[0756] Server: Select appropriate learning content (e.g., machine operation procedures, troubleshooting methods, etc.) according to the user's skill level and role.

[0757] Terminal: Provides selected content using an audio output device. For example, it might instruct, "First, press the power button. Then, press the start button to operate the machine."

[0758] Data storage and feedback

[0759] Terminal: Sends data collected during the interaction to the server.

[0760] Server: Receives data and saves it to the database, applying the feedback to the next interaction. This makes subsequent interactions more tailored to the user.

[0761] Examples of specific cases and prompt statements

[0762] As a concrete example, let's consider a scenario where a new employee, "Yamada," learns the basic operation of a new machine. In this case, the following prompt messages are used.

[0763] Example of a prompt

[0764] User: Yamada

[0765] Position: Operator

[0766] Skill level: beginner

[0767] User: Please tell me how to operate this machine.

[0768] Robot: "I'll teach you the basic operation of the machine. First, press the power button. Then, press the start button to start the machine."

[0769] In this way, the system of the present invention can utilize user information to provide personalized conversations and training content to each user.

[0770] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0771] Step 1: Register User Information

[0772] Terminal: Users (new recruits and employees) use a smartphone app or head-mounted display (HMD) to enter basic information such as name, job title, and skill level. This information is sent from the terminal to the server.

[0773] Input: User's basic information (name, job title, skill level)

[0774] Output: User information sent to the server

[0775] Step 2: Create an interaction profile

[0776] Server: Based on the received user information, it generates an individual dialogue profile. The dialogue profile includes dialogue scenarios based on the user's skill level and job title.

[0777] Input: User information

[0778] Output: Individual interaction profiles

[0779] Step 3: Providing the training program

[0780] Server: Based on the interaction profile, generates an appropriate training scenario and sends it to the terminal.

[0781] Input: Interaction Profile

[0782] Output: Training scenario

[0783] Step 4: Start the conversation

[0784] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[0785] Input: User's voice instructions

[0786] Output: Startup of speech recognition engine

[0787] Step 5: Voice Analysis

[0788] Terminal: Uses a speech recognition engine to convert user voice commands into text data and send it to the server.

[0789] Input: User's voice instructions

[0790] Output: Text data

[0791] Step 6: Generating dialogue content

[0792] Server: Based on the received text data, it generates the next dialogue content and sends it to the terminal.

[0793] Input: Text data, dialogue profile

[0794] Output: Dialogue content (training content)

[0795] Step 7: Instructions via voice and display

[0796] Terminal: Provides the user with the content of the conversation received from the server via voice output and display. For example, it might instruct the user, "First, press the power button. Then, press the start button to operate the machine."

[0797] Input: Dialogue content

[0798] Output: Audio and display

[0799] Step 8: Obtain and analyze user responses

[0800] Terminal: Analyzes user voice input in real time and sends the results to the server. For example, if a user asks, "Where is the power button?", the device analyzes the voice and converts it into text data.

[0801] Input: User voice input

[0802] Output: Analyzed text data

[0803] Step 9: Generating Feedback

[0804] Server: Based on the analyzed text data, it generates the next dialogue and sends it to the terminal. This provides personalized training content in real time.

[0805] Input: Analyzed text data, dialogue profile

[0806] Output: The following dialogue

[0807] Step 10: Data storage and feedback

[0808] Terminal: Sends data collected during the interaction to the server.

[0809] Server: Receives data and saves it to a database, which is then used in subsequent interactions. This data includes user responses, learning progress, and other information.

[0810] Input: Data collected during the conversation

[0811] Output: Saved to the database, feedback that will be reflected in the next interaction.

[0812] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0813] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles based on that information, and provides dialogue scenarios. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized dialogue and learning content.

[0814] User Information Registration

[0815] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0816] Terminal: Send this information to the server.

[0817] Server: Stores received information in a database and generates an interaction profile based on it.

[0818] Embedding an emotion engine

[0819] Device: The pet robot will be equipped with an emotion engine that recognizes emotions from the user's voice tone, facial expressions, and movements.

[0820] Server: Implement an algorithm that analyzes emotional data and identifies the user's current emotional state.

[0821] Creating and adapting dialogue profiles

[0822] Server: Generates dialogue profiles based on user information and emotional data. These dialogue profiles include personalized dialogue scenarios tailored to the user's age, interests, learning progress, and emotional state.

[0823] Terminal: Synchronizes conversation profiles with pet robots, allowing the system to provide optimal interactions for each user.

[0824] Initiating and facilitating the dialogue

[0825] User: The parent gives instructions via voice commands or the app, such as "Play with pet robot Taro."

[0826] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0827] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[0828] Acquiring and analyzing user responses

[0829] User: Taro replies, "Car."

[0830] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions.

[0831] Server: Analyzes response content and sentiment data to generate the next appropriate dialogue.

[0832] Providing content to promote learning

[0833] Server: Selects learning-enhancing content (English vocabulary, place names, etc.) based on the child's age, interests, and emotional state.

[0834] Terminal: The pet robot generates an encouraging response, saying, "The English word for red is 'Red.' Let's say it together, 'Red.' Sounds fun!"

[0835] Data storage and feedback

[0836] Terminal: Sends data collected during the conversation (Taro's responses, learning progress, emotional state, etc.) to the server.

[0837] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[0838] Terminal: Update and prepare a new interaction profile for the next session.

[0839] Specific example

[0840] For example, if a 3-year-old child named Taro is registered, and along with information that he likes cars, the emotion engine analyzes Taro's emotional state as "very interested." Based on this information, the pet robot asks, "Do you know what colors cars come in?" If Taro replies "red," and the emotion engine analyzes Taro's joy and excitement, the server generates a response such as, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for the content of the next conversation.

[0841] In this way, the system utilizes user information and emotional data to provide personalized dialogue and learning content, and can also apply feedback tailored to the user's emotional state.

[0842] The following describes the processing flow.

[0843] Step 1:

[0844] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[0845] Step 2:

[0846] Terminal: Sends the entered user information to the system server.

[0847] Step 3:

[0848] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[0849] Step 4:

[0850] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[0851] Step 5:

[0852] Terminal: The pet robot activates its built-in emotion engine to recognize the user's emotional state in real time based on their voice tone, facial expressions, and movements.

[0853] Step 6:

[0854] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[0855] Step 7:

[0856] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0857] Step 8:

[0858] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[0859] Step 9:

[0860] User: Taro replies, "I like cars."

[0861] Step 10:

[0862] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions (excitement, joy, etc.).

[0863] Step 11:

[0864] Server: Based on the response content and sentiment data, it generates the next appropriate dialogue.

[0865] Step 12:

[0866] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[0867] Step 13:

[0868] User: Taro replies "Red".

[0869] Step 14:

[0870] Terminal: Taro's response is analyzed by a speech recognition engine, and along with the content of the response, an emotion engine is used to analyze his emotion of joy.

[0871] Step 15:

[0872] Server: Based on this response and sentiment data, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[0873] Step 16:

[0874] Device: Provides selected learning content and says, "The English word for red is 'Red.' Let's say it together, 'Red.' Isn't that great!"

[0875] Step 17:

[0876] User: Taro repeats "Red".

[0877] Step 18:

[0878] Terminal: Recognizes Taro's pronunciation, records its accuracy and learning effectiveness, and simultaneously analyzes Taro's sense of accomplishment using an emotion engine.

[0879] Step 19:

[0880] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[0881] Step 20:

[0882] Server: Stores data collected during the interaction (Taro's responses, learning progress, emotional state, etc.) in a database.

[0883] Step 21:

[0884] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[0885] (Example 2)

[0886] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0887] Traditional dialogue systems have faced challenges in providing personalized dialogue and learning content due to a lack of individual dialogue profiles based on user information and insufficient emotion recognition. In particular, the lack of functionality to understand the user's emotional state and dynamically adjust dialogue content accordingly makes it difficult to improve the user experience and promote effective learning.

[0888] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0889] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, means for selecting dialogue scenarios based on the dialogue profiles, means for initiating and progressing a dialogue based on the selected dialogue scenario, means for acquiring and analyzing user responses during the dialogue, means for selecting and providing content to promote learning, means for saving data collected during the dialogue and applying the feedback to the next dialogue, means for analyzing emotional data using an emotion engine that recognizes the user's emotions, and means for dynamically adjusting the dialogue content based on the emotional data. This makes it possible to provide personalized dialogues and learning content that correspond to the user's emotional state.

[0890] "Means of registering user information" refers to means that include an interface for users to input their basic information (name, age, gender, interests, learning progress, etc.) and send it to the system.

[0891] "Means for creating individual dialogue profiles based on registered information" refers to a method for storing user-registered information in a database and generating optimal dialogue content for each user based on that information.

[0892] "Means for selecting dialogue scenarios based on dialogue profiles" refers to a means that has the function of automatically selecting an appropriate dialogue scenario from a large number of dialogue scenarios based on the created dialogue profile.

[0893] "Means for initiating and conducting a dialogue based on a selected dialogue scenario" refers to means for initiating a dialogue with the user according to a selected scenario and for smoothly conducting that dialogue.

[0894] "Means for acquiring and analyzing user responses during a conversation" refers to methods for collecting voice and text data obtained from users in real time during a conversation and analyzing that data to understand their responses.

[0895] "Means of selecting and providing content to promote learning" refers to means of selecting appropriate educational content based on the user's age, interests, and learning progress, and providing it to the user.

[0896] "A means of saving data collected during a conversation and applying the feedback to the next conversation" refers to a means of saving data obtained from the user during a conversation and applying that data as feedback for use in the next conversation.

[0897] "Means for analyzing emotional data using an emotion engine that recognizes user emotions" refers to a method for recognizing emotions from a user's voice tone, facial expressions, and actions using an emotion engine, and then analyzing that recognized data.

[0898] "Means for dynamically adjusting dialogue content based on emotional data" refers to means for adjusting and optimizing the content and progress of dialogue in real time based on analyzed emotional data.

[0899] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles, and provides dialogue scenarios based on emotional data. This system is implemented using the following hardware and software.

[0900] Hardware and software

[0901] Devices: Smartphones and tablets (including pet robots)

[0902] Servers: Cloud servers and database systems

[0903] Emotion engine: EmotionAI engine

[0904] Speech recognition engine: VoiceRec

[0905] System Overview

[0906] The system registers user information and creates an individual dialogue profile based on it. Based on the dialogue profile, the optimal dialogue scenario is selected, and the dialogue content is dynamically adjusted using an emotion engine that analyzes the user's emotional state. Appropriate content for learning is provided based on emotion and response data. Furthermore, data collected during the dialogue is used as feedback for subsequent dialogues.

[0907] User information registration and interaction profile creation

[0908] Users (e.g., parents) enter basic information about their children (name, age, gender, interests, learning progress, etc.) using a smartphone app or web interface. This information is sent to the server via the device's "information transmission module." The server receives this information and stores it in its user database. Based on the registered information, it then uses a dialogue profile generation module to create a dialogue profile.

[0909] Emotion engine integration and analysis

[0910] The terminal (pet robot) is equipped with an EmotionAI engine that recognizes emotions from the user's voice tone, facial expressions, and movements. The server analyzes the received emotion data and uses an emotion analysis algorithm to identify the user's emotional state.

[0911] Facilitating and adjusting the content of the dialogue

[0912] The user gives instructions via voice commands or a dedicated app, such as "Play with pet robot Taro." The device recognizes the voice command using the VoiceRec voice recognition engine and selects an appropriate dialogue scenario using the dialogue scenario selection module. The server receives the dialogue log in real time and dynamically adjusts the dialogue content.

[0913] Acquiring and analyzing user responses

[0914] During the conversation, when the user (i.e., the child) responds, the device analyzes the response using its VoiceRec speech recognition engine and analyzes the emotion using its EmotionAI engine. Based on this data, the server generates the next conversation and selects appropriate learning content.

[0915] Learning promotion and content provision

[0916] The server uses a "learning content selection module" to select educational content (such as words or place names in different languages) based on the child's age, interests, and emotional state to facilitate learning. The device uses a "dialogue generation module" and a "speech synthesis engine" to enable the pet robot to generate personalized responses.

[0917] Data storage and feedback

[0918] The device sends data collected during the interaction (responses, learning progress, emotional state, etc.) to the server. The server stores this data in an interaction log database and uses a feedback generation module to generate feedback to be reflected in the next interaction. The device updates its new interaction profile and prepares for the next session.

[0919] Specific example

[0920] For example, if a 3-year-old child named Taro is registered and the system determines that he likes cars and that the emotion engine is "excited," the pet robot might ask, "Do you know what colors cars come in?" If Taro replies "red" and the EmotionAI engine analyzes Taro's excitement, the server will generate a response like, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for future interactions.

[0921] Example of a prompt

[0922] An example of a prompt to input into the generative AI model is: "Based on user information and sentiment data, generate a conversational scenario about cars that a 3-year-old child would be interested in. Also, provide an appropriate response if the child is excited."

[0923] This system is expected to improve the user experience by providing personalized conversational and learning content.

[0924] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0925] Step 1:

[0926] User Information Registration

[0927] Users enter basic information about their child (name, age, gender, interests, learning progress, etc.) through a smartphone app or web interface.

[0928] Input: Basic information about the child

[0929] Data processing: Enter basic information into the input fields and click the submit button.

[0930] Output: Send input information

[0931] The terminal sends the information entered by the user to the server using an "information transmission module".

[0932] Step 2:

[0933] Creating a dialogue profile

[0934] The server saves the received information to the "user database".

[0935] Input: Received user information

[0936] Data processing: Saving operations to the database

[0937] Output: Base data for dialogue profile generation

[0938] The server generates an interaction profile using the "Interaction Profile Generation Module" based on the registration information. The interaction profile includes the user's basic information and interests.

[0939] Step 3:

[0940] Embedding an emotion engine

[0941] The device (pet robot) is equipped with an "EmotionAI engine" that recognizes emotions from the user's voice tone, facial expressions, and movements.

[0942] Input: User's voice tone, facial expressions, and gestures

[0943] Data processing: Audio analysis, image analysis, motion analysis

[0944] Output: Sentiment data

[0945] The server analyzes the received emotional data using an "emotion analysis algorithm" to identify the user's emotional state.

[0946] Step 4:

[0947] Initiating and facilitating the dialogue

[0948] Users can give commands via voice commands or a dedicated app, such as "Play with my pet robot, Taro."

[0949] Input: Parental voice commands

[0950] Data processing: Speech recognition and analysis

[0951] Output: Analysis result (instruction to "play with Taro")

[0952] The device uses the "VoiceRec" voice recognition engine to recognize voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[0953] The server receives dialogue logs in real time and dynamically adjusts the dialogue content using a "dialogue adjustment module".

[0954] Step 5:

[0955] Acquiring and analyzing user responses

[0956] The user (Taro) responds to the pet robot's question with "car".

[0957] Input: Taro's voice response

[0958] Data processing: Speech recognition and sentiment analysis

[0959] Output: Analysis results (response content and sentiment data)

[0960] The device analyzes Taro's responses using the "VoiceRec" voice recognition engine, and simultaneously analyzes his emotions using the "EmotionAI engine."

[0961] The server uses a "dialogue content analysis module" to analyze the response content and sentiment data, and then generates the next appropriate dialogue.

[0962] Step 6:

[0963] Providing content to promote learning

[0964] The server selects learning-enhancing content based on the user's age, interests, and emotional state.

[0965] Input: User age, interests, and sentiment data

[0966] Data processing: Execution of content selection algorithm

[0967] Output: Selected learning content

[0968] The device uses a "dialogue generation module" and a "speech synthesis engine" to generate learning content and motivating responses. For example, it might respond, "The English word for red is 'Red.' Let's say it together, red. Sounds fun!"

[0969] Step 7:

[0970] Data storage and feedback

[0971] The device sends the data collected during the conversation to the server.

[0972] Input: Data collected during the interaction (responses, learning progress, emotional state, etc.)

[0973] Data processing: Data transmission processing

[0974] Output: Sent data

[0975] The server saves the data in the "interaction log database" and uses the "feedback generation module" to apply the feedback to the next interaction.

[0976] The device updates its conversation profile with a new one and prepares for the next session.

[0977] (Application Example 2)

[0978] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0979] Traditional food delivery services have struggled to provide personalized menu recommendations and promotions tailored to users' emotions and preferences. This hindered the optimization of the user experience and the effective delivery of services. In addition, the lack of mechanisms to respond immediately to changes in users' emotions led to decreased convenience and satisfaction.

[0980] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0981] In this invention, the server includes means for acquiring and analyzing user responses and emotional data, means for selecting and providing personalized recommended menus and promotions based on the user's emotional state, and means for storing data collected during the interaction and applying the feedback to the next interaction in the food delivery service. This enables personalized interactions and menu provision that are tailored to the user's emotions and preferences, thereby improving the user experience.

[0982] "User information" refers to basic personal information such as the user's name, age, interests, and learning progress.

[0983] A "dialogue profile" refers to a dataset containing personalized dialogue scenarios based on user registration information and sentiment data.

[0984] A "dialogue scenario" refers to a scenario that defines the progression and content of the interaction with the user, selected based on the dialogue profile.

[0985] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice tone, facial expressions, and actions.

[0986] "Recommended menus" refer to personalized menus in food delivery services that are selected based on the user's emotional state and interests.

[0987] "Promotion" refers to sales promotion activities such as special discounts and campaigns offered based on the user's emotional state and interests.

[0988] A "food delivery service" refers to a service that delivers meals ordered by users.

[0989] "Feedback" refers to information collected during a conversation that is used to improve the next conversation, and it is data that contributes to system improvement and enhanced personalization.

[0990] "User experience" refers to the overall satisfaction and convenience that users feel when using a service.

[0991] This invention relates to a food delivery service that provides personalized menu recommendations and promotions based on user information and sentiment data. Embodiments thereof are described below.

[0992] Hardware and software used

[0993] Face recognition camera (e.g., Logitech C920)

[0994] This is a camera used to capture the user's facial expressions.

[0995] Server (e.g., Amazon Web Services EC2 instance)

[0996] This server stores and analyzes user information, sentiment data, and dialogue profiles.

[0997] Emotion recognition libraries (e.g., DeepFace)

[0998] This is a library that analyzes emotional data from a user's facial expressions.

[0999] HTTP request library (e.g., request)

[1000] This is a library for communicating user information, sentiment data, and other data with the server.

[1001] Program processing details

[1002] In this invention, three elements—a server, a terminal, and a user—work together to process information.

[1003] 1. User information registration

[1004] The user uses a smartphone app to enter basic information such as their name, age, interests, and learning progress. The device sends this information to the server via an HTTP POST request, and the server stores the received information in a database. This registers the user's basic information.

[1005] 2. Emotion recognition

[1006] The device uses a facial recognition camera to capture the user's facial expressions and analyzes the emotional data using the DeepFace library. The analyzed emotional data is sent to the server using an HTTP POST request, and the server stores this information in a database. This allows the user's emotional state to be monitored in real time.

[1007] 3. Providing personalized recommended menus

[1008] The server analyzes the user's basic information and sentiment data, and generates personalized recommendations and promotions based on that information. The device retrieves this information from the server using an HTTP GET request and displays it to the user. This provides personalized services tailored to the user's sentiments and interests.

[1009] Specific example

[1010] For example, suppose a user opens a food delivery app and faces the camera. If the emotion recognition engine recognizes the user's emotion as "joy," the server uses that emotion data to generate recommended menus of sushi dishes or special promotions that match the user's interests. The device then displays these recommended menus to the user, who can then order a meal based on them.

[1011] Example of a prompt

[1012] The following are examples of prompts for a generative AI model.

[1013] Design a food delivery application that analyzes user emotions and provides food recommendations based on those emotions. Use DeepFace for emotion recognition and include specific user scenarios and dialogue scenarios.

[1014] The above is a specific description of the embodiment for carrying out the invention. This system enables the provision of highly personalized services that respond to the user's emotions and preferences, thereby improving the user experience.

[1015] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1016] Step 1: Register User Information

[1017] Users enter basic information such as their name, age, interests, and learning progress using a smartphone app. This information is sent from the device to the server via an HTTP POST request. The server stores the received information in a database and generates a conversational profile for each user. This process registers the user's basic information in the database.

[1018] Input: User information such as name, age, interests, and learning progress.

[1019] Data processing: Convert user information to JSON format and construct an HTTP POST request.

[1020] Output: User information stored on the server, and generated interaction profile.

[1021] Step 2: Emotion Recognition

[1022] The device uses a facial recognition camera to capture the user's facial expressions. The captured image data is analyzed using the DeepFace library to identify dominant emotions (e.g., "joy," "sadness," etc.). The analyzed emotion data is sent from the device to the server via an HTTP POST request, and the server stores it in a database.

[1023] Input: Captured user's face image

[1024] Data processing: Sentiment analysis using DeepFace

[1025] Output: Emotional data stored on the server

[1026] Step 3: Generating personalized recommendation menus

[1027] The server analyzes registered user information and sentiment data to generate recommended menus and promotions based on the user's current state. This is done by combining menu and promotion information that has been previously stored in the database.

[1028] Input: User information, sentiment data

[1029] Data processing: Data analysis and menu selection using AI algorithms.

[1030] Output: Personalized menu recommendations and promotional information

[1031] Step 4: Serve the recommended menu

[1032] The device uses an HTTP GET request to retrieve personalized menu recommendations and promotional information from the server. Based on this information, the app displays menus and promotions that are suitable for the user.

[1033] Input: User's request

[1034] Data processing: Retrieving and displaying menu information sent from the server.

[1035] Output: Recommended menus and promotions presented to the user.

[1036] Step 5: Obtain and analyze user feedback

[1037] The user reacts to the presented menu or promotion, and the device captures that reaction again. This data is sent to the server using an HTTP POST request, where it is analyzed and stored.

[1038] Input: User response data

[1039] Data processing: Analysis of emotions and intentions using DeepFace and speech recognition engines.

[1040] Output: User response data stored on the server

[1041] Step 6: Applying Feedback

[1042] The server uses the emotional data and user response data collected during the conversation to apply feedback to the next conversation. This allows the next conversation scenario to be tailored to be more personalized.

[1043] Input: Collected sentiment data and reaction data

[1044] Data processing: Adjusting the next dialogue scenario using past data.

[1045] Output: Updated dialogue profile and next dialogue scenario

[1046] The above outlines the processing flow of the system program that implements the application example, and each step includes specific actions.

[1047] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1048] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1049] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1050] [Third Embodiment]

[1051] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1052] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1053] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1054] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1055] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1056] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1057] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1058] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1059] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1060] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1061] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1062] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1063] This invention relates to an interactive pet robot system that registers user information, creates a dialogue profile based on that information, and provides individually adapted dialogue scenarios. This system is designed to support children's learning and play, especially for parents and seniors, when parents are unable to attend to their children.

[1064] User Information Registration

[1065] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1066] Terminal: Send this information to the server.

[1067] Server: Stores received information in a database and generates individual interaction profiles.

[1068] Creating and adapting dialogue profiles

[1069] Server: The conversation profile created based on user information reflects age, interests, and learning progress, and includes individual conversation scenarios.

[1070] Terminal: Synchronize conversation profiles with pet robots, allowing the system to provide appropriate interactions for each user.

[1071] Initiating and facilitating the dialogue

[1072] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[1073] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. It says, "Hi Taro. What would you like to do today?"

[1074] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[1075] Acquiring and analyzing user responses

[1076] User: The child responds.

[1077] Device: The device analyzes the child's response using a speech recognition engine to understand its meaning. For example, if the child says "I like cars," it retrieves that information.

[1078] Server: Based on the analyzed information, the server generates the following dialogue, and the pet robot responds, "Do you know what colors cars come in?"

[1079] Providing content to promote learning

[1080] Server: Selects learning-enhancing content (e.g., English vocabulary, place names, etc.) based on the child's age and interests.

[1081] Terminal: The pet robot teaches, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1082] User: The child keeps repeating "red."

[1083] Device: Recognizes repeated content and records accuracy and learning effectiveness.

[1084] Data storage and feedback

[1085] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1086] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[1087] Terminal: Update and prepare a new interaction profile for the next session.

[1088] Specific example

[1089] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then says "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[1090] In this way, the system can utilize user information to provide personalized conversations and learning content tailored to each user.

[1091] The following describes the processing flow.

[1092] Step 1:

[1093] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1094] Step 2:

[1095] Terminal: Sends the entered user information to the system server.

[1096] Step 3:

[1097] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[1098] Step 4:

[1099] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[1100] Step 5:

[1101] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[1102] Step 6:

[1103] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1104] Step 7:

[1105] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[1106] Step 8:

[1107] User: Taro replies, "Car."

[1108] Step 9:

[1109] Terminal: The speech recognition engine analyzes Taro's response to understand its content.

[1110] Step 10:

[1111] Server: Analyzes the response and generates the next appropriate dialogue.

[1112] Step 11:

[1113] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[1114] Step 12:

[1115] User: Taro replies "Red".

[1116] Step 13:

[1117] Terminal: The speech recognition engine analyzes Taro's response again to verify its accuracy.

[1118] Step 14:

[1119] Server: Based on this response, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[1120] Step 15:

[1121] Device: Provides selected learning content and says, "The English word for red is 'Red'. Let's say it together, red."

[1122] Step 16:

[1123] User: Taro repeats "Red".

[1124] Step 17:

[1125] Terminal: Recognizes Taro's pronunciation and records its accuracy and learning effectiveness.

[1126] Step 18:

[1127] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[1128] Step 19:

[1129] Server: Stores data collected during the interaction (Taro's responses, learning progress, etc.) in a database.

[1130] Step 20:

[1131] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[1132] (Example 1)

[1133] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1134] One of the problems faced by modern families, particularly by parents and seniors, is the difficulty in adequately supporting children's learning and play when parents are unable to attend to them directly. Traditional educational tools and interactive learning toys have limited effectiveness because they cannot provide individualized dialogue and learning content tailored to the child's age and interests. Furthermore, the technology to monitor learning progress in real time during dialogue and incorporate that information into subsequent dialogues has been insufficient.

[1135] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1136] In this invention, the server includes means for registering user information, means for storing user information in a database and generating individual dialogue profiles, and means for synchronizing the generated dialogue profiles with a dialogue robot. This makes it possible to provide individual dialogue scenarios tailored to the user's age and interests in real time, and to effectively support learning progress.

[1137] "Means for registering user information" refers to devices or software that allow users to input and register basic information about their child and parent using a smartphone app or a dedicated web interface.

[1138] "Means for sending registered information to the server" refers to a communication device or software that has the role of sending information entered by the user to the server.

[1139] "Means for storing received information in a database and generating individual dialogue profiles" refers to a device or software that stores user information received by a server in a database and generates individual dialogue profiles based on that information.

[1140] "Means for synchronizing the generated dialogue profile with the dialogue robot" refers to a communication device or software that transmits the generated dialogue profile to the dialogue robot's memory, causing the robot to operate based on that profile.

[1141] "Means of initiating dialogue based on parental instructions" refers to a device or software that initiates a dialogue session when a parent gives instructions to a dialogue robot via voice or other interface.

[1142] "Means for selecting and implementing an appropriate dialogue scenario based on a dialogue profile" refers to a device or software that selects the optimal dialogue scenario based on the content of the dialogue profile and proceeds with the dialogue according to that scenario.

[1143] "Means for analyzing and understanding the content of a user's response using a speech recognition engine" refers to a device or software that uses a speech recognition engine to convert a user's response into text and understand its content.

[1144] "Means for generating the next dialogue based on analyzed information" refers to a device or software that uses a generative AI model to generate the content of the next dialogue based on information analyzed by a speech recognition engine.

[1145] "Means for selecting and providing learning-promoting content based on a child's age and interests" refers to a device or software that selects appropriate learning content according to a child's specific age and interests and provides it to the user.

[1146] "Means of sending data collected during a dialogue to a server and feeding it back into the next dialogue" refers to a device or software that sends detailed data collected during a dialogue to a server and processes it to reflect in the next dialogue profile.

[1147] This invention relates to an interactive pet robot system that performs a series of processes including user information registration, dialogue profile generation, initiation and progression of dialogue, acquisition and analysis of user responses, provision of content to promote learning, and data storage and feedback. This system is designed particularly for parents and seniors to support children's learning and play when parents are unable to attend to them.

[1148] User Information Registration

[1149] User: Use a smartphone app or a dedicated web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter information such as "Taro, 3 years old, boy, interest: cars."

[1150] Terminal: Sends user-entered information to the server. This process uses communication methods such as HTTP POST requests.

[1151] Server: Stores received information in a database (e.g., MySQL or PostgreSQL) and generates individual interaction profiles. A generative AI model (e.g., OpenAI GPT-4) is used for profile generation.

[1152] Creating and adapting dialogue profiles

[1153] Server: Generates a conversation profile based on user information. The profile reflects age, interests, and learning progress, and includes individual conversation scenarios.

[1154] Terminal: Synchronizes dialogue profiles with the conversational robot, enabling the system to provide appropriate dialogues for each user. Specifically, it downloads dialogue profiles from the server to the terminal and saves them to the robot's memory.

[1155] Initiating and facilitating the dialogue

[1156] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[1157] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[1158] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content. A NoSQL database (e.g., MongoDB) is used to store the logs.

[1159] Acquiring and analyzing user responses

[1160] User: The child responds. For example, the child says, "I like cars."

[1161] Device: The child's responses are analyzed using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to understand the content.

[1162] Server: Based on the analyzed information, it generates the following dialogue. For example, it generates the following question: "Do you know what colors cars come in?"

[1163] Providing content to promote learning

[1164] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools (e.g., NLTK or SpaCy) are used for content selection.

[1165] Device: The pet robot provides children with selected learning content. For example, it might teach them, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1166] User: The child keeps repeating "red."

[1167] Device: Analyzes repeated content using speech recognition and records accuracy and learning effectiveness.

[1168] Data storage and feedback

[1169] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1170] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[1171] Terminal: Update and prepare a new interaction profile for the next session. Specifically, download the updated profile from the server and synchronize it with the robot.

[1172] Specific example

[1173] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then mentions the color "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[1174] Example of a prompt

[1175] The following is an example of a prompt statement for a generative AI model:

[1176] User information:

[1177] Name: Taro

[1178] Age: 3 years old

[1179] Gender: Boy

[1180] Interests: Cars

[1181] Start conversation:

[1182] User: Playing with my pet robot, Taro

[1183] Robot: Hi Taro. What do you want to do today?

[1184] User: I like cars

[1185] Robot: Do you know what colors cars come in?

[1186] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1187] Step 1: Enter user information

[1188] User: Use a smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter "Taro, 3 years old, boy, interest: cars" into the app's form.

[1189] Input: Basic information entered by the user

[1190] Output: Input basic information data

[1191] Step 2: Submit User Information

[1192] Terminal: Sends the entered user information to the server using an HTTP POST request.

[1193] Input: Basic information data entered by the user

[1194] Output: Request data sent to the server

[1195] Step 3: Saving user information and generating a profile

[1196] Server: The server stores the received information in a database (e.g., MySQL or PostgreSQL) and uses a generative AI model (e.g., OpenAI GPT-4) to generate individual interaction profiles. This information includes prompts for the generative AI model.

[1197] Input: Request data sent to the server

[1198] Output: Saved database entries and generated interaction profiles

[1199] Specific operation: Use an "INSERT" SQL query to save information to the database, create a prompt statement for profile generation, and send it to the generating AI model.

[1200] Step 4: Synchronize conversation profiles

[1201] Terminal: Downloads the dialogue profile sent from the server and synchronizes it with the dialogue robot.

[1202] Input: Interaction profile sent from the server

[1203] Output: Dialogue profile synchronized with the conversational robot

[1204] Specific operation: Download the profile using an HTTP GET request and save it to the conversational robot's memory.

[1205] Step 5: Start the conversation

[1206] User: The parent gives voice commands to the conversational robot, saying, "Pet robot, play with Taro."

[1207] Input: Parental voice commands

[1208] Output: Trigger signal to start interaction

[1209] Specific operation: A microphone system for receiving voice commands recognizes the command and sends a trigger to the robot system to start the dialogue.

[1210] Step 6: Select and implement a dialogue scenario

[1211] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and starts the conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[1212] Input: Interaction Profile

[1213] Output: Selected dialogue scenario and its implementation

[1214] Specific operation: The robot selects the optimal scenario from the dialogue profiles and plays the scenario aloud.

[1215] Step 7: Obtain and analyze user responses

[1216] Device: The system analyzes the child's response using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, if the child says "I like cars," the audio is converted into text.

[1217] Input: Child's voice response

[1218] Output: Analyzed data in text format

[1219] Specific operation: Audio data is sent to the API and returned as text data.

[1220] Step 8: Generating the next dialogue

[1221] Server: Based on the analyzed text data, it generates the next dialogue using an AI model. For example, it generates the following question: "Do you know what colors cars come in?"

[1222] Input: Text-formatted analysis data

[1223] Output: The following dialogue

[1224] Specific operation: A prompt message is sent to the generative AI model, and the generated dialogue content is retrieved.

[1225] Step 9: Provide content to facilitate learning

[1226] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools are used to select content. For example, it might teach, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1227] Input: Interaction profile and user response data

[1228] Output: Selected learning content

[1229] Specific operation: The robot selects content from a database that matches the child's interests and teaches them using voice.

[1230] Step 10: Data storage and feedback

[1231] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1232] Input: Collected data

[1233] Output: Feedback data sent to the server

[1234] Specific operation: The collected data is sent to the server via an HTTP POST request.

[1235] Step 11: Save feedback data and update the next interaction profile.

[1236] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[1237] Input: Feedback data sent to the server

[1238] Output: Updated interaction profile

[1239] Specific actions: Record feedback data in a database and use a generative AI model to update the next interaction profile.

[1240] (Application Example 1)

[1241] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1242] Traditional interactive systems were designed to support user learning and play, but they were difficult to apply to new employee training in factories and other similar settings. These systems struggled with providing complex operational instructions and training content tailored to skill levels, resulting in limited practicality. Furthermore, they lacked the ability to adapt training content in real time or analyze user instructions using speech recognition technology. This highlighted the challenge of improving the efficiency of training in new environments.

[1243] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1244] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, and means for providing training programs based on the user information. This makes it possible to provide personalized training content according to the user's skill level and learning progress.

[1245] "Means of registering user information" refers to interfaces or mechanisms for inputting basic information such as a user's name, age, job title, and skill level, and saving it in a database.

[1246] "Means for creating dialogue profiles" refers to a mechanism that generates individual dialogue scenarios based on registered user information and sets up dialogue content tailored to the user.

[1247] "Means for selecting dialogue scenarios" refers to a mechanism that selects a scenario that fits the created dialogue profile and facilitates effective dialogue with the user.

[1248] "Means for initiating and conducting a dialogue" refers to a mechanism for initiating a dialogue with a user based on a selected dialogue scenario and continuing the dialogue while providing timely responses.

[1249] "Means for acquiring and analyzing user responses" refers to a mechanism that analyzes the voice and actions obtained from the user during a conversation and dynamically adjusts the content of the next conversation based on that analysis.

[1250] "Means for selecting and providing content to promote learning" refers to a mechanism for selecting and providing educational content (e.g., vocabulary and operating procedures) that is appropriate for the user's age, interests, and skill level.

[1251] "Means of saving data collected during a conversation and applying feedback to the next conversation" refers to a mechanism that records the user's responses and learning progress obtained during a conversation and reflects them in subsequent conversations.

[1252] "Means of providing training programs" refers to a mechanism for generating and providing suitable training content based on user information.

[1253] "A means of analyzing user instructions using a speech recognition engine and generating the next dialogue content" refers to a mechanism that uses speech recognition technology to convert user statements into text and generates the next response based on that text.

[1254] "Means for providing instructions and training content using voice output means" refers to a mechanism for providing text-based instructions and training content to the user as audio.

[1255] To implement this invention, a server, terminal, and user must work together to build a dialogue system. The system of this invention consists of the following steps: user information registration, dialogue profile creation, training program provision, dialogue initiation and progress, user response acquisition and analysis, learning promotion content provision, data storage and feedback. The details are described below.

[1256] User Information Registration

[1257] Terminal: Users enter their information (name, job title, skill level, etc.) using a smartphone app or head-mounted display (HMD). This information is sent to the server and stored in the database.

[1258] Server: Based on the received information, it generates individual interaction profiles. These interaction profiles are customized according to the user's skill level and role.

[1259] Training program provision

[1260] Server: Based on user information, generates an appropriate training scenario and sends it to the terminal.

[1261] Terminal: Displays and provides the received training scenario to the user. For example, if a new employee says, "Teach me how to operate this machine," the terminal will display instructions from the server in both audio and display format.

[1262] Initiating and facilitating the dialogue

[1263] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[1264] Terminal: Uses a speech recognition engine to analyze the user's words and sends them to the server.

[1265] Server: Based on the analysis results, it generates the next dialogue content and sends it to the terminal. The dialogue content is dynamically adjusted in real time.

[1266] Acquiring and analyzing user responses

[1267] Terminal: Analyzes the user's voice input in real time and sends the results to the server.

[1268] Server: Based on the received information, it generates the next dialogue content and sends it to the terminal. This provides personalized training content.

[1269] Providing learning-enhancing content

[1270] Server: Select appropriate learning content (e.g., machine operation procedures, troubleshooting methods, etc.) according to the user's skill level and role.

[1271] Terminal: Provides selected content using an audio output device. For example, it might instruct, "First, press the power button. Then, press the start button to operate the machine."

[1272] Data storage and feedback

[1273] Terminal: Sends data collected during the interaction to the server.

[1274] Server: Receives data and saves it to the database, applying the feedback to the next interaction. This makes subsequent interactions more tailored to the user.

[1275] Examples of specific cases and prompt statements

[1276] As a concrete example, let's consider a scenario where a new employee, "Yamada," learns the basic operation of a new machine. In this case, the following prompt messages are used.

[1277] Example of a prompt

[1278] User: Yamada

[1279] Position: Operator

[1280] Skill level: beginner

[1281] User: Please tell me how to operate this machine.

[1282] Robot: "I'll teach you the basic operation of the machine. First, press the power button. Then, press the start button to start the machine."

[1283] In this way, the system of the present invention can utilize user information to provide personalized conversations and training content to each user.

[1284] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1285] Step 1: Register User Information

[1286] Terminal: Users (new recruits and employees) use a smartphone app or head-mounted display (HMD) to enter basic information such as name, job title, and skill level. This information is sent from the terminal to the server.

[1287] Input: User's basic information (name, job title, skill level)

[1288] Output: User information sent to the server

[1289] Step 2: Create an interaction profile

[1290] Server: Based on the received user information, it generates an individual dialogue profile. The dialogue profile includes dialogue scenarios based on the user's skill level and job title.

[1291] Input: User information

[1292] Output: Individual interaction profiles

[1293] Step 3: Providing the training program

[1294] Server: Based on the interaction profile, generates an appropriate training scenario and sends it to the terminal.

[1295] Input: Interaction Profile

[1296] Output: Training scenario

[1297] Step 4: Start the conversation

[1298] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[1299] Input: User's voice instructions

[1300] Output: Startup of speech recognition engine

[1301] Step 5: Voice Analysis

[1302] Terminal: Uses a speech recognition engine to convert user voice commands into text data and send it to the server.

[1303] Input: User's voice instructions

[1304] Output: Text data

[1305] Step 6: Generating dialogue content

[1306] Server: Based on the received text data, it generates the next dialogue content and sends it to the terminal.

[1307] Input: Text data, dialogue profile

[1308] Output: Dialogue content (training content)

[1309] Step 7: Instructions via voice and display

[1310] Terminal: Provides the user with the content of the conversation received from the server via voice output and display. For example, it might instruct the user, "First, press the power button. Then, press the start button to operate the machine."

[1311] Input: Dialogue content

[1312] Output: Audio and display

[1313] Step 8: Obtain and analyze user responses

[1314] Terminal: Analyzes user voice input in real time and sends the results to the server. For example, if a user asks, "Where is the power button?", the device analyzes the voice and converts it into text data.

[1315] Input: User voice input

[1316] Output: Analyzed text data

[1317] Step 9: Generating Feedback

[1318] Server: Based on the analyzed text data, it generates the next dialogue and sends it to the terminal. This provides personalized training content in real time.

[1319] Input: Analyzed text data, dialogue profile

[1320] Output: The following dialogue

[1321] Step 10: Data storage and feedback

[1322] Terminal: Sends data collected during the interaction to the server.

[1323] Server: Receives data and saves it to a database, which is then used in subsequent interactions. This data includes user responses, learning progress, and other information.

[1324] Input: Data collected during the conversation

[1325] Output: Saved to the database, feedback that will be reflected in the next interaction.

[1326] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1327] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles based on that information, and provides dialogue scenarios. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized dialogue and learning content.

[1328] User Information Registration

[1329] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1330] Terminal: Send this information to the server.

[1331] Server: Stores received information in a database and generates an interaction profile based on it.

[1332] Embedding an emotion engine

[1333] Device: The pet robot will be equipped with an emotion engine that recognizes emotions from the user's voice tone, facial expressions, and movements.

[1334] Server: Implement an algorithm that analyzes emotional data and identifies the user's current emotional state.

[1335] Creating and adapting dialogue profiles

[1336] Server: Generates dialogue profiles based on user information and emotional data. These dialogue profiles include personalized dialogue scenarios tailored to the user's age, interests, learning progress, and emotional state.

[1337] Terminal: Synchronizes conversation profiles with pet robots, allowing the system to provide optimal interactions for each user.

[1338] Initiating and facilitating the dialogue

[1339] User: The parent gives instructions via voice commands or the app, such as "Play with pet robot Taro."

[1340] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1341] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[1342] Acquiring and analyzing user responses

[1343] User: Taro replies, "Car."

[1344] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions.

[1345] Server: Analyzes response content and sentiment data to generate the next appropriate dialogue.

[1346] Providing content to promote learning

[1347] Server: Selects learning-enhancing content (English vocabulary, place names, etc.) based on the child's age, interests, and emotional state.

[1348] Terminal: The pet robot generates an encouraging response, saying, "The English word for red is 'Red.' Let's say it together, 'Red.' Sounds fun!"

[1349] Data storage and feedback

[1350] Terminal: Sends data collected during the conversation (Taro's responses, learning progress, emotional state, etc.) to the server.

[1351] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[1352] Terminal: Update and prepare a new interaction profile for the next session.

[1353] Specific example

[1354] For example, if a 3-year-old child named Taro is registered, and along with information that he likes cars, the emotion engine analyzes Taro's emotional state as "very interested." Based on this information, the pet robot asks, "Do you know what colors cars come in?" If Taro replies "red," and the emotion engine analyzes Taro's joy and excitement, the server generates a response such as, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for the content of the next conversation.

[1355] In this way, the system utilizes user information and emotional data to provide personalized dialogue and learning content, and can also apply feedback tailored to the user's emotional state.

[1356] The following describes the processing flow.

[1357] Step 1:

[1358] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1359] Step 2:

[1360] Terminal: Sends the entered user information to the system server.

[1361] Step 3:

[1362] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[1363] Step 4:

[1364] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[1365] Step 5:

[1366] Terminal: The pet robot activates its built-in emotion engine to recognize the user's emotional state in real time based on their voice tone, facial expressions, and movements.

[1367] Step 6:

[1368] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[1369] Step 7:

[1370] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1371] Step 8:

[1372] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[1373] Step 9:

[1374] User: Taro replies, "I like cars."

[1375] Step 10:

[1376] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions (excitement, joy, etc.).

[1377] Step 11:

[1378] Server: Based on the response content and sentiment data, it generates the next appropriate dialogue.

[1379] Step 12:

[1380] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[1381] Step 13:

[1382] User: Taro replies "Red".

[1383] Step 14:

[1384] Terminal: Taro's response is analyzed by a speech recognition engine, and along with the content of the response, an emotion engine is used to analyze his emotion of joy.

[1385] Step 15:

[1386] Server: Based on this response and sentiment data, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[1387] Step 16:

[1388] Device: Provides selected learning content and says, "The English word for red is 'Red.' Let's say it together, 'Red.' Isn't that great!"

[1389] Step 17:

[1390] User: Taro repeats "Red".

[1391] Step 18:

[1392] Terminal: Recognizes Taro's pronunciation, records its accuracy and learning effectiveness, and simultaneously analyzes Taro's sense of accomplishment using an emotion engine.

[1393] Step 19:

[1394] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[1395] Step 20:

[1396] Server: Stores data collected during the interaction (Taro's responses, learning progress, emotional state, etc.) in a database.

[1397] Step 21:

[1398] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[1399] (Example 2)

[1400] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1401] Traditional dialogue systems have faced challenges in providing personalized dialogue and learning content due to a lack of individual dialogue profiles based on user information and insufficient emotion recognition. In particular, the lack of functionality to understand the user's emotional state and dynamically adjust dialogue content accordingly makes it difficult to improve the user experience and promote effective learning.

[1402] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1403] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, means for selecting dialogue scenarios based on the dialogue profiles, means for initiating and progressing a dialogue based on the selected dialogue scenario, means for acquiring and analyzing user responses during the dialogue, means for selecting and providing content to promote learning, means for saving data collected during the dialogue and applying the feedback to the next dialogue, means for analyzing emotional data using an emotion engine that recognizes the user's emotions, and means for dynamically adjusting the dialogue content based on the emotional data. This makes it possible to provide personalized dialogues and learning content that correspond to the user's emotional state.

[1404] "Means of registering user information" refers to means that include an interface for users to input their basic information (name, age, gender, interests, learning progress, etc.) and send it to the system.

[1405] "Means for creating individual dialogue profiles based on registered information" refers to a method for storing user-registered information in a database and generating optimal dialogue content for each user based on that information.

[1406] "Means for selecting dialogue scenarios based on dialogue profiles" refers to a means that has the function of automatically selecting an appropriate dialogue scenario from a large number of dialogue scenarios based on the created dialogue profile.

[1407] "Means for initiating and conducting a dialogue based on a selected dialogue scenario" refers to means for initiating a dialogue with the user according to a selected scenario and for smoothly conducting that dialogue.

[1408] "Means for acquiring and analyzing user responses during a conversation" refers to methods for collecting voice and text data obtained from users in real time during a conversation and analyzing that data to understand their responses.

[1409] "Means of selecting and providing content to promote learning" refers to means of selecting appropriate educational content based on the user's age, interests, and learning progress, and providing it to the user.

[1410] "A means of saving data collected during a conversation and applying the feedback to the next conversation" refers to a means of saving data obtained from the user during a conversation and applying that data as feedback for use in the next conversation.

[1411] "Means for analyzing emotional data using an emotion engine that recognizes user emotions" refers to a method for recognizing emotions from a user's voice tone, facial expressions, and actions using an emotion engine, and then analyzing that recognized data.

[1412] "Means for dynamically adjusting dialogue content based on emotional data" refers to means for adjusting and optimizing the content and progress of dialogue in real time based on analyzed emotional data.

[1413] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles, and provides dialogue scenarios based on emotional data. This system is implemented using the following hardware and software.

[1414] Hardware and software

[1415] Devices: Smartphones and tablets (including pet robots)

[1416] Servers: Cloud servers and database systems

[1417] Emotion engine: EmotionAI engine

[1418] Speech recognition engine: VoiceRec

[1419] System Overview

[1420] The system registers user information and creates an individual dialogue profile based on it. Based on the dialogue profile, the optimal dialogue scenario is selected, and the dialogue content is dynamically adjusted using an emotion engine that analyzes the user's emotional state. Appropriate content for learning is provided based on emotion and response data. Furthermore, data collected during the dialogue is used as feedback for subsequent dialogues.

[1421] User information registration and interaction profile creation

[1422] Users (e.g., parents) enter basic information about their children (name, age, gender, interests, learning progress, etc.) using a smartphone app or web interface. This information is sent to the server via the device's "information transmission module." The server receives this information and stores it in its user database. Based on the registered information, it then uses a dialogue profile generation module to create a dialogue profile.

[1423] Emotion engine integration and analysis

[1424] The terminal (pet robot) is equipped with an EmotionAI engine that recognizes emotions from the user's voice tone, facial expressions, and movements. The server analyzes the received emotion data and uses an emotion analysis algorithm to identify the user's emotional state.

[1425] Facilitating and adjusting the content of the dialogue

[1426] The user gives instructions via voice commands or a dedicated app, such as "Play with pet robot Taro." The device recognizes the voice command using the VoiceRec voice recognition engine and selects an appropriate dialogue scenario using the dialogue scenario selection module. The server receives the dialogue log in real time and dynamically adjusts the dialogue content.

[1427] Acquiring and analyzing user responses

[1428] During the conversation, when the user (i.e., the child) responds, the device analyzes the response using its VoiceRec speech recognition engine and analyzes the emotion using its EmotionAI engine. Based on this data, the server generates the next conversation and selects appropriate learning content.

[1429] Learning promotion and content provision

[1430] The server uses a "learning content selection module" to select educational content (such as words or place names in different languages) based on the child's age, interests, and emotional state to facilitate learning. The device uses a "dialogue generation module" and a "speech synthesis engine" to enable the pet robot to generate personalized responses.

[1431] Data storage and feedback

[1432] The device sends data collected during the interaction (responses, learning progress, emotional state, etc.) to the server. The server stores this data in an interaction log database and uses a feedback generation module to generate feedback to be reflected in the next interaction. The device updates its new interaction profile and prepares for the next session.

[1433] Specific example

[1434] For example, if a 3-year-old child named Taro is registered and the system determines that he likes cars and that the emotion engine is "excited," the pet robot might ask, "Do you know what colors cars come in?" If Taro replies "red" and the EmotionAI engine analyzes Taro's excitement, the server will generate a response like, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for future interactions.

[1435] Example of a prompt

[1436] An example of a prompt to input into the generative AI model is: "Based on user information and sentiment data, generate a conversational scenario about cars that a 3-year-old child would be interested in. Also, provide an appropriate response if the child is excited."

[1437] This system is expected to improve the user experience by providing personalized conversational and learning content.

[1438] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1439] Step 1:

[1440] User Information Registration

[1441] Users enter basic information about their child (name, age, gender, interests, learning progress, etc.) through a smartphone app or web interface.

[1442] Input: Basic information about the child

[1443] Data processing: Enter basic information into the input fields and click the submit button.

[1444] Output: Send input information

[1445] The terminal sends the information entered by the user to the server using an "information transmission module".

[1446] Step 2:

[1447] Creating a dialogue profile

[1448] The server saves the received information to the "user database".

[1449] Input: Received user information

[1450] Data processing: Saving operations to the database

[1451] Output: Base data for dialogue profile generation

[1452] The server generates an interaction profile using the "Interaction Profile Generation Module" based on the registration information. The interaction profile includes the user's basic information and interests.

[1453] Step 3:

[1454] Embedding an emotion engine

[1455] The device (pet robot) is equipped with an "EmotionAI engine" that recognizes emotions from the user's voice tone, facial expressions, and movements.

[1456] Input: User's voice tone, facial expressions, and gestures

[1457] Data processing: Audio analysis, image analysis, motion analysis

[1458] Output: Sentiment data

[1459] The server analyzes the received emotional data using an "emotion analysis algorithm" to identify the user's emotional state.

[1460] Step 4:

[1461] Initiating and facilitating the dialogue

[1462] Users can give commands via voice commands or a dedicated app, such as "Play with my pet robot, Taro."

[1463] Input: Parental voice commands

[1464] Data processing: Speech recognition and analysis

[1465] Output: Analysis result (instruction to "play with Taro")

[1466] The device uses the "VoiceRec" voice recognition engine to recognize voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1467] The server receives dialogue logs in real time and dynamically adjusts the dialogue content using a "dialogue adjustment module".

[1468] Step 5:

[1469] Acquiring and analyzing user responses

[1470] The user (Taro) responds to the pet robot's question with "car".

[1471] Input: Taro's voice response

[1472] Data processing: Speech recognition and sentiment analysis

[1473] Output: Analysis results (response content and sentiment data)

[1474] The device analyzes Taro's responses using the "VoiceRec" voice recognition engine, and simultaneously analyzes his emotions using the "EmotionAI engine."

[1475] The server uses a "dialogue content analysis module" to analyze the response content and sentiment data, and then generates the next appropriate dialogue.

[1476] Step 6:

[1477] Providing content to promote learning

[1478] The server selects learning-enhancing content based on the user's age, interests, and emotional state.

[1479] Input: User age, interests, and sentiment data

[1480] Data processing: Execution of content selection algorithm

[1481] Output: Selected learning content

[1482] The device uses a "dialogue generation module" and a "speech synthesis engine" to generate learning content and motivating responses. For example, it might respond, "The English word for red is 'Red.' Let's say it together, red. Sounds fun!"

[1483] Step 7:

[1484] Data storage and feedback

[1485] The device sends the data collected during the conversation to the server.

[1486] Input: Data collected during the interaction (responses, learning progress, emotional state, etc.)

[1487] Data processing: Data transmission processing

[1488] Output: Sent data

[1489] The server saves the data in the "interaction log database" and uses the "feedback generation module" to apply the feedback to the next interaction.

[1490] The device updates its conversation profile with a new one and prepares for the next session.

[1491] (Application Example 2)

[1492] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1493] Traditional food delivery services have struggled to provide personalized menu recommendations and promotions tailored to users' emotions and preferences. This hindered the optimization of the user experience and the effective delivery of services. In addition, the lack of mechanisms to respond immediately to changes in users' emotions led to decreased convenience and satisfaction.

[1494] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1495] In this invention, the server includes means for acquiring and analyzing user responses and emotional data, means for selecting and providing personalized recommended menus and promotions based on the user's emotional state, and means for storing data collected during the interaction and applying the feedback to the next interaction in the food delivery service. This enables personalized interactions and menu provision that are tailored to the user's emotions and preferences, thereby improving the user experience.

[1496] "User information" refers to basic personal information such as the user's name, age, interests, and learning progress.

[1497] A "dialogue profile" refers to a dataset containing personalized dialogue scenarios based on user registration information and sentiment data.

[1498] A "dialogue scenario" refers to a scenario that defines the progression and content of the interaction with the user, selected based on the dialogue profile.

[1499] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice tone, facial expressions, and actions.

[1500] "Recommended menus" refer to personalized menus in food delivery services that are selected based on the user's emotional state and interests.

[1501] "Promotion" refers to sales promotion activities such as special discounts and campaigns offered based on the user's emotional state and interests.

[1502] A "food delivery service" refers to a service that delivers meals ordered by users.

[1503] "Feedback" refers to information collected during a conversation that is used to improve the next conversation, and it is data that contributes to system improvement and enhanced personalization.

[1504] "User experience" refers to the overall satisfaction and convenience that users feel when using a service.

[1505] This invention relates to a food delivery service that provides personalized menu recommendations and promotions based on user information and sentiment data. Embodiments thereof are described below.

[1506] Hardware and software used

[1507] Face recognition camera (e.g., Logitech C920)

[1508] This is a camera used to capture the user's facial expressions.

[1509] Server (e.g., Amazon Web Services EC2 instance)

[1510] This server stores and analyzes user information, sentiment data, and dialogue profiles.

[1511] Emotion recognition libraries (e.g., DeepFace)

[1512] This is a library that analyzes emotional data from a user's facial expressions.

[1513] HTTP request library (e.g., request)

[1514] This is a library for communicating user information, sentiment data, and other data with the server.

[1515] Program processing details

[1516] In this invention, three elements—a server, a terminal, and a user—work together to process information.

[1517] 1. User information registration

[1518] The user uses a smartphone app to enter basic information such as their name, age, interests, and learning progress. The device sends this information to the server via an HTTP POST request, and the server stores the received information in a database. This registers the user's basic information.

[1519] 2. Emotion recognition

[1520] The device uses a facial recognition camera to capture the user's facial expressions and analyzes the emotional data using the DeepFace library. The analyzed emotional data is sent to the server using an HTTP POST request, and the server stores this information in a database. This allows the user's emotional state to be monitored in real time.

[1521] 3. Providing personalized recommended menus

[1522] The server analyzes the user's basic information and sentiment data, and generates personalized recommendations and promotions based on that information. The device retrieves this information from the server using an HTTP GET request and displays it to the user. This provides personalized services tailored to the user's sentiments and interests.

[1523] Specific example

[1524] For example, suppose a user opens a food delivery app and faces the camera. If the emotion recognition engine recognizes the user's emotion as "joy," the server uses that emotion data to generate recommended menus of sushi dishes or special promotions that match the user's interests. The device then displays these recommended menus to the user, who can then order a meal based on them.

[1525] Example of a prompt

[1526] The following are examples of prompts for a generative AI model.

[1527] Design a food delivery application that analyzes user emotions and provides food recommendations based on those emotions. Use DeepFace for emotion recognition and include specific user scenarios and dialogue scenarios.

[1528] The above is a specific description of the embodiment for carrying out the invention. This system enables the provision of highly personalized services that respond to the user's emotions and preferences, thereby improving the user experience.

[1529] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1530] Step 1: Register User Information

[1531] Users enter basic information such as their name, age, interests, and learning progress using a smartphone app. This information is sent from the device to the server via an HTTP POST request. The server stores the received information in a database and generates a conversational profile for each user. This process registers the user's basic information in the database.

[1532] Input: User information such as name, age, interests, and learning progress.

[1533] Data processing: Convert user information to JSON format and construct an HTTP POST request.

[1534] Output: User information stored on the server, and generated interaction profile.

[1535] Step 2: Emotion Recognition

[1536] The device uses a facial recognition camera to capture the user's facial expressions. The captured image data is analyzed using the DeepFace library to identify dominant emotions (e.g., "joy," "sadness," etc.). The analyzed emotion data is sent from the device to the server via an HTTP POST request, and the server stores it in a database.

[1537] Input: Captured user's face image

[1538] Data processing: Sentiment analysis using DeepFace

[1539] Output: Emotional data stored on the server

[1540] Step 3: Generating personalized recommendation menus

[1541] The server analyzes registered user information and sentiment data to generate recommended menus and promotions based on the user's current state. This is done by combining menu and promotion information that has been previously stored in the database.

[1542] Input: User information, sentiment data

[1543] Data processing: Data analysis and menu selection using AI algorithms.

[1544] Output: Personalized menu recommendations and promotional information

[1545] Step 4: Serve the recommended menu

[1546] The device uses an HTTP GET request to retrieve personalized menu recommendations and promotional information from the server. Based on this information, the app displays menus and promotions that are suitable for the user.

[1547] Input: User's request

[1548] Data processing: Retrieving and displaying menu information sent from the server.

[1549] Output: Recommended menus and promotions presented to the user.

[1550] Step 5: Obtain and analyze user feedback

[1551] The user reacts to the presented menu or promotion, and the device captures that reaction again. This data is sent to the server using an HTTP POST request, where it is analyzed and stored.

[1552] Input: User response data

[1553] Data processing: Analysis of emotions and intentions using DeepFace and speech recognition engines.

[1554] Output: User response data stored on the server

[1555] Step 6: Applying Feedback

[1556] The server uses the emotional data and user response data collected during the conversation to apply feedback to the next conversation. This allows the next conversation scenario to be tailored to be more personalized.

[1557] Input: Collected sentiment data and reaction data

[1558] Data processing: Adjusting the next dialogue scenario using past data.

[1559] Output: Updated dialogue profile and next dialogue scenario

[1560] The above outlines the processing flow of the system program that implements the application example, and each step includes specific actions.

[1561] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1562] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1563] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1564] [Fourth Embodiment]

[1565] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1566] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1567] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1568] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1569] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1570] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1571] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1572] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1573] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1574] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1575] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1576] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1577] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1578] This invention relates to an interactive pet robot system that registers user information, creates a dialogue profile based on that information, and provides individually adapted dialogue scenarios. This system is designed to support children's learning and play, especially for parents and seniors, when parents are unable to attend to their children.

[1579] User Information Registration

[1580] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1581] Terminal: Send this information to the server.

[1582] Server: Stores received information in a database and generates individual interaction profiles.

[1583] Creating and adapting dialogue profiles

[1584] Server: The conversation profile created based on user information reflects age, interests, and learning progress, and includes individual conversation scenarios.

[1585] Terminal: Synchronize conversation profiles with pet robots, allowing the system to provide appropriate interactions for each user.

[1586] Initiating and facilitating the dialogue

[1587] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[1588] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. It says, "Hi Taro. What would you like to do today?"

[1589] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[1590] Acquiring and analyzing user responses

[1591] User: The child responds.

[1592] Device: The device analyzes the child's response using a speech recognition engine to understand its meaning. For example, if the child says "I like cars," it retrieves that information.

[1593] Server: Based on the analyzed information, the server generates the following dialogue, and the pet robot responds, "Do you know what colors cars come in?"

[1594] Providing content to promote learning

[1595] Server: Selects learning-enhancing content (e.g., English vocabulary, place names, etc.) based on the child's age and interests.

[1596] Terminal: The pet robot teaches, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1597] User: The child keeps repeating "red."

[1598] Device: Recognizes repeated content and records accuracy and learning effectiveness.

[1599] Data storage and feedback

[1600] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1601] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[1602] Terminal: Update and prepare a new interaction profile for the next session.

[1603] Specific example

[1604] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then says "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[1605] In this way, the system can utilize user information to provide personalized conversations and learning content tailored to each user.

[1606] The following describes the processing flow.

[1607] Step 1:

[1608] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1609] Step 2:

[1610] Terminal: Sends the entered user information to the system server.

[1611] Step 3:

[1612] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[1613] Step 4:

[1614] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[1615] Step 5:

[1616] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[1617] Step 6:

[1618] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1619] Step 7:

[1620] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[1621] Step 8:

[1622] User: Taro replies, "Car."

[1623] Step 9:

[1624] Terminal: The speech recognition engine analyzes Taro's response to understand its content.

[1625] Step 10:

[1626] Server: Analyzes the response and generates the next appropriate dialogue.

[1627] Step 11:

[1628] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[1629] Step 12:

[1630] User: Taro replies "Red".

[1631] Step 13:

[1632] Terminal: The speech recognition engine analyzes Taro's response again to verify its accuracy.

[1633] Step 14:

[1634] Server: Based on this response, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[1635] Step 15:

[1636] Device: Provides selected learning content and says, "The English word for red is 'Red'. Let's say it together, red."

[1637] Step 16:

[1638] User: Taro repeats "Red".

[1639] Step 17:

[1640] Terminal: Recognizes Taro's pronunciation and records its accuracy and learning effectiveness.

[1641] Step 18:

[1642] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[1643] Step 19:

[1644] Server: Stores data collected during the interaction (Taro's responses, learning progress, etc.) in a database.

[1645] Step 20:

[1646] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[1647] (Example 1)

[1648] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1649] One of the problems faced by modern families, particularly by parents and seniors, is the difficulty in adequately supporting children's learning and play when parents are unable to attend to them directly. Traditional educational tools and interactive learning toys have limited effectiveness because they cannot provide individualized dialogue and learning content tailored to the child's age and interests. Furthermore, the technology to monitor learning progress in real time during dialogue and incorporate that information into subsequent dialogues has been insufficient.

[1650] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1651] In this invention, the server includes means for registering user information, means for storing user information in a database and generating individual dialogue profiles, and means for synchronizing the generated dialogue profiles with a dialogue robot. This makes it possible to provide individual dialogue scenarios tailored to the user's age and interests in real time, and to effectively support learning progress.

[1652] "Means for registering user information" refers to devices or software that allow users to input and register basic information about their child and parent using a smartphone app or a dedicated web interface.

[1653] "Means for sending registered information to the server" refers to a communication device or software that has the role of sending information entered by the user to the server.

[1654] "Means for storing received information in a database and generating individual dialogue profiles" refers to a device or software that stores user information received by a server in a database and generates individual dialogue profiles based on that information.

[1655] "Means for synchronizing the generated dialogue profile with the dialogue robot" refers to a communication device or software that transmits the generated dialogue profile to the dialogue robot's memory, causing the robot to operate based on that profile.

[1656] "Means of initiating dialogue based on parental instructions" refers to a device or software that initiates a dialogue session when a parent gives instructions to a dialogue robot via voice or other interface.

[1657] "Means for selecting and implementing an appropriate dialogue scenario based on a dialogue profile" refers to a device or software that selects the optimal dialogue scenario based on the content of the dialogue profile and proceeds with the dialogue according to that scenario.

[1658] "Means for analyzing and understanding the content of a user's response using a speech recognition engine" refers to a device or software that uses a speech recognition engine to convert a user's response into text and understand its content.

[1659] "Means for generating the next dialogue based on analyzed information" refers to a device or software that uses a generative AI model to generate the content of the next dialogue based on information analyzed by a speech recognition engine.

[1660] "Means for selecting and providing learning-promoting content based on a child's age and interests" refers to a device or software that selects appropriate learning content according to a child's specific age and interests and provides it to the user.

[1661] "Means of sending data collected during a dialogue to a server and feeding it back into the next dialogue" refers to a device or software that sends detailed data collected during a dialogue to a server and processes it to reflect in the next dialogue profile.

[1662] This invention relates to an interactive pet robot system that performs a series of processes including user information registration, dialogue profile generation, initiation and progression of dialogue, acquisition and analysis of user responses, provision of content to promote learning, and data storage and feedback. This system is designed particularly for parents and seniors to support children's learning and play when parents are unable to attend to them.

[1663] User Information Registration

[1664] User: Use a smartphone app or a dedicated web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter information such as "Taro, 3 years old, boy, interest: cars."

[1665] Terminal: Sends user-entered information to the server. This process uses communication methods such as HTTP POST requests.

[1666] Server: Stores received information in a database (e.g., MySQL or PostgreSQL) and generates individual interaction profiles. A generative AI model (e.g., OpenAI GPT-4) is used for profile generation.

[1667] Creating and adapting dialogue profiles

[1668] Server: Generates a conversation profile based on user information. The profile reflects age, interests, and learning progress, and includes individual conversation scenarios.

[1669] Terminal: Synchronizes dialogue profiles with the conversational robot, enabling the system to provide appropriate dialogues for each user. Specifically, it downloads dialogue profiles from the server to the terminal and saves them to the robot's memory.

[1670] Initiating and facilitating the dialogue

[1671] User: When the parent instructs the pet robot to "play with Taro," the pet robot will begin a conversation.

[1672] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and begins the actual conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[1673] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content. A NoSQL database (e.g., MongoDB) is used to store the logs.

[1674] Acquiring and analyzing user responses

[1675] User: The child responds. For example, the child says, "I like cars."

[1676] Device: The child's responses are analyzed using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) to understand the content.

[1677] Server: Based on the analyzed information, it generates the following dialogue. For example, it generates the following question: "Do you know what colors cars come in?"

[1678] Providing content to promote learning

[1679] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools (e.g., NLTK or SpaCy) are used for content selection.

[1680] Device: The pet robot provides children with selected learning content. For example, it might teach them, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1681] User: The child keeps repeating "red."

[1682] Device: Analyzes repeated content using speech recognition and records accuracy and learning effectiveness.

[1683] Data storage and feedback

[1684] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1685] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[1686] Terminal: Update and prepare a new interaction profile for the next session. Specifically, download the updated profile from the server and synchronize it with the robot.

[1687] Specific example

[1688] For example, a 3-year-old child named "Taro" is registered as a user, and his interest in cars is recorded. Based on this information, the pet robot speaks to Taro, saying, "You like cars, huh? Do you know what colors cars come in?" If Taro then mentions the color "red," the pet robot responds, "Red is called 'Red' in English. Let's say it together, red." This entire conversation is managed by the system and used as feedback for the next conversation.

[1689] Example of a prompt

[1690] The following is an example of a prompt statement for a generative AI model:

[1691] User information:

[1692] Name: Taro

[1693] Age: 3 years old

[1694] Gender: Boy

[1695] Interests: Cars

[1696] Start conversation:

[1697] User: Playing with my pet robot, Taro

[1698] Robot: Hi Taro. What do you want to do today?

[1699] User: I like cars

[1700] Robot: Do you know what colors cars come in?

[1701] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1702] Step 1: Enter user information

[1703] User: Use a smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.). For example, enter "Taro, 3 years old, boy, interest: cars" into the app's form.

[1704] Input: Basic information entered by the user

[1705] Output: Input basic information data

[1706] Step 2: Submit User Information

[1707] Terminal: Sends the entered user information to the server using an HTTP POST request.

[1708] Input: Basic information data entered by the user

[1709] Output: Request data sent to the server

[1710] Step 3: Saving user information and generating a profile

[1711] Server: The server stores the received information in a database (e.g., MySQL or PostgreSQL) and uses a generative AI model (e.g., OpenAI GPT-4) to generate individual interaction profiles. This information includes prompts for the generative AI model.

[1712] Input: Request data sent to the server

[1713] Output: Saved database entries and generated interaction profiles

[1714] Specific operation: Use an "INSERT" SQL query to save information to the database, create a prompt statement for profile generation, and send it to the generating AI model.

[1715] Step 4: Synchronize conversation profiles

[1716] Terminal: Downloads the dialogue profile sent from the server and synchronizes it with the dialogue robot.

[1717] Input: Interaction profile sent from the server

[1718] Output: Dialogue profile synchronized with the conversational robot

[1719] Specific operation: Download the profile using an HTTP GET request and save it to the conversational robot's memory.

[1720] Step 5: Start the conversation

[1721] User: The parent gives voice commands to the conversational robot, saying, "Pet robot, play with Taro."

[1722] Input: Parental voice commands

[1723] Output: Trigger signal to start interaction

[1724] Specific operation: A microphone system for receiving voice commands recognizes the command and sends a trigger to the robot system to start the dialogue.

[1725] Step 6: Select and implement a dialogue scenario

[1726] Terminal: Selects an appropriate dialogue scenario based on the dialogue profile and starts the conversation. For example, it might say, "Hi Taro. What do you want to do today?"

[1727] Input: Interaction Profile

[1728] Output: Selected dialogue scenario and its implementation

[1729] Specific operation: The robot selects the optimal scenario from the dialogue profiles and plays the scenario aloud.

[1730] Step 7: Obtain and analyze user responses

[1731] Device: The system analyzes the child's response using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) and converts it into text data. For example, if the child says "I like cars," the audio is converted into text.

[1732] Input: Child's voice response

[1733] Output: Analyzed data in text format

[1734] Specific operation: Audio data is sent to the API and returned as text data.

[1735] Step 8: Generating the next dialogue

[1736] Server: Based on the analyzed text data, it generates the next dialogue using an AI model. For example, it generates the following question: "Do you know what colors cars come in?"

[1737] Input: Text-formatted analysis data

[1738] Output: The following dialogue

[1739] Specific operation: A prompt message is sent to the generative AI model, and the generated dialogue content is retrieved.

[1740] Step 9: Provide content to facilitate learning

[1741] Server: Selects learning-enhancing content based on the child's age and interests. Natural language processing tools are used to select content. For example, it might teach, "Do you know what red is in English? It's 'Red'. Let's say it together. Red."

[1742] Input: Interaction profile and user response data

[1743] Output: Selected learning content

[1744] Specific operation: The robot selects content from a database that matches the child's interests and teaches them using voice.

[1745] Step 10: Data storage and feedback

[1746] Terminal: Sends data collected during the interaction (e.g., child's responses and learning progress) to the server.

[1747] Input: Collected data

[1748] Output: Feedback data sent to the server

[1749] Specific operation: The collected data is sent to the server via an HTTP POST request.

[1750] Step 11: Save feedback data and update the next interaction profile.

[1751] Server: Stores received data in a database (e.g., Amazon RDS) and generates feedback to be used in future interactions.

[1752] Input: Feedback data sent to the server

[1753] Output: Updated interaction profile

[1754] Specific actions: Record feedback data in a database and use a generative AI model to update the next interaction profile.

[1755] (Application Example 1)

[1756] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1757] Traditional interactive systems were designed to support user learning and play, but they were difficult to apply to new employee training in factories and other similar settings. These systems struggled with providing complex operational instructions and training content tailored to skill levels, resulting in limited practicality. Furthermore, they lacked the ability to adapt training content in real time or analyze user instructions using speech recognition technology. This highlighted the challenge of improving the efficiency of training in new environments.

[1758] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1759] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, and means for providing training programs based on the user information. This makes it possible to provide personalized training content according to the user's skill level and learning progress.

[1760] "Means of registering user information" refers to interfaces or mechanisms for inputting basic information such as a user's name, age, job title, and skill level, and saving it in a database.

[1761] "Means for creating dialogue profiles" refers to a mechanism that generates individual dialogue scenarios based on registered user information and sets up dialogue content tailored to the user.

[1762] "Means for selecting dialogue scenarios" refers to a mechanism that selects a scenario that fits the created dialogue profile and facilitates effective dialogue with the user.

[1763] "Means for initiating and conducting a dialogue" refers to a mechanism for initiating a dialogue with a user based on a selected dialogue scenario and continuing the dialogue while providing timely responses.

[1764] "Means for acquiring and analyzing user responses" refers to a mechanism that analyzes the voice and actions obtained from the user during a conversation and dynamically adjusts the content of the next conversation based on that analysis.

[1765] "Means for selecting and providing content to promote learning" refers to a mechanism for selecting and providing educational content (e.g., vocabulary and operating procedures) that is appropriate for the user's age, interests, and skill level.

[1766] "Means of saving data collected during a conversation and applying feedback to the next conversation" refers to a mechanism that records the user's responses and learning progress obtained during a conversation and reflects them in subsequent conversations.

[1767] "Means of providing training programs" refers to a mechanism for generating and providing suitable training content based on user information.

[1768] "A means of analyzing user instructions using a speech recognition engine and generating the next dialogue content" refers to a mechanism that uses speech recognition technology to convert user statements into text and generates the next response based on that text.

[1769] "Means for providing instructions and training content using voice output means" refers to a mechanism for providing text-based instructions and training content to the user as audio.

[1770] To implement this invention, a server, terminal, and user must work together to build a dialogue system. The system of this invention consists of the following steps: user information registration, dialogue profile creation, training program provision, dialogue initiation and progress, user response acquisition and analysis, learning promotion content provision, data storage and feedback. The details are described below.

[1771] User Information Registration

[1772] Terminal: Users enter their information (name, job title, skill level, etc.) using a smartphone app or head-mounted display (HMD). This information is sent to the server and stored in the database.

[1773] Server: Based on the received information, it generates individual interaction profiles. These interaction profiles are customized according to the user's skill level and role.

[1774] Training program provision

[1775] Server: Based on user information, generates an appropriate training scenario and sends it to the terminal.

[1776] Terminal: Displays and provides the received training scenario to the user. For example, if a new employee says, "Teach me how to operate this machine," the terminal will display instructions from the server in both audio and display format.

[1777] Initiating and facilitating the dialogue

[1778] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[1779] Terminal: Uses a speech recognition engine to analyze the user's words and sends them to the server.

[1780] Server: Based on the analysis results, it generates the next dialogue content and sends it to the terminal. The dialogue content is dynamically adjusted in real time.

[1781] Acquiring and analyzing user responses

[1782] Terminal: Analyzes the user's voice input in real time and sends the results to the server.

[1783] Server: Based on the received information, it generates the next dialogue content and sends it to the terminal. This provides personalized training content.

[1784] Providing learning-enhancing content

[1785] Server: Select appropriate learning content (e.g., machine operation procedures, troubleshooting methods, etc.) according to the user's skill level and role.

[1786] Terminal: Provides selected content using an audio output device. For example, it might instruct, "First, press the power button. Then, press the start button to operate the machine."

[1787] Data storage and feedback

[1788] Terminal: Sends data collected during the interaction to the server.

[1789] Server: Receives data and saves it to the database, applying the feedback to the next interaction. This makes subsequent interactions more tailored to the user.

[1790] Examples of specific cases and prompt statements

[1791] As a concrete example, let's consider a scenario where a new employee, "Yamada," learns the basic operation of a new machine. In this case, the following prompt messages are used.

[1792] Example of a prompt

[1793] User: Yamada

[1794] Position: Operator

[1795] Skill level: beginner

[1796] User: Please tell me how to operate this machine.

[1797] Robot: "I'll teach you the basic operation of the machine. First, press the power button. Then, press the start button to start the machine."

[1798] In this way, the system of the present invention can utilize user information to provide personalized conversations and training content to each user.

[1799] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1800] Step 1: Register User Information

[1801] Terminal: Users (new recruits and employees) use a smartphone app or head-mounted display (HMD) to enter basic information such as name, job title, and skill level. This information is sent from the terminal to the server.

[1802] Input: User's basic information (name, job title, skill level)

[1803] Output: User information sent to the server

[1804] Step 2: Create an interaction profile

[1805] Server: Based on the received user information, it generates an individual dialogue profile. The dialogue profile includes dialogue scenarios based on the user's skill level and job title.

[1806] Input: User information

[1807] Output: Individual interaction profiles

[1808] Step 3: Providing the training program

[1809] Server: Based on the interaction profile, generates an appropriate training scenario and sends it to the terminal.

[1810] Input: Interaction Profile

[1811] Output: Training scenario

[1812] Step 4: Start the conversation

[1813] User: The user initiates interaction by speaking to the device. For example, they might say, "Tell me how to operate this machine."

[1814] Input: User's voice instructions

[1815] Output: Startup of speech recognition engine

[1816] Step 5: Voice Analysis

[1817] Terminal: Uses a speech recognition engine to convert user voice commands into text data and send it to the server.

[1818] Input: User's voice instructions

[1819] Output: Text data

[1820] Step 6: Generating dialogue content

[1821] Server: Based on the received text data, it generates the next dialogue content and sends it to the terminal.

[1822] Input: Text data, dialogue profile

[1823] Output: Dialogue content (training content)

[1824] Step 7: Instructions via voice and display

[1825] Terminal: Provides the user with the content of the conversation received from the server via voice output and display. For example, it might instruct the user, "First, press the power button. Then, press the start button to operate the machine."

[1826] Input: Dialogue content

[1827] Output: Audio and display

[1828] Step 8: Obtain and analyze user responses

[1829] Terminal: Analyzes user voice input in real time and sends the results to the server. For example, if a user asks, "Where is the power button?", the device analyzes the voice and converts it into text data.

[1830] Input: User voice input

[1831] Output: Analyzed text data

[1832] Step 9: Generating Feedback

[1833] Server: Based on the analyzed text data, it generates the next dialogue and sends it to the terminal. This provides personalized training content in real time.

[1834] Input: Analyzed text data, dialogue profile

[1835] Output: The following dialogue

[1836] Step 10: Data storage and feedback

[1837] Terminal: Sends data collected during the interaction to the server.

[1838] Server: Receives data and saves it to a database, which is then used in subsequent interactions. This data includes user responses, learning progress, and other information.

[1839] Input: Data collected during the conversation

[1840] Output: Saved to the database, feedback that will be reflected in the next interaction.

[1841] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1842] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles based on that information, and provides dialogue scenarios. In particular, by combining it with an emotion engine that recognizes the user's emotions, it is possible to provide more personalized dialogue and learning content.

[1843] User Information Registration

[1844] User: Use a smartphone app or other interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1845] Terminal: Send this information to the server.

[1846] Server: Stores received information in a database and generates an interaction profile based on it.

[1847] Embedding an emotion engine

[1848] Device: The pet robot will be equipped with an emotion engine that recognizes emotions from the user's voice tone, facial expressions, and movements.

[1849] Server: Implement an algorithm that analyzes emotional data and identifies the user's current emotional state.

[1850] Creating and adapting dialogue profiles

[1851] Server: Generates dialogue profiles based on user information and emotional data. These dialogue profiles include personalized dialogue scenarios tailored to the user's age, interests, learning progress, and emotional state.

[1852] Terminal: Synchronizes conversation profiles with pet robots, allowing the system to provide optimal interactions for each user.

[1853] Initiating and facilitating the dialogue

[1854] User: The parent gives instructions via voice commands or the app, such as "Play with pet robot Taro."

[1855] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1856] Server: Receives dialogue logs in real time and dynamically adjusts the dialogue content.

[1857] Acquiring and analyzing user responses

[1858] User: Taro replies, "Car."

[1859] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions.

[1860] Server: Analyzes response content and sentiment data to generate the next appropriate dialogue.

[1861] Providing content to promote learning

[1862] Server: Selects learning-enhancing content (English vocabulary, place names, etc.) based on the child's age, interests, and emotional state.

[1863] Terminal: The pet robot generates an encouraging response, saying, "The English word for red is 'Red.' Let's say it together, 'Red.' Sounds fun!"

[1864] Data storage and feedback

[1865] Terminal: Sends data collected during the conversation (Taro's responses, learning progress, emotional state, etc.) to the server.

[1866] Server: Stores received data in the database and generates feedback to be used in the next interaction.

[1867] Terminal: Update and prepare a new interaction profile for the next session.

[1868] Specific example

[1869] For example, if a 3-year-old child named Taro is registered, and along with information that he likes cars, the emotion engine analyzes Taro's emotional state as "very interested." Based on this information, the pet robot asks, "Do you know what colors cars come in?" If Taro replies "red," and the emotion engine analyzes Taro's joy and excitement, the server generates a response such as, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for the content of the next conversation.

[1870] In this way, the system utilizes user information and emotional data to provide personalized dialogue and learning content, and can also apply feedback tailored to the user's emotional state.

[1871] The following describes the processing flow.

[1872] Step 1:

[1873] User: Use the smartphone app or web interface to enter basic information about the child and parent (name, age, gender, interests, learning progress, etc.).

[1874] Step 2:

[1875] Terminal: Sends the entered user information to the system server.

[1876] Step 3:

[1877] Server: Stores received user information in a database and generates individual interaction profiles based on that information.

[1878] Step 4:

[1879] Server: Synchronizes the generated dialogue profile to the pet robot terminal.

[1880] Step 5:

[1881] Terminal: The pet robot activates its built-in emotion engine to recognize the user's emotional state in real time based on their voice tone, facial expressions, and movements.

[1882] Step 6:

[1883] User: The parent gives commands via voice commands or the app, such as "Play with pet robot Taro."

[1884] Step 7:

[1885] Terminal: Recognizes voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1886] Step 8:

[1887] Terminal: Initiates a conversation and says, "Hi, Taro. What do you want to do today?"

[1888] Step 9:

[1889] User: Taro replies, "I like cars."

[1890] Step 10:

[1891] Terminal: The speech recognition engine analyzes Taro's responses, while the emotion engine simultaneously analyzes Taro's emotions (excitement, joy, etc.).

[1892] Step 11:

[1893] Server: Based on the response content and sentiment data, it generates the next appropriate dialogue.

[1894] Step 12:

[1895] Terminal: Based on the generated dialogue, it says, "Do you know what colors cars come in?"

[1896] Step 13:

[1897] User: Taro replies "Red".

[1898] Step 14:

[1899] Terminal: Taro's response is analyzed by a speech recognition engine, and along with the content of the response, an emotion engine is used to analyze his emotion of joy.

[1900] Step 15:

[1901] Server: Based on this response and sentiment data, select learning-enhancing content suitable for Taro (e.g., the English word "Red").

[1902] Step 16:

[1903] Device: Provides selected learning content and says, "The English word for red is 'Red.' Let's say it together, 'Red.' Isn't that great!"

[1904] Step 17:

[1905] User: Taro repeats "Red".

[1906] Step 18:

[1907] Terminal: Recognizes Taro's pronunciation, records its accuracy and learning effectiveness, and simultaneously analyzes Taro's sense of accomplishment using an emotion engine.

[1908] Step 19:

[1909] Terminal: Declares the end of the conversation and says, "That's all for today. Let's play again soon!"

[1910] Step 20:

[1911] Server: Stores data collected during the interaction (Taro's responses, learning progress, emotional state, etc.) in a database.

[1912] Step 21:

[1913] Server: Based on past dialogue data and user information, it generates feedback for the next dialogue and updates the dialogue profile.

[1914] (Example 2)

[1915] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1916] Traditional dialogue systems have faced challenges in providing personalized dialogue and learning content due to a lack of individual dialogue profiles based on user information and insufficient emotion recognition. In particular, the lack of functionality to understand the user's emotional state and dynamically adjust dialogue content accordingly makes it difficult to improve the user experience and promote effective learning.

[1917] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1918] In this invention, the server includes means for registering user information, means for creating individual dialogue profiles based on the registered information, means for selecting dialogue scenarios based on the dialogue profiles, means for initiating and progressing a dialogue based on the selected dialogue scenario, means for acquiring and analyzing user responses during the dialogue, means for selecting and providing content to promote learning, means for saving data collected during the dialogue and applying the feedback to the next dialogue, means for analyzing emotional data using an emotion engine that recognizes the user's emotions, and means for dynamically adjusting the dialogue content based on the emotional data. This makes it possible to provide personalized dialogues and learning content that correspond to the user's emotional state.

[1919] "Means of registering user information" refers to means that include an interface for users to input their basic information (name, age, gender, interests, learning progress, etc.) and send it to the system.

[1920] "Means for creating individual dialogue profiles based on registered information" refers to a method for storing user-registered information in a database and generating optimal dialogue content for each user based on that information.

[1921] "Means for selecting dialogue scenarios based on dialogue profiles" refers to a means that has the function of automatically selecting an appropriate dialogue scenario from a large number of dialogue scenarios based on the created dialogue profile.

[1922] "Means for initiating and conducting a dialogue based on a selected dialogue scenario" refers to means for initiating a dialogue with the user according to a selected scenario and for smoothly conducting that dialogue.

[1923] "Means for acquiring and analyzing user responses during a conversation" refers to methods for collecting voice and text data obtained from users in real time during a conversation and analyzing that data to understand their responses.

[1924] "Means of selecting and providing content to promote learning" refers to means of selecting appropriate educational content based on the user's age, interests, and learning progress, and providing it to the user.

[1925] "A means of saving data collected during a conversation and applying the feedback to the next conversation" refers to a means of saving data obtained from the user during a conversation and applying that data as feedback for use in the next conversation.

[1926] "Means for analyzing emotional data using an emotion engine that recognizes user emotions" refers to a method for recognizing emotions from a user's voice tone, facial expressions, and actions using an emotion engine, and then analyzing that recognized data.

[1927] "Means for dynamically adjusting dialogue content based on emotional data" refers to means for adjusting and optimizing the content and progress of dialogue in real time based on analyzed emotional data.

[1928] This invention relates to an interactive pet robot system that registers user information, creates individual dialogue profiles, and provides dialogue scenarios based on emotional data. This system is implemented using the following hardware and software.

[1929] Hardware and software

[1930] Devices: Smartphones and tablets (including pet robots)

[1931] Servers: Cloud servers and database systems

[1932] Emotion engine: EmotionAI engine

[1933] Speech recognition engine: VoiceRec

[1934] System Overview

[1935] The system registers user information and creates an individual dialogue profile based on it. Based on the dialogue profile, the optimal dialogue scenario is selected, and the dialogue content is dynamically adjusted using an emotion engine that analyzes the user's emotional state. Appropriate content for learning is provided based on emotion and response data. Furthermore, data collected during the dialogue is used as feedback for subsequent dialogues.

[1936] User information registration and interaction profile creation

[1937] Users (e.g., parents) enter basic information about their children (name, age, gender, interests, learning progress, etc.) using a smartphone app or web interface. This information is sent to the server via the device's "information transmission module." The server receives this information and stores it in its user database. Based on the registered information, it then uses a dialogue profile generation module to create a dialogue profile.

[1938] Emotion engine integration and analysis

[1939] The terminal (pet robot) is equipped with an EmotionAI engine that recognizes emotions from the user's voice tone, facial expressions, and movements. The server analyzes the received emotion data and uses an emotion analysis algorithm to identify the user's emotional state.

[1940] Facilitating and adjusting the content of the dialogue

[1941] The user gives instructions via voice commands or a dedicated app, such as "Play with pet robot Taro." The device recognizes the voice command using the VoiceRec voice recognition engine and selects an appropriate dialogue scenario using the dialogue scenario selection module. The server receives the dialogue log in real time and dynamically adjusts the dialogue content.

[1942] Acquiring and analyzing user responses

[1943] During the conversation, when the user (i.e., the child) responds, the device analyzes the response using its VoiceRec speech recognition engine and analyzes the emotion using its EmotionAI engine. Based on this data, the server generates the next conversation and selects appropriate learning content.

[1944] Learning promotion and content provision

[1945] The server uses a "learning content selection module" to select educational content (such as words or place names in different languages) based on the child's age, interests, and emotional state to facilitate learning. The device uses a "dialogue generation module" and a "speech synthesis engine" to enable the pet robot to generate personalized responses.

[1946] Data storage and feedback

[1947] The device sends data collected during the interaction (responses, learning progress, emotional state, etc.) to the server. The server stores this data in an interaction log database and uses a feedback generation module to generate feedback to be reflected in the next interaction. The device updates its new interaction profile and prepares for the next session.

[1948] Specific example

[1949] For example, if a 3-year-old child named Taro is registered and the system determines that he likes cars and that the emotion engine is "excited," the pet robot might ask, "Do you know what colors cars come in?" If Taro replies "red" and the EmotionAI engine analyzes Taro's excitement, the server will generate a response like, "Red is called 'Red' in English. Let's say it together, it's fun to learn!" This entire conversational flow is managed by the system and used as feedback for future interactions.

[1950] Example of a prompt

[1951] An example of a prompt to input into the generative AI model is: "Based on user information and sentiment data, generate a conversational scenario about cars that a 3-year-old child would be interested in. Also, provide an appropriate response if the child is excited."

[1952] This system is expected to improve the user experience by providing personalized conversational and learning content.

[1953] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1954] Step 1:

[1955] User Information Registration

[1956] Users enter basic information about their child (name, age, gender, interests, learning progress, etc.) through a smartphone app or web interface.

[1957] Input: Basic information about the child

[1958] Data processing: Enter basic information into the input fields and click the submit button.

[1959] Output: Send input information

[1960] The terminal sends the information entered by the user to the server using an "information transmission module".

[1961] Step 2:

[1962] Creating a dialogue profile

[1963] The server saves the received information to the "user database".

[1964] Input: Received user information

[1965] Data processing: Saving operations to the database

[1966] Output: Base data for dialogue profile generation

[1967] The server generates an interaction profile using the "Interaction Profile Generation Module" based on the registration information. The interaction profile includes the user's basic information and interests.

[1968] Step 3:

[1969] Embedding an emotion engine

[1970] The device (pet robot) is equipped with an "EmotionAI engine" that recognizes emotions from the user's voice tone, facial expressions, and movements.

[1971] Input: User's voice tone, facial expressions, and gestures

[1972] Data processing: Audio analysis, image analysis, motion analysis

[1973] Output: Sentiment data

[1974] The server analyzes the received emotional data using an "emotion analysis algorithm" to identify the user's emotional state.

[1975] Step 4:

[1976] Initiating and facilitating the dialogue

[1977] Users can give commands via voice commands or a dedicated app, such as "Play with my pet robot, Taro."

[1978] Input: Parental voice commands

[1979] Data processing: Speech recognition and analysis

[1980] Output: Analysis result (instruction to "play with Taro")

[1981] The device uses the "VoiceRec" voice recognition engine to recognize voice commands and selects an appropriate dialogue scenario based on Taro's registered dialogue profile.

[1982] The server receives dialogue logs in real time and dynamically adjusts the dialogue content using a "dialogue adjustment module".

[1983] Step 5:

[1984] Acquiring and analyzing user responses

[1985] The user (Taro) responds to the pet robot's question with "car".

[1986] Input: Taro's voice response

[1987] Data processing: Speech recognition and sentiment analysis

[1988] Output: Analysis results (response content and sentiment data)

[1989] The device analyzes Taro's responses using the "VoiceRec" voice recognition engine, and simultaneously analyzes his emotions using the "EmotionAI engine."

[1990] The server uses a "dialogue content analysis module" to analyze the response content and sentiment data, and then generates the next appropriate dialogue.

[1991] Step 6:

[1992] Providing content to promote learning

[1993] The server selects learning-enhancing content based on the user's age, interests, and emotional state.

[1994] Input: User age, interests, and sentiment data

[1995] Data processing: Execution of content selection algorithm

[1996] Output: Selected learning content

[1997] The device uses a "dialogue generation module" and a "speech synthesis engine" to generate learning content and motivating responses. For example, it might respond, "The English word for red is 'Red.' Let's say it together, red. Sounds fun!"

[1998] Step 7:

[1999] Data storage and feedback

[2000] The device sends the data collected during the conversation to the server.

[2001] Input: Data collected during the interaction (responses, learning progress, emotional state, etc.)

[2002] Data processing: Data transmission processing

[2003] Output: Sent data

[2004] The server saves the data in the "interaction log database" and uses the "feedback generation module" to apply the feedback to the next interaction.

[2005] The device updates its conversation profile with a new one and prepares for the next session.

[2006] (Application Example 2)

[2007] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2008] Traditional food delivery services have struggled to provide personalized menu recommendations and promotions tailored to users' emotions and preferences. This hindered the optimization of the user experience and the effective delivery of services. In addition, the lack of mechanisms to respond immediately to changes in users' emotions led to decreased convenience and satisfaction.

[2009] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[2010] In this invention, the server includes means for acquiring and analyzing user responses and emotional data, means for selecting and providing personalized recommended menus and promotions based on the user's emotional state, and means for storing data collected during the interaction and applying the feedback to the next interaction in the food delivery service. This enables personalized interactions and menu provision that are tailored to the user's emotions and preferences, thereby improving the user experience.

[2011] "User information" refers to basic personal information such as the user's name, age, interests, and learning progress.

[2012] A "dialogue profile" refers to a dataset containing personalized dialogue scenarios based on user registration information and sentiment data.

[2013] A "dialogue scenario" refers to a scenario that defines the progression and content of the interaction with the user, selected based on the dialogue profile.

[2014] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice tone, facial expressions, and actions.

[2015] "Recommended menus" refer to personalized menus in food delivery services that are selected based on the user's emotional state and interests.

[2016] "Promotion" refers to sales promotion activities such as special discounts and campaigns offered based on the user's emotional state and interests.

[2017] A "food delivery service" refers to a service that delivers meals ordered by users.

[2018] "Feedback" refers to information collected during a conversation that is used to improve the next conversation, and it is data that contributes to system improvement and enhanced personalization.

[2019] "User experience" refers to the overall satisfaction and convenience that users feel when using a service.

[2020] This invention relates to a food delivery service that provides personalized menu recommendations and promotions based on user information and sentiment data. Embodiments thereof are described below.

[2021] Hardware and software used

[2022] Face recognition camera (e.g., Logitech C920)

[2023] This is a camera used to capture the user's facial expressions.

[2024] Server (e.g., Amazon Web Services EC2 instance)

[2025] This server stores and analyzes user information, sentiment data, and dialogue profiles.

[2026] Emotion recognition libraries (e.g., DeepFace)

[2027] This is a library that analyzes emotional data from a user's facial expressions.

[2028] HTTP request library (e.g., request)

[2029] This is a library for communicating user information, sentiment data, and other data with the server.

[2030] Program processing details

[2031] In this invention, three elements—a server, a terminal, and a user—work together to process information.

[2032] 1. User information registration

[2033] The user uses a smartphone app to enter basic information such as their name, age, interests, and learning progress. The device sends this information to the server via an HTTP POST request, and the server stores the received information in a database. This registers the user's basic information.

[2034] 2. Emotion recognition

[2035] The device uses a facial recognition camera to capture the user's facial expressions and analyzes the emotional data using the DeepFace library. The analyzed emotional data is sent to the server using an HTTP POST request, and the server stores this information in a database. This allows the user's emotional state to be monitored in real time.

[2036] 3. Providing personalized recommended menus

[2037] The server analyzes the user's basic information and sentiment data, and generates personalized recommendations and promotions based on that information. The device retrieves this information from the server using an HTTP GET request and displays it to the user. This provides personalized services tailored to the user's sentiments and interests.

[2038] Specific example

[2039] For example, suppose a user opens a food delivery app and faces the camera. If the emotion recognition engine recognizes the user's emotion as "joy," the server uses that emotion data to generate recommended menus of sushi dishes or special promotions that match the user's interests. The device then displays these recommended menus to the user, who can then order a meal based on them.

[2040] Example of a prompt

[2041] The following are examples of prompts for a generative AI model.

[2042] Design a food delivery application that analyzes user emotions and provides food recommendations based on those emotions. Use DeepFace for emotion recognition and include specific user scenarios and dialogue scenarios.

[2043] The above is a specific description of the embodiment for carrying out the invention. This system enables the provision of highly personalized services that respond to the user's emotions and preferences, thereby improving the user experience.

[2044] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2045] Step 1: Register User Information

[2046] Users enter basic information such as their name, age, interests, and learning progress using a smartphone app. This information is sent from the device to the server via an HTTP POST request. The server stores the received information in a database and generates a conversational profile for each user. This process registers the user's basic information in the database.

[2047] Input: User information such as name, age, interests, and learning progress.

[2048] Data processing: Convert user information to JSON format and construct an HTTP POST request.

[2049] Output: User information stored on the server, and generated interaction profile.

[2050] Step 2: Emotion Recognition

[2051] The device uses a facial recognition camera to capture the user's facial expressions. The captured image data is analyzed using the DeepFace library to identify dominant emotions (e.g., "joy," "sadness," etc.). The analyzed emotion data is sent from the device to the server via an HTTP POST request, and the server stores it in a database.

[2052] Input: Captured user's face image

[2053] Data processing: Sentiment analysis using DeepFace

[2054] Output: Emotional data stored on the server

[2055] Step 3: Generating personalized recommendation menus

[2056] The server analyzes registered user information and sentiment data to generate recommended menus and promotions based on the user's current state. This is done by combining menu and promotion information that has been previously stored in the database.

[2057] Input: User information, sentiment data

[2058] Data processing: Data analysis and menu selection using AI algorithms.

[2059] Output: Personalized menu recommendations and promotional information

[2060] Step 4: Serve the recommended menu

[2061] The device uses an HTTP GET request to retrieve personalized menu recommendations and promotional information from the server. Based on this information, the app displays menus and promotions that are suitable for the user.

[2062] Input: User's request

[2063] Data processing: Retrieving and displaying menu information sent from the server.

[2064] Output: Recommended menus and promotions presented to the user.

[2065] Step 5: Obtain and analyze user feedback

[2066] The user reacts to the presented menu or promotion, and the device captures that reaction again. This data is sent to the server using an HTTP POST request, where it is analyzed and stored.

[2067] Input: User response data

[2068] Data processing: Analysis of emotions and intentions using DeepFace and speech recognition engines.

[2069] Output: User response data stored on the server

[2070] Step 6: Applying Feedback

[2071] The server uses the emotional data and user response data collected during the conversation to apply feedback to the next conversation. This allows the next conversation scenario to be tailored to be more personalized.

[2072] Input: Collected sentiment data and reaction data

[2073] Data processing: Adjusting the next dialogue scenario using past data.

[2074] Output: Updated dialogue profile and next dialogue scenario

[2075] The above outlines the processing flow of the system program that implements the application example, and each step includes specific actions.

[2076] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2077] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2078] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[2079] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2080] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[2081] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[2082] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[2083] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[2084] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[2085] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[2086] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[2087] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[2088] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[2089] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2090] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[2091] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[2092] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[2093] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[2094] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[2095] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[2096] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[2097] The following is further disclosed regarding the embodiments described above.

[2098] (Claim 1)

[2099] Means of registering user information,

[2100] A means of creating individual dialogue profiles based on registered information,

[2101] A means of selecting a dialogue scenario based on a dialogue profile,

[2102] A means of initiating and conducting a dialogue based on a selected dialogue scenario,

[2103] A means of obtaining and analyzing user responses during a conversation,

[2104] Means for selecting and providing content to promote learning,

[2105] A means of saving data collected during a conversation and applying the feedback to the next conversation,

[2106] A system that includes this.

[2107] (Claim 2)

[2108] The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, and learning progress.

[2109] (Claim 3)

[2110] The system according to claim 1, wherein the content for promoting learning includes words, place names, and other educational information in different languages.

[2111] "Example 1"

[2112] (Claim 1)

[2113] Means of registering user information,

[2114] A means of sending registered information to the server,

[2115] A means for storing received information in a database and generating individual dialogue profiles,

[2116] A means for synchronizing the generated dialogue profile with the dialogue robot,

[2117] Means of initiating a dialogue based on parental instructions,

[2118] A means of selecting and implementing an appropriate dialogue scenario based on a dialogue profile,

[2119] A means of analyzing the user's response using a speech recognition engine to understand its content,

[2120] A means of generating the next dialogue based on the analyzed information,

[2121] Means for selecting and providing learning-promoting content based on a child's age and interests,

[2122] A means of sending data collected during the conversation to a server and feeding it back into the next conversation,

[2123] A system that includes this.

[2124] (Claim 2)

[2125] The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, and learning progress.

[2126] (Claim 3)

[2127] The system according to claim 1, wherein the content for promoting learning includes words, place names, and other educational information in different languages.

[2128] "Application Example 1"

[2129] (Claim 1)

[2130] Means of registering user information,

[2131] A means of creating individual dialogue profiles based on registered information,

[2132] A means of selecting a dialogue scenario based on a dialogue profile,

[2133] A means of initiating and conducting a dialogue based on a selected dialogue scenario,

[2134] A means of obtaining and analyzing user responses during a conversation,

[2135] Means for selecting and providing content to promote learning,

[2136] A means of saving data collected during a conversation and applying the feedback to the next conversation,

[2137] A means of providing training programs based on user information,

[2138] A means of selecting and adapting training scenarios based on dialogue profiles through dialogue,

[2139] A means for analyzing user instructions using a speech recognition engine and generating the next dialogue content,

[2140] A means for providing instructions and training content using an audio output means,

[2141] A system that includes this.

[2142] (Claim 2)

[2143] The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, learning progress and skill level.

[2144] (Claim 3)

[2145] The system according to claim 1, wherein the learning-promoting content includes words, place names, and other educational information and training content in different languages.

[2146] "Example 2 of combining an emotion engine"

[2147] (Claim 1)

[2148] Means of registering user information,

[2149] A means of creating individual dialogue profiles based on registered information,

[2150] A means of selecting a dialogue scenario based on a dialogue profile,

[2151] A means of initiating and conducting a dialogue based on a selected dialogue scenario,

[2152] A means of obtaining and analyzing user responses during a conversation,

[2153] Means for selecting and providing content to promote learning,

[2154] A means of saving data collected during a conversation and applying the feedback to the next conversation,

[2155] A means of analyzing emotional data using an emotion engine that recognizes user emotions,

[2156] A means of dynamically adjusting dialogue content based on emotional data,

[2157] A system that includes this.

[2158] (Claim 2)

[2159] The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, and learning progress.

[2160] (Claim 3)

[2161] The system according to claim 1, wherein the content for promoting learning includes words, place names, and other educational information in different languages.

[2162] "Application example 2 when combining with an emotional engine"

[2163] (Claim 1)

[2164] Means of registering user information,

[2165] A means of creating individual dialogue profiles based on registered information,

[2166] A means of selecting a dialogue scenario based on a dialogue profile,

[2167] A means of initiating and conducting a dialogue based on a selected dialogue scenario,

[2168] A means of acquiring and analyzing user responses and sentiment data during a conversation,

[2169] A means of selecting and providing personalized recommended menus and promotions based on the user's emotional state,

[2170] A means of saving data collected during the conversation and applying the feedback to the next conversation in the food delivery service,

[2171] A system that includes this.

[2172] (Claim 2)

[2173] The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, learning progress and emotional state.

[2174] (Claim 3)

[2175] The system according to claim 1, wherein recommended menu items for a food delivery service are generated taking into account the user's emotional state. [Explanation of Symbols]

[2176] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means of registering user information, A means of creating individual dialogue profiles based on registered information, A means of selecting a dialogue scenario based on a dialogue profile, A means of initiating and conducting a dialogue based on a selected dialogue scenario, A means of obtaining and analyzing user responses during a conversation, Means for selecting and providing content to promote learning, A means of saving data collected during a conversation and applying the feedback to the next conversation, A system that includes this.

2. The system according to claim 1, wherein the conversation profile is personalized based on the user's age, interests, and learning progress.

3. The system according to claim 1, wherein the learning-enhancing content includes words, place names, and other educational information in different languages.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A