system

The system addresses the unnatural use of smartphones for conversation practice by integrating a conversation support app within a doll or stuffed animal, utilizing AI to generate natural responses, enhancing user engagement and effectiveness.

JP2026073179APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional methods for conversation practice and translation using smartphones are unnatural and embarrassing for users.

Method used

A system comprising a storage unit, analysis unit, and conversation support unit, where a smartphone with a conversation support app is installed inside a doll or stuffed animal, analyzing user input and generating appropriate responses using natural language processing and generative AI to facilitate natural conversations.

Benefits of technology

Enables users to engage in natural conversations and translations without the embarrassment associated with direct smartphone use, allowing for effective practice and interpretation in various situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073179000001_ABST
    Figure 2026073179000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to facilitate conversation practice and interpretation in a natural manner. [Solution] The system according to the embodiment comprises a storage unit, an analysis unit, and a conversation support unit. The storage unit holds a smartphone with a conversation support app installed inside a doll or stuffed animal. The analysis unit analyzes what the user says and generates an appropriate response. The conversation support unit provides the response generated by the analysis unit to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, when practicing conversation or translation, directly using a smartphone is unnatural, and there is a problem that it is embarrassing and difficult to use for the user.

[0005] The system according to the embodiment aims to perform conversation practice and translation in a natural form.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a storage unit, an analysis unit, and a conversation support unit. The storage unit holds a smartphone with a conversation support app installed inside a doll or stuffed animal. The analysis unit analyzes what the user says and generates an appropriate response. The conversation support unit provides the response generated by the analysis unit to the user. [Effects of the Invention]

[0007] The system according to this embodiment can perform conversation practice and interpretation in a natural manner. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communications among a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The conversation support system according to an embodiment of the present invention is a robot that supports conversations in various situations. This conversation support system is designed so that users can enjoy conversations in a natural way by placing a smartphone with a conversation support app installed inside a doll or stuffed animal. First, let's explain the case of practicing English conversation alone. Many people feel embarrassed to speak English into their smartphone, but this embarrassment is reduced when talking to a doll or stuffed animal. For example, if the user says "How are you?" to the doll, the smartphone app will give an appropriate response to the question. This allows the user to practice English conversation as if they were having a normal conversation. Also, even if someone sees it, the damage will be minimal, and they will only think, "What a kind person who likes dolls." Next, let's explain its use as an interpreter while traveling abroad. For example, when talking to a staff member at a supermarket checkout, if you place the doll between yourself and the staff member and say, "I'm at the supermarket checkout right now, can you interpret for me?", the smartphone app will mediate the conversation. Existing interpreter apps involve both parties peering at their smartphones, which is unnatural, but by talking to a doll, a natural three-way conversation can take place. Furthermore, let's discuss children's language learning. For young children just beginning to learn language, it's crucial to repeatedly listen and speak. By using a doll as a practice partner, children can freely talk to their favorite doll. For example, if a child asks the doll, "What is your name?", a smartphone app will provide an appropriate response. This allows children to learn language in a fun way. Thus, this invention provides a robot that supports conversations in various situations, realizing a mechanism that allows users to enjoy conversations in a natural way. As a result, the conversation support system allows users to enjoy conversations in a natural manner.

[0029] The conversation support system according to this embodiment comprises a storage unit, an analysis unit, and a conversation support unit. The storage unit holds a smartphone with a conversation support app installed inside a doll or stuffed animal. The analysis unit analyzes what the user says and generates an appropriate response. The analysis unit analyzes the user's utterance using, for example, natural language processing technology and generates an appropriate response. The analysis unit can also analyze the user's utterance and generate an appropriate response using a generation AI. For example, the analysis unit analyzes the user's utterance using a text generation AI (e.g., LLM) and generates an appropriate response. The analysis unit can also analyze the user's utterance and generate an appropriate response using a multimodal generation AI. The conversation support unit provides the response generated by the analysis unit to the user. The conversation support unit provides the response generated by the analysis unit in voice using, for example, speech synthesis technology. The conversation support unit can also provide the response generated by the analysis unit in voice using a generation AI. For example, the conversation support unit provides the response generated by the analysis unit in voice using a speech generation AI. Furthermore, the conversation support unit can also provide the responses generated by the analysis unit in text format using a text generation AI. This allows the conversation support system to enable users to enjoy conversations in a natural way. Some or all of the above-described processes in the analysis unit and conversation support unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the user's utterances as input and outputs appropriate responses. The conversation support unit can provide responses using an AI model that takes the responses generated by the analysis unit as input and outputs them in voice or text format.

[0030] The storage compartment holds a smartphone with a conversation-supporting app installed inside a doll or stuffed animal. Specifically, the smartphone is placed inside the doll or stuffed animal, out of sight from the outside. This allows the user to enjoy natural conversations with the doll or stuffed animal. The smartphone is secured in a dedicated holder or pocket and designed to prevent it from shifting during use. Furthermore, the materials and structure of the doll or stuffed animal are carefully designed to ensure that the smartphone's microphone and speaker function properly. For example, a mesh material that allows sound to pass through easily is used for the microphone, and an opening is provided for the speaker to prevent muffled sound. In addition, a low-power consumption mode is set to extend the smartphone's battery life. This allows the user to enjoy conversations for extended periods. The storage compartment is designed for easy smartphone charging and may include a USB port or wireless charging function. This allows the user to easily charge their smartphone without removing it from the device.

[0031] The analysis unit analyzes what the user says and generates an appropriate response. For example, the analysis unit uses natural language processing technology to analyze the user's utterances and generate an appropriate response. Specifically, it converts the user's utterances into text data using speech recognition technology and then analyzes that text data using natural language processing technology. The analysis unit can also use generative AI to analyze the user's utterances and generate an appropriate response. For example, the analysis unit uses text generation AI (e.g., LLM) to analyze the user's utterances and generate an appropriate response. LLM has learned from a large amount of text data and has the ability to understand context and generate natural responses. Furthermore, the analysis unit can also use multimodal generative AI to analyze the user's utterances and generate an appropriate response. Multimodal generative AI has the ability to integrate and analyze multiple data formats such as speech, text, and images, providing a richer conversational experience. For example, if a user speaks while showing a specific image, the content of the image can be analyzed and a relevant response can be generated. The analysis unit can also use sentiment analysis technology to understand the intent and emotions behind the user's utterances. This enables the generation of appropriate responses that respond to the user's emotions, resulting in more natural and engaging conversations.

[0032] The conversation support unit provides the user with responses generated by the analysis unit. For example, the conversation support unit provides the responses generated by the analysis unit in audio format using speech synthesis technology. Specifically, it uses speech synthesis technology to convert text data into natural-sounding speech for the user to hear. Speech synthesis technology can adjust the tone and speed of the voice according to the user's preferences. The conversation support unit can also provide the responses generated by the analysis unit in audio format using generative AI. For example, the conversation support unit uses speech generation AI to provide the responses generated by the analysis unit in audio format. Speech generation AI has the ability to generate more natural and human-like speech, providing the user with a more engaging conversational experience. Furthermore, the conversation support unit can also provide the responses generated by the analysis unit in text format using text generation AI. For example, if the user prefers a text response rather than an audio response, the conversation support unit displays the responses generated by the analysis unit in text format using text generation AI. This allows the conversation support system to enable users to enjoy conversations in a natural way. In addition, the conversation support unit can collect user feedback and continuously improve the accuracy and quality of its responses. For example, the system can provide a feature that allows users to rate responses, and adjust the response generation algorithm based on those ratings. This enables the conversation support unit to always provide the best possible response to the user, improving the conversational experience.

[0033] The conversation support unit includes an interpretation unit that provides interpretation functionality. The interpretation unit, for example, translates what the user says in real time and conveys it to the other party. The interpretation unit can also use generative AI to translate the user's statements and convey them to the other party. For example, the interpretation unit can use text generation AI to translate the user's statements and convey them to the other party. The interpretation unit can also use speech generation AI to translate the user's statements and convey them to the other party audibly. The interpretation unit supports multiple languages, such as English, Japanese, and Chinese. The interpretation unit analyzes the user's statements and translates them into the appropriate language. This enables interpretation during overseas travel by providing interpretation functionality. Some or all of the above-described processes in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform translation using an AI model that takes the user's statements as input and outputs the translation result.

[0034] The conversation support unit includes a learning unit that supports children's language learning. The learning unit, for example, analyzes what the child says and generates an appropriate response. The learning unit can also use generative AI to analyze the child's statements and generate an appropriate response. For example, the learning unit can use text generation AI to analyze the child's statements and generate an appropriate response. The learning unit can also use speech generation AI to analyze the child's statements and provide a voice response. The learning unit supports the progress of language learning based on what the child says. The learning unit analyzes the child's statements and provides appropriate feedback. This makes language learning more enjoyable by supporting the child's language learning. Some or all of the above-described processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform analysis using an AI model that takes the child's statements as input and outputs an appropriate response.

[0035] The storage unit includes a character unit that transforms favorite characters into robots through corporate collaborations. The character unit can transform, for example, anime characters, game characters, or corporate mascots into robots. The character unit can also generate character designs and movements using generative AI. For example, the character unit can generate character designs using image generation AI. It can also generate character movements using animation generation AI. The character unit customizes the character design and movements according to the user's preferences. This makes it easier to attract the user's interest by transforming their favorite characters into robots. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can perform generation using an AI model that takes the user's preferences as input and outputs character designs and movements.

[0036] The storage unit includes a dedicated equipment unit for developing specialized equipment. This unit develops, for example, specialized cases for storing smartphones and peripheral devices that interact with smartphones. The dedicated equipment unit can also design and manufacture specialized equipment using generative AI. For example, it can manufacture specialized cases using 3D printing technology. Furthermore, it can develop sensors and actuators that interact with smartphones. The dedicated equipment unit customizes the functions and designs of specialized equipment according to user needs. This expands the system's functionality through the development of specialized equipment. Some or all of the above-described processes in the dedicated equipment unit may be performed using AI, or not. For example, the dedicated equipment unit can develop equipment using an AI model that takes user needs as input and outputs the design and manufacture of specialized equipment.

[0037] The storage unit automatically adjusts the internal temperature and humidity to maintain the optimal operating environment for the smartphone. For example, if the internal temperature becomes high, the storage unit activates a cooling fan to lower the temperature. The storage unit can also use generative AI to monitor the internal temperature and humidity and make appropriate adjustments. For example, the storage unit monitors the internal environment using temperature and humidity sensors and controls the cooling fan and dehumidification function. The storage unit also constantly monitors the temperature and humidity with sensors and makes automatic adjustments to ensure they remain within an appropriate range. This improves system reliability by maintaining the optimal operating environment for the smartphone. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can make adjustments using an AI model that takes data from temperature and humidity sensors as input and outputs control of the cooling fan and dehumidification function.

[0038] The storage unit incorporates sensors to automatically adjust the position and orientation of the smartphone. For example, if the smartphone is tilted, the sensors will detect this and automatically return it to a horizontal position. The storage unit can also use generative AI to monitor the smartphone's position and orientation and make appropriate adjustments. For example, the storage unit can use acceleration sensors and gyroscopes to monitor the smartphone's position and orientation and control motors and actuators. The storage unit can also be equipped with a mechanism to secure the smartphone so that it does not move within the unit. This improves user convenience by automatically adjusting the smartphone's position and orientation. Some or all of the above processing in the storage unit may be performed using AI, or not. For example, the storage unit can perform adjustments using an AI model that takes data from acceleration sensors and gyroscopes as input and outputs control for motors and actuators.

[0039] The storage compartment can be equipped with a charging function to automatically charge a smartphone while it is stored inside. For example, the storage compartment can have a wireless charging function, so charging starts simply by placing the smartphone inside. The storage compartment can also use generative AI to optimize the timing and speed of charging. For example, the storage compartment can monitor the smartphone's battery level and start charging at the optimal time. The storage compartment can also have a USB port, allowing charging by connecting a cable. This eliminates the hassle of charging by automatically charging the smartphone while it is stored inside. Some or all of the above processes in the storage compartment may be performed using AI, for example, or not. For example, the storage compartment can optimize charging using an AI model that takes smartphone battery level data as input and outputs the timing and speed of charging.

[0040] The storage unit incorporates speakers and microphones to optimize audio input and output. For example, the storage unit may incorporate high-quality speakers to provide clear audio. The storage unit can also optimize audio input and output using generative AI. For example, the storage unit may incorporate a microphone with noise-canceling capabilities to remove ambient noise. The storage unit may also employ multiple microphones to accurately capture the user's voice. This optimizes audio input and output, improving the quality of conversation. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can optimize audio input and output using an AI model that takes audio data as input and outputs noise cancellation and audio optimization.

[0041] The analysis unit refers to the user's past conversation history and generates more appropriate responses. For example, the analysis unit generates appropriate responses based on phrases the user has used in the past. The analysis unit can also use generative AI to analyze the user's past conversation history and generate appropriate responses. For example, the analysis unit uses text generation AI to analyze the user's past conversation history and generate appropriate responses. The analysis unit can also use speech generation AI to analyze the user's past conversation history and provide responses in voice. The analysis unit extracts preferred topics based on the user's past conversation history and generates responses based on them. This allows for more natural conversations by referring to past conversation history. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the user's past conversation history data as input and outputs appropriate responses.

[0042] The analysis unit analyzes the user's pronunciation and intonation and provides advice for improving pronunciation. For example, the analysis unit analyzes the user's pronunciation and provides correct pronunciation in audio. The analysis unit can also use generative AI to analyze the user's pronunciation and intonation and provide appropriate advice. For example, the analysis unit uses speech generation AI to analyze the user's pronunciation and provides correct pronunciation in audio. The analysis unit can also use text generation AI to analyze the user's pronunciation and intonation and provide advice in text. The analysis unit analyzes the user's intonation and provides natural intonation in audio. This improves the effectiveness of language learning by providing advice for improving pronunciation and intonation. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes user pronunciation data as input and outputs advice for improving pronunciation.

[0043] The analysis unit analyzes the user's gestures and facial expressions to gain a deeper understanding of the conversation context. For example, the analysis unit analyzes the user's gestures to understand the intent of the conversation. The analysis unit can also use generative AI to analyze the user's gestures and facial expressions to understand the conversation context. For example, the analysis unit uses image generation AI to analyze the user's gestures to understand the intent of the conversation. The analysis unit can also use facial expression recognition technology to analyze the user's facial expressions to understand their emotions. The analysis unit gains a deeper understanding of the conversation context based on the user's gestures and facial expressions. This allows for a deeper understanding of the conversation context by analyzing gestures and facial expressions. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes user gesture and facial expression data as input and outputs an analysis for understanding the conversation context.

[0044] The analysis unit analyzes multiple languages ​​simultaneously to support multilingual conversations. For example, if a user uses multiple languages, the analysis unit analyzes them simultaneously and generates an appropriate response. The analysis unit can also use generative AI to analyze multiple languages ​​simultaneously and generate an appropriate response. For example, the analysis unit can use text generation AI to analyze multiple languages ​​and generate an appropriate response. Alternatively, the analysis unit can use speech generation AI to analyze multiple languages ​​and provide a voice response. If a user speaks in different languages, the analysis unit analyzes each language and generates a response. This enables multilingual conversations by analyzing multiple languages ​​simultaneously. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes multiple language data as input and outputs an appropriate response.

[0045] The conversation support unit learns the user's past conversation patterns to provide more natural conversations. For example, the conversation support unit learns the user's past conversation patterns and generates natural responses. The conversation support unit can also use generative AI to learn the user's past conversation patterns and generate appropriate responses. For example, the conversation support unit can use text generation AI to learn the user's past conversation patterns and generate appropriate responses. The conversation support unit can also use speech generation AI to learn the user's past conversation patterns and provide responses in voice. The conversation support unit learns the user's preferred topics and proceeds with the conversation based on them. This makes more natural conversations possible by learning past conversation patterns. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can learn using an AI model that takes the user's past conversation pattern data as input and outputs appropriate responses.

[0046] The conversation support unit suggests conversation topics based on the user's interests. For example, the conversation support unit learns the user's interests and suggests conversation topics based on them. The conversation support unit can also use generative AI to analyze the user's interests and suggest appropriate topics. For example, the conversation support unit can use text generation AI to analyze the user's interests and suggest appropriate topics. The conversation support unit can also use speech generation AI to analyze the user's interests and suggest topics in speech. The conversation support unit analyzes the user's past conversation history and suggests the most suitable topics. This makes conversations more enjoyable by suggesting topics based on the user's interests. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can make suggestions using an AI model that takes user interest data as input and outputs appropriate topics.

[0047] The conversation support unit analyzes the ambient sounds surrounding the user and provides appropriate conversation content. For example, if the surroundings are quiet, the conversation support unit provides detailed conversation content. The conversation support unit can also use generative AI to analyze ambient sounds and provide appropriate conversation content. For example, the conversation support unit uses speech generation AI to analyze ambient sounds and provide appropriate conversation content. The conversation support unit can also use noise cancellation technology to remove ambient noise and provide appropriate conversation content. If the surroundings are noisy, the conversation support unit provides concise conversation content. In this way, appropriate conversation content can be provided by analyzing ambient sounds. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can perform analysis using an AI model that takes ambient sound data as input and outputs appropriate conversation content.

[0048] The conversation support unit refers to the user's schedule and starts a conversation at an appropriate time. For example, the conversation support unit refers to the user's schedule and starts a conversation during an available time. The conversation support unit can also use generative AI to analyze the user's schedule and start a conversation at an appropriate time. For example, the conversation support unit uses schedule management AI to analyze the user's schedule and start a conversation at the optimal time. The conversation support unit can also start a conversation before an important appointment based on the user's schedule. This allows for more effective conversations by starting conversations based on the user's schedule. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can perform analysis using an AI model that takes user schedule data as input and outputs the timing of when to start a conversation.

[0049] The interpretation unit refers to the user's past interpretation history to provide a more appropriate interpretation. For example, the interpretation unit refers to the user's past interpretation history and selects appropriate expressions. The interpretation unit can also use generative AI to analyze the user's past interpretation history and provide an appropriate interpretation. For example, the interpretation unit uses text generation AI to analyze the user's past interpretation history and provide an appropriate interpretation. The interpretation unit can also use speech generation AI to analyze the user's past interpretation history and provide an interpretation in speech. The interpretation unit provides the optimal interpretation based on the user's past interpretation history. This makes it possible to provide a more appropriate interpretation by referring to past interpretation history. Some or all of the above processing in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform analysis using an AI model that takes the user's past interpretation history data as input and outputs an appropriate interpretation.

[0050] The interpretation unit incorporates region-specific expressions into the interpretation based on the user's geographical location information. For example, if the user is in a specific region, the interpretation unit will incorporate region-specific expressions into the interpretation. The interpretation unit can also use generative AI to analyze the user's geographical location information and incorporate appropriate expressions into the interpretation. For example, the interpretation unit can use location information analysis AI to analyze the user's geographical location information and incorporate region-specific expressions into the interpretation. Furthermore, if the user moves to a different region, the interpretation unit can also incorporate expressions from the new region into the interpretation. This allows for more natural interpretation by incorporating region-specific expressions. Some or all of the above processing in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform analysis using an AI model that takes the user's geographical location data as input and outputs appropriate expressions.

[0051] The learning unit refers to the user's past learning history and provides an optimal learning plan. For example, the learning unit refers to the user's past learning history and provides an optimal learning plan. The learning unit can also use generative AI to analyze the user's past learning history and provide an appropriate learning plan. For example, the learning unit uses text generation AI to analyze the user's past learning history and provide an appropriate learning plan. The learning unit can also use speech generation AI to analyze the user's past learning history and provide a learning plan in speech. The learning unit provides an individually customized learning plan based on the user's past learning history. This allows for the provision of a more effective learning plan by referring to past learning history. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform analysis using an AI model that takes the user's past learning history data as input and outputs an appropriate learning plan.

[0052] The learning unit suggests new learning topics based on the user's interests. For example, the learning unit learns the user's interests and suggests new learning topics based on them. The learning unit can also use generative AI to analyze the user's interests and suggest appropriate learning topics. For example, the learning unit can use text generation AI to analyze the user's interests and suggest appropriate learning topics. Alternatively, the learning unit can use speech generation AI to analyze the user's interests and suggest learning topics in speech. The learning unit analyzes the user's past learning history and suggests optimal new learning topics. This improves the user's motivation to learn by suggesting new learning topics based on their interests. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can make suggestions using an AI model that takes user interest data as input and outputs appropriate learning topics.

[0053] The character unit refers to the user's past character selection history and proposes the most suitable character. For example, the character unit refers to the user's past character selection history and proposes the most suitable character. The character unit can also use generative AI to analyze the user's past character selection history and propose an appropriate character. For example, the character unit uses text generation AI to analyze the user's past character selection history and propose an appropriate character. The character unit can also use image generation AI to analyze the user's past character selection history and propose a character as an image. The character unit proposes a character that is individually customized based on the user's past character selection history. In this way, the most suitable character is proposed to the user by referring to their past character selection history. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can perform analysis using an AI model that takes the user's past character selection history data as input and outputs an appropriate character.

[0054] The character department changes the character's costume and accessories according to the season and events. For example, the character department changes the character's costume according to the season. The character department can also design costumes and accessories according to the season and events using generative AI. For example, the character department designs seasonal costumes using image generation AI. The character department can also design character accessories to match events. The character department customizes the character's appearance to match specific events. This makes it easier to attract user interest by changing the character's appearance according to the season and events. Some or all of the above processing in the character department may be done using AI, for example, or not using AI. For example, the character department can design using an AI model that takes seasonal and event data as input and outputs character costumes and accessories.

[0055] The dedicated device unit refers to the user's past usage history and provides the optimal device settings. For example, the dedicated device unit refers to the user's past usage history and provides the optimal device settings. The dedicated device unit can also use generative AI to analyze the user's past usage history and provide appropriate device settings. For example, the dedicated device unit uses text generation AI to analyze the user's past usage history and provide appropriate device settings. The dedicated device unit also uses device information analysis AI to analyze the user's past usage history and provide optimal device settings. The dedicated device unit provides individually customized device settings based on the user's past usage history. This ensures that the optimal device settings are provided by referring to the past usage history. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform analysis using an AI model that takes the user's past usage history data as input and outputs appropriate device settings.

[0056] The dedicated device unit provides optimal device settings based on the user's device information. For example, the dedicated device unit refers to the user's device information and provides optimal device settings. The dedicated device unit can also analyze the user's device information using a generation AI and provide appropriate device settings. For example, the dedicated device unit analyzes the user's device information using a device information analysis AI and provides optimal device settings. Furthermore, the dedicated device unit can provide individually customized device settings based on the user's device information. This enables more effective operation by providing optimal device settings based on device information. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform analysis using an AI model that takes user device information data as input and outputs appropriate device settings.

[0057] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0058] The conversation support system can also learn from the user's past conversation history and suggest conversation topics based on the user's preferences and interests. For example, if the user has talked a lot about sports in the past, the system can suggest new sports-related topics. Similarly, if the user has talked about a particular movie or music, the system can suggest related topics. This enables conversations based on the user's interests, improving the quality of the conversation.

[0059] The conversation support system can also refer to the user's schedule and have the functionality to initiate conversations at appropriate times. For example, the system can refrain from initiating conversations during busy periods and start them at times when the user is free. It can also initiate conversations as reminders before important meetings or events to encourage preparation. This enables flexible conversations tailored to the user's schedule, improving user convenience.

[0060] The conversation support system can also be equipped with the ability to analyze the ambient sounds surrounding the user and provide appropriate conversation content. For example, if the surroundings are quiet, the system can provide detailed conversation content, and if the surroundings are noisy, it can provide concise conversation content. In addition, the system can use noise cancellation technology to remove ambient noise and provide conversation with clear voice. This enables appropriate conversation according to the surrounding environment, improving the quality of conversation.

[0061] The conversation support system can also be equipped with features to learn the user's past conversation patterns and provide more natural conversations. For example, it can generate appropriate responses based on phrases and topics the user has used in the past. It can also learn the user's preferred topics and guide the conversation accordingly. By referring to past conversation patterns, it enables more natural and consistent conversations, improving user satisfaction.

[0062] The conversation support system can also be equipped with features to analyze the user's pronunciation and intonation and provide advice for improving pronunciation. For example, it can analyze the user's pronunciation and provide correct pronunciation in audio. It can also analyze the user's intonation and provide natural intonation in audio. By providing advice on improving pronunciation and intonation, the effectiveness of language learning and the user's skills will improve.

[0063] The conversation support system can also include a feature that suggests conversation topics based on the user's interests. For example, it can learn the user's interests and suggest new conversation topics based on them. It can also analyze the user's past conversation history and suggest the most suitable topics. This enables conversations based on the user's interests, making conversations more enjoyable.

[0064] The following briefly describes the processing flow for example form 1.

[0065] Step 1: Place a smartphone with a conversation support app installed into the storage compartment, then place it inside the doll or stuffed animal. Step 2: The analysis unit analyzes what the user says and generates an appropriate response. The analysis unit analyzes the user's utterance using, for example, natural language processing technology or generative AI (e.g., text generation AI or multimodal generation AI) and generates an appropriate response. Step 3: The conversation support unit provides the user with the response generated by the analysis unit. The conversation support unit provides the response generated by the analysis unit in voice or text, for example, using speech synthesis technology or generative AI (e.g., speech generation AI or text generation AI).

[0066] (Example of form 2) The conversation support system according to an embodiment of the present invention is a robot that supports conversations in various situations. This conversation support system is designed so that users can enjoy conversations in a natural way by placing a smartphone with a conversation support app installed inside a doll or stuffed animal. First, let's explain the case of practicing English conversation alone. Many people feel embarrassed to speak English into their smartphone, but this embarrassment is reduced when talking to a doll or stuffed animal. For example, if the user says "How are you?" to the doll, the smartphone app will give an appropriate response to the question. This allows the user to practice English conversation as if they were having a normal conversation. Also, even if someone sees it, the damage will be minimal, and they will only think, "What a kind person who likes dolls." Next, let's explain its use as an interpreter while traveling abroad. For example, when talking to a staff member at a supermarket checkout, if you place the doll between yourself and the staff member and say, "I'm at the supermarket checkout right now, can you interpret for me?", the smartphone app will mediate the conversation. Existing interpreter apps involve both parties peering at their smartphones, which is unnatural, but by talking to a doll, a natural three-way conversation can take place. Furthermore, let's discuss children's language learning. For young children just beginning to learn language, it's crucial to repeatedly listen and speak. By using a doll as a practice partner, children can freely talk to their favorite doll. For example, if a child asks the doll, "What is your name?", a smartphone app will provide an appropriate response. This allows children to learn language in a fun way. Thus, this invention provides a robot that supports conversations in various situations, realizing a mechanism that allows users to enjoy conversations in a natural way. As a result, the conversation support system allows users to enjoy conversations in a natural manner.

[0067] The conversation support system according to this embodiment comprises a storage unit, an analysis unit, and a conversation support unit. The storage unit holds a smartphone with a conversation support app installed inside a doll or stuffed animal. The analysis unit analyzes what the user says and generates an appropriate response. The analysis unit analyzes the user's utterance using, for example, natural language processing technology and generates an appropriate response. The analysis unit can also analyze the user's utterance and generate an appropriate response using a generation AI. For example, the analysis unit analyzes the user's utterance using a text generation AI (e.g., LLM) and generates an appropriate response. The analysis unit can also analyze the user's utterance and generate an appropriate response using a multimodal generation AI. The conversation support unit provides the response generated by the analysis unit to the user. The conversation support unit provides the response generated by the analysis unit in voice using, for example, speech synthesis technology. The conversation support unit can also provide the response generated by the analysis unit in voice using a generation AI. For example, the conversation support unit provides the response generated by the analysis unit in voice using a speech generation AI. Furthermore, the conversation support unit can also provide the responses generated by the analysis unit in text format using a text generation AI. This allows the conversation support system to enable users to enjoy conversations in a natural way. Some or all of the above-described processes in the analysis unit and conversation support unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the user's utterances as input and outputs appropriate responses. The conversation support unit can provide responses using an AI model that takes the responses generated by the analysis unit as input and outputs them in voice or text format.

[0068] The storage compartment holds a smartphone with a conversation-supporting app installed inside a doll or stuffed animal. Specifically, the smartphone is placed inside the doll or stuffed animal, out of sight from the outside. This allows the user to enjoy natural conversations with the doll or stuffed animal. The smartphone is secured in a dedicated holder or pocket and designed to prevent it from shifting during use. Furthermore, the materials and structure of the doll or stuffed animal are carefully designed to ensure that the smartphone's microphone and speaker function properly. For example, a mesh material that allows sound to pass through easily is used for the microphone, and an opening is provided for the speaker to prevent muffled sound. In addition, a low-power consumption mode is set to extend the smartphone's battery life. This allows the user to enjoy conversations for extended periods. The storage compartment is designed for easy smartphone charging and may include a USB port or wireless charging function. This allows the user to easily charge their smartphone without removing it from the device.

[0069] The analysis unit analyzes what the user says and generates an appropriate response. For example, the analysis unit uses natural language processing technology to analyze the user's utterances and generate an appropriate response. Specifically, it converts the user's utterances into text data using speech recognition technology and then analyzes that text data using natural language processing technology. The analysis unit can also use generative AI to analyze the user's utterances and generate an appropriate response. For example, the analysis unit uses text generation AI (e.g., LLM) to analyze the user's utterances and generate an appropriate response. LLM has learned from a large amount of text data and has the ability to understand context and generate natural responses. Furthermore, the analysis unit can also use multimodal generative AI to analyze the user's utterances and generate an appropriate response. Multimodal generative AI has the ability to integrate and analyze multiple data formats such as speech, text, and images, providing a richer conversational experience. For example, if a user speaks while showing a specific image, the content of the image can be analyzed and a relevant response can be generated. The analysis unit can also use sentiment analysis technology to understand the intent and emotions behind the user's utterances. This enables the generation of appropriate responses that respond to the user's emotions, resulting in more natural and engaging conversations.

[0070] The conversation support unit provides the user with responses generated by the analysis unit. For example, the conversation support unit provides the responses generated by the analysis unit in audio format using speech synthesis technology. Specifically, it uses speech synthesis technology to convert text data into natural-sounding speech for the user to hear. Speech synthesis technology can adjust the tone and speed of the voice according to the user's preferences. The conversation support unit can also provide the responses generated by the analysis unit in audio format using generative AI. For example, the conversation support unit uses speech generation AI to provide the responses generated by the analysis unit in audio format. Speech generation AI has the ability to generate more natural and human-like speech, providing the user with a more engaging conversational experience. Furthermore, the conversation support unit can also provide the responses generated by the analysis unit in text format using text generation AI. For example, if the user prefers a text response rather than an audio response, the conversation support unit displays the responses generated by the analysis unit in text format using text generation AI. This allows the conversation support system to enable users to enjoy conversations in a natural way. In addition, the conversation support unit can collect user feedback and continuously improve the accuracy and quality of its responses. For example, the system can provide a feature that allows users to rate responses, and adjust the response generation algorithm based on those ratings. This enables the conversation support unit to always provide the best possible response to the user, improving the conversational experience.

[0071] The conversation support unit includes an interpretation unit that provides interpretation functionality. The interpretation unit, for example, translates what the user says in real time and conveys it to the other party. The interpretation unit can also use generative AI to translate the user's statements and convey them to the other party. For example, the interpretation unit can use text generation AI to translate the user's statements and convey them to the other party. The interpretation unit can also use speech generation AI to translate the user's statements and convey them to the other party audibly. The interpretation unit supports multiple languages, such as English, Japanese, and Chinese. The interpretation unit analyzes the user's statements and translates them into the appropriate language. This enables interpretation during overseas travel by providing interpretation functionality. Some or all of the above-described processes in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform translation using an AI model that takes the user's statements as input and outputs the translation result.

[0072] The conversation support unit includes a learning unit that supports children's language learning. The learning unit, for example, analyzes what the child says and generates an appropriate response. The learning unit can also use generative AI to analyze the child's statements and generate an appropriate response. For example, the learning unit can use text generation AI to analyze the child's statements and generate an appropriate response. The learning unit can also use speech generation AI to analyze the child's statements and provide a voice response. The learning unit supports the progress of language learning based on what the child says. The learning unit analyzes the child's statements and provides appropriate feedback. This makes language learning more enjoyable by supporting the child's language learning. Some or all of the above-described processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform analysis using an AI model that takes the child's statements as input and outputs an appropriate response.

[0073] The storage unit includes a character unit that transforms favorite characters into robots through corporate collaborations. The character unit can transform, for example, anime characters, game characters, or corporate mascots into robots. The character unit can also generate character designs and movements using generative AI. For example, the character unit can generate character designs using image generation AI. It can also generate character movements using animation generation AI. The character unit customizes the character design and movements according to the user's preferences. This makes it easier to attract the user's interest by transforming their favorite characters into robots. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can perform generation using an AI model that takes the user's preferences as input and outputs character designs and movements.

[0074] The storage unit includes a dedicated equipment unit for developing specialized equipment. This unit develops, for example, specialized cases for storing smartphones and peripheral devices that interact with smartphones. The dedicated equipment unit can also design and manufacture specialized equipment using generative AI. For example, it can manufacture specialized cases using 3D printing technology. Furthermore, it can develop sensors and actuators that interact with smartphones. The dedicated equipment unit customizes the functions and designs of specialized equipment according to user needs. This expands the system's functionality through the development of specialized equipment. Some or all of the above-described processes in the dedicated equipment unit may be performed using AI, or not. For example, the dedicated equipment unit can develop equipment using an AI model that takes user needs as input and outputs the design and manufacture of specialized equipment.

[0075] The storage unit estimates the user's emotions and adjusts the opening and closing timing based on the estimated emotions. For example, if the user is stressed, the storage unit will automatically open to make it easier to take out the smartphone. The storage unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the storage unit can estimate the user's emotions using facial recognition technology. The storage unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions, the storage unit adjusts the opening and closing timing. For example, if the user is relaxed, the storage unit will open and close slowly. If the user is in a hurry, the storage unit will open and close quickly. By adjusting the opening and closing timing of the storage unit according to the user's emotions, usability is improved. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can make adjustments using an AI model that takes user emotion data as input and outputs the opening and closing timing.

[0076] The storage unit automatically adjusts the internal temperature and humidity to maintain the optimal operating environment for the smartphone. For example, if the internal temperature becomes high, the storage unit activates a cooling fan to lower the temperature. The storage unit can also use generative AI to monitor the internal temperature and humidity and make appropriate adjustments. For example, the storage unit monitors the internal environment using temperature and humidity sensors and controls the cooling fan and dehumidification function. The storage unit also constantly monitors the temperature and humidity with sensors and makes automatic adjustments to ensure they remain within an appropriate range. This improves system reliability by maintaining the optimal operating environment for the smartphone. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can make adjustments using an AI model that takes data from temperature and humidity sensors as input and outputs control of the cooling fan and dehumidification function.

[0077] The storage unit incorporates sensors to automatically adjust the position and orientation of the smartphone. For example, if the smartphone is tilted, the sensors will detect this and automatically return it to a horizontal position. The storage unit can also use generative AI to monitor the smartphone's position and orientation and make appropriate adjustments. For example, the storage unit can use acceleration sensors and gyroscopes to monitor the smartphone's position and orientation and control motors and actuators. The storage unit can also be equipped with a mechanism to secure the smartphone so that it does not move within the unit. This improves user convenience by automatically adjusting the smartphone's position and orientation. Some or all of the above processing in the storage unit may be performed using AI, or not. For example, the storage unit can perform adjustments using an AI model that takes data from acceleration sensors and gyroscopes as input and outputs control for motors and actuators.

[0078] The storage unit estimates the user's emotions and changes its design and color based on the estimated emotions. For example, if the user is relaxed, the storage unit's color may change to blue or green. The storage unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the storage unit may estimate the user's emotions using facial recognition technology. The storage unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions, the storage unit changes its design and color. For example, if the user is excited, the storage unit's color may change to red or orange. If the user is sad, the storage unit's color may change to a calming color. By changing the design and color according to the user's emotions, user satisfaction is improved. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can be adjusted using an AI model that takes user emotion data as input and outputs changes to the design and color.

[0079] The storage compartment can be equipped with a charging function to automatically charge a smartphone while it is stored inside. For example, the storage compartment can have a wireless charging function, so charging starts simply by placing the smartphone inside. The storage compartment can also use generative AI to optimize the timing and speed of charging. For example, the storage compartment can monitor the smartphone's battery level and start charging at the optimal time. The storage compartment can also have a USB port, allowing charging by connecting a cable. This eliminates the hassle of charging by automatically charging the smartphone while it is stored inside. Some or all of the above processes in the storage compartment may be performed using AI, for example, or not. For example, the storage compartment can optimize charging using an AI model that takes smartphone battery level data as input and outputs the timing and speed of charging.

[0080] The storage unit incorporates speakers and microphones to optimize audio input and output. For example, the storage unit may incorporate high-quality speakers to provide clear audio. The storage unit can also optimize audio input and output using generative AI. For example, the storage unit may incorporate a microphone with noise-canceling capabilities to remove ambient noise. The storage unit may also employ multiple microphones to accurately capture the user's voice. This optimizes audio input and output, improving the quality of conversation. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can optimize audio input and output using an AI model that takes audio data as input and outputs noise cancellation and audio optimization.

[0081] The analysis unit estimates the user's emotions and adjusts the analysis algorithm based on the estimated emotions. For example, if the user is relaxed, the analysis unit adjusts the analysis algorithm to a relaxed pace. The analysis unit estimates the user's emotions using emotion estimation functions, such as an emotion engine or generative AI. For example, the analysis unit estimates the user's emotions using facial recognition technology. The analysis unit can also estimate the user's emotions using voice analysis technology. The analysis unit adjusts the analysis algorithm based on the estimated emotions of the user. For example, if the user is in a hurry, the analysis algorithm is adjusted to process quickly. If the user is excited, the analysis algorithm is adjusted to produce visually stimulating results. By adjusting the analysis algorithm according to the user's emotions, more appropriate analysis results can be obtained. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform adjustments using an AI model that takes user emotion data as input and outputs adjustments to the analysis algorithm.

[0082] The analysis unit refers to the user's past conversation history and generates more appropriate responses. For example, the analysis unit generates appropriate responses based on phrases the user has used in the past. The analysis unit can also use generative AI to analyze the user's past conversation history and generate appropriate responses. For example, the analysis unit uses text generation AI to analyze the user's past conversation history and generate appropriate responses. The analysis unit can also use speech generation AI to analyze the user's past conversation history and provide responses in voice. The analysis unit extracts preferred topics based on the user's past conversation history and generates responses based on them. This allows for more natural conversations by referring to past conversation history. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the user's past conversation history data as input and outputs appropriate responses.

[0083] The analysis unit analyzes the user's pronunciation and intonation and provides advice for improving pronunciation. For example, the analysis unit analyzes the user's pronunciation and provides correct pronunciation in audio. The analysis unit can also use generative AI to analyze the user's pronunciation and intonation and provide appropriate advice. For example, the analysis unit uses speech generation AI to analyze the user's pronunciation and provides correct pronunciation in audio. The analysis unit can also use text generation AI to analyze the user's pronunciation and intonation and provide advice in text. The analysis unit analyzes the user's intonation and provides natural intonation in audio. This improves the effectiveness of language learning by providing advice for improving pronunciation and intonation. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes user pronunciation data as input and outputs advice for improving pronunciation.

[0084] The analysis unit estimates the user's emotions and adjusts the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the analysis unit provides a simple and highly visible display method. The analysis unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the analysis unit estimates the user's emotions using facial recognition technology. The analysis unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions of the user, the analysis unit adjusts the display method of the analysis results. For example, if the user is relaxed, it provides a display method that includes detailed information. If the user is in a hurry, it provides a display method that focuses on the essentials. By adjusting the display method according to the user's emotions, visibility is improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform adjustments using an AI model that takes user emotion data as input and outputs adjustments to the display method.

[0085] The analysis unit analyzes the user's gestures and facial expressions to gain a deeper understanding of the conversation context. For example, the analysis unit analyzes the user's gestures to understand the intent of the conversation. The analysis unit can also use generative AI to analyze the user's gestures and facial expressions to understand the conversation context. For example, the analysis unit uses image generation AI to analyze the user's gestures to understand the intent of the conversation. The analysis unit can also use facial expression recognition technology to analyze the user's facial expressions to understand their emotions. The analysis unit gains a deeper understanding of the conversation context based on the user's gestures and facial expressions. This allows for a deeper understanding of the conversation context by analyzing gestures and facial expressions. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes user gesture and facial expression data as input and outputs an analysis for understanding the conversation context.

[0086] The analysis unit analyzes multiple languages ​​simultaneously to support multilingual conversations. For example, if a user uses multiple languages, the analysis unit analyzes them simultaneously and generates an appropriate response. The analysis unit can also use generative AI to analyze multiple languages ​​simultaneously and generate an appropriate response. For example, the analysis unit can use text generation AI to analyze multiple languages ​​and generate an appropriate response. Alternatively, the analysis unit can use speech generation AI to analyze multiple languages ​​and provide a voice response. If a user speaks in different languages, the analysis unit analyzes each language and generates a response. This enables multilingual conversations by analyzing multiple languages ​​simultaneously. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes multiple language data as input and outputs an appropriate response.

[0087] The conversation support unit estimates the user's emotions and adjusts the tone and content of its responses based on the estimated emotions. For example, if the user is relaxed, the conversation support unit will respond in a calm tone. The conversation support unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the conversation support unit can estimate the user's emotions using facial recognition technology. The conversation support unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions of the user, the conversation support unit adjusts the tone and content of its responses. For example, if the user is excited, it will respond in a lively tone. If the user is sad, it will respond in a gentle tone. This allows for more natural conversation by adjusting the tone and content of responses according to the user's emotions. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can perform adjustments using an AI model that takes user emotion data as input and outputs the tone and content of responses.

[0088] The conversation support unit learns the user's past conversation patterns to provide more natural conversations. For example, the conversation support unit learns the user's past conversation patterns and generates natural responses. The conversation support unit can also use generative AI to learn the user's past conversation patterns and generate appropriate responses. For example, the conversation support unit can use text generation AI to learn the user's past conversation patterns and generate appropriate responses. The conversation support unit can also use speech generation AI to learn the user's past conversation patterns and provide responses in voice. The conversation support unit learns the user's preferred topics and proceeds with the conversation based on them. This makes more natural conversations possible by learning past conversation patterns. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can learn using an AI model that takes the user's past conversation pattern data as input and outputs appropriate responses.

[0089] The conversation support unit suggests conversation topics based on the user's interests. For example, the conversation support unit learns the user's interests and suggests conversation topics based on them. The conversation support unit can also use generative AI to analyze the user's interests and suggest appropriate topics. For example, the conversation support unit can use text generation AI to analyze the user's interests and suggest appropriate topics. The conversation support unit can also use speech generation AI to analyze the user's interests and suggest topics in speech. The conversation support unit analyzes the user's past conversation history and suggests the most suitable topics. This makes conversations more enjoyable by suggesting topics based on the user's interests. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can make suggestions using an AI model that takes user interest data as input and outputs appropriate topics.

[0090] The conversation support unit estimates the user's emotions and adjusts the conversation pace based on the estimated emotions. For example, if the user is relaxed, the conversation support unit will proceed at a relaxed pace. The conversation support unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the conversation support unit can estimate the user's emotions using facial recognition technology. The conversation support unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions, the conversation support unit adjusts the conversation pace. For example, if the user is in a hurry, the conversation will proceed quickly. If the user is excited, the conversation will proceed at a lively pace. By adjusting the conversation pace according to the user's emotions, a more appropriate conversation becomes possible. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can make adjustments using an AI model that takes user emotion data as input and outputs the conversation pace.

[0091] The conversation support unit analyzes the ambient sounds surrounding the user and provides appropriate conversation content. For example, if the surroundings are quiet, the conversation support unit provides detailed conversation content. The conversation support unit can also use generative AI to analyze ambient sounds and provide appropriate conversation content. For example, the conversation support unit uses speech generation AI to analyze ambient sounds and provide appropriate conversation content. The conversation support unit can also use noise cancellation technology to remove ambient noise and provide appropriate conversation content. If the surroundings are noisy, the conversation support unit provides concise conversation content. In this way, appropriate conversation content can be provided by analyzing ambient sounds. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can perform analysis using an AI model that takes ambient sound data as input and outputs appropriate conversation content.

[0092] The conversation support unit refers to the user's schedule and starts a conversation at an appropriate time. For example, the conversation support unit refers to the user's schedule and starts a conversation during an available time. The conversation support unit can also use generative AI to analyze the user's schedule and start a conversation at an appropriate time. For example, the conversation support unit uses schedule management AI to analyze the user's schedule and start a conversation at the optimal time. The conversation support unit can also start a conversation before an important appointment based on the user's schedule. This allows for more effective conversations by starting conversations based on the user's schedule. Some or all of the above processing in the conversation support unit may be performed using AI, for example, or without AI. For example, the conversation support unit can perform analysis using an AI model that takes user schedule data as input and outputs the timing of when to start a conversation.

[0093] The interpreting unit estimates the user's emotions and adjusts the interpretation based on the estimated emotions. For example, if the user is relaxed, the interpreting unit will interpret in a calm tone. The interpreting unit estimates the user's emotions using emotion estimation functions, such as an emotion engine or generative AI. For example, the interpreting unit may estimate the user's emotions using facial recognition technology. The interpreting unit can also estimate the user's emotions using speech analysis technology. The interpreting unit adjusts the interpretation based on the estimated emotions of the user. For example, if the user is excited, the interpreting unit will interpret in a lively tone. If the user is sad, the interpreting unit will interpret in a gentle tone. By adjusting the interpretation according to the user's emotions, a more natural interpretation becomes possible. Some or all of the above processing in the interpreting unit may be performed using AI, for example, or without AI. For example, the interpreting unit can make adjustments using an AI model that takes user emotion data as input and outputs an interpretation expression.

[0094] The interpretation unit refers to the user's past interpretation history to provide a more appropriate interpretation. For example, the interpretation unit refers to the user's past interpretation history and selects appropriate expressions. The interpretation unit can also use generative AI to analyze the user's past interpretation history and provide an appropriate interpretation. For example, the interpretation unit uses text generation AI to analyze the user's past interpretation history and provide an appropriate interpretation. The interpretation unit can also use speech generation AI to analyze the user's past interpretation history and provide an interpretation in speech. The interpretation unit provides the optimal interpretation based on the user's past interpretation history. This makes it possible to provide a more appropriate interpretation by referring to past interpretation history. Some or all of the above processing in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform analysis using an AI model that takes the user's past interpretation history data as input and outputs an appropriate interpretation.

[0095] The interpretation unit estimates the user's emotions and determines the interpretation priority based on the estimated emotions. For example, if the user is in an urgent situation, the interpretation unit will set a high priority. The interpretation unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the interpretation unit will estimate the user's emotions using facial recognition technology. The interpretation unit can also estimate the user's emotions using speech analysis technology. Based on the estimated emotions, the interpretation unit determines the interpretation priority. For example, if the user is relaxed, the interpretation priority will be set low. If the user is excited, the interpretation priority will be set to medium. This allows for more effective interpretation by determining the interpretation priority according to the user's emotions. Some or all of the above processing in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can make decisions using an AI model that takes user emotion data as input and outputs interpretation priority.

[0096] The interpretation unit incorporates region-specific expressions into the interpretation based on the user's geographical location information. For example, if the user is in a specific region, the interpretation unit will incorporate region-specific expressions into the interpretation. The interpretation unit can also use generative AI to analyze the user's geographical location information and incorporate appropriate expressions into the interpretation. For example, the interpretation unit can use location information analysis AI to analyze the user's geographical location information and incorporate region-specific expressions into the interpretation. Furthermore, if the user moves to a different region, the interpretation unit can also incorporate expressions from the new region into the interpretation. This allows for more natural interpretation by incorporating region-specific expressions. Some or all of the above processing in the interpretation unit may be performed using AI, for example, or without AI. For example, the interpretation unit can perform analysis using an AI model that takes the user's geographical location data as input and outputs appropriate expressions.

[0097] The learning unit estimates the user's emotions and adjusts the learning content based on the estimated emotions. For example, if the user is relaxed, the learning unit provides learning content at a relaxed pace. The learning unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the learning unit estimates the user's emotions using facial recognition technology. The learning unit can also estimate the user's emotions using voice analysis technology. Based on the estimated emotions of the user, the learning unit adjusts the learning content. For example, if the user is in a hurry, it provides learning content quickly. If the user is excited, it provides visually stimulating learning content. By adjusting the learning content according to the user's emotions, more effective learning becomes possible. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform adjustments using an AI model that takes user emotion data as input and outputs adjusted learning content.

[0098] The learning unit refers to the user's past learning history and provides an optimal learning plan. For example, the learning unit refers to the user's past learning history and provides an optimal learning plan. The learning unit can also use generative AI to analyze the user's past learning history and provide an appropriate learning plan. For example, the learning unit uses text generation AI to analyze the user's past learning history and provide an appropriate learning plan. The learning unit can also use speech generation AI to analyze the user's past learning history and provide a learning plan in speech. The learning unit provides an individually customized learning plan based on the user's past learning history. This allows for the provision of a more effective learning plan by referring to past learning history. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can perform analysis using an AI model that takes the user's past learning history data as input and outputs an appropriate learning plan.

[0099] The learning unit estimates the user's emotions and adjusts the learning pace based on the estimated emotions. For example, if the user is relaxed, the learning unit will proceed at a relaxed pace. The learning unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the learning unit may estimate the user's emotions using facial recognition technology. The learning unit can also estimate the user's emotions using speech analysis technology. Based on the estimated emotions, the learning unit adjusts the learning pace. For example, if the user is in a hurry, the learning will proceed quickly. If the user is excited, the learning will proceed at a lively pace. By adjusting the learning pace according to the user's emotions, more effective learning becomes possible. Some or all of the above processing in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can make adjustments using an AI model that takes user emotion data as input and outputs the learning pace.

[0100] The learning unit suggests new learning topics based on the user's interests. For example, the learning unit learns the user's interests and suggests new learning topics based on them. The learning unit can also use generative AI to analyze the user's interests and suggest appropriate learning topics. For example, the learning unit can use text generation AI to analyze the user's interests and suggest appropriate learning topics. Alternatively, the learning unit can use speech generation AI to analyze the user's interests and suggest learning topics in speech. The learning unit analyzes the user's past learning history and suggests optimal new learning topics. This improves the user's motivation to learn by suggesting new learning topics based on their interests. Some or all of the above processes in the learning unit may be performed using AI, for example, or without AI. For example, the learning unit can make suggestions using an AI model that takes user interest data as input and outputs appropriate learning topics.

[0101] The character unit estimates the user's emotions and adjusts the character's facial expressions and movements based on the estimated emotions. For example, if the user is relaxed, the character will have a calm expression. The character unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the character unit estimates the user's emotions using facial recognition technology. The character unit can also estimate the user's emotions using voice analysis technology. The character unit adjusts the character's facial expressions and movements based on the estimated emotions. For example, if the user is excited, the character will move energetically. If the user is sad, the character will have a gentle expression. In this way, a more approachable character is provided by adjusting the character's facial expressions and movements according to the user's emotions. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can perform adjustments using an AI model that takes user emotion data as input and outputs the character's facial expressions and movements.

[0102] The character unit refers to the user's past character selection history and proposes the most suitable character. For example, the character unit refers to the user's past character selection history and proposes the most suitable character. The character unit can also use generative AI to analyze the user's past character selection history and propose an appropriate character. For example, the character unit uses text generation AI to analyze the user's past character selection history and propose an appropriate character. The character unit can also use image generation AI to analyze the user's past character selection history and propose a character as an image. The character unit proposes a character that is individually customized based on the user's past character selection history. In this way, the most suitable character is proposed to the user by referring to their past character selection history. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can perform analysis using an AI model that takes the user's past character selection history data as input and outputs an appropriate character.

[0103] The character unit estimates the user's emotions and changes the character's voice and speaking style based on the estimated emotions. For example, if the user is relaxed, the character unit will speak in a calm voice. The character unit estimates the user's emotions using emotion estimation functions, such as an emotion engine or generative AI. For example, the character unit can estimate the user's emotions using facial recognition technology. The character unit can also estimate the user's emotions using voice analysis technology. The character unit changes the character's voice and speaking style based on the estimated emotions of the user. For example, if the user is excited, the character will speak in a lively voice. If the user is sad, the character will speak in a gentle voice. In this way, a more approachable character is provided by changing the character's voice and speaking style according to the user's emotions. Some or all of the above processing in the character unit may be performed using AI, for example, or without AI. For example, the character unit can make adjustments using an AI model that takes user emotion data as input and outputs the character's voice and speaking style.

[0104] The character department changes the character's costume and accessories according to the season and events. For example, the character department changes the character's costume according to the season. The character department can also design costumes and accessories according to the season and events using generative AI. For example, the character department designs seasonal costumes using image generation AI. The character department can also design character accessories to match events. The character department customizes the character's appearance to match specific events. This makes it easier to attract user interest by changing the character's appearance according to the season and events. Some or all of the above processing in the character department may be done using AI, for example, or not using AI. For example, the character department can design using an AI model that takes seasonal and event data as input and outputs character costumes and accessories.

[0105] The dedicated device unit estimates the user's emotions and adjusts the operating mode of the dedicated device based on the estimated user emotions. For example, if the user is relaxed, the dedicated device unit switches to a calm operating mode. The dedicated device unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the dedicated device unit estimates the user's emotions using facial recognition technology. The dedicated device unit can also estimate the user's emotions using voice analysis technology. The dedicated device unit adjusts the operating mode of the dedicated device based on the estimated user emotions. For example, if the user is in a hurry, the dedicated device switches to a fast operating mode. If the user is excited, the dedicated device switches to a lively operating mode. By adjusting the operating mode of the dedicated device according to the user's emotions, more effective operation becomes possible. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform adjustments using an AI model that takes user emotion data as input and outputs the operating mode of the dedicated device.

[0106] The dedicated device unit refers to the user's past usage history and provides the optimal device settings. For example, the dedicated device unit refers to the user's past usage history and provides the optimal device settings. The dedicated device unit can also use generative AI to analyze the user's past usage history and provide appropriate device settings. For example, the dedicated device unit uses text generation AI to analyze the user's past usage history and provide appropriate device settings. The dedicated device unit also uses device information analysis AI to analyze the user's past usage history and provide optimal device settings. The dedicated device unit provides individually customized device settings based on the user's past usage history. This ensures that the optimal device settings are provided by referring to the past usage history. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform analysis using an AI model that takes the user's past usage history data as input and outputs appropriate device settings.

[0107] The dedicated device unit estimates the user's emotions and adjusts the operating procedure of the dedicated device based on the estimated user emotions. For example, if the user is relaxed, the dedicated device unit provides a calm operating procedure. The dedicated device unit estimates the user's emotions using an emotion estimation function, such as an emotion engine or generative AI. For example, the dedicated device unit estimates the user's emotions using facial recognition technology. The dedicated device unit can also estimate the user's emotions using voice analysis technology. Based on the estimated user emotions, the dedicated device unit adjusts the operating procedure of the dedicated device. For example, if the user is in a hurry, it provides a quick operating procedure. If the user is excited, it provides an energetic operating procedure. By adjusting the operating procedure according to the user's emotions, more effective operation becomes possible. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform adjustments using an AI model that takes user emotion data as input and outputs operating procedures.

[0108] The dedicated device unit provides optimal device settings based on the user's device information. For example, the dedicated device unit refers to the user's device information and provides optimal device settings. The dedicated device unit can also analyze the user's device information using a generation AI and provide appropriate device settings. For example, the dedicated device unit analyzes the user's device information using a device information analysis AI and provides optimal device settings. Furthermore, the dedicated device unit can provide individually customized device settings based on the user's device information. This enables more effective operation by providing optimal device settings based on device information. Some or all of the above processing in the dedicated device unit may be performed using AI, for example, or without AI. For example, the dedicated device unit can perform analysis using an AI model that takes user device information data as input and outputs appropriate device settings.

[0109] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0110] The conversation support system can also analyze the user's voice tone and speed to estimate their emotions. For example, if a user speaks quickly in a high-pitched voice, the system can estimate that the user is excited and respond in a calm tone. Conversely, if a user speaks slowly in a low-pitched voice, the system can estimate that the user is relaxed and respond in an energetic tone. This enables appropriate responses tailored to the user's emotions, resulting in more natural conversations.

[0111] The conversation support system can also learn from the user's past conversation history and suggest conversation topics based on the user's preferences and interests. For example, if the user has talked a lot about sports in the past, the system can suggest new sports-related topics. Similarly, if the user has talked about a particular movie or music, the system can suggest related topics. This enables conversations based on the user's interests, improving the quality of the conversation.

[0112] Conversation support systems can also analyze user gestures and facial expressions to gain a deeper understanding of the conversation context. For example, if a user is smiling while speaking, the system can determine that the user is enjoying the conversation and encourage them to continue. Conversely, if a user is frowning, the system can determine that the user is confused and provide additional explanations. This allows for a better understanding of the user's nonverbal communication and enables more appropriate responses.

[0113] The conversation support system can also refer to the user's schedule and have the functionality to initiate conversations at appropriate times. For example, the system can refrain from initiating conversations during busy periods and start them at times when the user is free. It can also initiate conversations as reminders before important meetings or events to encourage preparation. This enables flexible conversations tailored to the user's schedule, improving user convenience.

[0114] The conversation support system can also be equipped with the ability to analyze the ambient sounds surrounding the user and provide appropriate conversation content. For example, if the surroundings are quiet, the system can provide detailed conversation content, and if the surroundings are noisy, it can provide concise conversation content. In addition, the system can use noise cancellation technology to remove ambient noise and provide conversation with clear voice. This enables appropriate conversation according to the surrounding environment, improving the quality of conversation.

[0115] The conversation support system can also be equipped with the ability to estimate the user's emotions and adjust the pace of the conversation based on those emotions. For example, if the user is relaxed, the system can proceed at a leisurely pace; if the user is in a hurry, it can proceed quickly. Conversely, if the user is excited, it can proceed at a lively pace. This enables appropriate conversational pacing according to the user's emotions, resulting in a more natural conversation.

[0116] The conversation support system can also be equipped with features to learn the user's past conversation patterns and provide more natural conversations. For example, it can generate appropriate responses based on phrases and topics the user has used in the past. It can also learn the user's preferred topics and guide the conversation accordingly. By referring to past conversation patterns, it enables more natural and consistent conversations, improving user satisfaction.

[0117] The conversation support system can also be equipped with features to analyze the user's pronunciation and intonation and provide advice for improving pronunciation. For example, it can analyze the user's pronunciation and provide correct pronunciation in audio. It can also analyze the user's intonation and provide natural intonation in audio. By providing advice on improving pronunciation and intonation, the effectiveness of language learning and the user's skills will improve.

[0118] The conversation support system can also be equipped with the ability to estimate the user's emotions and adjust the tone and content of its responses based on those emotions. For example, if the user is relaxed, the system can respond in a calm tone; if the user is excited, it can respond in a lively tone; and if the user is sad, it can respond in a gentle tone. This enables appropriate responses that match the user's emotions, resulting in more natural conversations.

[0119] The conversation support system can also include a feature that suggests conversation topics based on the user's interests. For example, it can learn the user's interests and suggest new conversation topics based on them. It can also analyze the user's past conversation history and suggest the most suitable topics. This enables conversations based on the user's interests, making conversations more enjoyable.

[0120] The following briefly describes the processing flow for example form 2.

[0121] Step 1: Place a smartphone with a conversation support app installed into the storage compartment, then place it inside the doll or stuffed animal. Step 2: The analysis unit analyzes what the user says and generates an appropriate response. The analysis unit analyzes the user's utterance using, for example, natural language processing technology or generative AI (e.g., text generation AI or multimodal generation AI) and generates an appropriate response. Step 3: The conversation support unit provides the user with the response generated by the analysis unit. The conversation support unit provides the response generated by the analysis unit in voice or text, for example, using speech synthesis technology or generative AI (e.g., speech generation AI or text generation AI).

[0122] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0123] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0124] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0125] Each of the multiple elements described above, including the storage unit, analysis unit, and conversation support unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the storage unit is implemented as part of the smart device 14 and places a smartphone into a doll or stuffed animal. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the user's utterances and generates an appropriate response. The conversation support unit is implemented in the control unit 46A of the smart device 14 and provides the response generated by the analysis unit in voice or text. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0126] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0127] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0128] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0129] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0130] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0131] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0132] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0133] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0134] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0135] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0136] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0137] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0138] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0139] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0140] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0141] Each of the multiple elements described above, including the storage unit, analysis unit, and conversation support unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the storage unit is implemented as part of the smart glasses 214 and places a smartphone into a doll or stuffed animal. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12 and analyzes the user's utterances and generates an appropriate response. The conversation support unit is implemented by the control unit 46A of the smart glasses 214 and provides the response generated by the analysis unit in voice or text. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0142] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0143] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0144] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0145] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0146] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0147] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0148] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0149] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0150] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0151] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0152] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0153] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0154] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0155] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0156] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0157] Each of the multiple elements described above, including the storage unit, analysis unit, and conversation support unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the storage unit is implemented as part of the headset terminal 314 and places the smartphone into a doll or stuffed animal. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the user's utterances and generates an appropriate response. The conversation support unit is implemented in the control unit 46A of the headset terminal 314 and provides the response generated by the analysis unit in voice or text. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0158] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0159] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0160] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0161] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0162] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0163] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0164] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0165] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0166] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0167] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0168] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0169] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0170] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0171] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0172] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0173] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0174] Each of the multiple elements described above, including the storage unit, analysis unit, and conversation support unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the storage unit is implemented as part of the robot 414 and places a smartphone into a doll or stuffed animal. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the user's statements and generates an appropriate response. The conversation support unit is implemented, for example, by the control unit 46A of the robot 414, which provides the response generated by the analysis unit in voice or text. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0175] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0176] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0177] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0178] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0179] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0180] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0181] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0182] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0183] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0184] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0185] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0186] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0187] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0188] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0189] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0190] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0191] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0192] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0193] (Note 1) A storage compartment for dolls or stuffed animals where a smartphone with a conversation support app installed can be placed, An analysis unit that analyzes what the user says and generates an appropriate response, The system includes a conversation support unit that provides the user with a response generated by the analysis unit. A system characterized by the following features. (Note 2) The aforementioned conversation support unit, It has an interpreting department that provides interpretation services. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned conversation support unit, It has a learning section that supports children's language learning. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned storage compartment is The company has a character division that creates robots based on favorite characters through corporate collaborations. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned storage compartment is It has a dedicated equipment division that develops specialized equipment. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned storage compartment is The system estimates the user's emotions and adjusts the opening and closing timing of the storage compartment based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned storage compartment is It automatically adjusts the internal temperature and humidity to maintain the optimal operating environment for your smartphone. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned storage compartment is By installing sensors, the position and orientation of the smartphone are automatically adjusted. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned storage compartment is The system estimates the user's emotions and changes the design and color of the storage compartments based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned storage compartment is Added a charging function that automatically charges smartphones while they are stored inside. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned storage compartment is It has built-in speakers and microphones to optimize audio input and output. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, It estimates the user's emotions and adjusts the analysis algorithm based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, Referencing the user's past conversation history generates more appropriate responses. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, It analyzes the user's pronunciation and intonation and provides advice on how to improve their pronunciation. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, By analyzing the user's gestures and facial expressions, we gain a deeper understanding of the conversational context. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, It analyzes multiple languages ​​simultaneously and supports multilingual conversations. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned conversation support unit, It estimates the user's emotions and adjusts the tone and content of the response based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned conversation support unit, It learns the user's past conversation patterns to provide more natural conversations. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned conversation support unit, Suggest conversation topics based on the user's interests and preferences. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned conversation support unit, It estimates the user's emotions and adjusts the conversation speed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned conversation support unit, It analyzes ambient sounds around the user and provides appropriate conversation content. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned conversation support unit, Refer to the user's schedule and start a conversation at the appropriate time. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned interpretation department, The system estimates the user's emotions and adjusts the translation's expression based on those estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 25) The aforementioned interpretation department, Referencing the user's past interpreting history provides more appropriate interpretations. The system described in Appendix 2, characterized by the features described herein. (Note 26) The aforementioned interpretation department, The system estimates the user's emotions and determines the priority of interpretation based on those estimated emotions. The system described in Appendix 2, characterized by the features described herein. (Note 27) The aforementioned interpretation department, Based on the user's geographical location information, region-specific expressions are reflected in the interpretation. The system described in Appendix 2, characterized by the features described herein. (Note 28) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning content based on the estimated user emotions. The system described in Appendix 3, characterized by the features described herein. (Note 29) The aforementioned learning unit, It provides an optimal learning plan by referring to the user's past learning history. The system described in Appendix 3, characterized by the features described herein. (Note 30) The aforementioned learning unit, It estimates the user's emotions and adjusts the learning rate based on the estimated user emotions. The system described in Appendix 3, characterized by the features described herein. (Note 31) The aforementioned learning unit, Based on the user's interests, we suggest new learning topics. The system described in Appendix 3, characterized by the features described herein. (Note 32) The aforementioned character section is It estimates the user's emotions and adjusts the character's facial expressions and movements based on those estimated emotions. The system described in Appendix 4, characterized by the features described herein. (Note 33) The aforementioned character section is The system refers to the user's past character selection history and suggests the most suitable character. The system described in Appendix 4, characterized by the features described herein. (Note 34) The aforementioned character section is It estimates the user's emotions and changes the character's voice and speaking style based on those estimated emotions. The system described in Appendix 4, characterized by the features described herein. (Note 35) The aforementioned character section is Change the character's costume and accessories according to the season and events. The system described in Appendix 4, characterized by the features described herein. (Note 36) The aforementioned dedicated equipment unit is It estimates the user's emotions and adjusts the operating mode of the dedicated device based on the estimated user emotions. The system described in Appendix 5, characterized by the features described herein. (Note 37) The aforementioned dedicated equipment unit is Referencing the user's past usage history provides optimal device settings. The system described in Appendix 5, characterized by the features described herein. (Note 38) The aforementioned dedicated equipment unit is It estimates the user's emotions and adjusts the operating procedures of the dedicated device based on the estimated user emotions. The system described in Appendix 5, characterized by the features described herein. (Note 39) The aforementioned dedicated equipment unit is Based on the user's device information, we provide optimal device settings. The system described in Appendix 5, characterized by the features described herein. [Explanation of Symbols]

[0194] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A storage compartment for dolls or stuffed animals where a smartphone with a conversation support app installed can be placed, An analysis unit that analyzes what the user says and generates an appropriate response, The system includes a conversation support unit that provides the user with a response generated by the analysis unit. A system characterized by the following features.

2. The aforementioned conversation support unit, It has an interpreting department that provides interpretation services. The system according to feature 1.

3. The aforementioned conversation support unit, It has a learning section that supports children's language learning. The system according to feature 1.

4. The aforementioned storage compartment is The company has a character division that creates robots based on favorite characters through corporate collaborations. The system according to feature 1.

5. The aforementioned storage compartment is It has a dedicated equipment division that develops specialized equipment. The system according to feature 1.

6. The aforementioned storage compartment is The system estimates the user's emotions and adjusts the opening and closing timing of the storage compartment based on the estimated user emotions. The system according to feature 1.

7. The aforementioned storage compartment is It automatically adjusts the internal temperature and humidity to maintain the optimal operating environment for your smartphone. The system according to feature 1.

8. The aforementioned storage compartment is By installing sensors, the position and orientation of the smartphone are automatically adjusted. The system according to feature 1.

9. The aforementioned storage compartment is The system estimates the user's emotions and changes the design and color of the storage compartment based on the estimated emotions. The system according to feature 1.

10. The aforementioned storage compartment is Added a charging function that automatically charges smartphones while they are stored inside. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A