system

An AI-powered conversational trainer system addresses the challenge of busy caregivers by providing consistent language teaching and assessment, enhancing children's language development and offering parents feedback.

JP2026073132APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Caregivers face challenges in supporting children's language development when they are busy, and maintaining a consistent teaching method is difficult.

Method used

An AI-powered conversational trainer system that includes a prompt setting unit, response analysis unit, and conversation generation unit to provide consistent language teaching and assessment, even when caregivers are busy.

Benefits of technology

Supports children's language development by providing regular, consistent language teaching and assessment, ensuring accurate word usage and emotional expression learning, and offering parents specific feedback on their child's progress.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073132000001_ABST
    Figure 2026073132000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to support a child's language development and provide a consistent teaching method, even when parents are busy. [Solution] The system according to the embodiment comprises a prompt setting unit, a response analysis unit, a conversation generation unit, and a provision unit. The prompt setting unit sets prompts to speak to the child. The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. The conversation generation unit generates the next conversation based on the response analyzed by the response analysis unit. The provision unit provides the conversation generated by the conversation generation unit to the child.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] ,

[0006] , , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003] <00XXXXX16>

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there are problems that it is difficult for a caregiver to support a child's language development when the caregiver is busy, and it is difficult to maintain a consistent teaching method.

[0005] The system according to the embodiment aims to support a child's language development even when the caregiver is busy and provide a consistent teaching method.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a prompt setting unit, a response analysis unit, a conversation generation unit, and a provision unit. The prompt setting unit sets prompts to speak to the child. The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. The conversation generation unit generates the next conversation based on the response analyzed by the response analysis unit. The provision unit provides the conversation generated by the conversation generation unit to the child. [Effects of the Invention]

[0007] The system according to this embodiment can support a child's language development and provide a consistent teaching method, even when parents are busy. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) An AI-powered conversational trainer according to an embodiment of the present invention is a system for infants who have just begun to speak and their guardians. This system allows children to regularly learn new words even when their guardians are busy, by having the AI ​​speak to them periodically and progressing the conversation based on their responses. The AI ​​also consistently teaches accurate word usage, synonyms, antonyms, nuances, and emotional expressions. Furthermore, the AI ​​analyzes the child's statements and responses and regularly reports on their vocabulary, comprehension, and speaking ability development levels. This allows guardians to receive specific feedback and advice to guide their next steps. For example, the AI ​​can set prompts when speaking to the child, such as "What did you do today?" or "What's your favorite food?". When the child answers, the AI ​​analyzes their response and generates further questions or comments. This allows children to learn new words in a natural conversational flow. Next, the AI ​​consistently teaches word usage. For example, when teaching the word "dog," it provides specific examples such as, "A dog is an animal. Dogs bark." The AI ​​also teaches the difference between "dog" and "cat," as well as synonyms and antonyms for the word "dog." This allows children to deeply understand the meaning and usage of words. Furthermore, the AI ​​analyzes the child's speech and responses and assesses their developmental level. For example, if a child says "dog," the AI ​​analyzes the pronunciation and context to assess whether it is pronounced correctly and used in an appropriate context. This assessment is regularly provided to parents as a report, which includes specific feedback and advice. Parents can use this report to plan the next steps. This ensures that children have regular opportunities to learn new words, even when their parents are busy. In addition, teaching words in a consistent manner prevents confusion and improves the child's language comprehension. Moreover, the developmental level assessment and feedback allow parents to specifically understand their child's language development and provide appropriate support. In this way, the AI ​​conversational trainer can effectively support the language development of young children and provide specific feedback to parents.

[0029] The AI ​​conversational trainer according to this embodiment comprises a prompt setting unit, a response analysis unit, a conversation generation unit, and a provision unit. The prompt setting unit sets prompts to speak to the child. The prompt setting unit sets, for example, specific questions and comments to be asked when speaking to the child. For example, the prompt setting unit can set questions such as "What did you do today?" or "What's your favorite food?" The prompt setting unit can also set questions and comments that are appropriate for the child's age and interests. For example, the prompt setting unit can set simple questions for toddlers and more complex questions for slightly older children. The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. The response analysis unit analyzes, for example, the child's pronunciation and context, and evaluates whether the pronunciation is accurate and whether the word is used in an appropriate context. For example, if the child says "dog," the response analysis unit can analyze the pronunciation and context, and evaluate whether the pronunciation is accurate and whether the word is used in an appropriate context. The response analysis unit can also analyze the child's response in real time and provide data for generating the next conversation. The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. The conversation generation unit generates the next question or comment based on the child's response, for example. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, the conversation generation unit can generate the next conversation based on a topic the child has shown interest in. The provision unit provides the conversation generated by the conversation generation unit to the child. The provision unit provides the generated conversation to the child in the form of voice or text, for example. For example, the provision unit can provide the child with the generated conversation in voice using speech synthesis technology. The provision unit can also display the generated conversation as text so that the child can read it. This allows the AI ​​conversational trainer according to the embodiment to teach language to the child in a consistent manner, assess their progress, and provide feedback to the parent.

[0030] The prompt setting unit sets prompts to use when speaking to children. For example, it sets specific questions and comments to ask children. Specifically, the prompt setting unit collects and analyzes the child's profile information in advance to set questions and comments appropriate to the child's age and interests. For example, the prompt setting unit generates appropriate prompts based on information such as the child's age, gender, topics of interest, and past conversation history. For toddlers, simple questions such as "What did you do today?" or "What's your favorite food?" can be set, while for slightly older children, more complex questions such as "What book have you read recently?" or "What are your dreams for the future?" can be set. The prompt setting unit can also set questions appropriate to the season or event. For example, during Christmas, a question such as "What would you like to ask Santa for?" can be set, and during summer vacation, a question such as "What are your plans for summer vacation?" can be set. This allows the prompt setting unit to flexibly set prompts that will capture the child's interest and promote natural conversation. Furthermore, the prompt setting unit also has a function to evaluate the effectiveness of the set prompts and modify them as needed. For example, if a child's response to a particular prompt is unsatisfactory, the prompt can be modified and changed to a more appropriate question or comment. This allows the prompt setting unit to continue providing prompts that facilitate more effective conversations with children.

[0031] The response analysis unit analyzes the child's responses based on prompts set by the prompt setting unit. For example, the response analysis unit analyzes the child's pronunciation and context to evaluate whether the pronunciation is accurate and whether the words are used in an appropriate context. Specifically, the response analysis unit uses speech recognition technology to convert the child's utterances into text data and then analyzes that text data. For example, if the child says "dog," the unit analyzes the speech data at the phoneme level to evaluate whether the pronunciation is accurate and compares it to standard pronunciation. It also uses contextual analysis technology to evaluate whether the child's utterances are used in an appropriate context. For example, if the child says "dog," the unit analyzes the conversation before and after the utterance to determine whether the word "dog" is used in an appropriate context. Furthermore, the response analysis unit can analyze the child's responses in real time and provide data for generating the next conversation. For example, if the child says "dog," it provides data to the conversation generation unit to generate the next question or comment based on that response. This allows the response analysis unit to accurately evaluate the child's pronunciation and context and provide data to facilitate the next conversation. In addition, the response analysis unit can accumulate the child's responses and build a database to evaluate long-term growth. For example, the system records changes in a child's pronunciation and context over time and analyzes growth trends. This allows the response analysis unit to continuously evaluate the child's language development and provide feedback to parents.

[0032] The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. For example, the conversation generation unit generates the next question or comment based on the child's response. Specifically, the conversation generation unit uses natural language processing technology to generate appropriate conversations in response to the child's response. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, if the child says "I like dogs," the conversation generation unit can generate the question, "What do you like about dogs?" to pique the child's interest. Furthermore, the conversation generation unit can adjust the tone and style of the conversation based on the child's response. For example, it can speak in a gentle tone to toddlers and generate more complex conversations for slightly older children. In this way, the conversation generation unit can generate appropriate conversations according to the child's age and interests, promoting natural conversation. In addition, the conversation generation unit has a function to evaluate the effectiveness of the generated conversations and improve the conversation generation algorithm as needed. For example, if the child's response to a particular conversation pattern is unfavorable, it can modify that pattern to generate more appropriate conversations. This allows the conversation generation unit to continuously provide conversations that facilitate more effective communication with children.

[0033] The delivery unit provides children with conversations generated by the conversation generation unit. The delivery unit provides children with generated conversations in the form of audio or text. Specifically, the delivery unit can provide children with audio conversations generated using speech synthesis technology. For example, the delivery unit uses a speech synthesis engine with natural pronunciation and intonation to convert generated conversations into audio in real time. The delivery unit can also display the generated conversations as text so that children can read them. For example, the delivery unit can display the generated conversations on the screen of a tablet or smartphone so that children can visually confirm them. Furthermore, the delivery unit can adjust the method of delivery (audio or text) according to the child's response. For example, if the child prefers audio conversations, it can provide them in audio; if the child prefers text conversations, it can provide them in text. This allows the delivery unit to provide conversations in the most optimal way according to the child's preferences, promoting natural communication. In addition, the delivery unit has a function to evaluate the effectiveness of the delivered conversations and improve the delivery method as needed. For example, if a child's response to a particular delivery method is poor, it can modify that method and change to a more appropriate delivery method. This allows the delivery unit to continue providing conversations that facilitate more effective communication with children.

[0034] The prompt setting unit can set specific questions and comments to be asked when speaking to a child. For example, the prompt setting unit can set questions such as, "What did you do today?" or "What's your favorite food?" The prompt setting unit can also set questions and comments that are appropriate for the child's age and interests. For example, the prompt setting unit can set simple questions for toddlers and more complex questions for slightly older children. This promotes language learning by providing children with appropriate questions and comments. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or not using AI. For example, the prompt setting unit can set questions and comments using an AI model that generates prompts based on the child's age and interests.

[0035] The response analysis unit can analyze a child's pronunciation and context to evaluate whether they are pronouncing words accurately and using them in an appropriate context. For example, if a child says "dog," the response analysis unit can analyze the pronunciation and context to evaluate whether they are pronouncing words accurately and using them in an appropriate context. The response analysis unit can also analyze a child's response in real time and provide data to generate the next conversation. This supports language learning by evaluating the appropriateness of the child's pronunciation and context. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can take a child's pronunciation data as input and use an AI model to evaluate the accuracy of pronunciation.

[0036] The conversation generation unit can generate the next question or comment based on the child's response. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, the conversation generation unit can generate the next conversation based on a topic the child has shown interest in. This maintains a natural flow of conversation by generating appropriate conversations that respond to the child's response. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate a conversation using an AI model that takes the child's response data as input and generates the next question or comment.

[0037] The service provider can provide the generated conversations to the child. The service provider can provide the generated conversations to the child in the form of voice or text. For example, the service provider can provide the child with voice generated conversations using speech synthesis technology. The service provider can also display the generated conversations as text so that the child can read them. This promotes language learning by providing the child with generated conversations. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take generated conversation data as input and provide the conversation in voice using an AI model that performs speech synthesis.

[0038] The response analysis unit can analyze a child's statements and responses to evaluate their growth level in vocabulary, comprehension, and speaking ability. For example, if a child says "dog," the response analysis unit can analyze the pronunciation and context to evaluate whether it is pronounced correctly and used in an appropriate context. The response analysis unit can also periodically analyze a child's statements and responses to evaluate their growth level. This allows for the provision of appropriate feedback and advice by evaluating the child's growth level. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can use a child's statement data as input and evaluate their growth level using an AI model that evaluates vocabulary and comprehension.

[0039] The service provider can provide parents with a report of the assessment results of their child's growth level. For example, the service provider can provide parents with a report of the assessment results of their child's growth level. For example, the service provider can create a detailed report of the assessment results and provide it to the parents. The service provider can also include specific feedback and advice based on the assessment results in the report. For example, the service provider can assess the growth level of a child's vocabulary and comprehension and report the results to the parents. This allows parents to understand their child's growth level and provide appropriate support. Some or all of the above processing in the service provider may be performed using AI, or not. For example, the service provider can use an AI model that takes growth level assessment data as input and generates a report to provide parents with the assessment results.

[0040] The prompt setting unit can analyze a child's past response history and select the optimal prompt. For example, the prompt setting unit can set prompts based on topics the child has shown interest in in the past. The prompt setting unit can also prioritize question formats that the child has found easy to answer in the past. For example, the prompt setting unit can set prompts to avoid questions that the child has found difficult in the past. This promotes effective language learning by providing optimal prompts based on past response history. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's past response data as input and selects the optimal prompt.

[0041] The prompt setting unit can filter prompts based on the child's current interests and concerns. For example, the prompt setting unit can set questions related to characters the child has recently become interested in. It can also set prompts based on television programs the child has recently watched. For example, the prompt setting unit can set questions related to toys the child has recently been playing with. This enhances the child's motivation to learn by providing prompts based on their interests and concerns. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's interest data as input and filters the prompts.

[0042] The prompt setting unit can prioritize setting highly relevant prompts by considering the child's geographical location when setting prompts. For example, the prompt setting unit can set questions related to parks if the child is in a park. It can also set questions related to household items if the child is at home. For example, it can set questions related to school if the child is at school. This makes it easier to capture the child's interest by providing prompts based on geographical location information. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's geographical location data as input and selects highly relevant prompts.

[0043] The prompt setting unit can analyze a child's social media activity and set relevant prompts when setting prompts. For example, the prompt setting unit can set questions related to posts the child has recently "liked." The prompt setting unit can also set prompts based on what the child has recently commented on. For example, the prompt setting unit can set questions related to accounts the child follows. This makes it easier to capture the child's interest by providing prompts based on their social media activity. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's social media activity data as input and selects relevant prompts.

[0044] The reaction analysis unit can detect subtle changes in a child's pronunciation during reaction analysis, thereby improving the accuracy of the analysis. For example, the reaction analysis unit can detect changes in a child's pronunciation at the phoneme level. The reaction analysis unit can also analyze changes in the rhythm and intonation of a child's pronunciation. For example, the reaction analysis unit can detect changes in the volume and speed of a child's pronunciation. This improves the accuracy of the analysis by detecting subtle changes in pronunciation. Some or all of the above-described processes in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can take a child's pronunciation data as input and perform pronunciation analysis using an AI model that detects subtle changes.

[0045] The response analysis unit can refer to past conversation history to gain a deeper understanding of the context of a child's statements during response analysis. For example, the response analysis unit can analyze the context based on what the child has said in the past. The response analysis unit can also refer to patterns of words the child has used in the past. For example, the response analysis unit can analyze the context by considering the interests and concerns the child has shown in the past. By referring to past conversation history, the understanding of the context is deepened and the accuracy of the analysis is improved. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can take the child's past conversation data as input and perform utterance analysis using an AI model that understands context.

[0046] The reaction analysis unit can perform analysis while considering the child's geographical background. For example, the reaction analysis unit can consider the dialect and accent of the area where the child lives. It can also consider the educational policies of the school the child attends. For example, the reaction analysis unit can consider the cultural background of places the child frequently visits. This improves the accuracy of the analysis by considering geographical background. Some or all of the above processing in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can analyze reactions using an AI model that takes the child's geographical background data as input.

[0047] The reaction analysis unit can improve the accuracy of its analysis by referring to relevant literature on children during reaction analysis. For example, the reaction analysis unit can refer to the content of a picture book that a child is reading during the analysis. The reaction analysis unit can also refer to the content of educational materials that a child is learning from during the analysis. For example, the reaction analysis unit can refer to the content of an educational program that a child is watching during the analysis. By referring to relevant literature, the accuracy of the analysis is improved. Some or all of the above processing in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can take data on relevant literature on children as input and analyze the reaction using an AI model that performs the analysis.

[0048] The conversation generation unit can generate the most suitable conversation by analyzing the child's past responses during conversation generation. For example, the conversation generation unit can generate conversations based on topics the child has shown interest in in the past. The conversation generation unit can also prioritize question formats that the child found easier to answer in the past. For example, the conversation generation unit can generate conversations that avoid questions the child has found difficult in the past. This promotes effective language learning by generating the most suitable conversation based on past responses. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate conversations using an AI model that takes the child's past response data as input and generates the most suitable conversation.

[0049] The conversation generation unit can adjust the difficulty level of conversations based on the child's current learning level when generating conversations. For example, if the child is at a beginner level, the conversation generation unit can generate conversations using simple words and phrases. If the child is at an intermediate level, the conversation generation unit can generate conversations using slightly more complex sentences. If the child is at an advanced level, the conversation generation unit can generate conversations using difficult words and phrases. This promotes effective language learning by generating conversations appropriate to the learning level. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can take the child's learning level data as input and generate conversations using an AI model that adjusts the difficulty level of conversations.

[0050] The conversation generation unit can prioritize conversations based on the child's submission timing when generating conversations. For example, the conversation generation unit can prioritize generating conversations related to assignments recently submitted by the child. The conversation generation unit can also generate conversations based on the content of assignments previously submitted by the child. For example, the conversation generation unit can generate conversations related to assignments the child plans to submit in the future. This promotes effective language learning by generating conversations based on submission timing. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can take the child's submission timing data as input and generate conversations using an AI model that determines conversation priorities.

[0051] The conversation generation unit can adjust the order of conversations based on the child's relevance during conversation generation. For example, the conversation generation unit can prioritize generating conversations related to topics the child has shown interest in. The conversation generation unit can also generate conversations related to what the child has learned in the past. For example, the conversation generation unit can generate conversations related to what the child will learn in the future. This promotes effective language learning by generating conversations based on relevance. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate conversations using an AI model that takes the child's relevance data as input and adjusts the order of conversations.

[0052] The service provider can select the optimal service delivery method by referring to the child's past response history when providing a conversation. For example, the service provider can provide a conversation based on topics the child has shown interest in in the past. The service provider can also prioritize question formats that the child has found easy to answer in the past. For example, the service provider can provide a conversation that avoids questions that the child has found difficult in the past. This promotes effective language learning by selecting a service delivery method based on past response history. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide a conversation using an AI model that takes the child's past response data as input and selects the optimal service delivery method.

[0053] The service provider can adjust the timing of conversation delivery based on the child's current life circumstances. For example, the service provider can deliver a conversation immediately after the child returns home from school. It can also deliver a conversation after the child has finished eating. For example, the service provider can deliver a conversation when the child is relaxed before going to bed. By adjusting the timing of delivery based on the child's life circumstances, effective language learning is promoted. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take the child's life circumstances data as input and deliver a conversation using an AI model that adjusts the timing of delivery.

[0054] The service provider can select the optimal delivery method when providing conversations, taking into account the child's geographical location. For example, if the child is in a park, the service provider can provide conversations related to the park. If the child is at home, the service provider can provide conversations related to household items. If the child is at school, the service provider can provide conversations related to school. This promotes effective language learning by selecting a delivery method based on geographical location information. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide conversations using an AI model that takes the child's geographical location data as input and selects the optimal delivery method.

[0055] The service provider can analyze a child's social media activity and suggest a means of providing conversations. For example, the service provider can provide conversations related to posts the child has recently "liked." The service provider can also provide conversations based on the content the child has recently commented on. For example, the service provider can provide conversations related to accounts the child follows. This promotes effective language learning by suggesting means of providing conversations based on social media activity. Some or all of the processing described above in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take a child's social media activity data as input and provide conversations using an AI model that suggests means of providing conversations.

[0056] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0057] An AI-powered speech trainer can be equipped with a pronunciation analysis unit that detects subtle changes in a child's pronunciation and supports pronunciation improvement. The pronunciation analysis unit can, for example, detect changes in a child's pronunciation at the phoneme level and evaluate the accuracy of the pronunciation. It can also analyze changes in the rhythm and intonation of a child's pronunciation and provide appropriate feedback. For example, if a child says "dog," the pronunciation analysis unit can detect changes in the stress and speed of the pronunciation and teach the child how to pronounce it correctly. This allows for the detection of subtle changes in a child's pronunciation and supports pronunciation improvement, thereby facilitating language learning. Some or all of the above-described processes in the pronunciation analysis unit may be performed using AI, or without AI. For example, the pronunciation analysis unit can take a child's pronunciation data as input and perform pronunciation analysis using an AI model that detects subtle changes.

[0058] An AI-powered conversational trainer can be equipped with a progress display unit that visualizes a child's learning progress. This unit can, for example, display the child's vocabulary and comprehension growth in graphs or charts. It can also visually show the child's pronunciation improvement. For instance, it could display the number of new words the child has learned or a list of words they can now pronounce correctly. This allows parents to quickly grasp their child's learning progress and provide appropriate support. Some or all of the above-described processes in the progress display unit may be performed using AI, or not. For example, the progress display unit can take the child's learning data as input and display progress using an AI model that visualizes the progress.

[0059] An AI-powered conversational trainer can be equipped with a reward system to enhance children's motivation to learn. The reward system could, for example, award points or badges when a child correctly pronounces a new word or uses it in an appropriate context. The reward system could also provide special rewards when a child achieves certain goals. For example, the reward system could provide a digital sticker when a child learns 10 new words. This can increase children's motivation and encourage continuous learning. Some or all of the processes described above in the reward system may be performed using AI, or not. For example, the reward system could use an AI model that takes the child's learning data as input and determines the reward to provide it.

[0060] An AI-powered learning trainer may be equipped with an environment adjustment unit to optimize the child's learning environment. This unit can, for example, adjust the ambient sound and light conditions while the child is learning. It can also provide background music or noise cancellation to create an environment conducive to concentration. For instance, it can play relaxing music while the child is learning. This optimizes the child's learning environment and supports effective learning. Some or all of the above processing in the environment adjustment unit may be performed using AI, or without AI. For example, the environment adjustment unit can take the child's environmental data as input and adjust the environment using an AI model that provides the optimal environment.

[0061] An AI-powered learning trainer can be equipped with a learning plan suggestion unit that analyzes a child's learning history and proposes an optimal learning plan. For example, the learning plan suggestion unit can analyze a child's past learning history and suggest what they should learn next. Furthermore, the learning plan suggestion unit can provide individually customized learning plans based on the child's learning pace and level of understanding. For instance, it can suggest a plan that focuses on areas where the child struggles. This allows for effective learning support by providing an optimal learning plan based on the child's learning history. Some or all of the above-described processes in the learning plan suggestion unit may be performed using AI, or not. For example, the learning plan suggestion unit can use an AI model that takes the child's learning history data as input and proposes an optimal learning plan to provide a learning plan.

[0062] The AI ​​conversational trainer may include a sharing section for sharing children's learning outcomes. This sharing section can, for example, share children's learning outcomes with parents and teachers. It can also generate reports on children's learning progress and evaluation results and provide them to relevant parties. For example, the sharing section can send parents a list of new words and phrases their child has learned. This allows parents and teachers to provide appropriate support by sharing children's learning outcomes. Some or all of the above processing in the sharing section may be performed using AI, or not. For example, the sharing section can use an AI model that takes children's learning data as input and shares the outcomes to provide the results.

[0063] The following briefly describes the processing flow for example form 1.

[0064] Step 1: The prompt setting section allows you to set prompts to speak to the child. Specifically, you set questions and comments to ask the child, selecting content that is appropriate for the child's age and interests. For example, you can set questions such as "What did you do today?" or "What's your favorite food?". Step 2: The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. Specifically, it analyzes the child's pronunciation and context to evaluate whether the pronunciation is accurate and whether it is used in an appropriate context. For example, if the child says "dog," the unit analyzes the pronunciation and context to evaluate whether the pronunciation is accurate and whether it is used in an appropriate context. Step 3: The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. Specifically, it generates the next questions and comments based on the child's responses and adjusts the content of the conversation. For example, if the child says "dog," it can generate the next question, "What kind of animal is a dog?" Step 4: The provider unit provides the conversation generated by the conversation generation unit to the child. Specifically, it provides the generated conversation to the child in audio or text format. For example, it may provide the conversation generated using speech synthesis technology as audio, or display it as text so that the child can read it.

[0065] (Example of form 2) An AI-powered conversational trainer according to an embodiment of the present invention is a system for infants who have just begun to speak and their guardians. This system allows children to regularly learn new words even when their guardians are busy, by having the AI ​​speak to them periodically and progressing the conversation based on their responses. The AI ​​also consistently teaches accurate word usage, synonyms, antonyms, nuances, and emotional expressions. Furthermore, the AI ​​analyzes the child's statements and responses and regularly reports on their vocabulary, comprehension, and speaking ability development levels. This allows guardians to receive specific feedback and advice to guide their next steps. For example, the AI ​​can set prompts when speaking to the child, such as "What did you do today?" or "What's your favorite food?". When the child answers, the AI ​​analyzes their response and generates further questions or comments. This allows children to learn new words in a natural conversational flow. Next, the AI ​​consistently teaches word usage. For example, when teaching the word "dog," it provides specific examples such as, "A dog is an animal. Dogs bark." The AI ​​also teaches the difference between "dog" and "cat," as well as synonyms and antonyms for the word "dog." This allows children to deeply understand the meaning and usage of words. Furthermore, the AI ​​analyzes the child's speech and responses and assesses their developmental level. For example, if a child says "dog," the AI ​​analyzes the pronunciation and context to assess whether it is pronounced correctly and used in an appropriate context. This assessment is regularly provided to parents as a report, which includes specific feedback and advice. Parents can use this report to plan the next steps. This ensures that children have regular opportunities to learn new words, even when their parents are busy. In addition, teaching words in a consistent manner prevents confusion and improves the child's language comprehension. Moreover, the developmental level assessment and feedback allow parents to specifically understand their child's language development and provide appropriate support. In this way, the AI ​​conversational trainer can effectively support the language development of young children and provide specific feedback to parents.

[0066] The AI ​​conversational trainer according to this embodiment comprises a prompt setting unit, a response analysis unit, a conversation generation unit, and a provision unit. The prompt setting unit sets prompts to speak to the child. The prompt setting unit sets, for example, specific questions and comments to be asked when speaking to the child. For example, the prompt setting unit can set questions such as "What did you do today?" or "What's your favorite food?" The prompt setting unit can also set questions and comments that are appropriate for the child's age and interests. For example, the prompt setting unit can set simple questions for toddlers and more complex questions for slightly older children. The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. The response analysis unit analyzes, for example, the child's pronunciation and context, and evaluates whether the pronunciation is accurate and whether the word is used in an appropriate context. For example, if the child says "dog," the response analysis unit can analyze the pronunciation and context, and evaluate whether the pronunciation is accurate and whether the word is used in an appropriate context. The response analysis unit can also analyze the child's response in real time and provide data for generating the next conversation. The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. The conversation generation unit generates the next question or comment based on the child's response, for example. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, the conversation generation unit can generate the next conversation based on a topic the child has shown interest in. The provision unit provides the conversation generated by the conversation generation unit to the child. The provision unit provides the generated conversation to the child in the form of voice or text, for example. For example, the provision unit can provide the child with the generated conversation in voice using speech synthesis technology. The provision unit can also display the generated conversation as text so that the child can read it. This allows the AI ​​conversational trainer according to the embodiment to teach language to the child in a consistent manner, assess their progress, and provide feedback to the parent.

[0067] The prompt setting unit sets prompts to use when speaking to children. For example, it sets specific questions and comments to ask children. Specifically, the prompt setting unit collects and analyzes the child's profile information in advance to set questions and comments appropriate to the child's age and interests. For example, the prompt setting unit generates appropriate prompts based on information such as the child's age, gender, topics of interest, and past conversation history. For toddlers, simple questions such as "What did you do today?" or "What's your favorite food?" can be set, while for slightly older children, more complex questions such as "What book have you read recently?" or "What are your dreams for the future?" can be set. The prompt setting unit can also set questions appropriate to the season or event. For example, during Christmas, a question such as "What would you like to ask Santa for?" can be set, and during summer vacation, a question such as "What are your plans for summer vacation?" can be set. This allows the prompt setting unit to flexibly set prompts that will capture the child's interest and promote natural conversation. Furthermore, the prompt setting unit also has a function to evaluate the effectiveness of the set prompts and modify them as needed. For example, if a child's response to a particular prompt is unsatisfactory, the prompt can be modified and changed to a more appropriate question or comment. This allows the prompt setting unit to continue providing prompts that facilitate more effective conversations with children.

[0068] The response analysis unit analyzes the child's responses based on prompts set by the prompt setting unit. For example, the response analysis unit analyzes the child's pronunciation and context to evaluate whether the pronunciation is accurate and whether the words are used in an appropriate context. Specifically, the response analysis unit uses speech recognition technology to convert the child's utterances into text data and then analyzes that text data. For example, if the child says "dog," the unit analyzes the speech data at the phoneme level to evaluate whether the pronunciation is accurate and compares it to standard pronunciation. It also uses contextual analysis technology to evaluate whether the child's utterances are used in an appropriate context. For example, if the child says "dog," the unit analyzes the conversation before and after the utterance to determine whether the word "dog" is used in an appropriate context. Furthermore, the response analysis unit can analyze the child's responses in real time and provide data for generating the next conversation. For example, if the child says "dog," it provides data to the conversation generation unit to generate the next question or comment based on that response. This allows the response analysis unit to accurately evaluate the child's pronunciation and context and provide data to facilitate the next conversation. In addition, the response analysis unit can accumulate the child's responses and build a database to evaluate long-term growth. For example, the system records changes in a child's pronunciation and context over time and analyzes growth trends. This allows the response analysis unit to continuously evaluate the child's language development and provide feedback to parents.

[0069] The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. For example, the conversation generation unit generates the next question or comment based on the child's response. Specifically, the conversation generation unit uses natural language processing technology to generate appropriate conversations in response to the child's response. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, if the child says "I like dogs," the conversation generation unit can generate the question, "What do you like about dogs?" to pique the child's interest. Furthermore, the conversation generation unit can adjust the tone and style of the conversation based on the child's response. For example, it can speak in a gentle tone to toddlers and generate more complex conversations for slightly older children. In this way, the conversation generation unit can generate appropriate conversations according to the child's age and interests, promoting natural conversation. In addition, the conversation generation unit has a function to evaluate the effectiveness of the generated conversations and improve the conversation generation algorithm as needed. For example, if the child's response to a particular conversation pattern is unfavorable, it can modify that pattern to generate more appropriate conversations. This allows the conversation generation unit to continuously provide conversations that facilitate more effective communication with children.

[0070] The delivery unit provides children with conversations generated by the conversation generation unit. The delivery unit provides children with generated conversations in the form of audio or text. Specifically, the delivery unit can provide children with audio conversations generated using speech synthesis technology. For example, the delivery unit uses a speech synthesis engine with natural pronunciation and intonation to convert generated conversations into audio in real time. The delivery unit can also display the generated conversations as text so that children can read them. For example, the delivery unit can display the generated conversations on the screen of a tablet or smartphone so that children can visually confirm them. Furthermore, the delivery unit can adjust the method of delivery (audio or text) according to the child's response. For example, if the child prefers audio conversations, it can provide them in audio; if the child prefers text conversations, it can provide them in text. This allows the delivery unit to provide conversations in the most optimal way according to the child's preferences, promoting natural communication. In addition, the delivery unit has a function to evaluate the effectiveness of the delivered conversations and improve the delivery method as needed. For example, if a child's response to a particular delivery method is poor, it can modify that method and change to a more appropriate delivery method. This allows the delivery unit to continue providing conversations that facilitate more effective communication with children.

[0071] The prompt setting unit can set specific questions and comments to be asked when speaking to a child. For example, the prompt setting unit can set questions such as, "What did you do today?" or "What's your favorite food?" The prompt setting unit can also set questions and comments that are appropriate for the child's age and interests. For example, the prompt setting unit can set simple questions for toddlers and more complex questions for slightly older children. This promotes language learning by providing children with appropriate questions and comments. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or not using AI. For example, the prompt setting unit can set questions and comments using an AI model that generates prompts based on the child's age and interests.

[0072] The response analysis unit can analyze a child's pronunciation and context to evaluate whether they are pronouncing words accurately and using them in an appropriate context. For example, if a child says "dog," the response analysis unit can analyze the pronunciation and context to evaluate whether they are pronouncing words accurately and using them in an appropriate context. The response analysis unit can also analyze a child's response in real time and provide data to generate the next conversation. This supports language learning by evaluating the appropriateness of the child's pronunciation and context. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can take a child's pronunciation data as input and use an AI model to evaluate the accuracy of pronunciation.

[0073] The conversation generation unit can generate the next question or comment based on the child's response. For example, if the child says "dog," the conversation generation unit can generate the next question, "What kind of animal is a dog?" The conversation generation unit can also adjust the content of the conversation according to the child's response. For example, the conversation generation unit can generate the next conversation based on a topic the child has shown interest in. This maintains a natural flow of conversation by generating appropriate conversations that respond to the child's response. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate a conversation using an AI model that takes the child's response data as input and generates the next question or comment.

[0074] The service provider can provide the generated conversations to the child. The service provider can provide the generated conversations to the child in the form of voice or text. For example, the service provider can provide the child with voice generated conversations using speech synthesis technology. The service provider can also display the generated conversations as text so that the child can read them. This promotes language learning by providing the child with generated conversations. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take generated conversation data as input and provide the conversation in voice using an AI model that performs speech synthesis.

[0075] The response analysis unit can analyze a child's statements and responses to evaluate their growth level in vocabulary, comprehension, and speaking ability. For example, if a child says "dog," the response analysis unit can analyze the pronunciation and context to evaluate whether it is pronounced correctly and used in an appropriate context. The response analysis unit can also periodically analyze a child's statements and responses to evaluate their growth level. This allows for the provision of appropriate feedback and advice by evaluating the child's growth level. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can use a child's statement data as input and evaluate their growth level using an AI model that evaluates vocabulary and comprehension.

[0076] The service provider can provide parents with a report of the assessment results of their child's growth level. For example, the service provider can provide parents with a report of the assessment results of their child's growth level. For example, the service provider can create a detailed report of the assessment results and provide it to the parents. The service provider can also include specific feedback and advice based on the assessment results in the report. For example, the service provider can assess the growth level of a child's vocabulary and comprehension and report the results to the parents. This allows parents to understand their child's growth level and provide appropriate support. Some or all of the above processing in the service provider may be performed using AI, or not. For example, the service provider can use an AI model that takes growth level assessment data as input and generates a report to provide parents with the assessment results.

[0077] The prompt setting unit can estimate the child's emotions and adjust the content of the prompt based on the estimated emotions. For example, if the child is excited, the prompt setting unit can set a prompt that speaks in a calm tone. If the child is tired, the prompt setting unit can set a simple and short question. For example, if the child is having fun, the prompt setting unit can set an engaging question. This promotes more effective language learning by providing prompts that are appropriate to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes child emotion data as input and adjusts the content of the prompts.

[0078] The prompt setting unit can analyze a child's past response history and select the optimal prompt. For example, the prompt setting unit can set prompts based on topics the child has shown interest in in the past. The prompt setting unit can also prioritize question formats that the child has found easy to answer in the past. For example, the prompt setting unit can set prompts to avoid questions that the child has found difficult in the past. This promotes effective language learning by providing optimal prompts based on past response history. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's past response data as input and selects the optimal prompt.

[0079] The prompt setting unit can filter prompts based on the child's current interests and concerns. For example, the prompt setting unit can set questions related to characters the child has recently become interested in. It can also set prompts based on television programs the child has recently watched. For example, the prompt setting unit can set questions related to toys the child has recently been playing with. This enhances the child's motivation to learn by providing prompts based on their interests and concerns. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's interest data as input and filters the prompts.

[0080] The prompt setting unit can estimate a child's emotions and determine prompt priorities based on the estimated emotions. For example, if a child is feeling anxious, the prompt setting unit can prioritize reassuring questions. If a child is excited, the prompt setting unit can prioritize questions that pique their interest. If a child is tired, the prompt setting unit can prioritize simple questions. This promotes effective language learning by determining prompt priorities according to the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the prompt setting unit may be performed using AI, or not. For example, the prompt setting unit can set prompts using an AI model that takes child emotion data as input and determines prompt priorities.

[0081] The prompt setting unit can prioritize setting highly relevant prompts by considering the child's geographical location when setting prompts. For example, the prompt setting unit can set questions related to parks if the child is in a park. It can also set questions related to household items if the child is at home. For example, it can set questions related to school if the child is at school. This makes it easier to capture the child's interest by providing prompts based on geographical location information. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's geographical location data as input and selects highly relevant prompts.

[0082] The prompt setting unit can analyze a child's social media activity and set relevant prompts when setting prompts. For example, the prompt setting unit can set questions related to posts the child has recently "liked." The prompt setting unit can also set prompts based on what the child has recently commented on. For example, the prompt setting unit can set questions related to accounts the child follows. This makes it easier to capture the child's interest by providing prompts based on their social media activity. Some or all of the above processing in the prompt setting unit may be performed using AI, for example, or without AI. For example, the prompt setting unit can set prompts using an AI model that takes the child's social media activity data as input and selects relevant prompts.

[0083] The reaction analysis unit can estimate a child's emotions and adjust the criteria for reaction analysis based on the estimated emotions. For example, if a child is excited, the reaction analysis unit can prioritize the speed of the reaction. If a child is tired, the reaction analysis unit can prioritize the accuracy of the reaction. If a child is having fun, the reaction analysis unit can prioritize the diversity of the reaction. This improves the accuracy of the analysis by performing reaction analysis according to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can take child emotion data as input and analyze the reaction using an AI model that adjusts the criteria for reaction analysis.

[0084] The reaction analysis unit can detect subtle changes in a child's pronunciation during reaction analysis, thereby improving the accuracy of the analysis. For example, the reaction analysis unit can detect changes in a child's pronunciation at the phoneme level. The reaction analysis unit can also analyze changes in the rhythm and intonation of a child's pronunciation. For example, the reaction analysis unit can detect changes in the volume and speed of a child's pronunciation. This improves the accuracy of the analysis by detecting subtle changes in pronunciation. Some or all of the above-described processes in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can take a child's pronunciation data as input and perform pronunciation analysis using an AI model that detects subtle changes.

[0085] The response analysis unit can refer to past conversation history to gain a deeper understanding of the context of a child's statements during response analysis. For example, the response analysis unit can analyze the context based on what the child has said in the past. The response analysis unit can also refer to patterns of words the child has used in the past. For example, the response analysis unit can analyze the context by considering the interests and concerns the child has shown in the past. By referring to past conversation history, the understanding of the context is deepened and the accuracy of the analysis is improved. Some or all of the above processing in the response analysis unit may be performed using AI, for example, or without AI. For example, the response analysis unit can take the child's past conversation data as input and perform utterance analysis using an AI model that understands context.

[0086] The reaction analysis unit can estimate a child's emotions and adjust the order in which the reaction analysis results are displayed based on the estimated emotions. For example, if a child is feeling anxious, the reaction analysis unit can prioritize displaying results that provide reassurance. If a child is excited, the reaction analysis unit can prioritize displaying results that attract interest. If a child is tired, the reaction analysis unit can prioritize displaying simple results. This helps in understanding the analysis results by displaying them in an order that corresponds to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can analyze reactions using an AI model that takes child emotion data as input and adjusts the order in which the results are displayed.

[0087] The reaction analysis unit can perform analysis while considering the child's geographical background. For example, the reaction analysis unit can consider the dialect and accent of the area where the child lives. It can also consider the educational policies of the school the child attends. For example, the reaction analysis unit can consider the cultural background of places the child frequently visits. This improves the accuracy of the analysis by considering geographical background. Some or all of the above processing in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can analyze reactions using an AI model that takes the child's geographical background data as input.

[0088] The reaction analysis unit can improve the accuracy of its analysis by referring to relevant literature on children during reaction analysis. For example, the reaction analysis unit can refer to the content of a picture book that a child is reading during the analysis. The reaction analysis unit can also refer to the content of educational materials that a child is learning from during the analysis. For example, the reaction analysis unit can refer to the content of an educational program that a child is watching during the analysis. By referring to relevant literature, the accuracy of the analysis is improved. Some or all of the above processing in the reaction analysis unit may be performed using AI, for example, or without AI. For example, the reaction analysis unit can take data on relevant literature on children as input and analyze the reaction using an AI model that performs the analysis.

[0089] The conversation generation unit can estimate a child's emotions and adjust the content of the next conversation based on the estimated emotions. For example, if the child is excited, the conversation generation unit can generate the next conversation in a calm tone. Also, if the child is tired, the conversation generation unit can generate a simple and short conversation. For example, if the child is having fun, the conversation generation unit can generate an engaging conversation. This promotes more effective language learning by generating conversations that respond to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate conversations using an AI model that takes child emotion data as input and adjusts the content of the conversation.

[0090] The conversation generation unit can generate the most suitable conversation by analyzing the child's past responses during conversation generation. For example, the conversation generation unit can generate conversations based on topics the child has shown interest in in the past. The conversation generation unit can also prioritize question formats that the child found easier to answer in the past. For example, the conversation generation unit can generate conversations that avoid questions the child has found difficult in the past. This promotes effective language learning by generating the most suitable conversation based on past responses. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate conversations using an AI model that takes the child's past response data as input and generates the most suitable conversation.

[0091] The conversation generation unit can adjust the difficulty level of conversations based on the child's current learning level when generating conversations. For example, if the child is at a beginner level, the conversation generation unit can generate conversations using simple words and phrases. If the child is at an intermediate level, the conversation generation unit can generate conversations using slightly more complex sentences. If the child is at an advanced level, the conversation generation unit can generate conversations using difficult words and phrases. This promotes effective language learning by generating conversations appropriate to the learning level. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can take the child's learning level data as input and generate conversations using an AI model that adjusts the difficulty level of conversations.

[0092] The conversation generation unit can estimate a child's emotions and adjust the length of the conversation based on the estimated emotions. For example, if the child is excited, the conversation generation unit can generate a short, to-the-point conversation. If the child is relaxed, the conversation generation unit can generate a longer conversation with more detailed explanations. If the child is tired, the conversation generation unit can generate a simple, short conversation. This promotes effective language learning by adjusting the length of the conversation according to the child's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate a conversation using an AI model that takes child emotion data as input and adjusts the length of the conversation.

[0093] The conversation generation unit can prioritize conversations based on the child's submission timing when generating conversations. For example, the conversation generation unit can prioritize generating conversations related to assignments recently submitted by the child. The conversation generation unit can also generate conversations based on the content of assignments previously submitted by the child. For example, the conversation generation unit can generate conversations related to assignments the child plans to submit in the future. This promotes effective language learning by generating conversations based on submission timing. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can take the child's submission timing data as input and generate conversations using an AI model that determines conversation priorities.

[0094] The conversation generation unit can adjust the order of conversations based on the child's relevance during conversation generation. For example, the conversation generation unit can prioritize generating conversations related to topics the child has shown interest in. The conversation generation unit can also generate conversations related to what the child has learned in the past. For example, the conversation generation unit can generate conversations related to what the child will learn in the future. This promotes effective language learning by generating conversations based on relevance. Some or all of the above processing in the conversation generation unit may be performed using AI, for example, or without AI. For example, the conversation generation unit can generate conversations using an AI model that takes the child's relevance data as input and adjusts the order of conversations.

[0095] The service provider can estimate a child's emotions and adjust the way it delivers conversation based on those estimated emotions. For example, if a child is excited, the service provider can deliver a conversation in a calm tone. If a child is tired, the service provider can deliver a simple and short conversation. If a child is having fun, the service provider can deliver an engaging conversation. This adjusts the delivery method according to the child's emotions, thereby promoting effective language learning. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take child emotion data as input and deliver conversation using an AI model that adjusts the delivery method.

[0096] The service provider can select the optimal service delivery method by referring to the child's past response history when providing a conversation. For example, the service provider can provide a conversation based on topics the child has shown interest in in the past. The service provider can also prioritize question formats that the child has found easy to answer in the past. For example, the service provider can provide a conversation that avoids questions that the child has found difficult in the past. This promotes effective language learning by selecting a service delivery method based on past response history. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide a conversation using an AI model that takes the child's past response data as input and selects the optimal service delivery method.

[0097] The service provider can adjust the timing of conversation delivery based on the child's current life circumstances. For example, the service provider can deliver a conversation immediately after the child returns home from school. It can also deliver a conversation after the child has finished eating. For example, the service provider can deliver a conversation when the child is relaxed before going to bed. By adjusting the timing of delivery based on the child's life circumstances, effective language learning is promoted. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take the child's life circumstances data as input and deliver a conversation using an AI model that adjusts the timing of delivery.

[0098] The service provider can estimate a child's emotions and determine the order in which to deliver conversations based on the estimated emotions. For example, if a child is feeling anxious, the service provider can prioritize delivering conversations that provide reassurance. If a child is excited, the service provider can prioritize delivering conversations that pique their interest. If a child is tired, the service provider can prioritize delivering simple conversations. This promotes effective language learning by determining the order of delivery according to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take child emotion data as input and deliver conversations using an AI model that determines the order of delivery.

[0099] The service provider can select the optimal delivery method when providing conversations, taking into account the child's geographical location. For example, if the child is in a park, the service provider can provide conversations related to the park. If the child is at home, the service provider can provide conversations related to household items. If the child is at school, the service provider can provide conversations related to school. This promotes effective language learning by selecting a delivery method based on geographical location information. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide conversations using an AI model that takes the child's geographical location data as input and selects the optimal delivery method.

[0100] The service provider can analyze a child's social media activity and suggest a means of providing conversations. For example, the service provider can provide conversations related to posts the child has recently "liked." The service provider can also provide conversations based on the content the child has recently commented on. For example, the service provider can provide conversations related to accounts the child follows. This promotes effective language learning by suggesting means of providing conversations based on social media activity. Some or all of the processing described above in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take a child's social media activity data as input and provide conversations using an AI model that suggests means of providing conversations.

[0101] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0102] An AI-powered speech trainer can be equipped with a pronunciation analysis unit that detects subtle changes in a child's pronunciation and supports pronunciation improvement. The pronunciation analysis unit can, for example, detect changes in a child's pronunciation at the phoneme level and evaluate the accuracy of the pronunciation. It can also analyze changes in the rhythm and intonation of a child's pronunciation and provide appropriate feedback. For example, if a child says "dog," the pronunciation analysis unit can detect changes in the stress and speed of the pronunciation and teach the child how to pronounce it correctly. This allows for the detection of subtle changes in a child's pronunciation and supports pronunciation improvement, thereby facilitating language learning. Some or all of the above-described processes in the pronunciation analysis unit may be performed using AI, or without AI. For example, the pronunciation analysis unit can take a child's pronunciation data as input and perform pronunciation analysis using an AI model that detects subtle changes.

[0103] An AI conversational trainer may include a conversation tone adjustment unit that estimates a child's emotions and adjusts the tone of conversation based on the estimated emotions. For example, if the child is excited, the conversation tone adjustment unit can speak in a calm tone. It can also speak in a gentle tone if the child is tired. For example, if the conversation tone adjustment unit is having fun, it can speak in a bright tone. This promotes more effective language learning by providing conversations in a tone that matches the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the conversation tone adjustment unit may be performed using AI, for example, or without AI. For example, the conversation tone adjustment unit can take child emotion data as input and provide conversations using an AI model that adjusts the tone.

[0104] An AI-powered conversational trainer can be equipped with a progress display unit that visualizes a child's learning progress. This unit can, for example, display the child's vocabulary and comprehension growth in graphs or charts. It can also visually show the child's pronunciation improvement. For instance, it could display the number of new words the child has learned or a list of words they can now pronounce correctly. This allows parents to quickly grasp their child's learning progress and provide appropriate support. Some or all of the above-described processes in the progress display unit may be performed using AI, or not. For example, the progress display unit can take the child's learning data as input and display progress using an AI model that visualizes the progress.

[0105] An AI-powered conversational trainer can be equipped with a reward system to enhance children's motivation to learn. The reward system could, for example, award points or badges when a child correctly pronounces a new word or uses it in an appropriate context. The reward system could also provide special rewards when a child achieves certain goals. For example, the reward system could provide a digital sticker when a child learns 10 new words. This can increase children's motivation and encourage continuous learning. Some or all of the processes described above in the reward system may be performed using AI, or not. For example, the reward system could use an AI model that takes the child's learning data as input and determines the reward to provide it.

[0106] An AI-powered learning trainer may be equipped with an environment adjustment unit to optimize the child's learning environment. This unit can, for example, adjust the ambient sound and light conditions while the child is learning. It can also provide background music or noise cancellation to create an environment conducive to concentration. For instance, it can play relaxing music while the child is learning. This optimizes the child's learning environment and supports effective learning. Some or all of the above processing in the environment adjustment unit may be performed using AI, or without AI. For example, the environment adjustment unit can take the child's environmental data as input and adjust the environment using an AI model that provides the optimal environment.

[0107] An AI conversational trainer may include a learning content adjustment unit that estimates a child's emotions and adjusts the learning content based on the estimated emotions. For example, if the child is excited, the learning content adjustment unit can provide calming learning content. If the child is tired, it can also provide simple and short learning content. For example, if the child is having fun, the learning content adjustment unit can provide engaging learning content. This promotes more effective learning by providing learning content that matches the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the learning content adjustment unit may be performed using AI, for example, or without AI. For example, the learning content adjustment unit can take the child's emotion data as input and provide learning content using an AI model that adjusts the learning content.

[0108] An AI-powered learning trainer can be equipped with a learning plan suggestion unit that analyzes a child's learning history and proposes an optimal learning plan. For example, the learning plan suggestion unit can analyze a child's past learning history and suggest what they should learn next. Furthermore, the learning plan suggestion unit can provide individually customized learning plans based on the child's learning pace and level of understanding. For instance, it can suggest a plan that focuses on areas where the child struggles. This allows for effective learning support by providing an optimal learning plan based on the child's learning history. Some or all of the above-described processes in the learning plan suggestion unit may be performed using AI, or not. For example, the learning plan suggestion unit can use an AI model that takes the child's learning history data as input and proposes an optimal learning plan to provide a learning plan.

[0109] An AI-powered learning trainer may include a pace adjustment unit that estimates a child's emotions and adjusts the learning pace based on the estimated emotions. For example, if the child is excited, the pace adjustment unit can proceed with learning at a slower pace. If the child is tired, it can also proceed with learning in shorter sessions. For example, if the child is enjoying themselves, the pace adjustment unit can proceed with learning at a normal pace. This promotes more effective learning by providing learning at a pace that matches the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the pace adjustment unit may be performed using AI, for example, or without AI. For example, the pace adjustment unit can provide learning using an AI model that takes the child's emotion data as input and adjusts the pace.

[0110] The AI ​​conversational trainer may include a sharing section for sharing children's learning outcomes. This sharing section can, for example, share children's learning outcomes with parents and teachers. It can also generate reports on children's learning progress and evaluation results and provide them to relevant parties. For example, the sharing section can send parents a list of new words and phrases their child has learned. This allows parents and teachers to provide appropriate support by sharing children's learning outcomes. Some or all of the above processing in the sharing section may be performed using AI, or not. For example, the sharing section can use an AI model that takes children's learning data as input and shares the outcomes to provide the results.

[0111] An AI-powered conversational trainer may include a feedback adjustment unit that estimates a child's emotions and adjusts learning feedback based on the estimated emotions. For example, the feedback adjustment unit can provide positive feedback if the child is excited, or gentle feedback if the child is tired. For example, the feedback adjustment unit can provide encouraging feedback if the child is having fun. This promotes more effective learning by providing feedback that is appropriate to the child's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback adjustment unit may be performed using AI, for example, or without AI. For example, the feedback adjustment unit can take the child's emotion data as input and provide feedback using an AI model that adjusts the feedback.

[0112] The following briefly describes the processing flow for example form 2.

[0113] Step 1: The prompt setting section allows you to set prompts to speak to the child. Specifically, you set questions and comments to ask the child, selecting content that is appropriate for the child's age and interests. For example, you can set questions such as "What did you do today?" or "What's your favorite food?". Step 2: The response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit. Specifically, it analyzes the child's pronunciation and context to evaluate whether the pronunciation is accurate and whether it is used in an appropriate context. For example, if the child says "dog," the unit analyzes the pronunciation and context to evaluate whether the pronunciation is accurate and whether it is used in an appropriate context. Step 3: The conversation generation unit generates the next conversation based on the responses analyzed by the response analysis unit. Specifically, it generates the next questions and comments based on the child's responses and adjusts the content of the conversation. For example, if the child says "dog," it can generate the next question, "What kind of animal is a dog?" Step 4: The provider unit provides the conversation generated by the conversation generation unit to the child. Specifically, it provides the generated conversation to the child in audio or text format. For example, it may provide the conversation generated using speech synthesis technology as audio, or display it as text so that the child can read it.

[0114] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0115] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0116] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0117] Each of the multiple elements described above, including the prompt setting unit, response analysis unit, conversation generation unit, and provision unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the prompt setting unit is implemented by the control unit 46A of the smart device 14 and sets specific questions and comments to be spoken to the child. The response analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes the child's pronunciation and context. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12 and generates the next conversation. The provision unit is implemented by the control unit 46A of the smart device 14 and provides the generated conversation to the child in voice or text. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0118] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0119] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0120] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0121] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0122] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0123] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0124] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0125] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0126] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0127] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0128] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0129] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0130] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0131] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0132] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0133] Each of the multiple elements described above, including the prompt setting unit, response analysis unit, conversation generation unit, and provision unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the prompt setting unit is implemented by the control unit 46A of the smart glasses 214 and sets specific questions and comments to be spoken to the child. The response analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes the child's pronunciation and context. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12 and generates the next conversation. The provision unit is implemented by the control unit 46A of the smart glasses 214 and provides the generated conversation to the child in voice or text. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0134] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0135] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0136] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0137] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0138] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0140] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0141] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0142] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0143] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0144] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0145] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0146] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0147] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0148] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0149] Each of the multiple elements described above, including the prompt setting unit, response analysis unit, conversation generation unit, and provision unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the prompt setting unit is implemented by the control unit 46A of the headset terminal 314 and sets specific questions and comments to be spoken to the child. The response analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes the child's pronunciation and context. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12 and generates the next conversation. The provision unit is implemented by the control unit 46A of the headset terminal 314 and provides the generated conversation to the child in voice or text. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0150] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0151] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0152] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0153] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0154] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0155] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0156] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0157] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0158] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0159] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0160] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0161] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0162] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0163] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0164] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0165] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0166] Each of the multiple elements described above, including the prompt setting unit, response analysis unit, conversation generation unit, and provision unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the prompt setting unit is implemented by the control unit 46A of the robot 414 and sets specific questions and comments to be spoken to the child. The response analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes the child's pronunciation and context. The conversation generation unit is implemented by the specific processing unit 290 of the data processing unit 12 and generates the next conversation. The provision unit is implemented by the control unit 46A of the robot 414 and provides the generated conversation to the child in voice or text. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0167] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0168] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0169] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0170] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0171] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0172] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0173] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0174] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0175] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0176] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0177] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0178] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0179] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0180] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0181] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0182] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0183] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0184] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0185] (Note 1) A prompt setting section for setting prompts to speak to children, A response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit, A conversation generation unit generates the next conversation based on the reaction analyzed by the reaction analysis unit, The system includes a providing unit that provides the conversation generated by the conversation generation unit to the child. A system characterized by the following features. (Note 2) The prompt setting unit is, Set specific questions and comments to use when talking to children. The system described in Appendix 1, characterized by the features described herein. (Note 3) The reaction analysis unit is The system analyzes children's pronunciation and context to evaluate whether they are pronouncing words accurately and using them in appropriate contexts. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned conversation generation unit, Based on the child's response, generate the following questions and comments. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned supply unit is, Provide the generated conversation to the child. The system described in Appendix 1, characterized by the features described herein. (Note 6) The reaction analysis unit is We analyze children's statements and responses to assess their growth levels in vocabulary, comprehension, and speaking skills. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned supply unit is, The results of the growth level assessment will be provided to parents as a report. The system described in Appendix 1, characterized by the features described herein. (Note 8) The prompt setting unit is, The system estimates the child's emotions and adjusts the content of the prompts based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The prompt setting unit is, Analyze the child's past response history to select the optimal prompt. The system described in Appendix 1, characterized by the features described herein. (Note 10) The prompt setting unit is, When setting prompts, filter based on the child's current interests. The system described in Appendix 1, characterized by the features described herein. (Note 11) The prompt setting unit is, The system estimates the child's emotions and prioritizes prompts based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The prompt setting unit is, When setting prompts, the system prioritizes relevant prompts by considering the child's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 13) The prompt setting unit is, When setting prompts, analyze your child's social media activity and set relevant prompts. The system described in Appendix 1, characterized by the features described herein. (Note 14) The reaction analysis unit is We estimate the child's emotions and adjust the criteria for response analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The reaction analysis unit is During response analysis, subtle changes in children's pronunciation are detected to improve the accuracy of the analysis. The system described in Appendix 1, characterized by the features described herein. (Note 16) The reaction analysis unit is During response analysis, we refer to past conversation history to gain a deeper understanding of the context of the child's statements. The system described in Appendix 1, characterized by the features described herein. (Note 17) The reaction analysis unit is The system estimates the child's emotions and adjusts the order in which the response analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The reaction analysis unit is When analyzing responses, the geographical background of the children should be taken into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 19) The reaction analysis unit is When analyzing responses, we refer to relevant literature on children to improve the accuracy of the analysis. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned conversation generation unit, The system estimates the child's emotions and adjusts the content of the next conversation based on those estimates. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned conversation generation unit, When generating conversations, the system analyzes the child's past responses to create the most optimal dialogue. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned conversation generation unit, When generating conversations, adjust the difficulty level of the conversation based on the child's current learning level. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned conversation generation unit, The system estimates the child's emotions and adjusts the length of the conversation based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned conversation generation unit, When generating conversations, prioritize conversations based on when the child submitted their contributions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned conversation generation unit, When generating conversations, the order of conversations is adjusted based on the children's relevance. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned supply unit is, The system estimates the child's emotions and adjusts the way conversations are delivered based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned supply unit is, When providing conversation, the optimal method of delivery is selected by referring to the child's past response history. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned supply unit is, When providing conversations, the timing of the conversations will be adjusted based on the child's current living situation. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned supply unit is, The system estimates the child's emotions and determines the order in which conversations are presented based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned supply unit is, When providing conversations, the optimal method of delivery is selected considering the child's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned supply unit is, When providing conversational support, we analyze children's social media activity and suggest ways to provide support. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0186] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A prompt setting section for setting prompts to speak to children, A response analysis unit analyzes the child's response based on the prompt set by the prompt setting unit, A conversation generation unit generates the next conversation based on the reaction analyzed by the reaction analysis unit, The system includes a providing unit that provides the conversation generated by the conversation generation unit to the child. A system characterized by the following features.

2. The prompt setting unit is, Set specific questions and comments to use when talking to children. The system according to feature 1.

3. The reaction analysis unit is, The system analyzes children's pronunciation and context to evaluate whether they are pronouncing words accurately and using them in appropriate contexts. The system according to feature 1.

4. The aforementioned conversation generation unit, Based on the child's response, generate the following questions and comments. The system according to feature 1.

5. The aforementioned supply unit is, Provide the generated conversation to the child. The system according to feature 1.

6. The reaction analysis unit is, We analyze children's statements and responses to assess their growth levels in vocabulary, comprehension, and speaking skills. The system according to feature 1.

7. The aforementioned supply unit is, The results of the growth level assessment will be provided to parents as a report. The system according to feature 1.

8. The prompt setting unit is, The system estimates the child's emotions and adjusts the content of the prompts based on those emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A