system

The system addresses the lack of customization in learning tools by using a reception, analysis, and generation framework to provide personalized language learning experiences, effectively improving user engagement and learning outcomes.

JP2026072841APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing learning tools fail to provide customized learning experiences based on user input information, leading to inefficiencies and user dissatisfaction.

Method used

A system comprising a reception unit, analysis unit, and generation unit that receives, analyzes, and generates personalized learning content tailored to user inputs, using natural language processing, speech recognition, and image analysis to provide grammatical, pronunciation, and video-based learning.

Benefits of technology

Enables efficient and personalized language learning by correcting grammatical and pronunciation errors, generating tailored content, and proposing optimal learning plans, addressing user-specific challenges and preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072841000001_ABST
    Figure 2026072841000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to provide a customized learning tool based on user input information. [Solution] The system according to the embodiment comprises a reception unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives user input. The analysis unit analyzes the information received by the reception unit. The generation unit generates output based on the information analyzed by the analysis unit. The provision unit provides the output generated by the generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , , ,

[0005] , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, a customized learning tool based on user input information has not been sufficiently provided, and there is room for improvement.

[0005] The system according to an embodiment aims to provide a customized learning tool based on user input information.

Means for Solving the Problems

[0006] The system according to an embodiment includes a reception unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives user input. The analysis unit analyzes the information received by the reception unit. The generation unit generates an output based on the information analyzed by the analysis unit. The provision unit provides the output generated by the generation unit. [Effects of the Invention]

[0007] The system according to this embodiment can provide a customized learning tool based on user input information. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The learning tool according to an embodiment of the present invention is a tool that aims to improve conversational skills and abilities in languages ​​such as English and Chinese using generative AI. This learning tool uses a language learning function to provide correction and guidance on compositions. Next, it uses speech recognition and generation functions to conduct conversation lessons and pronunciation practice. Furthermore, the generative AI generates videos based on videos and topics tailored to the user's preferences and tastes, providing video-based learning. Based on these functions, it is a learning tool that can be customized to suit the user. For example, when a user inputs a composition, the generative AI analyzes the composition and points out grammatical and vocabulary errors. For example, if a user inputs "I went to the park," the generative AI instructs the user to correct "goed" to "went." This allows the user to learn correct grammar. Next, when a user speaks into the microphone, the generative AI analyzes the audio and points out pronunciation errors. For example, if a user pronounces "apple" as "aple," the generative AI shows the correct pronunciation of "apple" and encourages the user to practice. This allows the user to acquire correct pronunciation. Furthermore, the AI ​​generates videos based on the user's preferences and interests, providing video-based learning. For example, if a user enjoys cooking, the AI ​​will generate English videos related to cooking. Users can learn English expressions and vocabulary while watching these videos. This allows users to learn while maintaining interest. Based on these functions, it is a learning tool that can be customized to the user. The AI ​​analyzes the user's learning history and preferences and proposes the optimal learning plan. For example, if a user has difficulty with English pronunciation, the AI ​​will propose a plan that focuses on pronunciation practice. This allows users to learn efficiently using a learning method that suits them. This tool is extremely useful for people who want to learn languages ​​such as English or Chinese. It solves problems such as not knowing which learning service is right for them because there are too many options, getting bored with repetitive learning, not knowing what to do specifically, and feeling anxious about speaking with native speakers from the start. By providing lessons tailored to individual types and preferences, and lessons that match individual goals, users can efficiently improve their language skills.This allows learning tools to efficiently support users in their language learning.

[0029] The learning tool according to this embodiment comprises a reception unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may, for example, use a keyboard or touchscreen to receive text input. The reception unit may also use a microphone to receive voice input. Furthermore, the reception unit may use a camera to receive image input. For example, the reception unit receives text entered by the user using a keyboard. The reception unit may also receive voice entered by the user using a microphone. Furthermore, the reception unit may also receive images entered by the user using a camera. The analysis unit analyzes the information received by the reception unit. The analysis unit may, for example, use methods such as text analysis, voice analysis, and image analysis, but is not limited to these. For example, the analysis unit may use natural language processing technology to perform text analysis. The analysis unit may also use speech recognition technology to perform voice analysis. Furthermore, the analysis unit may also use image recognition technology to perform image analysis. For example, the analysis unit uses natural language processing technology to analyze grammatical and vocabulary errors in text. The analysis unit can also use speech recognition technology to analyze pronunciation errors in speech. Furthermore, the analysis unit can use image recognition technology to analyze the content of images. The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in formats such as text output, audio output, and image output, but is not limited to these examples. For example, the generation unit uses text generation AI to generate text output. The generation unit can also use speech generation AI to generate audio output. Furthermore, the generation unit can also use image generation AI to generate image output. For example, the generation unit uses text generation AI to summarize a user's writing. The generation unit can also use speech generation AI to correct the user's pronunciation. Furthermore, the generation unit can use image generation AI to generate videos tailored to the user's preferences. The provisioning unit provides the output generated by the generation unit.The provider delivers output through, for example, a web application or a mobile application, but is not limited to such examples. For example, the provider displays text output to the user through a web application. The provider can also play audio output to the user through a mobile application. Furthermore, the provider can stream video output. For example, the provider displays generated text to the user through a web application. The provider can also play generated audio to the user through a mobile application. Furthermore, the provider can stream generated video. In this way, the learning tool according to the embodiment can support efficient learning by analyzing user input and providing generated output.

[0030] The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. For example, the reception unit can use a keyboard or touchscreen to receive text input. Specifically, keyboards can be physical or software keyboards, and touchscreens can use different technologies such as capacitive or resistive touchscreens. The reception unit can also use a microphone to receive voice input. Microphones can be of various types, such as those with noise-canceling capabilities or directional microphones. Furthermore, the reception unit can use a camera to receive image input. Cameras can capture not only still images but also videos, and their performance, such as resolution and frame rate, varies. For example, the reception unit can receive text entered by the user using a keyboard. Text input is used for various purposes such as writing documents or entering questions. The reception unit can also receive voice input from the user using a microphone. Voice input is used for pronunciation practice or entering voice commands. Furthermore, the reception unit can also receive images entered by the user using a camera. Image input is used for purposes such as handwritten character recognition and object identification. This allows the reception unit to receive user information through various input methods and provide the data to the subsequent analysis unit.

[0031] The analysis unit analyzes the information received by the reception unit. The analysis unit uses, but is not limited to, methods such as text analysis, speech analysis, and image analysis. For example, the analysis unit uses natural language processing techniques to perform text analysis. Natural language processing techniques include morphological analysis, grammatical analysis, and semantic analysis, which enable the detection of grammatical and vocabulary errors in text. The analysis unit can also use speech recognition techniques to perform speech analysis. Speech recognition techniques include speech-to-text conversion using acoustic and language models, and pronunciation error detection. Furthermore, the analysis unit can also use image recognition techniques to perform image analysis. Image recognition techniques include object detection, face recognition, and handwriting recognition, which enable detailed analysis of image content. For example, the analysis unit uses natural language processing techniques to analyze grammatical and vocabulary errors in text. Specifically, it performs morphological analysis to identify the part of speech and meaning of words and detect grammatical errors. The analysis unit can also use speech recognition techniques to analyze pronunciation errors in speech. The system inputs audio data into an acoustic model to detect errors at the phoneme level and identify areas for improvement in pronunciation. Furthermore, the analysis unit can also analyze image content using image recognition technology. For example, it can recognize handwritten characters, analyze their shape and arrangement, and convert them into text data. This allows the analysis unit to analyze the diverse information received in detail and provide data to the subsequent generation unit.

[0032] The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in various forms, such as text output, audio output, and image output, but is not limited to these examples. For example, the generation unit uses a text generation AI to generate text output. The text generation AI uses natural language generation technology to generate text based on user input. For example, it can summarize a composition entered by the user or generate answers to questions. The generation unit can also use an audio generation AI to generate audio output. The audio generation AI uses speech synthesis technology to convert text data into natural-sounding speech. For example, it can generate audio with corrected pronunciation of the user or read text aloud. Furthermore, the generation unit can also use an image generation AI to generate image output. The image generation AI uses a generation model to generate images and videos tailored to the user's preferences. For example, it can generate relevant images based on keywords entered by the user or create videos on a specific theme. This allows the generation unit to generate output in various formats based on the analyzed information and provide data to the subsequent supply unit.

[0033] The provider provides the output generated by the generator. The provider provides the output through, for example, a web application or a mobile application. For example, the provider displays text output to the user through a web application. The web application runs on a browser and provides an easily accessible interface for the user. The provider can also play audio output to the user through a mobile application. The mobile application runs on a smartphone or tablet and can be used on a device that is easy for the user to carry around. Furthermore, the provider can also stream video output. Streaming is a technology that plays video in real time, allowing the user to watch the video over an internet connection. For example, the provider displays generated text to the user through a web application. The user can open a browser, access the web application, and view the generated text. The provider can also play generated audio to the user through a mobile application. The user can open the application on a smartphone or tablet and listen to the generated audio. Furthermore, the provider can also stream generated video. The user can watch the video in real time over an internet connection and review the learning content. This allows the output provider to deliver the generated output to the user in various ways, supporting efficient learning.

[0034] The analytics unit can analyze user preferences and tastes. For example, the analytics unit can identify user preferences and tastes based on survey results and past behavioral history. For instance, the analytics unit can analyze a user's past video viewing history to identify their preferences. It can also analyze the content of text entered by users in the past to identify their tastes. Furthermore, the analytics unit can analyze a user's past search history to identify their areas of interest. For example, the analytics unit can analyze the genres of videos a user has watched in the past to identify their preferences. It can also analyze the topics of text entered by users in the past to identify their tastes. Furthermore, the analytics unit can analyze keywords from a user's past search history to identify their areas of interest. This allows for the provision of more appropriate learning content based on user preferences and tastes.

[0035] The generation unit can analyze the user's pronunciation. For example, the generation unit can analyze the sound waveform and identify phonemes. For instance, it can analyze the waveform of the sound the user pronounces and identify pronunciation errors. It can also identify the phonemes of the sound the user pronounces and show the correct pronunciation. Furthermore, the generation unit can analyze the user's pronunciation and suggest pronunciation practice methods. This improves the accuracy of pronunciation practice by analyzing the user's pronunciation.

[0036] The generation unit can analyze the user's writing. For example, the generation unit can perform grammar checks and content evaluations. For instance, it can check the grammar of the user's input and point out errors. It can also evaluate the content of the user's input and suggest areas for improvement. Furthermore, the generation unit can analyze the user's writing and provide feedback on its revisions. This allows for writing feedback by analyzing the user's writing.

[0037] The service provider can provide users with the most suitable learning plan. For example, the service provider can create a learning plan based on the user's learning goals and progress. For instance, the service provider can set goals to be achieved based on the user's learning goals. Furthermore, the service provider can suggest appropriate learning content based on the user's learning progress. In addition, the service provider can customize the learning plan according to the user's individual needs. This allows for efficient learning by providing users with the most suitable learning plan.

[0038] The service provider can analyze the user's learning history and propose an optimal learning plan. For example, the service provider creates a learning plan based on past learning content and learning outcomes. For example, the service provider analyzes the user's past learning content and proposes a relevant learning plan. The service provider can also propose an appropriate learning method based on the user's learning outcomes. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. For example, the service provider analyzes the user's past learning content and proposes a relevant learning plan. The service provider can also propose an appropriate learning method based on the user's learning outcomes. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. As a result, learning effectiveness is improved by proposing an optimal learning plan based on the user's learning history.

[0039] The reception desk can analyze the user's past input history and select the optimal input method. For example, the reception desk can prioritize suggesting input methods (voice, text, etc.) that the user has frequently used in the past. The reception desk can also suggest similar input methods based on the user's past input. Furthermore, the reception desk can suggest input methods suitable for specific time periods based on the user's past input history. For example, the reception desk can suggest similar input methods based on the user's past input. Furthermore, the reception desk can suggest input methods suitable for specific time periods based on the user's past input history. This enables efficient input by selecting the optimal input method based on the user's past input history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI.

[0040] The reception unit can filter input based on the user's current learning status and areas of interest. For example, the reception unit can accept only input related to the topic the user is currently learning. The reception unit can also prioritize accepting relevant input based on the user's areas of interest. Furthermore, the reception unit can accept input of appropriate difficulty level according to the user's learning progress. For example, the reception unit can prioritize accepting relevant input based on the user's areas of interest. The reception unit can also accept input of appropriate difficulty level according to the user's learning progress. This allows for the acceptance of more appropriate input by filtering input based on the user's learning status and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0041] The reception unit can prioritize accepting inputs that are highly relevant to the user, taking into account the user's geographical location. For example, if the user is in a specific region, the reception unit will prioritize accepting inputs related to that region. The reception unit can also prioritize accepting inputs related to the user's travel destination if the user is traveling. Furthermore, if the user is at home, the reception unit can prioritize accepting inputs related to their home. This allows for the acceptance of more appropriate inputs by prioritizing highly relevant inputs based on the user's geographical location. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0042] The reception unit can analyze the user's social media activity and accept relevant input when receiving input. For example, the reception unit can accept relevant input based on what the user has shared on social media. The reception unit can also accept relevant input based on the activity of the user's social media followers and friends. Furthermore, the reception unit can also accept relevant input based on the content of the user's social media posts. For example, the reception unit can accept relevant input based on the activity of the user's social media followers and friends. Furthermore, the reception unit can also accept relevant input based on the content of the user's social media posts. This allows for the acceptance of more appropriate input by accepting relevant input based on the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0043] The analysis unit can adjust the level of detail of the analysis based on the importance of the information during the analysis. For example, the analysis unit performs a detailed analysis on important information. The analysis unit can also perform a concise analysis on general information. Furthermore, the analysis unit can adjust the level of detail of the analysis based on the user's level of interest. For example, the analysis unit performs a concise analysis on general information. Furthermore, the analysis unit can adjust the level of detail of the analysis based on the user's level of interest. By adjusting the level of detail of the analysis based on the importance of the information, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0044] The analysis unit can apply different analysis algorithms depending on the category of information during analysis. For example, the analysis unit can use a specialized algorithm for grammatical analysis. The analysis unit can also use a speech recognition algorithm for pronunciation analysis. Furthermore, the analysis unit can use a recommendation system for analyzing preferences and tastes. By applying different analysis algorithms depending on the category of information, more appropriate analysis results can be provided. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0045] The analysis unit can determine the priority of analysis based on the timing of information submission during the analysis. For example, the analysis unit may prioritize the analysis of the most recent information. The analysis unit may also prioritize the analysis of information with an approaching submission deadline. Furthermore, the analysis unit may also determine the priority of analysis based on the user's schedule. For example, the analysis unit may prioritize the analysis of information with an approaching submission deadline. The analysis unit may also determine the priority of analysis based on the timing of information submission, thereby providing more appropriate analysis results. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0046] The analysis unit can adjust the order of analysis based on the relevance of the information during the analysis. For example, the analysis unit prioritizes the analysis of highly relevant information. The analysis unit can also adjust the order of analysis based on the user's level of interest. Furthermore, the analysis unit can adjust the order of analysis based on the importance of the information. For example, the analysis unit adjusts the order of analysis based on the user's level of interest. Furthermore, the analysis unit can adjust the order of analysis based on the importance of the information. By adjusting the order of analysis based on the relevance of the information, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0047] The generation unit can adjust the level of detail of the generated output based on the importance of the information during generation. For example, the generation unit generates detailed output for important information. The generation unit can also generate concise output for general information. Furthermore, the generation unit can adjust the level of detail of the generated output based on the user's level of interest. For example, the generation unit generates concise output for general information. The generation unit can also adjust the level of detail of the generated output based on the user's level of interest. By adjusting the level of detail of the generated output based on the importance of the information, a more appropriate output can be provided. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0048] The generation unit can apply different generation algorithms depending on the category of information during generation. For example, the generation unit can use a specialized algorithm for grammar generation. The generation unit can also use a speech recognition algorithm for pronunciation generation. Furthermore, the generation unit can use a recommendation system for generating preferences and tastes. By applying different generation algorithms depending on the category of information, it is possible to provide more appropriate output. Some or all of the above-described processes in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0049] The generation unit can determine the generation priority based on the information submission timing during generation. For example, the generation unit can prioritize generating the latest information. The generation unit can also prioritize generating information with an approaching submission deadline. Furthermore, the generation unit can also determine the generation priority based on the user's schedule. For example, the generation unit can prioritize generating information with an approaching submission deadline. The generation unit can also determine the generation priority based on the user's schedule. This allows for the provision of more appropriate output by determining the generation priority based on the information submission timing. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0050] The generation unit can adjust the order of generation based on the relevance of the information during generation. For example, the generation unit can prioritize the generation of highly relevant information. The generation unit can also adjust the order of generation based on the user's level of interest. Furthermore, the generation unit can also adjust the order of generation based on the importance of the information. By adjusting the order of generation based on the relevance of the information, a more appropriate output can be provided. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0051] The delivery unit can select the optimal delivery method by referring to the user's past learning history at the time of delivery. For example, the delivery unit provides relevant output based on what the user has learned in the past. For example, the delivery unit provides relevant output based on what the user has learned in the past. The delivery unit can also suggest the optimal learning method from the user's learning history. Furthermore, the delivery unit can provide appropriate output according to the user's learning progress. For example, the delivery unit suggests the optimal learning method from the user's learning history. Furthermore, the delivery unit can also provide appropriate output according to the user's learning progress. This allows for the provision of more appropriate output by selecting the optimal delivery method based on the user's past learning history. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without using AI.

[0052] The delivery unit can select the optimal delivery method at the time of delivery, taking into account the user's device information. For example, if the user is using a smartphone, the delivery unit can provide output that matches the screen size. For example, if the user is using a smartphone, the delivery unit can provide output that matches the screen size. The delivery unit can also provide output optimized for a larger screen if the user is using a tablet. Furthermore, if the user is using a smartwatch, the delivery unit can provide concise and highly visible output. For example, if the user is using a tablet, the delivery unit can provide output optimized for a larger screen. Furthermore, if the user is using a smartwatch, the delivery unit can also provide concise and highly visible output. This allows for the provision of more appropriate output by selecting the optimal delivery method based on the user's device information. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI.

[0053] The service provider can analyze the user's learning history and propose an optimal learning plan at the time of delivery. For example, the service provider can propose a relevant learning plan based on the content the user has previously learned. The service provider can also propose an optimal learning method based on the user's learning history. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. For example, the service provider can propose an optimal learning method based on the user's learning history. The service provider can also propose an appropriate learning plan according to the user's learning progress. In this way, by proposing an optimal learning plan based on the user's learning history, more appropriate learning can be provided. Some or all of the above processing in the service provider may be performed using AI, for example, or without using AI.

[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0055] The reception desk can analyze the user's past input history when receiving user input and suggest the most suitable input method. For example, it can prioritize suggesting input methods that the user has frequently used in the past (such as voice or text). It can also suggest similar input methods based on the user's past input content. Furthermore, it can suggest input methods suitable for specific time periods based on the user's past input history. This enables efficient input by selecting the optimal input method based on the user's past input history.

[0056] The analytics unit can perform more detailed analyses of user preferences and tastes based on the user's past behavior history and survey results. For example, it can analyze the user's past video viewing history to identify their preferences. It can also analyze the content of text the user has entered in the past to identify their preferences. Furthermore, it can analyze the user's past search history to identify their areas of interest. This allows for the provision of more appropriate learning content based on the user's preferences and tastes.

[0057] The generation unit, when analyzing a user's writing, not only performs grammatical checks and evaluates content, but can also provide more personalized feedback based on the user's past writing history. For example, it can identify grammatical errors the user has made in the past and provide advice to prevent similar errors. It can also evaluate the content of the user's past writing and indicate areas for improvement. Furthermore, it can provide more specific editing guidance based on the user's writing style and tone. In this way, analyzing the user's writing enables more personalized editing guidance.

[0058] The service provider analyzes the user's learning history and proposes an optimal learning plan, customizing the plan based on the user's learning progress and goals. For example, if a user has difficulty with a particular skill, the service provider can propose a learning plan that focuses on strengthening that skill. It can also suggest learning content of appropriate difficulty level according to the user's learning progress. Furthermore, it can set goals to be achieved based on the user's learning objectives and propose a learning plan to achieve them. This allows for efficient learning by proposing an optimal learning plan based on the user's learning history.

[0059] The service provider can analyze a user's learning history and propose an optimal learning plan, taking into account the user's geographical location to suggest highly relevant learning content. For example, if a user is in a specific region, it can prioritize suggesting learning content related to that region. If a user is traveling, it can also suggest learning content related to their travel destination. Furthermore, if a user is at home, it can suggest learning content related to their home. This allows for more appropriate learning by suggesting highly relevant learning content based on the user's geographical location.

[0060] The following briefly describes the processing flow for example form 1.

[0061] Step 1: The reception desk receives user input. User input includes text input, voice input, and image input. For example, the reception desk accepts text input using a keyboard or touchscreen, voice input using a microphone, and image input using a camera. Step 2: The analysis unit analyzes the information received by the reception unit. The analysis unit uses methods such as text analysis, speech analysis, and image analysis. For example, it uses natural language processing technology to analyze grammatical and vocabulary errors in text, speech recognition technology to analyze pronunciation errors in speech, and image recognition technology to analyze the content of images. Step 3: The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in formats such as text output, audio output, and image output. For example, it can use a text generation AI to summarize the user's writing, an audio generation AI to correct the user's pronunciation, and an image generation AI to generate a video tailored to the user's preferences. Step 4: The provider unit provides the output generated by the generator unit. The provider unit provides the output through a web application or a mobile application. For example, it may display generated text to the user through a web application, play generated audio to the user through a mobile application, or stream generated video.

[0062] (Example of form 2) The learning tool according to an embodiment of the present invention is a tool that aims to improve conversational skills and abilities in languages ​​such as English and Chinese using generative AI. This learning tool uses a language learning function to provide correction and guidance on compositions. Next, it uses speech recognition and generation functions to conduct conversation lessons and pronunciation practice. Furthermore, the generative AI generates videos based on videos and topics tailored to the user's preferences and tastes, providing video-based learning. Based on these functions, it is a learning tool that can be customized to suit the user. For example, when a user inputs a composition, the generative AI analyzes the composition and points out grammatical and vocabulary errors. For example, if a user inputs "I went to the park," the generative AI instructs the user to correct "goed" to "went." This allows the user to learn correct grammar. Next, when a user speaks into the microphone, the generative AI analyzes the audio and points out pronunciation errors. For example, if a user pronounces "apple" as "aple," the generative AI shows the correct pronunciation of "apple" and encourages the user to practice. This allows the user to acquire correct pronunciation. Furthermore, the AI ​​generates videos based on the user's preferences and interests, providing video-based learning. For example, if a user enjoys cooking, the AI ​​will generate English videos related to cooking. Users can learn English expressions and vocabulary while watching these videos. This allows users to learn while maintaining interest. Based on these functions, it is a learning tool that can be customized to the user. The AI ​​analyzes the user's learning history and preferences and proposes the optimal learning plan. For example, if a user has difficulty with English pronunciation, the AI ​​will propose a plan that focuses on pronunciation practice. This allows users to learn efficiently using a learning method that suits them. This tool is extremely useful for people who want to learn languages ​​such as English or Chinese. It solves problems such as not knowing which learning service is right for them because there are too many options, getting bored with repetitive learning, not knowing what to do specifically, and feeling anxious about speaking with native speakers from the start. By providing lessons tailored to individual types and preferences, and lessons that match individual goals, users can efficiently improve their language skills.This allows learning tools to efficiently support users in their language learning.

[0063] The learning tool according to this embodiment comprises a reception unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. The reception unit may, for example, use a keyboard or touchscreen to receive text input. The reception unit may also use a microphone to receive voice input. Furthermore, the reception unit may use a camera to receive image input. For example, the reception unit receives text entered by the user using a keyboard. The reception unit may also receive voice entered by the user using a microphone. Furthermore, the reception unit may also receive images entered by the user using a camera. The analysis unit analyzes the information received by the reception unit. The analysis unit may, for example, use methods such as text analysis, voice analysis, and image analysis, but is not limited to these. For example, the analysis unit may use natural language processing technology to perform text analysis. The analysis unit may also use speech recognition technology to perform voice analysis. Furthermore, the analysis unit may also use image recognition technology to perform image analysis. For example, the analysis unit uses natural language processing technology to analyze grammatical and vocabulary errors in text. The analysis unit can also use speech recognition technology to analyze pronunciation errors in speech. Furthermore, the analysis unit can use image recognition technology to analyze the content of images. The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in formats such as text output, audio output, and image output, but is not limited to these examples. For example, the generation unit uses text generation AI to generate text output. The generation unit can also use speech generation AI to generate audio output. Furthermore, the generation unit can also use image generation AI to generate image output. For example, the generation unit uses text generation AI to summarize a user's writing. The generation unit can also use speech generation AI to correct the user's pronunciation. Furthermore, the generation unit can use image generation AI to generate videos tailored to the user's preferences. The provisioning unit provides the output generated by the generation unit.The provider delivers output through, for example, a web application or a mobile application, but is not limited to such examples. For example, the provider displays text output to the user through a web application. The provider can also play audio output to the user through a mobile application. Furthermore, the provider can stream video output. For example, the provider displays generated text to the user through a web application. The provider can also play generated audio to the user through a mobile application. Furthermore, the provider can stream generated video. In this way, the learning tool according to the embodiment can support efficient learning by analyzing user input and providing generated output.

[0064] The reception unit receives user input. User input includes, but is not limited to, text input, voice input, and image input. For example, the reception unit can use a keyboard or touchscreen to receive text input. Specifically, keyboards can be physical or software keyboards, and touchscreens can use different technologies such as capacitive or resistive touchscreens. The reception unit can also use a microphone to receive voice input. Microphones can be of various types, such as those with noise-canceling capabilities or directional microphones. Furthermore, the reception unit can use a camera to receive image input. Cameras can capture not only still images but also videos, and their performance, such as resolution and frame rate, varies. For example, the reception unit can receive text entered by the user using a keyboard. Text input is used for various purposes such as writing documents or entering questions. The reception unit can also receive voice input from the user using a microphone. Voice input is used for pronunciation practice or entering voice commands. Furthermore, the reception unit can also receive images entered by the user using a camera. Image input is used for purposes such as handwritten character recognition and object identification. This allows the reception unit to receive user information through various input methods and provide the data to the subsequent analysis unit.

[0065] The analysis unit analyzes the information received by the reception unit. The analysis unit uses, but is not limited to, methods such as text analysis, speech analysis, and image analysis. For example, the analysis unit uses natural language processing techniques to perform text analysis. Natural language processing techniques include morphological analysis, grammatical analysis, and semantic analysis, which enable the detection of grammatical and vocabulary errors in text. The analysis unit can also use speech recognition techniques to perform speech analysis. Speech recognition techniques include speech-to-text conversion using acoustic and language models, and pronunciation error detection. Furthermore, the analysis unit can also use image recognition techniques to perform image analysis. Image recognition techniques include object detection, face recognition, and handwriting recognition, which enable detailed analysis of image content. For example, the analysis unit uses natural language processing techniques to analyze grammatical and vocabulary errors in text. Specifically, it performs morphological analysis to identify the part of speech and meaning of words and detect grammatical errors. The analysis unit can also use speech recognition techniques to analyze pronunciation errors in speech. The system inputs audio data into an acoustic model to detect errors at the phoneme level and identify areas for improvement in pronunciation. Furthermore, the analysis unit can also analyze image content using image recognition technology. For example, it can recognize handwritten characters, analyze their shape and arrangement, and convert them into text data. This allows the analysis unit to analyze the diverse information received in detail and provide data to the subsequent generation unit.

[0066] The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in various forms, such as text output, audio output, and image output, but is not limited to these examples. For example, the generation unit uses a text generation AI to generate text output. The text generation AI uses natural language generation technology to generate text based on user input. For example, it can summarize a composition entered by the user or generate answers to questions. The generation unit can also use an audio generation AI to generate audio output. The audio generation AI uses speech synthesis technology to convert text data into natural-sounding speech. For example, it can generate audio with corrected pronunciation of the user or read text aloud. Furthermore, the generation unit can also use an image generation AI to generate image output. The image generation AI uses a generation model to generate images and videos tailored to the user's preferences. For example, it can generate relevant images based on keywords entered by the user or create videos on a specific theme. This allows the generation unit to generate output in various formats based on the analyzed information and provide data to the subsequent supply unit.

[0067] The provider provides the output generated by the generator. The provider provides the output through, for example, a web application or a mobile application. For example, the provider displays text output to the user through a web application. The web application runs on a browser and provides an easily accessible interface for the user. The provider can also play audio output to the user through a mobile application. The mobile application runs on a smartphone or tablet and can be used on a device that is easy for the user to carry around. Furthermore, the provider can also stream video output. Streaming is a technology that plays video in real time, allowing the user to watch the video over an internet connection. For example, the provider displays generated text to the user through a web application. The user can open a browser, access the web application, and view the generated text. The provider can also play generated audio to the user through a mobile application. The user can open the application on a smartphone or tablet and listen to the generated audio. Furthermore, the provider can also stream generated video. The user can watch the video in real time over an internet connection and review the learning content. This allows the output provider to deliver the generated output to the user in various ways, supporting efficient learning.

[0068] The analytics unit can analyze user preferences and tastes. For example, the analytics unit can identify user preferences and tastes based on survey results and past behavioral history. For instance, the analytics unit can analyze a user's past video viewing history to identify their preferences. It can also analyze the content of text entered by users in the past to identify their tastes. Furthermore, the analytics unit can analyze a user's past search history to identify their areas of interest. For example, the analytics unit can analyze the genres of videos a user has watched in the past to identify their preferences. It can also analyze the topics of text entered by users in the past to identify their tastes. Furthermore, the analytics unit can analyze keywords from a user's past search history to identify their areas of interest. This allows for the provision of more appropriate learning content based on user preferences and tastes.

[0069] The generation unit can analyze the user's pronunciation. For example, the generation unit can analyze the sound waveform and identify phonemes. For instance, it can analyze the waveform of the sound the user pronounces and identify pronunciation errors. It can also identify the phonemes of the sound the user pronounces and show the correct pronunciation. Furthermore, the generation unit can analyze the user's pronunciation and suggest pronunciation practice methods. This improves the accuracy of pronunciation practice by analyzing the user's pronunciation.

[0070] The generation unit can analyze the user's writing. For example, the generation unit can perform grammar checks and content evaluations. For instance, it can check the grammar of the user's input and point out errors. It can also evaluate the content of the user's input and suggest areas for improvement. Furthermore, the generation unit can analyze the user's writing and provide feedback on its revisions. This allows for writing feedback by analyzing the user's writing.

[0071] The service provider can provide users with the most suitable learning plan. For example, the service provider can create a learning plan based on the user's learning goals and progress. For instance, the service provider can set goals to be achieved based on the user's learning goals. Furthermore, the service provider can suggest appropriate learning content based on the user's learning progress. In addition, the service provider can customize the learning plan according to the user's individual needs. This allows for efficient learning by providing users with the most suitable learning plan.

[0072] The service provider can analyze the user's learning history and propose an optimal learning plan. For example, the service provider creates a learning plan based on past learning content and learning outcomes. For example, the service provider analyzes the user's past learning content and proposes a relevant learning plan. The service provider can also propose an appropriate learning method based on the user's learning outcomes. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. For example, the service provider analyzes the user's past learning content and proposes a relevant learning plan. The service provider can also propose an appropriate learning method based on the user's learning outcomes. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. As a result, learning effectiveness is improved by proposing an optimal learning plan based on the user's learning history.

[0073] The reception system can estimate the user's emotions and adjust the timing of input acceptance based on the estimated emotions. For example, if the user is feeling stressed, the reception system can prompt for input at a time when the user can relax. The reception system can also accept input immediately if the user is concentrating. Furthermore, if the user is tired, the reception system can prompt for input after a break. For example, if the reception system is concentrating, it can accept input immediately. Furthermore, if the user is tired, it can prompt for input after a break. By adjusting the timing of input acceptance according to the user's emotions, input can be accepted at a more appropriate time. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0074] The reception desk can analyze the user's past input history and select the optimal input method. For example, the reception desk can prioritize suggesting input methods (voice, text, etc.) that the user has frequently used in the past. The reception desk can also suggest similar input methods based on the user's past input. Furthermore, the reception desk can suggest input methods suitable for specific time periods based on the user's past input history. For example, the reception desk can suggest similar input methods based on the user's past input. Furthermore, the reception desk can suggest input methods suitable for specific time periods based on the user's past input history. This enables efficient input by selecting the optimal input method based on the user's past input history. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI.

[0075] The reception unit can filter input based on the user's current learning status and areas of interest. For example, the reception unit can accept only input related to the topic the user is currently learning. The reception unit can also prioritize accepting relevant input based on the user's areas of interest. Furthermore, the reception unit can accept input of appropriate difficulty level according to the user's learning progress. For example, the reception unit can prioritize accepting relevant input based on the user's areas of interest. The reception unit can also accept input of appropriate difficulty level according to the user's learning progress. This allows for the acceptance of more appropriate input by filtering input based on the user's learning status and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0076] The reception system can estimate the user's emotions and determine the priority of input to be received based on the estimated emotions. For example, if the user is nervous, the reception system will prioritize simple inputs. The reception system can also prioritize detailed inputs if the user is relaxed. Furthermore, if the user is in a hurry, the reception system can prioritize inputs that can be processed quickly. This allows for more appropriate input to be received by prioritizing inputs according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0077] The reception unit can prioritize accepting inputs that are highly relevant to the user, taking into account the user's geographical location. For example, if the user is in a specific region, the reception unit will prioritize accepting inputs related to that region. The reception unit can also prioritize accepting inputs related to the user's travel destination if the user is traveling. Furthermore, if the user is at home, the reception unit can prioritize accepting inputs related to their home. This allows for the acceptance of more appropriate inputs by prioritizing highly relevant inputs based on the user's geographical location. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0078] The reception unit can analyze the user's social media activity and accept relevant input when receiving input. For example, the reception unit can accept relevant input based on what the user has shared on social media. The reception unit can also accept relevant input based on the activity of the user's social media followers and friends. Furthermore, the reception unit can also accept relevant input based on the content of the user's social media posts. For example, the reception unit can accept relevant input based on the activity of the user's social media followers and friends. Furthermore, the reception unit can also accept relevant input based on the content of the user's social media posts. This allows for the acceptance of more appropriate input by accepting relevant input based on the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI.

[0079] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the user is relaxed, the analysis unit can provide detailed analysis results. For example, if the user is relaxed, the analysis unit can provide detailed analysis results. The analysis unit can also provide concise analysis results if the user is tense. Furthermore, if the user is excited, the analysis unit can provide visually appealing analysis results. For example, if the user is tense, the analysis unit can provide concise analysis results. Furthermore, if the user is excited, the analysis unit can also provide visually appealing analysis results. This allows for the provision of more appropriate analysis results by adjusting the presentation of the analysis according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0080] The analysis unit can adjust the level of detail of the analysis based on the importance of the information during the analysis. For example, the analysis unit performs a detailed analysis on important information. The analysis unit can also perform a concise analysis on general information. Furthermore, the analysis unit can adjust the level of detail of the analysis based on the user's level of interest. For example, the analysis unit performs a concise analysis on general information. Furthermore, the analysis unit can adjust the level of detail of the analysis based on the user's level of interest. By adjusting the level of detail of the analysis based on the importance of the information, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0081] The analysis unit can apply different analysis algorithms depending on the category of information during analysis. For example, the analysis unit can use a specialized algorithm for grammatical analysis. The analysis unit can also use a speech recognition algorithm for pronunciation analysis. Furthermore, the analysis unit can use a recommendation system for analyzing preferences and tastes. By applying different analysis algorithms depending on the category of information, more appropriate analysis results can be provided. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0082] The analysis unit can estimate the user's emotions and adjust the length of the analysis based on the estimated emotions. For example, if the user is in a hurry, the analysis unit will provide a short analysis result. The analysis unit can also provide a detailed analysis result if the user is relaxed. Furthermore, if the user is excited, the analysis unit can provide a visually appealing analysis result. For example, if the user is relaxed, the analysis unit will provide a detailed analysis result. The analysis unit can also provide a visually appealing analysis result if the user is excited. By adjusting the length of the analysis according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0083] The analysis unit can determine the priority of analysis based on the timing of information submission during the analysis. For example, the analysis unit may prioritize the analysis of the most recent information. The analysis unit may also prioritize the analysis of information with an approaching submission deadline. Furthermore, the analysis unit may also determine the priority of analysis based on the user's schedule. For example, the analysis unit may prioritize the analysis of information with an approaching submission deadline. The analysis unit may also determine the priority of analysis based on the timing of information submission, thereby providing more appropriate analysis results. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0084] The analysis unit can adjust the order of analysis based on the relevance of the information during the analysis. For example, the analysis unit prioritizes the analysis of highly relevant information. The analysis unit can also adjust the order of analysis based on the user's level of interest. Furthermore, the analysis unit can adjust the order of analysis based on the importance of the information. For example, the analysis unit adjusts the order of analysis based on the user's level of interest. Furthermore, the analysis unit can adjust the order of analysis based on the importance of the information. By adjusting the order of analysis based on the relevance of the information, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI.

[0085] The generation unit can estimate the user's emotions and adjust the way it expresses the generated output based on the estimated user emotions. For example, if the user is relaxed, the generation unit can generate detailed output. For example, if the user is relaxed, the generation unit can generate detailed output. The generation unit can also generate concise output if the user is tense. Furthermore, if the user is excited, the generation unit can generate visually appealing output. For example, if the user is tense, the generation unit can generate concise output. Furthermore, if the user is excited, the generation unit can also generate visually appealing output. This allows for the provision of more appropriate output by adjusting the way the output is expressed according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0086] The generation unit can adjust the level of detail of the generated output based on the importance of the information during generation. For example, the generation unit generates detailed output for important information. The generation unit can also generate concise output for general information. Furthermore, the generation unit can adjust the level of detail of the generated output based on the user's level of interest. For example, the generation unit generates concise output for general information. The generation unit can also adjust the level of detail of the generated output based on the user's level of interest. By adjusting the level of detail of the generated output based on the importance of the information, a more appropriate output can be provided. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0087] The generation unit can apply different generation algorithms depending on the category of information during generation. For example, the generation unit can use a specialized algorithm for grammar generation. The generation unit can also use a speech recognition algorithm for pronunciation generation. Furthermore, the generation unit can use a recommendation system for generating preferences and tastes. By applying different generation algorithms depending on the category of information, it is possible to provide more appropriate output. Some or all of the above-described processes in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0088] The generation unit can estimate the user's emotions and adjust the length of the output it generates based on the estimated emotions. For example, if the user is in a hurry, the generation unit will generate a short output. The generation unit can also generate a detailed output if the user is relaxed. Furthermore, if the user is excited, the generation unit can generate a visually appealing output. For example, if the user is relaxed, the generation unit will generate a detailed output. Furthermore, if the user is excited, the generation unit can also generate a visually appealing output. This allows for the provision of more appropriate output by adjusting the length of the output according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generation AI. Generation AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0089] The generation unit can determine the generation priority based on the information submission timing during generation. For example, the generation unit can prioritize generating the latest information. The generation unit can also prioritize generating information with an approaching submission deadline. Furthermore, the generation unit can also determine the generation priority based on the user's schedule. For example, the generation unit can prioritize generating information with an approaching submission deadline. The generation unit can also determine the generation priority based on the user's schedule. This allows for the provision of more appropriate output by determining the generation priority based on the information submission timing. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0090] The generation unit can adjust the order of generation based on the relevance of the information during generation. For example, the generation unit can prioritize the generation of highly relevant information. The generation unit can also adjust the order of generation based on the user's level of interest. Furthermore, the generation unit can also adjust the order of generation based on the importance of the information. By adjusting the order of generation based on the relevance of the information, a more appropriate output can be provided. Some or all of the above processing in the generation unit may be performed using, for example, a generation AI, or without using a generation AI.

[0091] The output unit can estimate the user's emotions and adjust the way it presents the output based on the estimated emotions. For example, if the user is relaxed, the output unit can provide detailed output. For example, if the user is relaxed, the output unit can provide detailed output. The output unit can also provide concise output if the user is tense. Furthermore, if the user is excited, the output unit can provide visually appealing output. For example, if the user is tense, the output unit can provide concise output. Furthermore, if the user is excited, the output unit can also provide visually appealing output. This allows for the provision of more appropriate output by adjusting the way the output is presented according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0092] The delivery unit can select the optimal delivery method by referring to the user's past learning history at the time of delivery. For example, the delivery unit provides relevant output based on what the user has learned in the past. For example, the delivery unit provides relevant output based on what the user has learned in the past. The delivery unit can also suggest the optimal learning method from the user's learning history. Furthermore, the delivery unit can provide appropriate output according to the user's learning progress. For example, the delivery unit suggests the optimal learning method from the user's learning history. Furthermore, the delivery unit can also provide appropriate output according to the user's learning progress. This allows for the provision of more appropriate output by selecting the optimal delivery method based on the user's past learning history. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without using AI.

[0093] The service provider can estimate the user's emotions and determine the priority of the output to be provided based on the estimated emotions. For example, if the user is nervous, the service provider may prioritize providing simple output. The service provider may also prioritize providing detailed output if the user is relaxed. Furthermore, if the user is in a hurry, the service provider may prioritize providing output that can be processed quickly. This allows for the provision of more appropriate output by prioritizing output according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0094] The delivery unit can select the optimal delivery method at the time of delivery, taking into account the user's device information. For example, if the user is using a smartphone, the delivery unit can provide output that matches the screen size. For example, if the user is using a smartphone, the delivery unit can provide output that matches the screen size. The delivery unit can also provide output optimized for a larger screen if the user is using a tablet. Furthermore, if the user is using a smartwatch, the delivery unit can provide concise and highly visible output. For example, if the user is using a tablet, the delivery unit can provide output optimized for a larger screen. Furthermore, if the user is using a smartwatch, the delivery unit can also provide concise and highly visible output. This allows for the provision of more appropriate output by selecting the optimal delivery method based on the user's device information. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI.

[0095] The service provider can analyze the user's learning history and propose an optimal learning plan at the time of delivery. For example, the service provider can propose a relevant learning plan based on the content the user has previously learned. The service provider can also propose an optimal learning method based on the user's learning history. Furthermore, the service provider can propose an appropriate learning plan according to the user's learning progress. For example, the service provider can propose an optimal learning method based on the user's learning history. The service provider can also propose an appropriate learning plan according to the user's learning progress. In this way, by proposing an optimal learning plan based on the user's learning history, more appropriate learning can be provided. Some or all of the above processing in the service provider may be performed using AI, for example, or without using AI.

[0096] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0097] The reception desk can estimate the user's current mood and emotions when receiving user input, and suggest input methods based on those estimates. For example, if the user is stressed, the reception desk can suggest a simple input method that helps them relax. If the user is focused, it can also suggest a more detailed input method. Furthermore, if the user is tired, it can display a message encouraging them to take a break. By providing the optimal input method according to the user's emotions, it can support more effective learning.

[0098] The analysis unit can estimate the user's emotions when analyzing user input and adjust the depth and detail of the analysis based on the estimated emotions. For example, if the user is relaxed, it can provide detailed analysis results. If the user is tense, it can provide concise analysis results. Furthermore, if the user is excited, it can provide visually appealing analysis results. This allows for more appropriate feedback by providing analysis results that match the user's emotions.

[0099] The generation unit can analyze the user's pronunciation, estimate the user's emotions, and suggest pronunciation practice methods based on those emotions. For example, if the user is nervous, it can suggest simple pronunciation practice. If the user is relaxed, it can suggest more detailed pronunciation practice. Furthermore, if the user is excited, it can suggest visually engaging pronunciation practice. By providing pronunciation practice tailored to the user's emotions, it can support more effective learning.

[0100] The system can estimate the user's emotions when providing a learning plan and adjust the plan's content based on those emotions. For example, if the user is feeling stressed, it can suggest a learning plan that helps them relax. If the user is focused, it can suggest a more detailed learning plan. Furthermore, if the user is tired, it can suggest a learning plan that includes breaks. By providing an optimal learning plan tailored to the user's emotions, the system can support more effective learning.

[0101] The system analyzes the user's learning history and, when proposing the optimal learning plan, can estimate the user's emotions and prioritize the learning plan based on those emotions. For example, if the user is feeling stressed, it can prioritize simple learning plans. Conversely, if the user is relaxed, it can prioritize detailed learning plans. Furthermore, if the user is in a hurry, it can prioritize learning plans that can be completed quickly. By providing the optimal learning plan tailored to the user's emotions, it can support more effective learning.

[0102] The reception desk can analyze the user's past input history when receiving user input and suggest the most suitable input method. For example, it can prioritize suggesting input methods that the user has frequently used in the past (such as voice or text). It can also suggest similar input methods based on the user's past input content. Furthermore, it can suggest input methods suitable for specific time periods based on the user's past input history. This enables efficient input by selecting the optimal input method based on the user's past input history.

[0103] The analytics unit can perform more detailed analyses of user preferences and tastes based on the user's past behavior history and survey results. For example, it can analyze the user's past video viewing history to identify their preferences. It can also analyze the content of text the user has entered in the past to identify their preferences. Furthermore, it can analyze the user's past search history to identify their areas of interest. This allows for the provision of more appropriate learning content based on the user's preferences and tastes.

[0104] The generation unit, when analyzing a user's writing, not only performs grammatical checks and evaluates content, but can also provide more personalized feedback based on the user's past writing history. For example, it can identify grammatical errors the user has made in the past and provide advice to prevent similar errors. It can also evaluate the content of the user's past writing and indicate areas for improvement. Furthermore, it can provide more specific editing guidance based on the user's writing style and tone. In this way, analyzing the user's writing enables more personalized editing guidance.

[0105] The service provider analyzes the user's learning history and proposes an optimal learning plan, customizing the plan based on the user's learning progress and goals. For example, if a user has difficulty with a particular skill, the service provider can propose a learning plan that focuses on strengthening that skill. It can also suggest learning content of appropriate difficulty level according to the user's learning progress. Furthermore, it can set goals to be achieved based on the user's learning objectives and propose a learning plan to achieve them. This allows for efficient learning by proposing an optimal learning plan based on the user's learning history.

[0106] The service provider can analyze a user's learning history and propose an optimal learning plan, taking into account the user's geographical location to suggest highly relevant learning content. For example, if a user is in a specific region, it can prioritize suggesting learning content related to that region. If a user is traveling, it can also suggest learning content related to their travel destination. Furthermore, if a user is at home, it can suggest learning content related to their home. This allows for more appropriate learning by suggesting highly relevant learning content based on the user's geographical location.

[0107] The following briefly describes the processing flow for example form 2.

[0108] Step 1: The reception desk receives user input. User input includes text input, voice input, and image input. For example, the reception desk accepts text input using a keyboard or touchscreen, voice input using a microphone, and image input using a camera. Step 2: The analysis unit analyzes the information received by the reception unit. The analysis unit uses methods such as text analysis, speech analysis, and image analysis. For example, it uses natural language processing technology to analyze grammatical and vocabulary errors in text, speech recognition technology to analyze pronunciation errors in speech, and image recognition technology to analyze the content of images. Step 3: The generation unit generates output based on the information analyzed by the analysis unit. The generation unit generates output in formats such as text output, audio output, and image output. For example, it can use a text generation AI to summarize the user's writing, an audio generation AI to correct the user's pronunciation, and an image generation AI to generate a video tailored to the user's preferences. Step 4: The provider unit provides the output generated by the generator unit. The provider unit provides the output through a web application or a mobile application. For example, it may display generated text to the user through a web application, play generated audio to the user through a mobile application, or stream generated video.

[0109] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0110] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0111] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0112] Each of the multiple elements described above, including the reception unit, analysis unit, generation unit, and provision unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives user input using the keyboard, touchscreen, microphone, and camera of the smart device 14. The analysis unit performs text analysis, speech analysis, and image analysis using the specific processing unit 290 of the data processing unit 12. The generation unit generates output using the text generation AI, speech generation AI, and image generation AI using the specific processing unit 290 of the data processing unit 12. The provision unit provides the generated output to the user through the control unit 46A of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0113] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0114] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0116] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0120] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0123] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0125] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0127] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0128] Each of the multiple elements described above, including the reception unit, analysis unit, generation unit, and provision unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives user input using the microphone and camera of the smart glasses 214. The analysis unit performs text analysis, speech analysis, and image analysis using the specific processing unit 290 of the data processing unit 12. The generation unit generates output using the text generation AI, speech generation AI, and image generation AI using the specific processing unit 290 of the data processing unit 12. The provision unit provides the generated output to the user through the control unit 46A of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0129] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0130] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0132] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0136] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0139] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0141] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0143] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0144] Each of the multiple elements described above, including the reception unit, analysis unit, generation unit, and provision unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives user input using the microphone and camera of the headset terminal 314. The analysis unit performs text analysis, speech analysis, and image analysis using the specific processing unit 290 of the data processing unit 12. The generation unit generates output using the text generation AI, speech generation AI, and image generation AI using the specific processing unit 290 of the data processing unit 12. The provision unit provides the generated output to the user through the control unit 46A of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0145] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0146] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0152] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0153] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0154] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0155] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0156] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0157] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0158] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0159] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0160] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0161] Each of the multiple elements described above, including the reception unit, analysis unit, generation unit, and provision unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives user input using the microphone and camera of the robot 414. The analysis unit performs text analysis, speech analysis, and image analysis using the specific processing unit 290 of the data processing unit 12. The generation unit generates output using the text generation AI, speech generation AI, and image generation AI using the specific processing unit 290 of the data processing unit 12. The provision unit provides the generated output to the user through the control unit 46A of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0162] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0163] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0164] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0165] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0166] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0167] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0169] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0170] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0172] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0173] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0174] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0175] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0176] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0177] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0178] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0179] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0180] (Note 1) A reception area that receives user input, An analysis unit that analyzes the information received by the reception unit, A generation unit that generates output based on the information analyzed by the analysis unit, The system comprises a providing unit that provides the output generated by the generating unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, Analyze user preferences and tastes. The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is Analyze the user's pronunciation. The system described in Appendix 1, characterized by the features described herein. (Note 4) The generating unit is Analyze the user's writing. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned supply unit is, We provide users with the most suitable learning plan. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned supply unit is, It analyzes the user's learning history and proposes the optimal learning plan. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving input, filtering is performed based on the user's current learning status and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is It estimates the user's emotions and determines the priority of input to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and accepts relevant input. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the representation of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, adjust the level of detail based on the importance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, It estimates the user's emotions and adjusts the length of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During the analysis, the priority of the analysis is determined based on when the information was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is It estimates the user's emotions and adjusts how the output generated based on those estimated emotions is represented. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is During generation, adjust the level of detail based on the importance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is During generation, different generation algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is It estimates the user's emotions and adjusts the length of the output generated based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is During generation, the generation priority is determined based on when the information was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is During generation, the order of generation is adjusted based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned supply unit is, It estimates the user's emotions and adjusts how the output is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned supply unit is, When providing the content, the system will refer to the user's past learning history to select the most suitable delivery method. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned supply unit is, It estimates the user's emotions and determines the priority of the output to be provided based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned supply unit is, When providing the service, the optimal delivery method will be selected, taking into account the user's device information. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned supply unit is, Upon delivery, the system analyzes the user's learning history and proposes the optimal learning plan. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0181] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception area that receives user input, An analysis unit that analyzes the information received by the reception unit, A generation unit that generates output based on the information analyzed by the analysis unit, The system comprises a providing unit that provides the output generated by the generating unit. A system characterized by the following features.

2. The aforementioned analysis unit, Analyze user preferences and tastes. The system according to feature 1.

3. The generating unit is Analyze the user's pronunciation. The system according to feature 1.

4. The generating unit is Analyze the user's writing. The system according to feature 1.

5. The aforementioned supply unit is, We provide users with the most suitable learning plan. The system according to feature 1.

6. The aforementioned supply unit is, It analyzes the user's learning history and proposes the optimal learning plan. The system according to feature 1.

7. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system according to feature 1.

8. The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A