system

The system addresses the challenge of maintaining motivation and adapting to individual learning progress by generating personalized, visually engaging, and emotionally responsive educational content, enhancing learning efficiency and motivation.

JP2026069003APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional learning methods struggle to maintain children's motivation and adapt to individual learning progress, especially in high-stakes examinations like junior high school entrance exams, lacking efficient and personalized educational systems that integrate visual aids and emotional feedback.

Method used

A system that collects data from an educational database, generates quiz questions automatically, integrates images and diagrams, and provides interactive learning experiences tailored to individual progress, using a generative AI model to adjust content based on user responses and emotional analysis.

Benefits of technology

Enhances learning efficiency and motivation by providing personalized, visually engaging, and emotionally responsive educational content, allowing users to learn at their own pace and adapt to their unique learning needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069003000001_ABST
    Figure 2026069003000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of accessing an educational database and obtaining data for the subject being studied, A means of analyzing acquired data and automatically generating quiz questions, A means for selecting images or diagrams related to the generated quiz questions and integrating them visually, A means of sending integrated quizzes and visuals to the user's terminal, A means of collecting and analyzing user responses, A means of suggesting the next learning content based on the user's learning progress, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0003] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] As the number of examinees taking the junior high school entrance examination increases, efficient learning has become an important factor in success or failure. In conventional learning methods, it is difficult for children to maintain their motivation for studying, and it is also difficult to adjust learning according to individual progress. Against this background, there is a demand for a system that allows children to study efficiently while having fun.

Means for Solving the Problems

[0005] This invention provides a means for collecting data on target subjects from an educational database and automatically generating quiz questions through analysis. It also integrates images and diagrams related to the generated quiz questions and delivers them to the user's terminal as visually appealing content. Furthermore, it promotes an interactive learning experience through an interface incorporating manga-like elements. By analyzing the user's answer data and suggesting the next learning content, it enables customized learning tailored to each individual's learning progress.

[0006] An "educational database" is a collection of information that stores and allows for the searching and extraction of text, images, and diagrams related to learning.

[0007] "Methods for automatically generating quiz questions" refers to a function that automatically creates quizzes in the form of multiple-choice questions, fill-in-the-blank questions, etc., based on analyzed data.

[0008] "Visual integration" refers to a function that combines generated quiz questions with related images and diagrams to create a visually integrated display.

[0009] "Means of sending to user terminals" refers to the function of sending generated quizzes and visuals from the server to the user's computer or mobile device.

[0010] "Means for collecting and analyzing user responses" refers to a function that collects data from user responses and analyzes it to evaluate learning progress and understanding.

[0011] "A means of suggesting the next learning content based on learning progress" refers to a function that recommends the next learning items and quiz difficulty levels to tackle, taking into account the user's previous learning evaluations. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] To implement this invention, first, an educational database is prepared. The database stores text information, images, charts, and other data related to the subjects being studied, such as Japanese language, mathematics, science, and social studies.

[0034] The server connects to the aforementioned educational database and retrieves data for the subject selected by the user. The retrieved data is then analyzed using natural language processing and image recognition technologies. This analysis process extracts important keywords and concepts, which are then used to automatically generate quiz questions. These quiz questions are created in various formats, including multiple-choice and fill-in-the-blank questions.

[0035] When a quiz is generated, the server picks relevant images and diagrams from a database and creates a manga-style visual. This visual content is designed to allow users to intuitively understand the learning material.

[0036] The generated quiz and visual content are sent from the server to the device. The device receives this data and displays it to the user in an interactive format. The user enters their answers to the presented questions, and these answers are sent to the server in real time.

[0037] The server analyzes the data submitted by the user and evaluates their learning progress and weaknesses. Based on the evaluation results, it adjusts the next learning content and quiz difficulty level, and proposes an appropriate learning plan for the user. This entire process allows users to learn efficiently at their own pace.

[0038] As a concrete example, when generating a quiz on fraction calculations in arithmetic, the server retrieves rules and example problems for fraction calculations from a database and creates multiple-choice questions that allow the user to practice fraction addition. In addition, diagrams and charts that help visualize fractions are displayed in the explanation of the problems, allowing the user to deepen their understanding of the problems. In this way, the present invention is implemented with the aim of maximizing the learning effect of the user.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The server accesses an educational database and retrieves text, images, and diagrams for the subject selected by the user. This data is prepared as material necessary for generating quizzes.

[0042] Step 2:

[0043] The server applies natural language processing to the acquired text data to analyze important keywords and concepts. During this analysis, it extracts content that should be included in the questions and related information.

[0044] Step 3:

[0045] The server automatically generates quiz questions based on the analysis results. These questions are structured as multiple-choice and fill-in-the-blank questions and include multiple answer choices and different formats to suit different learning styles.

[0046] Step 4:

[0047] The server selects images and diagrams related to the generated quiz from the database and integrates them as visual content, creating manga-like elements that aid intuitive understanding.

[0048] Step 5:

[0049] The server sends the generated quiz and visual content to the user's terminal. At this point, communication stability and data integrity are checked.

[0050] Step 6:

[0051] The terminal displays the received quiz and visuals as a user interface, and adjusts the operation to allow the user to easily input answers.

[0052] Step 7:

[0053] Users enter their answers to the presented quiz by selecting from multiple-choice options or providing written answers. This input is sent from the terminal to the server in real time.

[0054] Step 8:

[0055] The server analyzes the user's answers and calculates performance data, including correctness ratings and answer time. This allows the user's learning progress to be evaluated.

[0056] Step 9:

[0057] Based on the analysis results, the server adjusts the next learning content and quiz difficulty level, proposing a personalized learning plan to the user. This ensures that content optimized for individual learning needs is provided.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] Conventional learning support systems have not adequately customized their systems to suit each user's learning progress and level of understanding, making it difficult to achieve effective individualized learning. Furthermore, visual learning aids are limited, lacking opportunities for learners to intuitively grasp the learning content. Additionally, the provision of learning plans is fixed, lacking dynamic adjustments that adapt to the user's learning situation.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for accessing an educational information storage device and acquiring information from the learning target information area, means for analyzing the acquired information and automatically generating test questions, and means for selecting visual media or figures related to the generated test questions and visually integrating them. This makes it possible to realize effective individualized learning according to each user's learning progress, provide visually intuitive learning support, and dynamically adjust the learning plan.

[0063] An "educational information storage device" is a system that stores learning materials and education-related data and provides them in a format that can be accessed as needed.

[0064] A "learning target information domain" is a part of a database that contains learning materials related to a specific subject or theme.

[0065] "Exam questions" are questions or tasks designed for educational purposes and used to assess a user's knowledge and understanding.

[0066] "Visual media" refers to visual information materials, including images and diagrams, that serve to visually support learning.

[0067] A "dynamic information display device" is a display device that provides a visual interface that users can use interactively during learning, and that adapts in real time.

[0068] A "generative AI model" is a type of artificial intelligence that uses machine learning techniques to analyze data and generate new data or perform pattern recognition.

[0069] A "server" is a computer system that shares data over a network and processes information in response to requests from clients.

[0070] A "user device" is a device used by a user to receive, manipulate, and display information, and generally includes personal computers and smartphones.

[0071] To implement this invention, the server first connects to an educational information storage device and retrieves necessary data from the learning target information area. This data includes text data, images, and figures related to basic subjects such as Japanese language, mathematics, science, and social studies. The server analyzes the retrieved data using natural language processing and image recognition technologies. Through this analysis, important keywords and concepts are extracted, and test questions are created based on them.

[0072] The system automatically generates test questions according to the user's learning progress and provides them in various formats, such as multiple-choice and fill-in-the-blank questions. For example, when generating questions about fraction calculations in arithmetic, the server retrieves rules and examples of fraction calculations from a database and creates questions that allow the user to practice fraction addition. Furthermore, it selects relevant visual aids and shapes and uses a generative AI model to create manga-style visual content.

[0073] The generated test questions and visual content are sent from the server to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The user answers the questions through this interface, and the data is sent to the server in real time. The server analyzes the user's answers, provides feedback based on learning progress and understanding, and suggests the next learning content as needed.

[0074] As a concrete example, by inputting a prompt message such as "Generate a quiz about fraction addition" into an AI model, the server can automatically generate questions to deepen the user's understanding. In this way, users can learn efficiently at their own pace.

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The server connects to the educational information storage device and searches the learning target information area specified by the user to retrieve the necessary data. This process involves sending search queries and retrieving learning resources such as text data, images, and diagrams. The input is the user's subject selection, and the output is material data related to the specified subject.

[0078] Step 2:

[0079] The server analyzes the acquired data using natural language processing and image recognition technologies. Specifically, it extracts key keywords and concepts from text data using natural language processing and recognizes visual information from images. The input is the acquired training material, and the output is the extracted keywords and identified visual information.

[0080] Step 3:

[0081] The server automatically generates test questions based on keywords and concepts obtained through analysis. At this stage, a generation AI model is used to create questions of varying difficulty levels according to the user's level of understanding and learning progress. The input is the extracted keywords, and the output is the generated test questions. Specifically, multiple-choice and fill-in-the-blank questions are generated in a variety of formats.

[0082] Step 4:

[0083] The server selects visual media and figures related to the generated exam questions and processes them to integrate them visually. Using a generative AI model, it automatically generates manga-style visual content. This process includes retrieving appropriate images from a database and creating visual content based on them. The input is the generated exam questions and associated images, and the output is the integrated visual content.

[0084] Step 5:

[0085] The server sends integrated exam questions and visual content to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The display takes an intuitive form that the user can easily operate. The input is the integrated data sent from the server, and the output is the learning interface presented to the user.

[0086] Step 6:

[0087] The user enters their answers to the displayed test questions and sends those answers to the server via their device. The server receives this answer data, determines whether it is correct or incorrect, and generates feedback. The input is the user's answer data, and the output is the evaluation result and feedback.

[0088] Step 7:

[0089] The server uses a generative AI model to evaluate the user's learning progress and understanding based on their answer data, and to propose the next learning plan. It analyzes each user's learning pattern and automatically suggests the appropriate next step. The input is the user's answer history, and the output is a customized learning plan.

[0090] (Application Example 1)

[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0092] Modern learning methods demand educational materials that enable efficient understanding of information and its usage. Furthermore, improving the customer experience in physical stores requires a system that allows customers to learn about product background information while enjoying the experience. However, traditional methods make it difficult to directly link information about individual products to learning, posing challenges to efficient learning promotion and improved customer engagement.

[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0094] In this invention, the server includes means for accessing an educational database and acquiring information, means for analyzing the acquired information and automatically generating multiple-choice questions, and means for reading product identification information, receiving related information, and displaying it to the user. This allows users to learn background knowledge about products in an enjoyable way while receiving visual feedback within a physical store, thereby deepening their understanding of the products and improving the customer experience.

[0095] An "educational database" is an information aggregation system that stores information related to the subject being studied.

[0096] An "information processing terminal" is a device that allows users to send, receive, and display data through an interface.

[0097] "Visual elements" refer to expressive techniques that facilitate user understanding through visual information such as images and diagrams.

[0098] "Product identification information" refers to information used to distinguish a specific item, such as an ID code or barcode.

[0099] A "multiple-choice question" is a type of question where the user is presented with options and asked to select the correct answer.

[0100] "Proficiency level" is an indicator that shows the extent to which a user can understand and apply information.

[0101] To implement this invention, it is first necessary to prepare an educational database. This database contains information about the subject of learning, and the server using it retrieves information from this database. The server also receives product-related information through a QR code (registered trademark) reader and product identification information. When a user scans a QR code with an information processing terminal such as a smartphone, a request is sent to the server.

[0102] The server analyzes the acquired information using natural language processing libraries (such as NLTK and Transformers) implemented in Python, and then automatically generates multiple-choice questions. These questions utilize a generative AI model and include relevant options and information using prompts.

[0103] The information processing terminal presents the user with multiple-choice questions and visual elements sent from the server. These visual elements include images and diagrams related to the products and are visually integrated using image processing libraries such as OpenCV.

[0104] Users answer the displayed multiple-choice questions, and their answers are sent back to the server. The server analyzes the answers, determines the user's level of proficiency, and suggests the next learning steps and information. This allows users to deepen their background knowledge of products in an enjoyable way while in a physical store.

[0105] As a concrete example, in a certain tea shop, when a user scans a QR code related to the production method of green tea, a multiple-choice question such as "What production method is used to make this green tea?" is displayed. If the answer is correct, a reward that can be used on the next purchase is given. In this way, it is possible to promote learning in a physical store.

[0106] An example of a prompt to input into a generative AI model might be: "Generate a quiz to learn about the production methods of green tea. Also, include some interesting trivia about the history of green tea."

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The user scans the QR code attached to the product using the QR code reader on their smartphone. The input is the information from the QR code, and the output is the identification information of the acquired product. The device sends this identification information to the server.

[0110] Step 2:

[0111] The server retrieves relevant information from an educational database based on the received product identification information. The input is the product identification information, and the output is text information related to the product. The server analyzes this information using a natural language processing library and generates a multiple-choice question.

[0112] Step 3:

[0113] The server uses a generative AI model to generate prompts from acquired text information and automatically generate multiple-choice questions based on those prompts. The input is text information related to a product, and the output is a multiple-choice question. Specifically, it generates a prompt such as "Generate a quiz to learn about the production method of green tea," and then creates the question and answer choices.

[0114] Step 4:

[0115] The server retrieves images and diagrams related to the multiple-choice questions from a database and integrates the visual content using an image processing library. The input is the visual information associated with the multiple-choice questions, and the output is the visually integrated quiz content.

[0116] Step 5:

[0117] The server sends integrated quiz content to the device. The device displays this to the user, facilitating interactive learning. The input is the visually integrated quiz content, and the output is the display on the user's device.

[0118] Step 6:

[0119] The user answers the presented multiple-choice questions, and the terminal sends the answer information to the server. The input is the user's answer, and the output is the answer information sent to the server.

[0120] Step 7:

[0121] The server analyzes the user's answers and evaluates their proficiency level. The input is the user's answers, and the output is the user's proficiency evaluation. Based on this evaluation, the server generates data to suggest the next learning steps and information.

[0122] Step 8:

[0123] The server adjusts the next learning content based on the evaluation results and generates bonus information that can be used on the next visit. The input is the user's proficiency evaluation result, and the output is the next learning content and bonus information.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] This invention is a system that incorporates an emotion engine that recognizes the user's emotions, in addition to existing mechanisms that use an educational database to collect information on subjects to be studied and generate quiz questions.

[0126] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is analyzed, and quiz questions are automatically generated. The quiz questions are integrated with relevant images and diagrams, which are presented to the user in a visually appealing format.

[0127] The device receives quizzes and provides an easy-to-use interface for the user. It also features a camera and microphone to collect the user's facial expressions and voice tone. This data is sent to an emotion engine for real-time analysis.

[0128] The emotion engine analyzes the acquired user data to identify their emotional state. For example, it determines whether the user is confused, focused, or enjoying themselves. This result is then fed back to the server.

[0129] Based on the results of emotion recognition, the server adjusts the quiz content and difficulty level. If the user is confused, the questions can be simplified and supplementary explanations added. Conversely, if the user is focused, the difficulty level can be increased to present more challenging questions. This process makes it possible to maximize the learning effect for each individual.

[0130] For example, if a user is struggling with a difficult math problem, the emotion engine recognizes this as a state of confusion. As a result, the server breaks down the problem into simpler steps and presents a quiz with added hints to help solve it gradually. Conversely, if the user shows an expression of enjoyment, the server appropriately increases the difficulty level to provide a further challenge. In this way, the present invention provides an optimal learning experience that is tailored to the user's emotions.

[0131] The following describes the processing flow.

[0132] Step 1:

[0133] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is used as the basic material for generating quizzes.

[0134] Step 2:

[0135] The server analyzes the acquired data and automatically generates quiz questions using natural language processing. These questions include various formats, such as multiple-choice and fill-in-the-blank questions.

[0136] Step 3:

[0137] The server selects images and diagrams related to the generated quiz and visually integrates them to create visually appealing content.

[0138] Step 4:

[0139] The server sends the integrated quiz and visuals to the user's device, verifying data integrity and communication stability during this process.

[0140] Step 5:

[0141] The device displays the received quiz and visuals as an interactive user interface, and adjusts it to be easy for the user to operate.

[0142] Step 6:

[0143] The user answers the presented quiz, and their facial expressions and voice tone are recorded in real time through the camera and microphone built into the device.

[0144] Step 7:

[0145] The device transmits the user's facial expression data and voice data to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0146] Step 8:

[0147] The server receives feedback from the emotion engine and adjusts the quiz content and difficulty level according to the user's emotions. For example, if the user is confused, the questions are made easier and additional information is provided.

[0148] Step 9:

[0149] The server then resends the newly adjusted quiz questions to the device, allowing the user to continue learning. In this way, the learning experience is optimized based on the user's individual emotions.

[0150] (Example 2)

[0151] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0152] In modern education systems, it is difficult to consider the emotions and learning progress of individual learners in real time, and general educational content is not optimized for all learners. Furthermore, providing content in a uniform manner without understanding the emotional state of learners results in a failure to maximize learning effectiveness. In addition, there is a need to provide appropriate feedback based on emotions, but the means to achieve this are insufficient.

[0153] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0154] In this invention, the server includes means for accessing an educational database and acquiring information on the subject to be studied; means for analyzing the acquired information and automatically generating quiz questions using a generative AI model; and means for selecting images or charts related to the generated quiz questions and integrating them as visual elements. This makes it possible to analyze the learner's emotional state and adjust appropriate feedback and learning content in real time according to their emotions, thereby providing a learning experience optimized for each individual learner.

[0155] An "educational database" is a source of information for accumulating and managing information related to the subjects being studied, and is a database that is referenced to support users' learning activities.

[0156] A "generative AI model" is a model that uses artificial intelligence algorithms to analyze data and automatically generate specific tasks or processes, in this case, quiz questions.

[0157] A "quiz question" is an automatically generated question format intended for evaluation and feedback to learners, and is a means of measuring their level of understanding of the learning material.

[0158] "Images or diagrams" are visual elements used to provide learners with information visually and are integrated to complement quiz questions.

[0159] An "emotion analysis engine" is a system or tool that identifies and analyzes a learner's emotional state based on data from the user's facial expressions and voice.

[0160] "Visual integration" refers to selecting images or diagrams related to the generated quiz questions and integrating them for visual display.

[0161] "Real-time" means that data acquisition and analysis are performed instantly without delay, which enables immediate feedback tailored to the user's situation.

[0162] An "interactive learning experience" refers to a form of learning in which learners can acquire knowledge effectively and efficiently by interacting with the system in a two-way manner.

[0163] This invention is an educational support system incorporating a generative AI model and an emotion analysis engine. Specific embodiments are described below.

[0164] The server is responsible for accessing the educational database and retrieving information on the subject the user has selected to study. This information includes subject-related data and past exam questions. The server analyzes the retrieved information and uses a generative AI model to create quiz questions. This model learns from a large amount of data and automatically generates question formats suitable for the user.

[0165] The generated quiz questions are visually integrated by selecting relevant images and diagrams. This makes the quiz questions more intuitive and easier to understand. The integrated content is sent to the device and presented in an interactive interface for the user.

[0166] The device uses its camera and microphone to collect the user's facial expressions and voice tone in real time as they answer quizzes. This process is crucial for accurately capturing the user's reactions and improving the personalized learning experience.

[0167] The collected data is analyzed by an emotion analysis engine. This engine detects the user's emotional state based on facial recognition technology. For example, it can identify whether the user is confused, focused, or enjoying themselves. The analysis results are fed back to the server, forming the basis for adjusting the content and difficulty level of quiz questions in real time.

[0168] For example, if a user is faced with a complex math problem and the emotion analysis engine detects confusion from the user's facial expression, the server incorporates step-by-step hints into the quiz to make solving the problem easier. Conversely, if the analysis indicates the user is enjoying the problem, the server increases the difficulty level, providing further challenges to facilitate learning.

[0169] As an example of a prompt, the AI ​​generator can be input with the following message: "Based on the user's emotional state, please suggest the type of problem to present next." In this way, the present invention realizes effective learning support.

[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0171] Step 1:

[0172] The user selects the subject they wish to study through the interface on their device. This selected subject information is then sent to the server.

[0173] Step 2:

[0174] The server accesses an educational database based on the received subject information. It retrieves information related to the selected subject from the database. This retrieved information includes teaching materials, past exam questions, and related literature. This data is then used as input for analysis.

[0175] Step 3:

[0176] The server analyzes data obtained from an educational database and automatically generates quiz questions using a generative AI model. The analysis includes selecting appropriate topics and adjusting the question format and difficulty level. The resulting quiz questions are then output.

[0177] Step 4:

[0178] The server selects images or diagrams related to the generated quiz questions and integrates them as visual elements. This process aims to output a visualized integrated question, and the data is processed so that the user can understand it intuitively.

[0179] Step 5:

[0180] The server sends the integrated quiz questions to the terminal. The terminal displays the received quiz to the user in an easy-to-use format. Specifically, the question text, answer choices, images, and diagrams are arranged on the screen.

[0181] Step 6:

[0182] Users answer quizzes through their devices. During this process, the devices use their cameras and microphones to collect the user's facial expressions and voice tone in real time. This collected data is then sent to an emotion analysis engine.

[0183] Step 7:

[0184] The emotion analysis engine analyzes the received user data. This analysis identifies the user's emotional state, such as whether they are confused, focused, or enjoying themselves. The results of this analysis are then sent back to the server.

[0185] Step 8:

[0186] The server adjusts the quiz content and difficulty based on the results from the sentiment analysis engine. For example, if confusion is detected, the difficulty of the question is lowered or hints are added. This generates new quiz questions that are appropriate to the situation and sends them back to the terminal.

[0187] This series of steps provides an appropriate learning environment that takes user emotions into consideration.

[0188] (Application Example 2)

[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0190] In education, providing personalized learning experiences tailored to individual learners is crucial. However, traditional learning systems have struggled to provide appropriate learning content that takes into account the emotional state of users, posing challenges to motivating learners and promoting effective understanding. Furthermore, creating a visually appealing and interactive learning environment has also been difficult.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0192] In this invention, the server includes means for accessing an educational knowledge base and obtaining information on subjects to be studied; means for recognizing the user's emotional state and adjusting the content and difficulty level of quizzes in real time; and means for incorporating entertainment elements into the visual user interface to promote an interactive and emotionally sensitive learning experience. This makes it possible to provide optimal learning content tailored to the emotions of each individual learner, thereby improving learner comprehension and increasing motivation to learn.

[0193] An "educational knowledge base" is a database that systematically stores information about the subjects being studied, and serves as a foundation for providing quiz questions and related materials.

[0194] A "subject to be studied" refers to a specific educational field or theme that learners are expected to understand and master.

[0195] A "quiz-style assignment" is a question designed to allow learners to confirm their knowledge and deepen their understanding by answering it.

[0196] "Visual media" refers to digital content such as images and videos that contain visual information, and is intended to complement and enhance learning content.

[0197] A "user terminal" refers to a device that learners use to access and interact with learning content via the internet.

[0198] "User's emotional state" refers to the emotional state a learner exhibits during learning, and includes factors such as the degree of concentration and stress level.

[0199] "User learning progress" refers to an indicator that shows the extent to which a learner has acquired the target knowledge.

[0200] "Entertainment elements" refer to entertaining designs and content that aim to capture learners' interest and make the learning experience more enjoyable.

[0201] The system for realizing this invention is configured as an educational platform.

[0202] The server first connects to an educational knowledge base to retrieve information about the subject being studied. This retrieved information is then analyzed by a data processing algorithm to generate quiz-style questions. Next, relevant visual media are selected for the generated quiz and integrated into it. Finally, the integrated quiz is transmitted to the user's terminal via the internet.

[0203] The terminal has the function of receiving and collecting the user's answers. The terminal also has an integrated camera and microphone, which transmits the user's facial expressions and voice to the emotion engine. The emotion engine uses the image analysis library OpenCV and the deep learning framework TENSORFLOW® to analyze the user's emotional state in real time. Based on the analysis results, the server adjusts the quiz content and difficulty level as needed and retransmits the answers.

[0204] Users answer quizzes provided by the device and receive feedback tailored to their emotional state during the process. The device stores the user's learning progress and provides advice to help them in their next learning step.

[0205] As a concrete example, let's assume a learner is tackling a challenging history quiz. Initially, the learner shows a surprised expression, so the server immediately lowers the difficulty of the quiz and provides additional explanations. If emotional analysis indicates that the learner is enjoying themselves, the server provides feedback such as preparing a more challenging quiz next time.

[0206] An example of a prompt message might be, "Estimate the user's emotions from their facial expressions and voice, and generate optimal learning feedback."

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The server accesses an educational knowledge base to retrieve information related to the subject being studied. In this process, the server executes database queries using the subject ID as input. As a result, it outputs text, images, and videos related to the subject.

[0210] Step 2:

[0211] The server generates quiz-style questions based on the acquired information. Here, a generative AI model is used to analyze the relevant information using natural language processing to form the quiz questions and answer choices. The output is a set of quiz questions.

[0212] Step 3:

[0213] The server integrates appropriate visual media into the generated quiz. This involves using image processing algorithms to combine images and videos with the question text. A quiz with visual media is then generated.

[0214] Step 4:

[0215] The terminal displays quizzes received from the server to the user. The terminal's input is a quiz dataset from the server, and its output is a visual display for the user. This is implemented using a user interface library.

[0216] Step 5:

[0217] The user answers a quiz displayed on the terminal, and the terminal sends the answer to the server. The input for this step is the user's choices, and the output is the answer data packet.

[0218] Step 6:

[0219] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits them to the emotion engine. The input here is the user's real-time facial and voice data, and the output is a data stream to the emotion engine.

[0220] Step 7:

[0221] The emotion engine analyzes the acquired data to estimate the user's emotional state. Inputs are facial expressions and voice data, and output is the estimated emotional category. Specific operations include on-the-spot data analysis using TensorFlow.

[0222] Step 8:

[0223] The server receives the results of the emotion analysis and adjusts the quiz content and difficulty level. The output is a newly adjusted quiz set.

[0224] Step 9:

[0225] The device then presents the adjusted quiz data to the user again and provides interactive feedback. The input is the adjusted quiz, and the output is the result of the re-presentation to the user.

[0226] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0227] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0228] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0229] [Second Embodiment]

[0230] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0231] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0232] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0233] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0234] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0235] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0236] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0237] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0238] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0239] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0240] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0241] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0242] To implement this invention, first, an educational database is prepared. The database stores text information, images, charts, and other data related to the subjects being studied, such as Japanese language, mathematics, science, and social studies.

[0243] The server connects to the aforementioned educational database and retrieves data for the subject selected by the user. The retrieved data is then analyzed using natural language processing and image recognition technologies. This analysis process extracts important keywords and concepts, which are then used to automatically generate quiz questions. These quiz questions are created in various formats, including multiple-choice and fill-in-the-blank questions.

[0244] When a quiz is generated, the server picks relevant images and diagrams from a database and creates a manga-style visual. This visual content is designed to allow users to intuitively understand the learning material.

[0245] The generated quiz and visual content are sent from the server to the device. The device receives this data and displays it to the user in an interactive format. The user enters their answers to the presented questions, and these answers are sent to the server in real time.

[0246] The server analyzes the data submitted by the user and evaluates their learning progress and weaknesses. Based on the evaluation results, it adjusts the next learning content and quiz difficulty level, and proposes an appropriate learning plan for the user. This entire process allows users to learn efficiently at their own pace.

[0247] As a concrete example, when generating a quiz on fraction calculations in arithmetic, the server retrieves rules and example problems for fraction calculations from a database and creates multiple-choice questions that allow the user to practice fraction addition. In addition, diagrams and charts that help visualize fractions are displayed in the explanation of the problems, allowing the user to deepen their understanding of the problems. In this way, the present invention is implemented with the aim of maximizing the learning effect of the user.

[0248] The following describes the processing flow.

[0249] Step 1:

[0250] The server accesses an educational database and retrieves text, images, and diagrams for the subject selected by the user. This data is prepared as material necessary for generating quizzes.

[0251] Step 2:

[0252] The server applies natural language processing to the acquired text data to analyze important keywords and concepts. During this analysis, it extracts content that should be included in the questions and related information.

[0253] Step 3:

[0254] The server automatically generates quiz questions based on the analysis results. These questions are structured as multiple-choice and fill-in-the-blank questions and include multiple answer choices and different formats to suit different learning styles.

[0255] Step 4:

[0256] The server selects images and diagrams related to the generated quiz from the database and integrates them as visual content, creating manga-like elements that aid intuitive understanding.

[0257] Step 5:

[0258] The server sends the generated quiz and visual content to the user's terminal. At this point, communication stability and data integrity are checked.

[0259] Step 6:

[0260] The terminal displays the received quiz and visuals as a user interface, and adjusts the operation to allow the user to easily input answers.

[0261] Step 7:

[0262] Users enter their answers to the presented quiz by selecting from multiple-choice options or providing written answers. This input is sent from the terminal to the server in real time.

[0263] Step 8:

[0264] The server analyzes the user's answers and calculates performance data, including correctness ratings and answer time. This allows the user's learning progress to be evaluated.

[0265] Step 9:

[0266] Based on the analysis results, the server adjusts the next learning content and quiz difficulty level, proposing a personalized learning plan to the user. This ensures that content optimized for individual learning needs is provided.

[0267] (Example 1)

[0268] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0269] Conventional learning support systems have not adequately customized their systems to suit each user's learning progress and level of understanding, making it difficult to achieve effective individualized learning. Furthermore, visual learning aids are limited, lacking opportunities for learners to intuitively grasp the learning content. Additionally, the provision of learning plans is fixed, lacking dynamic adjustments that adapt to the user's learning situation.

[0270] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0271] In this invention, the server includes means for accessing an educational information storage device and acquiring information from the learning target information area, means for analyzing the acquired information and automatically generating test questions, and means for selecting visual media or figures related to the generated test questions and visually integrating them. This makes it possible to realize effective individualized learning according to each user's learning progress, provide visually intuitive learning support, and dynamically adjust the learning plan.

[0272] An "educational information storage device" is a system that stores learning materials and education-related data and provides them in a format that can be accessed as needed.

[0273] A "learning target information domain" is a part of a database that contains learning materials related to a specific subject or theme.

[0274] "Exam questions" are questions or tasks designed for educational purposes and used to assess a user's knowledge and understanding.

[0275] "Visual media" refers to visual information materials, including images and diagrams, that serve to visually support learning.

[0276] A "dynamic information display device" is a display device that provides a visual interface that users can use interactively during learning, and that adapts in real time.

[0277] A "generative AI model" is a type of artificial intelligence that uses machine learning techniques to analyze data and generate new data or perform pattern recognition.

[0278] A "server" is a computer system that shares data over a network and processes information in response to requests from clients.

[0279] A "user device" is a device used by a user to receive, manipulate, and display information, and generally includes personal computers and smartphones.

[0280] To implement this invention, the server first connects to an educational information storage device and retrieves necessary data from the learning target information area. This data includes text data, images, and figures related to basic subjects such as Japanese language, mathematics, science, and social studies. The server analyzes the retrieved data using natural language processing and image recognition technologies. Through this analysis, important keywords and concepts are extracted, and test questions are created based on them.

[0281] Test questions are automatically generated according to the user's learning progress and provided in various forms such as multiple-choice questions and fill-in-the-blank questions. For example, when generating questions related to fraction calculations in arithmetic, the server retrieves the rules and examples of fraction calculations from the database and creates questions for the user to practice adding fractions. Furthermore, relevant visual media and figures are selected, and manga-style visual content is created using a generative AI model.

[0282] The generated test questions and visual content are sent from the server to the terminal. The terminal receives this data and displays it to the user via an interactive user interface. The user answers the questions through this interface, and the data is sent to the server in real time. The server analyzes the user's answers, provides feedback based on the learning progress and understanding level, and proposes the next learning content if necessary.

[0283] As a specific example, by inputting a prompt sentence such as "Please generate a quiz on adding fractions" into the generative AI model by the server, questions for deepening the user's understanding can be automatically generated. In this way, the user can efficiently deepen their learning at their own pace.

[0284] The flow of the specific process in Example 1 will be described using FIG. 11.

[0285] Step 1:

[0286] The server connects to the educational information storage device, searches for the learning target information area specified by the user, and obtains the necessary data. In this process, a search query is sent to retrieve learning resources such as text data, images, and figures. The input is the user's subject selection, and the output is the material data related to the specified subject.

[0287] Step 2:

[0288] The server analyzes the acquired data using natural language processing and image recognition technologies. Specifically, it extracts key keywords and concepts from text data using natural language processing and recognizes visual information from images. The input is the acquired training material, and the output is the extracted keywords and identified visual information.

[0289] Step 3:

[0290] The server automatically generates test questions based on keywords and concepts obtained through analysis. At this stage, a generation AI model is used to create questions of varying difficulty levels according to the user's level of understanding and learning progress. The input is the extracted keywords, and the output is the generated test questions. Specifically, multiple-choice and fill-in-the-blank questions are generated in a variety of formats.

[0291] Step 4:

[0292] The server selects visual media and figures related to the generated exam questions and processes them to integrate them visually. Using a generative AI model, it automatically generates manga-style visual content. This process includes retrieving appropriate images from a database and creating visual content based on them. The input is the generated exam questions and associated images, and the output is the integrated visual content.

[0293] Step 5:

[0294] The server sends integrated exam questions and visual content to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The display takes an intuitive form that the user can easily operate. The input is the integrated data sent from the server, and the output is the learning interface presented to the user.

[0295] Step 6:

[0296] The user enters their answers to the displayed test questions and sends those answers to the server via their device. The server receives this answer data, determines whether it is correct or incorrect, and generates feedback. The input is the user's answer data, and the output is the evaluation result and feedback.

[0297] Step 7:

[0298] The server uses a generative AI model to evaluate the user's learning progress and understanding based on their answer data, and to propose the next learning plan. It analyzes each user's learning pattern and automatically suggests the appropriate next step. The input is the user's answer history, and the output is a customized learning plan.

[0299] (Application Example 1)

[0300] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0301] Modern learning methods demand educational materials that enable efficient understanding of information and its usage. Furthermore, improving the customer experience in physical stores requires a system that allows customers to learn about product background information while enjoying the experience. However, traditional methods make it difficult to directly link information about individual products to learning, posing challenges to efficient learning promotion and improved customer engagement.

[0302] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0303] In this invention, the server includes means for accessing an educational database and acquiring information, means for analyzing the acquired information and automatically generating multiple-choice questions, and means for reading product identification information, receiving related information, and displaying it to the user. This allows users to learn background knowledge about products in an enjoyable way while receiving visual feedback within a physical store, thereby deepening their understanding of the products and improving the customer experience.

[0304] An "educational database" is an information aggregation system that stores information related to learning targets.

[0305] An "information processing terminal" is a device for a user to transmit and receive and display data through an interface.

[0306] A "visual element" is an expression method that promotes user understanding through visual information such as images and charts.

[0307] "Product identification information" refers to information indicating information for distinguishing a specific article, such as an ID code or a barcode.

[0308] A "selection question" is a type of question that presents options to a user and asks the user to select the correct answer.

[0309] "Proficiency level" is an index indicating the level to which a user can understand and apply information.

[0310] To implement this invention, first, an educational database needs to be prepared. This database contains information related to learning targets, and the server to be used acquires information from this database. The server also receives information related to products through a QR code reader or product identification information. When a user scans a QR code with an information processing terminal such as a smartphone, a request is sent to the server.

[0311] The server analyzes the acquired information using a natural language processing library (such as NLTK or Transformers) implemented in Python, etc., and further automatically generates selection questions. These selection questions utilize a generation AI model and contain related options and information using prompt sentences.

[0312] The information processing terminal presents the user with multiple-choice questions and visual elements sent from the server. These visual elements include images and diagrams related to the products and are visually integrated using image processing libraries such as OpenCV.

[0313] Users answer the displayed multiple-choice questions, and their answers are sent back to the server. The server analyzes the answers, determines the user's level of proficiency, and suggests the next learning steps and information. This allows users to deepen their background knowledge of products in an enjoyable way while in a physical store.

[0314] As a concrete example, in a certain tea shop, when a user scans a QR code related to the production method of green tea, a multiple-choice question such as "What production method is used to make this green tea?" is displayed. If the answer is correct, a reward that can be used on the next purchase is given. In this way, it is possible to promote learning in a physical store.

[0315] An example of a prompt to input into a generative AI model might be: "Generate a quiz to learn about the production methods of green tea. Also, include some interesting trivia about the history of green tea."

[0316] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0317] Step 1:

[0318] The user scans the QR code attached to the product using the QR code reader on their smartphone. The input is the information from the QR code, and the output is the identification information of the acquired product. The device sends this identification information to the server.

[0319] Step 2:

[0320] The server retrieves relevant information from an educational database based on the received product identification information. The input is the product identification information, and the output is text information related to the product. The server analyzes this information using a natural language processing library and generates a multiple-choice question.

[0321] Step 3:

[0322] The server uses a generative AI model to generate prompts from acquired text information and automatically generate multiple-choice questions based on those prompts. The input is text information related to a product, and the output is a multiple-choice question. Specifically, it generates a prompt such as "Generate a quiz to learn about the production method of green tea," and then creates the question and answer choices.

[0323] Step 4:

[0324] The server retrieves images and diagrams related to the multiple-choice questions from a database and integrates the visual content using an image processing library. The input is the visual information associated with the multiple-choice questions, and the output is the visually integrated quiz content.

[0325] Step 5:

[0326] The server sends integrated quiz content to the device. The device displays this to the user, facilitating interactive learning. The input is the visually integrated quiz content, and the output is the display on the user's device.

[0327] Step 6:

[0328] The user answers the presented multiple-choice questions, and the terminal sends the answer information to the server. The input is the user's answer, and the output is the answer information sent to the server.

[0329] Step 7:

[0330] The server analyzes the user's answers and evaluates their proficiency level. The input is the user's answers, and the output is the user's proficiency evaluation. Based on this evaluation, the server generates data to suggest the next learning steps and information.

[0331] Step 8:

[0332] The server adjusts the next learning content based on the evaluation results and generates bonus information that can be used on the next visit. The input is the user's proficiency evaluation result, and the output is the next learning content and bonus information.

[0333] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0334] This invention is a system that incorporates an emotion engine that recognizes the user's emotions, in addition to existing mechanisms that use an educational database to collect information on subjects to be studied and generate quiz questions.

[0335] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is analyzed, and quiz questions are automatically generated. The quiz questions are integrated with relevant images and diagrams, which are presented to the user in a visually appealing format.

[0336] The device receives quizzes and provides an easy-to-use interface for the user. It also features a camera and microphone to collect the user's facial expressions and voice tone. This data is sent to an emotion engine for real-time analysis.

[0337] The emotion engine analyzes the acquired user data to identify their emotional state. For example, it determines whether the user is confused, focused, or enjoying themselves. This result is then fed back to the server.

[0338] Based on the results of emotion recognition, the server adjusts the quiz content and difficulty level. If the user is confused, the questions can be simplified and supplementary explanations added. Conversely, if the user is focused, the difficulty level can be increased to present more challenging questions. This process makes it possible to maximize the learning effect for each individual.

[0339] For example, if a user is struggling with a difficult math problem, the emotion engine recognizes this as a state of confusion. As a result, the server breaks down the problem into simpler steps and presents a quiz with added hints to help solve it gradually. Conversely, if the user shows an expression of enjoyment, the server appropriately increases the difficulty level to provide a further challenge. In this way, the present invention provides an optimal learning experience that is tailored to the user's emotions.

[0340] The following describes the processing flow.

[0341] Step 1:

[0342] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is used as the basic material for generating quizzes.

[0343] Step 2:

[0344] The server analyzes the acquired data and automatically generates quiz questions using natural language processing. These questions include various formats, such as multiple-choice and fill-in-the-blank questions.

[0345] Step 3:

[0346] The server selects images and diagrams related to the generated quiz and visually integrates them to create visually appealing content.

[0347] Step 4:

[0348] The server sends the integrated quiz and visuals to the user's device, verifying data integrity and communication stability during this process.

[0349] Step 5:

[0350] The device displays the received quiz and visuals as an interactive user interface, and adjusts it to be easy for the user to operate.

[0351] Step 6:

[0352] The user answers the presented quiz, and their facial expressions and voice tone are recorded in real time through the camera and microphone built into the device.

[0353] Step 7:

[0354] The device transmits the user's facial expression data and voice data to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0355] Step 8:

[0356] The server receives feedback from the emotion engine and adjusts the quiz content and difficulty level according to the user's emotions. For example, if the user is confused, the questions are made easier and additional information is provided.

[0357] Step 9:

[0358] The server then resends the newly adjusted quiz questions to the device, allowing the user to continue learning. In this way, the learning experience is optimized based on the user's individual emotions.

[0359] (Example 2)

[0360] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0361] In modern education systems, it is difficult to consider the emotions and learning progress of individual learners in real time, and general educational content is not optimized for all learners. Furthermore, providing content in a uniform manner without understanding the emotional state of learners results in a failure to maximize learning effectiveness. In addition, there is a need to provide appropriate feedback based on emotions, but the means to achieve this are insufficient.

[0362] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0363] In this invention, the server includes means for accessing an educational database and acquiring information on the subject to be studied; means for analyzing the acquired information and automatically generating quiz questions using a generative AI model; and means for selecting images or charts related to the generated quiz questions and integrating them as visual elements. This makes it possible to analyze the learner's emotional state and adjust appropriate feedback and learning content in real time according to their emotions, thereby providing a learning experience optimized for each individual learner.

[0364] An "educational database" is a source of information for accumulating and managing information related to the subjects being studied, and is a database that is referenced to support users' learning activities.

[0365] A "generative AI model" is a model that uses artificial intelligence algorithms to analyze data and automatically generate specific tasks or processes, in this case, quiz questions.

[0366] A "quiz question" is an automatically generated question format intended for evaluation and feedback to learners, and is a means of measuring their level of understanding of the learning material.

[0367] "Images or diagrams" are visual elements used to provide learners with information visually and are integrated to complement quiz questions.

[0368] An "emotion analysis engine" is a system or tool that identifies and analyzes a learner's emotional state based on data from the user's facial expressions and voice.

[0369] "Visual integration" refers to selecting images or diagrams related to the generated quiz questions and integrating them for visual display.

[0370] "Real-time" means that data acquisition and analysis are performed instantly without delay, which enables immediate feedback tailored to the user's situation.

[0371] An "interactive learning experience" refers to a form of learning in which learners can acquire knowledge effectively and efficiently by interacting with the system in a two-way manner.

[0372] This invention is an educational support system incorporating a generative AI model and an emotion analysis engine. Specific embodiments are described below.

[0373] The server is responsible for accessing the educational database and retrieving information on the subject the user has selected to study. This information includes subject-related data and past exam questions. The server analyzes the retrieved information and uses a generative AI model to create quiz questions. This model learns from a large amount of data and automatically generates question formats suitable for the user.

[0374] The generated quiz questions are visually integrated by selecting relevant images and diagrams. This makes the quiz questions more intuitive and easier to understand. The integrated content is sent to the device and presented in an interactive interface for the user.

[0375] The device uses its camera and microphone to collect the user's facial expressions and voice tone in real time as they answer quizzes. This process is crucial for accurately capturing the user's reactions and improving the personalized learning experience.

[0376] The collected data is analyzed by an emotion analysis engine. This engine detects the user's emotional state based on facial recognition technology. For example, it can identify whether the user is confused, focused, or enjoying themselves. The analysis results are fed back to the server, forming the basis for adjusting the content and difficulty level of quiz questions in real time.

[0377] For example, if a user is faced with a complex math problem and the emotion analysis engine detects confusion from the user's facial expression, the server incorporates step-by-step hints into the quiz to make solving the problem easier. Conversely, if the analysis indicates the user is enjoying the problem, the server increases the difficulty level, providing further challenges to facilitate learning.

[0378] As an example of a prompt, the AI ​​generator can be input with the following message: "Based on the user's emotional state, please suggest the type of problem to present next." In this way, the present invention realizes effective learning support.

[0379] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0380] Step 1:

[0381] The user selects the subject they wish to study through the interface on their device. This selected subject information is then sent to the server.

[0382] Step 2:

[0383] The server accesses an educational database based on the received subject information. It retrieves information related to the selected subject from the database. This retrieved information includes teaching materials, past exam questions, and related literature. This data is then used as input for analysis.

[0384] Step 3:

[0385] The server analyzes data obtained from an educational database and automatically generates quiz questions using a generative AI model. The analysis includes selecting appropriate topics and adjusting the question format and difficulty level. The resulting quiz questions are then output.

[0386] Step 4:

[0387] The server selects images or diagrams related to the generated quiz questions and integrates them as visual elements. This process aims to output a visualized integrated question, and the data is processed so that the user can understand it intuitively.

[0388] Step 5:

[0389] The server sends the integrated quiz questions to the terminal. The terminal displays the received quiz to the user in an easy-to-use format. Specifically, the question text, answer choices, images, and diagrams are arranged on the screen.

[0390] Step 6:

[0391] Users answer quizzes through their devices. During this process, the devices use their cameras and microphones to collect the user's facial expressions and voice tone in real time. This collected data is then sent to an emotion analysis engine.

[0392] Step 7:

[0393] The emotion analysis engine analyzes the received user data. This analysis identifies the user's emotional state, such as whether they are confused, focused, or enjoying themselves. The results of this analysis are then sent back to the server.

[0394] Step 8:

[0395] The server adjusts the quiz content and difficulty based on the results from the sentiment analysis engine. For example, if confusion is detected, the difficulty of the question is lowered or hints are added. This generates new quiz questions that are appropriate to the situation and sends them back to the terminal.

[0396] This series of steps provides an appropriate learning environment that takes user emotions into consideration.

[0397] (Application Example 2)

[0398] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0399] In education, providing personalized learning experiences tailored to individual learners is crucial. However, traditional learning systems have struggled to provide appropriate learning content that takes into account the emotional state of users, posing challenges to motivating learners and promoting effective understanding. Furthermore, creating a visually appealing and interactive learning environment has also been difficult.

[0400] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0401] In this invention, the server includes means for accessing an educational knowledge base and obtaining information on subjects to be studied; means for recognizing the user's emotional state and adjusting the content and difficulty level of quizzes in real time; and means for incorporating entertainment elements into the visual user interface to promote an interactive and emotionally sensitive learning experience. This makes it possible to provide optimal learning content tailored to the emotions of each individual learner, thereby improving learner comprehension and increasing motivation to learn.

[0402] An "educational knowledge base" is a database that systematically stores information about the subjects being studied, and serves as a foundation for providing quiz questions and related materials.

[0403] A "subject to be studied" refers to a specific educational field or theme that learners are expected to understand and master.

[0404] A "quiz-style assignment" is a question designed to allow learners to confirm their knowledge and deepen their understanding by answering it.

[0405] "Visual media" refers to digital content such as images and videos that contain visual information, and is intended to complement and enhance learning content.

[0406] A "user terminal" refers to a device that learners use to access and interact with learning content via the internet.

[0407] "User's emotional state" refers to the emotional state a learner exhibits during learning, and includes factors such as the degree of concentration and stress level.

[0408] "User learning progress" refers to an indicator that shows the extent to which a learner has acquired the target knowledge.

[0409] "Entertainment elements" refer to entertaining designs and content that aim to capture learners' interest and make the learning experience more enjoyable.

[0410] The system for realizing this invention is configured as an educational platform.

[0411] The server first connects to an educational knowledge base to retrieve information about the subject being studied. This retrieved information is then analyzed by a data processing algorithm to generate quiz-style questions. Next, relevant visual media are selected for the generated quiz and integrated into it. Finally, the integrated quiz is transmitted to the user's terminal via the internet.

[0412] The terminal has the functionality to receive and collect user answers. It also has an integrated camera and microphone, which transmits the user's facial expressions and voice to the emotion engine. The emotion engine uses the image analysis library OpenCV and the deep learning framework TensorFlow to analyze the user's emotional state in real time. Based on the analysis results, the server adjusts the quiz content and difficulty level as needed and retransmits the answers.

[0413] Users answer quizzes provided by the device and receive feedback tailored to their emotional state during the process. The device stores the user's learning progress and provides advice to help them in their next learning step.

[0414] As a concrete example, let's assume a learner is tackling a challenging history quiz. Initially, the learner shows a surprised expression, so the server immediately lowers the difficulty of the quiz and provides additional explanations. If emotional analysis indicates that the learner is enjoying themselves, the server provides feedback such as preparing a more challenging quiz next time.

[0415] An example of a prompt message might be, "Estimate the user's emotions from their facial expressions and voice, and generate optimal learning feedback."

[0416] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0417] Step 1:

[0418] The server accesses an educational knowledge base to retrieve information related to the subject being studied. In this process, the server executes database queries using the subject ID as input. As a result, it outputs text, images, and videos related to the subject.

[0419] Step 2:

[0420] The server generates quiz-style questions based on the acquired information. Here, a generative AI model is used to analyze the relevant information using natural language processing to form the quiz questions and answer choices. The output is a set of quiz questions.

[0421] Step 3:

[0422] The server integrates appropriate visual media into the generated quiz. This involves using image processing algorithms to combine images and videos with the question text. A quiz with visual media is then generated.

[0423] Step 4:

[0424] The terminal displays quizzes received from the server to the user. The terminal's input is a quiz dataset from the server, and its output is a visual display for the user. This is implemented using a user interface library.

[0425] Step 5:

[0426] The user answers a quiz displayed on the terminal, and the terminal sends the answer to the server. The input for this step is the user's choices, and the output is the answer data packet.

[0427] Step 6:

[0428] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits them to the emotion engine. The input here is the user's real-time facial and voice data, and the output is a data stream to the emotion engine.

[0429] Step 7:

[0430] The emotion engine analyzes the acquired data to estimate the user's emotional state. Inputs are facial expressions and voice data, and output is the estimated emotional category. Specific operations include on-the-spot data analysis using TensorFlow.

[0431] Step 8:

[0432] The server receives the results of the emotion analysis and adjusts the quiz content and difficulty level. The output is a newly adjusted quiz set.

[0433] Step 9:

[0434] The device then presents the adjusted quiz data to the user again and provides interactive feedback. The input is the adjusted quiz, and the output is the result of the re-presentation to the user.

[0435] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0436] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0437] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0438] [Third Embodiment]

[0439] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0440] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0441] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0442] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0443] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0445] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0446] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0449] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0451] To implement this invention, first, an educational database is prepared. The database stores text information, images, charts, and other data related to the subjects being studied, such as Japanese language, mathematics, science, and social studies.

[0452] The server connects to the aforementioned educational database and retrieves data for the subject selected by the user. The retrieved data is then analyzed using natural language processing and image recognition technologies. This analysis process extracts important keywords and concepts, which are then used to automatically generate quiz questions. These quiz questions are created in various formats, including multiple-choice and fill-in-the-blank questions.

[0453] When a quiz is generated, the server picks relevant images and diagrams from a database and creates a manga-style visual. This visual content is designed to allow users to intuitively understand the learning material.

[0454] The generated quiz and visual content are sent from the server to the device. The device receives this data and displays it to the user in an interactive format. The user enters their answers to the presented questions, and these answers are sent to the server in real time.

[0455] The server analyzes the data submitted by the user and evaluates their learning progress and weaknesses. Based on the evaluation results, it adjusts the next learning content and quiz difficulty level, and proposes an appropriate learning plan for the user. This entire process allows users to learn efficiently at their own pace.

[0456] As a concrete example, when generating a quiz on fraction calculations in arithmetic, the server retrieves rules and example problems for fraction calculations from a database and creates multiple-choice questions that allow the user to practice fraction addition. In addition, diagrams and charts that help visualize fractions are displayed in the explanation of the problems, allowing the user to deepen their understanding of the problems. In this way, the present invention is implemented with the aim of maximizing the learning effect of the user.

[0457] The following describes the processing flow.

[0458] Step 1:

[0459] The server accesses an educational database and retrieves text, images, and diagrams for the subject selected by the user. This data is prepared as material necessary for generating quizzes.

[0460] Step 2:

[0461] The server applies natural language processing to the acquired text data to analyze important keywords and concepts. During this analysis, it extracts content that should be included in the questions and related information.

[0462] Step 3:

[0463] The server automatically generates quiz questions based on the analysis results. These questions are structured as multiple-choice and fill-in-the-blank questions and include multiple answer choices and different formats to suit different learning styles.

[0464] Step 4:

[0465] The server selects images and diagrams related to the generated quiz from the database and integrates them as visual content, creating manga-like elements that aid intuitive understanding.

[0466] Step 5:

[0467] The server sends the generated quiz and visual content to the user's terminal. At this point, communication stability and data integrity are checked.

[0468] Step 6:

[0469] The terminal displays the received quiz and visuals as a user interface, and adjusts the operation to allow the user to easily input answers.

[0470] Step 7:

[0471] Users enter their answers to the presented quiz by selecting from multiple-choice options or providing written answers. This input is sent from the terminal to the server in real time.

[0472] Step 8:

[0473] The server analyzes the user's answers and calculates performance data, including correctness ratings and answer time. This allows the user's learning progress to be evaluated.

[0474] Step 9:

[0475] Based on the analysis results, the server adjusts the next learning content and quiz difficulty level, proposing a personalized learning plan to the user. This ensures that content optimized for individual learning needs is provided.

[0476] (Example 1)

[0477] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0478] Conventional learning support systems have not adequately customized their systems to suit each user's learning progress and level of understanding, making it difficult to achieve effective individualized learning. Furthermore, visual learning aids are limited, lacking opportunities for learners to intuitively grasp the learning content. Additionally, the provision of learning plans is fixed, lacking dynamic adjustments that adapt to the user's learning situation.

[0479] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0480] In this invention, the server includes means for accessing an educational information storage device and acquiring information from the learning target information area, means for analyzing the acquired information and automatically generating test questions, and means for selecting visual media or figures related to the generated test questions and visually integrating them. This makes it possible to realize effective individualized learning according to each user's learning progress, provide visually intuitive learning support, and dynamically adjust the learning plan.

[0481] An "educational information storage device" is a system that stores learning materials and education-related data and provides them in a format that can be accessed as needed.

[0482] A "learning target information domain" is a part of a database that contains learning materials related to a specific subject or theme.

[0483] "Exam questions" are questions or tasks designed for educational purposes and used to assess a user's knowledge and understanding.

[0484] "Visual media" refers to visual information materials, including images and diagrams, that serve to visually support learning.

[0485] A "dynamic information display device" is a display device that provides a visual interface that users can use interactively during learning, and that adapts in real time.

[0486] A "generative AI model" is a type of artificial intelligence that uses machine learning techniques to analyze data and generate new data or perform pattern recognition.

[0487] A "server" is a computer system that shares data over a network and processes information in response to requests from clients.

[0488] A "user device" is a device used by a user to receive, manipulate, and display information, and generally includes personal computers and smartphones.

[0489] To implement this invention, the server first connects to an educational information storage device and retrieves necessary data from the learning target information area. This data includes text data, images, and figures related to basic subjects such as Japanese language, mathematics, science, and social studies. The server analyzes the retrieved data using natural language processing and image recognition technologies. Through this analysis, important keywords and concepts are extracted, and test questions are created based on them.

[0490] The system automatically generates test questions according to the user's learning progress and provides them in various formats, such as multiple-choice and fill-in-the-blank questions. For example, when generating questions about fraction calculations in arithmetic, the server retrieves rules and examples of fraction calculations from a database and creates questions that allow the user to practice fraction addition. Furthermore, it selects relevant visual aids and shapes and uses a generative AI model to create manga-style visual content.

[0491] The generated test questions and visual content are sent from the server to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The user answers the questions through this interface, and the data is sent to the server in real time. The server analyzes the user's answers, provides feedback based on learning progress and understanding, and suggests the next learning content as needed.

[0492] As a concrete example, by inputting a prompt message such as "Generate a quiz about fraction addition" into an AI model, the server can automatically generate questions to deepen the user's understanding. In this way, users can learn efficiently at their own pace.

[0493] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0494] Step 1:

[0495] The server connects to the educational information storage device and searches the learning target information area specified by the user to retrieve the necessary data. This process involves sending search queries and retrieving learning resources such as text data, images, and diagrams. The input is the user's subject selection, and the output is material data related to the specified subject.

[0496] Step 2:

[0497] The server analyzes the acquired data using natural language processing and image recognition technologies. Specifically, it extracts key keywords and concepts from text data using natural language processing and recognizes visual information from images. The input is the acquired training material, and the output is the extracted keywords and identified visual information.

[0498] Step 3:

[0499] The server automatically generates test questions based on keywords and concepts obtained through analysis. At this stage, a generation AI model is used to create questions of varying difficulty levels according to the user's level of understanding and learning progress. The input is the extracted keywords, and the output is the generated test questions. Specifically, multiple-choice and fill-in-the-blank questions are generated in a variety of formats.

[0500] Step 4:

[0501] The server selects visual media and figures related to the generated exam questions and processes them to integrate them visually. Using a generative AI model, it automatically generates manga-style visual content. This process includes retrieving appropriate images from a database and creating visual content based on them. The input is the generated exam questions and associated images, and the output is the integrated visual content.

[0502] Step 5:

[0503] The server sends integrated exam questions and visual content to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The display takes an intuitive form that the user can easily operate. The input is the integrated data sent from the server, and the output is the learning interface presented to the user.

[0504] Step 6:

[0505] The user enters their answers to the displayed test questions and sends those answers to the server via their device. The server receives this answer data, determines whether it is correct or incorrect, and generates feedback. The input is the user's answer data, and the output is the evaluation result and feedback.

[0506] Step 7:

[0507] The server uses a generative AI model to evaluate the user's learning progress and understanding based on their answer data, and to propose the next learning plan. It analyzes each user's learning pattern and automatically suggests the appropriate next step. The input is the user's answer history, and the output is a customized learning plan.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] Modern learning methods demand educational materials that enable efficient understanding of information and its usage. Furthermore, improving the customer experience in physical stores requires a system that allows customers to learn about product background information while enjoying the experience. However, traditional methods make it difficult to directly link information about individual products to learning, posing challenges to efficient learning promotion and improved customer engagement.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes means for accessing an educational database and acquiring information, means for analyzing the acquired information and automatically generating multiple-choice questions, and means for reading product identification information, receiving related information, and displaying it to the user. This allows users to learn background knowledge about products in an enjoyable way while receiving visual feedback within a physical store, thereby deepening their understanding of the products and improving the customer experience.

[0513] An "educational database" is an information aggregation system that stores information related to the subject being studied.

[0514] An "information processing terminal" is a device that allows users to send, receive, and display data through an interface.

[0515] "Visual elements" refer to expressive techniques that facilitate user understanding through visual information such as images and diagrams.

[0516] "Product identification information" refers to information used to distinguish a specific item, such as an ID code or barcode.

[0517] A "multiple-choice question" is a type of question where the user is presented with options and asked to select the correct answer.

[0518] "Proficiency level" is an indicator that shows the extent to which a user can understand and apply information.

[0519] To implement this invention, it is first necessary to prepare an educational database. This database contains information about the subject of learning, and the server using it retrieves information from this database. The server also receives product-related information through a QR code reader or product identification information. When a user scans a QR code with an information processing terminal such as a smartphone, a request is sent to the server.

[0520] The server analyzes the acquired information using natural language processing libraries (such as NLTK and Transformers) implemented in Python, and then automatically generates multiple-choice questions. These questions utilize a generative AI model and include relevant options and information using prompts.

[0521] The information processing terminal presents the user with multiple-choice questions and visual elements sent from the server. These visual elements include images and diagrams related to the products and are visually integrated using image processing libraries such as OpenCV.

[0522] Users answer the displayed multiple-choice questions, and their answers are sent back to the server. The server analyzes the answers, determines the user's level of proficiency, and suggests the next learning steps and information. This allows users to deepen their background knowledge of products in an enjoyable way while in a physical store.

[0523] As a concrete example, in a certain tea shop, when a user scans a QR code related to the production method of green tea, a multiple-choice question such as "What production method is used to make this green tea?" is displayed. If the answer is correct, a reward that can be used on the next purchase is given. In this way, it is possible to promote learning in a physical store.

[0524] An example of a prompt to input into a generative AI model might be: "Generate a quiz to learn about the production methods of green tea. Also, include some interesting trivia about the history of green tea."

[0525] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0526] Step 1:

[0527] The user scans the QR code attached to the product using the QR code reader on their smartphone. The input is the information from the QR code, and the output is the identification information of the acquired product. The device sends this identification information to the server.

[0528] Step 2:

[0529] The server retrieves relevant information from an educational database based on the received product identification information. The input is the product identification information, and the output is text information related to the product. The server analyzes this information using a natural language processing library and generates a multiple-choice question.

[0530] Step 3:

[0531] The server uses a generative AI model to generate prompts from acquired text information and automatically generate multiple-choice questions based on those prompts. The input is text information related to a product, and the output is a multiple-choice question. Specifically, it generates a prompt such as "Generate a quiz to learn about the production method of green tea," and then creates the question and answer choices.

[0532] Step 4:

[0533] The server retrieves images and diagrams related to the multiple-choice questions from a database and integrates the visual content using an image processing library. The input is the visual information associated with the multiple-choice questions, and the output is the visually integrated quiz content.

[0534] Step 5:

[0535] The server sends integrated quiz content to the device. The device displays this to the user, facilitating interactive learning. The input is the visually integrated quiz content, and the output is the display on the user's device.

[0536] Step 6:

[0537] The user answers the presented multiple-choice questions, and the terminal sends the answer information to the server. The input is the user's answer, and the output is the answer information sent to the server.

[0538] Step 7:

[0539] The server analyzes the user's answers and evaluates their proficiency level. The input is the user's answers, and the output is the user's proficiency evaluation. Based on this evaluation, the server generates data to suggest the next learning steps and information.

[0540] Step 8:

[0541] The server adjusts the next learning content based on the evaluation results and generates bonus information that can be used on the next visit. The input is the user's proficiency evaluation result, and the output is the next learning content and bonus information.

[0542] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0543] This invention is a system that incorporates an emotion engine that recognizes the user's emotions, in addition to existing mechanisms that use an educational database to collect information on subjects to be studied and generate quiz questions.

[0544] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is analyzed, and quiz questions are automatically generated. The quiz questions are integrated with relevant images and diagrams, which are presented to the user in a visually appealing format.

[0545] The device receives quizzes and provides an easy-to-use interface for the user. It also features a camera and microphone to collect the user's facial expressions and voice tone. This data is sent to an emotion engine for real-time analysis.

[0546] The emotion engine analyzes the acquired user data to identify their emotional state. For example, it determines whether the user is confused, focused, or enjoying themselves. This result is then fed back to the server.

[0547] Based on the results of emotion recognition, the server adjusts the quiz content and difficulty level. If the user is confused, the questions can be simplified and supplementary explanations added. Conversely, if the user is focused, the difficulty level can be increased to present more challenging questions. This process makes it possible to maximize the learning effect for each individual.

[0548] For example, if a user is struggling with a difficult math problem, the emotion engine recognizes this as a state of confusion. As a result, the server breaks down the problem into simpler steps and presents a quiz with added hints to help solve it gradually. Conversely, if the user shows an expression of enjoyment, the server appropriately increases the difficulty level to provide a further challenge. In this way, the present invention provides an optimal learning experience that is tailored to the user's emotions.

[0549] The following describes the processing flow.

[0550] Step 1:

[0551] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is used as the basic material for generating quizzes.

[0552] Step 2:

[0553] The server analyzes the acquired data and automatically generates quiz questions using natural language processing. These questions include various formats, such as multiple-choice and fill-in-the-blank questions.

[0554] Step 3:

[0555] The server selects images and diagrams related to the generated quiz and visually integrates them to create visually appealing content.

[0556] Step 4:

[0557] The server sends the integrated quiz and visuals to the user's device, verifying data integrity and communication stability during this process.

[0558] Step 5:

[0559] The device displays the received quiz and visuals as an interactive user interface, and adjusts it to be easy for the user to operate.

[0560] Step 6:

[0561] The user answers the presented quiz, and their facial expressions and voice tone are recorded in real time through the camera and microphone built into the device.

[0562] Step 7:

[0563] The device transmits the user's facial expression data and voice data to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0564] Step 8:

[0565] The server receives feedback from the emotion engine and adjusts the quiz content and difficulty level according to the user's emotions. For example, if the user is confused, the questions are made easier and additional information is provided.

[0566] Step 9:

[0567] The server then resends the newly adjusted quiz questions to the device, allowing the user to continue learning. In this way, the learning experience is optimized based on the user's individual emotions.

[0568] (Example 2)

[0569] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0570] In modern education systems, it is difficult to consider the emotions and learning progress of individual learners in real time, and general educational content is not optimized for all learners. Furthermore, providing content in a uniform manner without understanding the emotional state of learners results in a failure to maximize learning effectiveness. In addition, there is a need to provide appropriate feedback based on emotions, but the means to achieve this are insufficient.

[0571] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0572] In this invention, the server includes means for accessing an educational database and acquiring information on the subject to be studied; means for analyzing the acquired information and automatically generating quiz questions using a generative AI model; and means for selecting images or charts related to the generated quiz questions and integrating them as visual elements. This makes it possible to analyze the learner's emotional state and adjust appropriate feedback and learning content in real time according to their emotions, thereby providing a learning experience optimized for each individual learner.

[0573] An "educational database" is a source of information for accumulating and managing information related to the subjects being studied, and is a database that is referenced to support users' learning activities.

[0574] A "generative AI model" is a model that uses artificial intelligence algorithms to analyze data and automatically generate specific tasks or processes, in this case, quiz questions.

[0575] A "quiz question" is an automatically generated question format intended for evaluation and feedback to learners, and is a means of measuring their level of understanding of the learning material.

[0576] "Images or diagrams" are visual elements used to provide learners with information visually and are integrated to complement quiz questions.

[0577] An "emotion analysis engine" is a system or tool that identifies and analyzes a learner's emotional state based on data from the user's facial expressions and voice.

[0578] "Visual integration" refers to selecting images or diagrams related to the generated quiz questions and integrating them for visual display.

[0579] "Real-time" means that data acquisition and analysis are performed instantly without delay, which enables immediate feedback tailored to the user's situation.

[0580] An "interactive learning experience" refers to a form of learning in which learners can acquire knowledge effectively and efficiently by interacting with the system in a two-way manner.

[0581] This invention is an educational support system incorporating a generative AI model and an emotion analysis engine. Specific embodiments are described below.

[0582] The server is responsible for accessing the educational database and retrieving information on the subject the user has selected to study. This information includes subject-related data and past exam questions. The server analyzes the retrieved information and uses a generative AI model to create quiz questions. This model learns from a large amount of data and automatically generates question formats suitable for the user.

[0583] The generated quiz questions are visually integrated by selecting relevant images and diagrams. This makes the quiz questions more intuitive and easier to understand. The integrated content is sent to the device and presented in an interactive interface for the user.

[0584] The device uses its camera and microphone to collect the user's facial expressions and voice tone in real time as they answer quizzes. This process is crucial for accurately capturing the user's reactions and improving the personalized learning experience.

[0585] The collected data is analyzed by an emotion analysis engine. This engine detects the user's emotional state based on facial recognition technology. For example, it can identify whether the user is confused, focused, or enjoying themselves. The analysis results are fed back to the server, forming the basis for adjusting the content and difficulty level of quiz questions in real time.

[0586] For example, if a user is faced with a complex math problem and the emotion analysis engine detects confusion from the user's facial expression, the server incorporates step-by-step hints into the quiz to make solving the problem easier. Conversely, if the analysis indicates the user is enjoying the problem, the server increases the difficulty level, providing further challenges to facilitate learning.

[0587] As an example of a prompt, the AI ​​generator can be input with the following message: "Based on the user's emotional state, please suggest the type of problem to present next." In this way, the present invention realizes effective learning support.

[0588] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0589] Step 1:

[0590] The user selects the subject they wish to study through the interface on their device. This selected subject information is then sent to the server.

[0591] Step 2:

[0592] The server accesses an educational database based on the received subject information. It retrieves information related to the selected subject from the database. This retrieved information includes teaching materials, past exam questions, and related literature. This data is then used as input for analysis.

[0593] Step 3:

[0594] The server analyzes data obtained from an educational database and automatically generates quiz questions using a generative AI model. The analysis includes selecting appropriate topics and adjusting the question format and difficulty level. The resulting quiz questions are then output.

[0595] Step 4:

[0596] The server selects images or diagrams related to the generated quiz questions and integrates them as visual elements. This process aims to output a visualized integrated question, and the data is processed so that the user can understand it intuitively.

[0597] Step 5:

[0598] The server sends the integrated quiz questions to the terminal. The terminal displays the received quiz to the user in an easy-to-use format. Specifically, the question text, answer choices, images, and diagrams are arranged on the screen.

[0599] Step 6:

[0600] Users answer quizzes through their devices. During this process, the devices use their cameras and microphones to collect the user's facial expressions and voice tone in real time. This collected data is then sent to an emotion analysis engine.

[0601] Step 7:

[0602] The emotion analysis engine analyzes the received user data. This analysis identifies the user's emotional state, such as whether they are confused, focused, or enjoying themselves. The results of this analysis are then sent back to the server.

[0603] Step 8:

[0604] The server adjusts the quiz content and difficulty based on the results from the sentiment analysis engine. For example, if confusion is detected, the difficulty of the question is lowered or hints are added. This generates new quiz questions that are appropriate to the situation and sends them back to the terminal.

[0605] This series of steps provides an appropriate learning environment that takes user emotions into consideration.

[0606] (Application Example 2)

[0607] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0608] In education, providing personalized learning experiences tailored to individual learners is crucial. However, traditional learning systems have struggled to provide appropriate learning content that takes into account the emotional state of users, posing challenges to motivating learners and promoting effective understanding. Furthermore, creating a visually appealing and interactive learning environment has also been difficult.

[0609] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0610] In this invention, the server includes means for accessing an educational knowledge base and obtaining information on subjects to be studied; means for recognizing the user's emotional state and adjusting the content and difficulty level of quizzes in real time; and means for incorporating entertainment elements into the visual user interface to promote an interactive and emotionally sensitive learning experience. This makes it possible to provide optimal learning content tailored to the emotions of each individual learner, thereby improving learner comprehension and increasing motivation to learn.

[0611] An "educational knowledge base" is a database that systematically stores information about the subjects being studied, and serves as a foundation for providing quiz questions and related materials.

[0612] A "subject to be studied" refers to a specific educational field or theme that learners are expected to understand and master.

[0613] A "quiz-style assignment" is a question designed to allow learners to confirm their knowledge and deepen their understanding by answering it.

[0614] "Visual media" refers to digital content such as images and videos that contain visual information, and is intended to complement and enhance learning content.

[0615] A "user terminal" refers to a device that learners use to access and interact with learning content via the internet.

[0616] "User's emotional state" refers to the emotional state a learner exhibits during learning, and includes factors such as the degree of concentration and stress level.

[0617] "User learning progress" refers to an indicator that shows the extent to which a learner has acquired the target knowledge.

[0618] "Entertainment elements" refer to entertaining designs and content that aim to capture learners' interest and make the learning experience more enjoyable.

[0619] The system for realizing this invention is configured as an educational platform.

[0620] The server first connects to an educational knowledge base to retrieve information about the subject being studied. This retrieved information is then analyzed by a data processing algorithm to generate quiz-style questions. Next, relevant visual media are selected for the generated quiz and integrated into it. Finally, the integrated quiz is transmitted to the user's terminal via the internet.

[0621] The terminal has the functionality to receive and collect user answers. It also has an integrated camera and microphone, which transmits the user's facial expressions and voice to the emotion engine. The emotion engine uses the image analysis library OpenCV and the deep learning framework TensorFlow to analyze the user's emotional state in real time. Based on the analysis results, the server adjusts the quiz content and difficulty level as needed and retransmits the answers.

[0622] Users answer quizzes provided by the device and receive feedback tailored to their emotional state during the process. The device stores the user's learning progress and provides advice to help them in their next learning step.

[0623] As a concrete example, let's assume a learner is tackling a challenging history quiz. Initially, the learner shows a surprised expression, so the server immediately lowers the difficulty of the quiz and provides additional explanations. If emotional analysis indicates that the learner is enjoying themselves, the server provides feedback such as preparing a more challenging quiz next time.

[0624] An example of a prompt message might be, "Estimate the user's emotions from their facial expressions and voice, and generate optimal learning feedback."

[0625] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0626] Step 1:

[0627] The server accesses an educational knowledge base to retrieve information related to the subject being studied. In this process, the server executes database queries using the subject ID as input. As a result, it outputs text, images, and videos related to the subject.

[0628] Step 2:

[0629] The server generates quiz-style questions based on the acquired information. Here, a generative AI model is used to analyze the relevant information using natural language processing to form the quiz questions and answer choices. The output is a set of quiz questions.

[0630] Step 3:

[0631] The server integrates appropriate visual media into the generated quiz. This involves using image processing algorithms to combine images and videos with the question text. A quiz with visual media is then generated.

[0632] Step 4:

[0633] The terminal displays quizzes received from the server to the user. The terminal's input is a quiz dataset from the server, and its output is a visual display for the user. This is implemented using a user interface library.

[0634] Step 5:

[0635] The user answers a quiz displayed on the terminal, and the terminal sends the answer to the server. The input for this step is the user's choices, and the output is the answer data packet.

[0636] Step 6:

[0637] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits them to the emotion engine. The input here is the user's real-time facial and voice data, and the output is a data stream to the emotion engine.

[0638] Step 7:

[0639] The emotion engine analyzes the acquired data to estimate the user's emotional state. Inputs are facial expressions and voice data, and output is the estimated emotional category. Specific operations include on-the-spot data analysis using TensorFlow.

[0640] Step 8:

[0641] The server receives the results of the emotion analysis and adjusts the quiz content and difficulty level. The output is a newly adjusted quiz set.

[0642] Step 9:

[0643] The device then presents the adjusted quiz data to the user again and provides interactive feedback. The input is the adjusted quiz, and the output is the result of the re-presentation to the user.

[0644] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0645] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0646] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0647] [Fourth Embodiment]

[0648] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0649] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0650] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0651] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0652] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0653] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0654] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0655] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0656] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0657] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0658] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0659] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0660] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0661] To implement this invention, first, an educational database is prepared. The database stores text information, images, charts, and other data related to the subjects being studied, such as Japanese language, mathematics, science, and social studies.

[0662] The server connects to the aforementioned educational database and retrieves data for the subject selected by the user. The retrieved data is then analyzed using natural language processing and image recognition technologies. This analysis process extracts important keywords and concepts, which are then used to automatically generate quiz questions. These quiz questions are created in various formats, including multiple-choice and fill-in-the-blank questions.

[0663] When a quiz is generated, the server picks relevant images and diagrams from a database and creates a manga-style visual. This visual content is designed to allow users to intuitively understand the learning material.

[0664] The generated quiz and visual content are sent from the server to the device. The device receives this data and displays it to the user in an interactive format. The user enters their answers to the presented questions, and these answers are sent to the server in real time.

[0665] The server analyzes the data submitted by the user and evaluates their learning progress and weaknesses. Based on the evaluation results, it adjusts the next learning content and quiz difficulty level, and proposes an appropriate learning plan for the user. This entire process allows users to learn efficiently at their own pace.

[0666] As a concrete example, when generating a quiz on fraction calculations in arithmetic, the server retrieves rules and example problems for fraction calculations from a database and creates multiple-choice questions that allow the user to practice fraction addition. In addition, diagrams and charts that help visualize fractions are displayed in the explanation of the problems, allowing the user to deepen their understanding of the problems. In this way, the present invention is implemented with the aim of maximizing the learning effect of the user.

[0667] The following describes the processing flow.

[0668] Step 1:

[0669] The server accesses an educational database and retrieves text, images, and diagrams for the subject selected by the user. This data is prepared as material necessary for generating quizzes.

[0670] Step 2:

[0671] The server applies natural language processing to the acquired text data to analyze important keywords and concepts. During this analysis, it extracts content that should be included in the questions and related information.

[0672] Step 3:

[0673] The server automatically generates quiz questions based on the analysis results. These questions are structured as multiple-choice and fill-in-the-blank questions and include multiple answer choices and different formats to suit different learning styles.

[0674] Step 4:

[0675] The server selects images and diagrams related to the generated quiz from the database and integrates them as visual content, creating manga-like elements that aid intuitive understanding.

[0676] Step 5:

[0677] The server sends the generated quiz and visual content to the user's terminal. At this point, communication stability and data integrity are checked.

[0678] Step 6:

[0679] The terminal displays the received quiz and visuals as a user interface, and adjusts the operation to allow the user to easily input answers.

[0680] Step 7:

[0681] Users enter their answers to the presented quiz by selecting from multiple-choice options or providing written answers. This input is sent from the terminal to the server in real time.

[0682] Step 8:

[0683] The server analyzes the user's answers and calculates performance data, including correctness ratings and answer time. This allows the user's learning progress to be evaluated.

[0684] Step 9:

[0685] Based on the analysis results, the server adjusts the next learning content and quiz difficulty level, proposing a personalized learning plan to the user. This ensures that content optimized for individual learning needs is provided.

[0686] (Example 1)

[0687] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0688] Conventional learning support systems have not adequately customized their systems to suit each user's learning progress and level of understanding, making it difficult to achieve effective individualized learning. Furthermore, visual learning aids are limited, lacking opportunities for learners to intuitively grasp the learning content. Additionally, the provision of learning plans is fixed, lacking dynamic adjustments that adapt to the user's learning situation.

[0689] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0690] In this invention, the server includes means for accessing an educational information storage device and acquiring information from the learning target information area, means for analyzing the acquired information and automatically generating test questions, and means for selecting visual media or figures related to the generated test questions and visually integrating them. This makes it possible to realize effective individualized learning according to each user's learning progress, provide visually intuitive learning support, and dynamically adjust the learning plan.

[0691] An "educational information storage device" is a system that stores learning materials and education-related data and provides them in a format that can be accessed as needed.

[0692] A "learning target information domain" is a part of a database that contains learning materials related to a specific subject or theme.

[0693] "Exam questions" are questions or tasks designed for educational purposes and used to assess a user's knowledge and understanding.

[0694] "Visual media" refers to visual information materials, including images and diagrams, that serve to visually support learning.

[0695] A "dynamic information display device" is a display device that provides a visual interface that users can use interactively during learning, and that adapts in real time.

[0696] A "generative AI model" is a type of artificial intelligence that uses machine learning techniques to analyze data and generate new data or perform pattern recognition.

[0697] A "server" is a computer system that shares data over a network and processes information in response to requests from clients.

[0698] A "user device" is a device used by a user to receive, manipulate, and display information, and generally includes personal computers and smartphones.

[0699] To implement this invention, the server first connects to an educational information storage device and retrieves necessary data from the learning target information area. This data includes text data, images, and figures related to basic subjects such as Japanese language, mathematics, science, and social studies. The server analyzes the retrieved data using natural language processing and image recognition technologies. Through this analysis, important keywords and concepts are extracted, and test questions are created based on them.

[0700] The system automatically generates test questions according to the user's learning progress and provides them in various formats, such as multiple-choice and fill-in-the-blank questions. For example, when generating questions about fraction calculations in arithmetic, the server retrieves rules and examples of fraction calculations from a database and creates questions that allow the user to practice fraction addition. Furthermore, it selects relevant visual aids and shapes and uses a generative AI model to create manga-style visual content.

[0701] The generated test questions and visual content are sent from the server to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The user answers the questions through this interface, and the data is sent to the server in real time. The server analyzes the user's answers, provides feedback based on learning progress and understanding, and suggests the next learning content as needed.

[0702] As a concrete example, by inputting a prompt message such as "Generate a quiz about fraction addition" into an AI model, the server can automatically generate questions to deepen the user's understanding. In this way, users can learn efficiently at their own pace.

[0703] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0704] Step 1:

[0705] The server connects to the educational information storage device and searches the learning target information area specified by the user to retrieve the necessary data. This process involves sending search queries and retrieving learning resources such as text data, images, and diagrams. The input is the user's subject selection, and the output is material data related to the specified subject.

[0706] Step 2:

[0707] The server analyzes the acquired data using natural language processing and image recognition technologies. Specifically, it extracts key keywords and concepts from text data using natural language processing and recognizes visual information from images. The input is the acquired training material, and the output is the extracted keywords and identified visual information.

[0708] Step 3:

[0709] The server automatically generates test questions based on keywords and concepts obtained through analysis. At this stage, a generation AI model is used to create questions of varying difficulty levels according to the user's level of understanding and learning progress. The input is the extracted keywords, and the output is the generated test questions. Specifically, multiple-choice and fill-in-the-blank questions are generated in a variety of formats.

[0710] Step 4:

[0711] The server selects visual media and figures related to the generated exam questions and processes them to integrate them visually. Using a generative AI model, it automatically generates manga-style visual content. This process includes retrieving appropriate images from a database and creating visual content based on them. The input is the generated exam questions and associated images, and the output is the integrated visual content.

[0712] Step 5:

[0713] The server sends integrated exam questions and visual content to the terminal. The terminal receives this data and displays it to the user through an interactive user interface. The display takes an intuitive form that the user can easily operate. The input is the integrated data sent from the server, and the output is the learning interface presented to the user.

[0714] Step 6:

[0715] The user enters their answers to the displayed test questions and sends those answers to the server via their device. The server receives this answer data, determines whether it is correct or incorrect, and generates feedback. The input is the user's answer data, and the output is the evaluation result and feedback.

[0716] Step 7:

[0717] The server uses a generative AI model to evaluate the user's learning progress and understanding based on their answer data, and to propose the next learning plan. It analyzes each user's learning pattern and automatically suggests the appropriate next step. The input is the user's answer history, and the output is a customized learning plan.

[0718] (Application Example 1)

[0719] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0720] Modern learning methods demand educational materials that enable efficient understanding of information and its usage. Furthermore, improving the customer experience in physical stores requires a system that allows customers to learn about product background information while enjoying the experience. However, traditional methods make it difficult to directly link information about individual products to learning, posing challenges to efficient learning promotion and improved customer engagement.

[0721] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0722] In this invention, the server includes means for accessing an educational database and acquiring information, means for analyzing the acquired information and automatically generating multiple-choice questions, and means for reading product identification information, receiving related information, and displaying it to the user. This allows users to learn background knowledge about products in an enjoyable way while receiving visual feedback within a physical store, thereby deepening their understanding of the products and improving the customer experience.

[0723] An "educational database" is an information aggregation system that stores information related to the subject being studied.

[0724] An "information processing terminal" is a device that allows users to send, receive, and display data through an interface.

[0725] "Visual elements" refer to expressive techniques that facilitate user understanding through visual information such as images and diagrams.

[0726] "Product identification information" refers to information used to distinguish a specific item, such as an ID code or barcode.

[0727] A "multiple-choice question" is a type of question where the user is presented with options and asked to select the correct answer.

[0728] "Proficiency level" is an indicator that shows the extent to which a user can understand and apply information.

[0729] To implement this invention, it is first necessary to prepare an educational database. This database contains information about the subject of learning, and the server using it retrieves information from this database. The server also receives product-related information through a QR code reader or product identification information. When a user scans a QR code with an information processing terminal such as a smartphone, a request is sent to the server.

[0730] The server analyzes the acquired information using natural language processing libraries (such as NLTK and Transformers) implemented in Python, and then automatically generates multiple-choice questions. These questions utilize a generative AI model and include relevant options and information using prompts.

[0731] The information processing terminal presents the user with multiple-choice questions and visual elements sent from the server. These visual elements include images and diagrams related to the products and are visually integrated using image processing libraries such as OpenCV.

[0732] Users answer the displayed multiple-choice questions, and their answers are sent back to the server. The server analyzes the answers, determines the user's level of proficiency, and suggests the next learning steps and information. This allows users to deepen their background knowledge of products in an enjoyable way while in a physical store.

[0733] As a concrete example, in a certain tea shop, when a user scans a QR code related to the production method of green tea, a multiple-choice question such as "What production method is used to make this green tea?" is displayed. If the answer is correct, a reward that can be used on the next purchase is given. In this way, it is possible to promote learning in a physical store.

[0734] An example of a prompt to input into a generative AI model might be: "Generate a quiz to learn about the production methods of green tea. Also, include some interesting trivia about the history of green tea."

[0735] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0736] Step 1:

[0737] The user scans the QR code attached to the product using the QR code reader on their smartphone. The input is the information from the QR code, and the output is the identification information of the acquired product. The device sends this identification information to the server.

[0738] Step 2:

[0739] The server retrieves relevant information from an educational database based on the received product identification information. The input is the product identification information, and the output is text information related to the product. The server analyzes this information using a natural language processing library and generates a multiple-choice question.

[0740] Step 3:

[0741] The server uses a generative AI model to generate prompts from acquired text information and automatically generate multiple-choice questions based on those prompts. The input is text information related to a product, and the output is a multiple-choice question. Specifically, it generates a prompt such as "Generate a quiz to learn about the production method of green tea," and then creates the question and answer choices.

[0742] Step 4:

[0743] The server retrieves images and diagrams related to the multiple-choice questions from a database and integrates the visual content using an image processing library. The input is the visual information associated with the multiple-choice questions, and the output is the visually integrated quiz content.

[0744] Step 5:

[0745] The server sends integrated quiz content to the device. The device displays this to the user, facilitating interactive learning. The input is the visually integrated quiz content, and the output is the display on the user's device.

[0746] Step 6:

[0747] The user answers the presented multiple-choice questions, and the terminal sends the answer information to the server. The input is the user's answer, and the output is the answer information sent to the server.

[0748] Step 7:

[0749] The server analyzes the user's answers and evaluates their proficiency level. The input is the user's answers, and the output is the user's proficiency evaluation. Based on this evaluation, the server generates data to suggest the next learning steps and information.

[0750] Step 8:

[0751] The server adjusts the next learning content based on the evaluation results and generates bonus information that can be used on the next visit. The input is the user's proficiency evaluation result, and the output is the next learning content and bonus information.

[0752] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0753] This invention is a system that incorporates an emotion engine that recognizes the user's emotions, in addition to existing mechanisms that use an educational database to collect information on subjects to be studied and generate quiz questions.

[0754] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is analyzed, and quiz questions are automatically generated. The quiz questions are integrated with relevant images and diagrams, which are presented to the user in a visually appealing format.

[0755] The device receives quizzes and provides an easy-to-use interface for the user. It also features a camera and microphone to collect the user's facial expressions and voice tone. This data is sent to an emotion engine for real-time analysis.

[0756] The emotion engine analyzes the acquired user data to identify their emotional state. For example, it determines whether the user is confused, focused, or enjoying themselves. This result is then fed back to the server.

[0757] Based on the results of emotion recognition, the server adjusts the quiz content and difficulty level. If the user is confused, the questions can be simplified and supplementary explanations added. Conversely, if the user is focused, the difficulty level can be increased to present more challenging questions. This process makes it possible to maximize the learning effect for each individual.

[0758] For example, if a user is struggling with a difficult math problem, the emotion engine recognizes this as a state of confusion. As a result, the server breaks down the problem into simpler steps and presents a quiz with added hints to help solve it gradually. Conversely, if the user shows an expression of enjoyment, the server appropriately increases the difficulty level to provide a further challenge. In this way, the present invention provides an optimal learning experience that is tailored to the user's emotions.

[0759] The following describes the processing flow.

[0760] Step 1:

[0761] The server accesses an educational database and retrieves data related to the subject selected by the user. This data is used as the basic material for generating quizzes.

[0762] Step 2:

[0763] The server analyzes the acquired data and automatically generates quiz questions using natural language processing. These questions include various formats, such as multiple-choice and fill-in-the-blank questions.

[0764] Step 3:

[0765] The server selects images and diagrams related to the generated quiz and visually integrates them to create visually appealing content.

[0766] Step 4:

[0767] The server sends the integrated quiz and visuals to the user's device, verifying data integrity and communication stability during this process.

[0768] Step 5:

[0769] The device displays the received quiz and visuals as an interactive user interface, and adjusts it to be easy for the user to operate.

[0770] Step 6:

[0771] The user answers the presented quiz, and their facial expressions and voice tone are recorded in real time through the camera and microphone built into the device.

[0772] Step 7:

[0773] The device transmits the user's facial expression data and voice data to the emotion engine. The emotion engine analyzes this data to identify the user's emotional state.

[0774] Step 8:

[0775] The server receives feedback from the emotion engine and adjusts the quiz content and difficulty level according to the user's emotions. For example, if the user is confused, the questions are made easier and additional information is provided.

[0776] Step 9:

[0777] The server then resends the newly adjusted quiz questions to the device, allowing the user to continue learning. In this way, the learning experience is optimized based on the user's individual emotions.

[0778] (Example 2)

[0779] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0780] In modern education systems, it is difficult to consider the emotions and learning progress of individual learners in real time, and general educational content is not optimized for all learners. Furthermore, providing content in a uniform manner without understanding the emotional state of learners results in a failure to maximize learning effectiveness. In addition, there is a need to provide appropriate feedback based on emotions, but the means to achieve this are insufficient.

[0781] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0782] In this invention, the server includes means for accessing an educational database and acquiring information on the subject to be studied; means for analyzing the acquired information and automatically generating quiz questions using a generative AI model; and means for selecting images or charts related to the generated quiz questions and integrating them as visual elements. This makes it possible to analyze the learner's emotional state and adjust appropriate feedback and learning content in real time according to their emotions, thereby providing a learning experience optimized for each individual learner.

[0783] An "educational database" is a source of information for accumulating and managing information related to the subjects being studied, and is a database that is referenced to support users' learning activities.

[0784] A "generative AI model" is a model that uses artificial intelligence algorithms to analyze data and automatically generate specific tasks or processes, in this case, quiz questions.

[0785] A "quiz question" is an automatically generated question format intended for evaluation and feedback to learners, and is a means of measuring their level of understanding of the learning material.

[0786] "Images or diagrams" are visual elements used to provide learners with information visually and are integrated to complement quiz questions.

[0787] An "emotion analysis engine" is a system or tool that identifies and analyzes a learner's emotional state based on data from the user's facial expressions and voice.

[0788] "Visual integration" refers to selecting images or diagrams related to the generated quiz questions and integrating them for visual display.

[0789] "Real-time" means that data acquisition and analysis are performed instantly without delay, which enables immediate feedback tailored to the user's situation.

[0790] An "interactive learning experience" refers to a form of learning in which learners can acquire knowledge effectively and efficiently by interacting with the system in a two-way manner.

[0791] This invention is an educational support system incorporating a generative AI model and an emotion analysis engine. Specific embodiments are described below.

[0792] The server is responsible for accessing the educational database and retrieving information on the subject the user has selected to study. This information includes subject-related data and past exam questions. The server analyzes the retrieved information and uses a generative AI model to create quiz questions. This model learns from a large amount of data and automatically generates question formats suitable for the user.

[0793] The generated quiz questions are visually integrated by selecting relevant images and diagrams. This makes the quiz questions more intuitive and easier to understand. The integrated content is sent to the device and presented in an interactive interface for the user.

[0794] The device uses its camera and microphone to collect the user's facial expressions and voice tone in real time as they answer quizzes. This process is crucial for accurately capturing the user's reactions and improving the personalized learning experience.

[0795] The collected data is analyzed by an emotion analysis engine. This engine detects the user's emotional state based on facial recognition technology. For example, it can identify whether the user is confused, focused, or enjoying themselves. The analysis results are fed back to the server, forming the basis for adjusting the content and difficulty level of quiz questions in real time.

[0796] For example, if a user is faced with a complex math problem and the emotion analysis engine detects confusion from the user's facial expression, the server incorporates step-by-step hints into the quiz to make solving the problem easier. Conversely, if the analysis indicates the user is enjoying the problem, the server increases the difficulty level, providing further challenges to facilitate learning.

[0797] As an example of a prompt, the AI ​​generator can be input with the following message: "Based on the user's emotional state, please suggest the type of problem to present next." In this way, the present invention realizes effective learning support.

[0798] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0799] Step 1:

[0800] The user selects the subject they wish to study through the interface on their device. This selected subject information is then sent to the server.

[0801] Step 2:

[0802] The server accesses an educational database based on the received subject information. It retrieves information related to the selected subject from the database. This retrieved information includes teaching materials, past exam questions, and related literature. This data is then used as input for analysis.

[0803] Step 3:

[0804] The server analyzes data obtained from an educational database and automatically generates quiz questions using a generative AI model. The analysis includes selecting appropriate topics and adjusting the question format and difficulty level. The resulting quiz questions are then output.

[0805] Step 4:

[0806] The server selects images or diagrams related to the generated quiz questions and integrates them as visual elements. This process aims to output a visualized integrated question, and the data is processed so that the user can understand it intuitively.

[0807] Step 5:

[0808] The server sends the integrated quiz questions to the terminal. The terminal displays the received quiz to the user in an easy-to-use format. Specifically, the question text, answer choices, images, and diagrams are arranged on the screen.

[0809] Step 6:

[0810] Users answer quizzes through their devices. During this process, the devices use their cameras and microphones to collect the user's facial expressions and voice tone in real time. This collected data is then sent to an emotion analysis engine.

[0811] Step 7:

[0812] The emotion analysis engine analyzes the received user data. This analysis identifies the user's emotional state, such as whether they are confused, focused, or enjoying themselves. The results of this analysis are then sent back to the server.

[0813] Step 8:

[0814] The server adjusts the quiz content and difficulty based on the results from the sentiment analysis engine. For example, if confusion is detected, the difficulty of the question is lowered or hints are added. This generates new quiz questions that are appropriate to the situation and sends them back to the terminal.

[0815] This series of steps provides an appropriate learning environment that takes user emotions into consideration.

[0816] (Application Example 2)

[0817] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0818] In education, providing personalized learning experiences tailored to individual learners is crucial. However, traditional learning systems have struggled to provide appropriate learning content that takes into account the emotional state of users, posing challenges to motivating learners and promoting effective understanding. Furthermore, creating a visually appealing and interactive learning environment has also been difficult.

[0819] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0820] In this invention, the server includes means for accessing an educational knowledge base and obtaining information on subjects to be studied; means for recognizing the user's emotional state and adjusting the content and difficulty level of quizzes in real time; and means for incorporating entertainment elements into the visual user interface to promote an interactive and emotionally sensitive learning experience. This makes it possible to provide optimal learning content tailored to the emotions of each individual learner, thereby improving learner comprehension and increasing motivation to learn.

[0821] An "educational knowledge base" is a database that systematically stores information about the subjects being studied, and serves as a foundation for providing quiz questions and related materials.

[0822] A "subject to be studied" refers to a specific educational field or theme that learners are expected to understand and master.

[0823] A "quiz-style assignment" is a question designed to allow learners to confirm their knowledge and deepen their understanding by answering it.

[0824] "Visual media" refers to digital content such as images and videos that contain visual information, and is intended to complement and enhance learning content.

[0825] A "user terminal" refers to a device that learners use to access and interact with learning content via the internet.

[0826] "User's emotional state" refers to the emotional state a learner exhibits during learning, and includes factors such as the degree of concentration and stress level.

[0827] "User learning progress" refers to an indicator that shows the extent to which a learner has acquired the target knowledge.

[0828] "Entertainment elements" refer to entertaining designs and content that aim to capture learners' interest and make the learning experience more enjoyable.

[0829] The system for realizing this invention is configured as an educational platform.

[0830] The server first connects to an educational knowledge base to retrieve information about the subject being studied. This retrieved information is then analyzed by a data processing algorithm to generate quiz-style questions. Next, relevant visual media are selected for the generated quiz and integrated into it. Finally, the integrated quiz is transmitted to the user's terminal via the internet.

[0831] The terminal has the functionality to receive and collect user answers. It also has an integrated camera and microphone, which transmits the user's facial expressions and voice to the emotion engine. The emotion engine uses the image analysis library OpenCV and the deep learning framework TensorFlow to analyze the user's emotional state in real time. Based on the analysis results, the server adjusts the quiz content and difficulty level as needed and retransmits the answers.

[0832] Users answer quizzes provided by the device and receive feedback tailored to their emotional state during the process. The device stores the user's learning progress and provides advice to help them in their next learning step.

[0833] As a concrete example, let's assume a learner is tackling a challenging history quiz. Initially, the learner shows a surprised expression, so the server immediately lowers the difficulty of the quiz and provides additional explanations. If emotional analysis indicates that the learner is enjoying themselves, the server provides feedback such as preparing a more challenging quiz next time.

[0834] An example of a prompt message might be, "Estimate the user's emotions from their facial expressions and voice, and generate optimal learning feedback."

[0835] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0836] Step 1:

[0837] The server accesses an educational knowledge base to retrieve information related to the subject being studied. In this process, the server executes database queries using the subject ID as input. As a result, it outputs text, images, and videos related to the subject.

[0838] Step 2:

[0839] The server generates quiz-style questions based on the acquired information. Here, a generative AI model is used to analyze the relevant information using natural language processing to form the quiz questions and answer choices. The output is a set of quiz questions.

[0840] Step 3:

[0841] The server integrates appropriate visual media into the generated quiz. This involves using image processing algorithms to combine images and videos with the question text. A quiz with visual media is then generated.

[0842] Step 4:

[0843] The terminal displays quizzes received from the server to the user. The terminal's input is a quiz dataset from the server, and its output is a visual display for the user. This is implemented using a user interface library.

[0844] Step 5:

[0845] The user answers a quiz displayed on the terminal, and the terminal sends the answer to the server. The input for this step is the user's choices, and the output is the answer data packet.

[0846] Step 6:

[0847] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits them to the emotion engine. The input here is the user's real-time facial and voice data, and the output is a data stream to the emotion engine.

[0848] Step 7:

[0849] The emotion engine analyzes the acquired data to estimate the user's emotional state. Inputs are facial expressions and voice data, and output is the estimated emotional category. Specific operations include on-the-spot data analysis using TensorFlow.

[0850] Step 8:

[0851] The server receives the results of the emotion analysis and adjusts the quiz content and difficulty level. The output is a newly adjusted quiz set.

[0852] Step 9:

[0853] The device then presents the adjusted quiz data to the user again and provides interactive feedback. The input is the adjusted quiz, and the output is the result of the re-presentation to the user.

[0854] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0855] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0856] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0857] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0858] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0859] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0860] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0861] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0862] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0863] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0864] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0865] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0866] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0867] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0868] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0869] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0870] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0871] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0872] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0873] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0874] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0875] The following is further disclosed regarding the embodiments described above.

[0876] (Claim 1)

[0877] A means of accessing an educational database and obtaining data for the subject being studied,

[0878] A means of analyzing acquired data and automatically generating quiz questions,

[0879] A means for selecting images or diagrams related to the generated quiz questions and integrating them visually,

[0880] A means of sending integrated quizzes and visuals to the user's terminal,

[0881] A means of collecting and analyzing user responses,

[0882] A means of suggesting the next learning content based on the user's learning progress,

[0883] A system that includes this.

[0884] (Claim 2)

[0885] The system according to claim 1, which adjusts the difficulty level of quiz questions according to the user's learning progress based on the analyzed data.

[0886] (Claim 3)

[0887] The system according to claim 1, which provides a visual user interface that includes manga-like elements and promotes an interactive learning experience.

[0888] "Example 1"

[0889] (Claim 1)

[0890] A means for accessing an educational information storage device and obtaining information from the learning target information area,

[0891] A means of analyzing acquired information and automatically generating test questions,

[0892] A means for selecting and visually integrating visual media or figures related to the generated test questions,

[0893] Means for transmitting integrated tests and visuals to the user's device,

[0894] A means of collecting and analyzing user responses,

[0895] A means of suggesting the next learning content based on the user's learning progress,

[0896] A means of adjusting a user-specific learning plan using a generative AI model based on the analysis results,

[0897] A system that includes this.

[0898] (Claim 2)

[0899] The system according to claim 1, which adjusts the difficulty level of the test questions according to the user's learning progress based on the analyzed information.

[0900] (Claim 3)

[0901] The system according to claim 1, which provides a dynamic information display device that includes visual cartoon elements and facilitates an interactive learning experience.

[0902] "Application Example 1"

[0903] (Claim 1)

[0904] Means of accessing educational databases and obtaining information,

[0905] A means of analyzing acquired information and automatically generating multiple-choice questions,

[0906] A means for selecting and visually integrating visual elements related to the generated multiple-choice questions,

[0907] A means for transmitting integrated multiple-choice questions and visuals to an information processing terminal,

[0908] A means of collecting and analyzing user responses,

[0909] A means of suggesting the following information based on the user's level of proficiency,

[0910] A means of reading product identification information, receiving related information, and displaying it to the user,

[0911] A system that includes this.

[0912] (Claim 2)

[0913] The system according to claim 1, which adjusts the difficulty level of multiple-choice questions according to the user's proficiency level based on the analyzed data.

[0914] (Claim 3)

[0915] The system according to claim 1, which provides a user interface that includes visual elements and promotes an interactive experience.

[0916] "Example 2 of combining an emotion engine"

[0917] (Claim 1)

[0918] A means of accessing educational databases and obtaining information on subjects to be studied,

[0919] A means of analyzing acquired information and automatically generating quiz questions using a generative AI model,

[0920] A means for selecting images or diagrams related to the generated quiz questions and integrating them as visual elements,

[0921] A means for sending and displaying integrated quizzes and visual elements on the user's terminal,

[0922] In addition to the user's answers, the system uses cameras and microphones to collect data on the user's facial expressions and voice,

[0923] A means of analyzing collected user data with an emotion analysis engine to identify the user's emotional state,

[0924] A means of adjusting the content and difficulty level of quiz questions in real time according to the user's emotional state,

[0925] A system that includes this.

[0926] (Claim 2)

[0927] The system according to claim 1, which adjusts the difficulty level of quiz questions based on the emotional state of the analyzed user.

[0928] (Claim 3)

[0929] The system according to claim 1, which provides feedback that responds to the user's emotions and maximizes the interactive learning experience.

[0930] "Application example 2 of combining emotional engines"

[0931] (Claim 1)

[0932] A means of accessing an educational knowledge base and obtaining information on the subject being studied,

[0933] A means of analyzing acquired information and automatically generating quiz-style questions,

[0934] A means for selecting visual information related to the generated quiz and integrating it into a visual medium,

[0935] A means for transmitting an integrated quiz and visual media to the user's terminal,

[0936] A means of collecting and analyzing user responses,

[0937] A means of recognizing the user's emotional state and adjusting the quiz content and difficulty level in real time,

[0938] A means of suggesting the next learning content based on the user's learning progress,

[0939] A system that includes this.

[0940] (Claim 2)

[0941] The system according to claim 1, which adjusts the difficulty level of a quiz according to the user's learning and emotional state using the results of emotional analysis based on the analyzed information.

[0942] (Claim 3)

[0943] The system according to claim 1, which includes entertainment elements in the visual user interface to promote an interactive and emotionally conscious learning experience. [Explanation of Symbols]

[0944] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of accessing an educational database and obtaining data for the subject being studied, A means of analyzing acquired data and automatically generating quiz questions, A means for selecting images or diagrams related to the generated quiz questions and integrating them visually, A means of sending integrated quizzes and visuals to the user's terminal, A means of collecting and analyzing user responses, A means of suggesting the next learning content based on the user's learning progress, A system that includes this.

2. The system according to claim 1, which adjusts the difficulty level of quiz questions according to the user's learning progress based on the analyzed data.

3. The system according to claim 1, which provides a visual user interface that includes manga-like elements and promotes an interactive learning experience.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A