system
The system addresses the challenge of personalized learning by creating a virtual reality environment with 3D models and real-time feedback, improving educational engagement and career guidance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Existing educational systems struggle to provide content that is tailored to individual learning needs, making it difficult for learners to understand abstract concepts and lack personalized career guidance, leading to disinterest and ineffective learning.
A system that generates a virtual reality environment using 3D models and supplementary materials based on user interests and knowledge, analyzes voice inputs, and provides real-time feedback and career suggestions.
Enables an individually optimized learning experience by providing immersive and interactive education, enhancing understanding and offering personalized career guidance.
Smart Images

Figure 2026070162000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In existing educational services, there are problems that abstract concepts are difficult to understand and it is difficult to provide content according to individual learning needs. As a result, learners are not interested and effective learning is hindered. Also, since future career guidance remains at the level of general information provision, optimal career paths have not been proposed for individual learners.
Means for Solving the Problems
[0005] This invention provides a system that selects 3D models and supplementary materials according to the learning theme based on the user's interests and existing knowledge, and generates a virtual reality environment in real time. It analyzes the user's voice input using natural language processing technology and presents appropriate information to aid the learner's understanding. Furthermore, it provides a means for recording and analyzing learning behavior and providing feedback and future career path suggestions based on individual comprehension levels.
[0006] A "user" is an individual who uses a system to learn.
[0007] A "virtual reality environment" is a 3D environment created using computer technology that provides an experience similar to reality.
[0008] A "3D model" refers to a visual representation of an object in three-dimensional space that can be observed and manipulated by the user.
[0009] "Supplementary materials" refer to additional information and data provided to help understand the learning content, and can be in various formats such as text, audio, and video.
[0010] A "voice command" is a voice-based input method used by users to request actions or information from a system.
[0011] "Natural language processing technology" is a technology that enables computers to understand and process the language that humans use in everyday life.
[0012] "Feedback" refers to the advice and evaluation that a system provides regarding a learner's understanding and behavior.
[0013] "Career guidance" refers to recommending future occupations and academic fields based on the user's learning progress and interests.
[0014] "Real-time" refers to the time interval during which processing and action are taken immediately after an event occurs.
[0015] A "learning theme" refers to a specific field or topic that learners are interested in and want to deepen their understanding of.
Brief Explanation of Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiment for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention provides an educational platform that generates a virtual reality environment according to the user's interests and learning needs, thereby supporting knowledge acquisition. This system provides a individually optimized learning experience through the user's interactive operation.
[0038] Specifically, users first input their interests and goals to access the system. Once the device receives this information, it sends it to the server, and content preparation based on the selected learning theme begins. The server creates the selected 3D models and supplementary materials, sends them to the device, and sets up the virtual reality environment. The user then uses their smartphone to put on a VR headset and experience immersive learning.
[0039] When a user asks a question to the system using voice or gestures, the server receives the information and analyzes the question using natural language processing technology. The result is generated as an appropriate response and presented to the user via the terminal along with relevant information. In this process, the server performs data processing in real time to help the user resolve their questions. Furthermore, the data generated during learning is accumulated by the server, and the user's learning progress is evaluated based on this data.
[0040] As a concrete example, consider a case where a user becomes interested in "plant photosynthesis" and selects this topic through the system. In this case, the server prepares 3D models of plant cells and visualized materials of the photosynthetic process and sends them to the terminal. The user can observe these in a VR environment and learn about the mechanism of photosynthesis through experience. If the user asks aloud, "What are the final products of photosynthesis?", the server processes the question immediately and provides an answer such as, "Glucose and oxygen."
[0041] In this way, the invention provides an environment in which users can proactively advance their learning, and further contributes to learners' future planning by providing career advice based on information collected by the server. This invention aims to improve the quality of education by providing an individually optimized learning environment.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The user accesses the system and logs in. Once logged in, the user proceeds to a screen where they can select learning topics of interest.
[0045] Step 2:
[0046] The device retrieves the user's selected learning theme information and sends it to the server. The theme includes specific fields or topics that the user wants to learn about.
[0047] Step 3:
[0048] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected from the database. The user's interests and existing knowledge level are also considered during the selection process.
[0049] Step 4:
[0050] The server sends selected learning content to the device. The sent content includes 3D models and text materials related to the theme.
[0051] Step 5:
[0052] The device uses the received content to prepare the virtual reality environment, enabling the user to begin the experience using a VR headset. The user becomes immersed in the virtual reality environment and learns interactively.
[0053] Step 6:
[0054] Users can ask the system questions using voice and gestures during the learning process. For example, questions like, "What is the theory behind this phenomenon?"
[0055] Step 7:
[0056] The server receives voice input from the user and analyzes the input using natural language processing techniques. It then generates an appropriate response and prepares relevant 3D objects and information.
[0057] Step 8:
[0058] The server sends the generated responses and information to the terminal and presents them to the user. This allows the user to receive answers to their questions in real time.
[0059] Step 9:
[0060] The server records the user's learning behavior. This data includes information such as what content was spent on and what questions were asked.
[0061] Step 10:
[0062] The server analyzes the collected data to evaluate the user's understanding and progress. Based on these results, it generates and provides individually optimized feedback to the user.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] Traditional education systems have struggled to provide an optimal learning experience tailored to each user's interests and learning style. Furthermore, they lacked immediate answers to user questions during learning and suggestions for future directions based on learning progress. This has resulted in a challenge in providing effective learning support.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes means for inputting user information and objectives, means for selecting a three-dimensional model and additional materials according to the learning target, and means for analyzing the user's characteristics and suggesting future choices. This makes it possible to provide individually optimized real-time learning support and an environment that promotes the user's proactive learning.
[0068] A "user" refers to a person who uses a system to gain a learning experience.
[0069] "Information" refers to data that users input to record their interests and learning objectives.
[0070] A "three-dimensional model" refers to a digital model generated and provided by a system to visualize objects or processes in three dimensions.
[0071] "Additional materials" refer to documents and diagrams that provide supplementary information in digital format related to the learning theme selected by the user.
[0072] "Voice instructions" refer to instructions or questions that users input into the system using their voice.
[0073] "Action" refers to a method of inputting information into a system using physical movements such as gestures.
[0074] "Related information" refers to information related to the learning content that the system generates in response to user instructions or questions.
[0075] "Immediate" refers to the timing of responses that provide information without delay in response to user actions or inquiries.
[0076] "Learning activity" refers to a series of actions taken by a user to acquire knowledge and skills using a system.
[0077] "Progress" refers to the state indicating how far the user's learning has progressed.
[0078] A "generative artificial intelligence model" refers to artificial intelligence technology used to analyze a user's learning patterns and generate a learning experience tailored to their individual needs.
[0079] "Characteristics" refer to elements that shape an individual's learning style, such as a user's interests, learning ability, and habits.
[0080] "Choice" refers to future activities or directions that the system suggests to the user based on its analysis results.
[0081] This invention provides an educational platform that generates a virtual reality environment based on the user's interests and learning needs, and assists in knowledge acquisition. This system enables a individually optimized learning experience through user interaction.
[0082] The user first inputs their interests and learning objectives through a terminal. The terminal sends this information to a server, which generates appropriate 3D models and supplementary materials based on the input. This generation process utilizes a generative AI model, efficiently generating data using prompts.
[0083] As a concrete example, consider a case where a user expresses interest in "plant photosynthesis." Based on this information, the server generates a three-dimensional model of plant cells and materials explaining the photosynthetic process, and sends this data to the user's device. The user then uses their smartphone to wear a VR headset and experience an immersive learning environment. Within this virtual reality environment, the user can understand the mechanism of photosynthesis through visual information.
[0084] Furthermore, if a user asks a question via voice, such as "What are the final products of photosynthesis?", the server uses natural language processing technology to analyze the question. Based on the analysis, the server generates an appropriate answer, such as "glucose and oxygen," and immediately presents it to the user via the terminal. This process allows users to resolve their questions in real time and deepen their learning.
[0085] User learning activities are recorded on the server and used for progress evaluation. The accumulated data is analyzed using a generative AI model, and feedback is provided to further improve individual learning experiences. An example of a prompt might be, "Prepare a 3D model and supplementary materials needed if the user says they want to learn about plant photosynthesis."
[0086] In this way, the system provides users with an optimized learning environment, contributing to an improvement in the quality of learning.
[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0088] Step 1:
[0089] The user inputs their interests and learning objectives into the device. For example, by inputting "photosynthesis in plants," the learning theme is transmitted to the server. This information forms the basis for generating learning content based on the user's interests. The device sends this input to the server as digital data.
[0090] Step 2:
[0091] The server generates content using a generative AI model based on the received user information. In this process, three-dimensional models and supplementary materials related to the input theme are selected and generated based on prompt text. This ensures that the most suitable visual materials and explanations are prepared for the learning target. The generated content is processed as digital data on the server and sent to the terminal.
[0092] Step 3:
[0093] The terminal receives data sent from the server and sets up the VR environment. It prepares the application on the terminal to display the 3D model correctly. The user puts on a VR headset and begins the learning experience in virtual reality.
[0094] Step 4:
[0095] Users can ask questions and give instructions using voice and gestures within the VR environment. For example, they can ask a question by voice, such as, "What is the final product of photosynthesis?" The device converts the voice information into text data and sends it to the server.
[0096] Step 5:
[0097] The server analyzes the received input data using natural language processing techniques. Based on the user's question, the server generates relevant and accurate response data. This response is intended to aid the user's understanding and is immediately sent to the terminal.
[0098] Step 6:
[0099] The terminal receives response data from the server and presents it to the user visually or audibly. The user then deepens their understanding of the material based on the presented information. This enables interactive and individually optimized learning.
[0100] Step 7:
[0101] The server records the user's actions and questions during learning and stores them in a database. This stored data is used to analyze the user's learning progress and provide feedback. Through this data, users can receive continuous learning improvement and progress evaluation.
[0102] (Application Example 1)
[0103] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0104] Modern education demands providing learning environments tailored to the individual needs and interests of each learner. However, general educational content tends to be delivered in a uniform format, making it difficult to customize according to the learner's level of understanding and interests. Furthermore, limited opportunities for learning through real-world experiences mean that knowledge is absorbed only passively and in a fixed manner. It is necessary to improve the quality of education by utilizing virtual environments to provide optimal learning experiences for each individual learner.
[0105] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0106] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting information models and supplementary data based on the learning theme, and means for transmitting the selected content to the user information processing device and preparing a virtual environment. This enables learning in an individually optimized virtual environment.
[0107] A "user" is a learner who uses the system to access educational content.
[0108] "Interests and existing knowledge" refers to a user's interest in a particular field and the information they have already acquired.
[0109] An "information model" refers to three-dimensional data or virtual objects created based on a learning theme, which aids in visual understanding.
[0110] "Supplemental data" refers to information such as text, images, and videos that complement the information model and deepen the understanding of the learning material.
[0111] A "user information processing device" is a general term for information technology terminals used by users that enable access to a virtual environment.
[0112] A "virtual environment" is a simulation space constructed using digital technology, where users can learn through immersive experiences.
[0113] "Situation analysis" refers to interpreting voice commands and gestures made by the user and processing information accordingly.
[0114] "Relevant information" refers to learning-enhancing answers and additional data provided based on the user's questions and actions.
[0115] "AI technology" refers to technologies that use algorithms and programs related to artificial intelligence to imitate human intelligent behavior.
[0116] "Analyzing voice questions and generating information" refers to the process of receiving a user's voice inquiry, analyzing it, and creating an appropriate response.
[0117] In this invention, users access an interface using a smart device and select a theme based on their interests and learning needs. The server searches for relevant information models and supplementary data based on the data transmitted by the user and transmits the selected content to the user's information processing device. The virtual environment is built using Unity software, and users can obtain an immersive learning experience through it via a head-mounted display or similar device.
[0118] When a user asks a question to the system using voice commands or gestures, the server uses the Google® Cloud Natural Language API to analyze the speech. Based on the analysis, AI technology generates relevant information and answers, which are immediately presented to the user. This process is designed to allow users to resolve their questions in real time.
[0119] For example, if the educational theme is "the structure of historical buildings," the server retrieves and transmits an information model of ancient Roman architecture from the system. The user explores this model in a VR environment and asks a voice question such as, "What are the features of this building?", to which the AI responds with appropriate information based on the situation. An example of a prompt is as follows:
[0120] Examples of prompts for a generative AI model:
[0121] "Please describe the structure and features of this ancient building in detail."
[0122] In this way, the present invention provides an individually optimized learning space, enabling users to learn proactively and effectively.
[0123] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0124] Step 1:
[0125] Users access the interface via their smart devices and select learning themes that match their interests and learning needs. The input consists of data about the user's interests and themes, which is sent from the device to the server.
[0126] Step 2:
[0127] The server searches for relevant information models and supplementary data based on the user's theme information received. This process involves matching the database and determining the selected content. The output is the information model and supplementary data that are sent.
[0128] Step 3:
[0129] The selected content is sent from the server to the user's information processing device. The terminal processes the received data and generates a 3D view to prepare the virtual reality environment. The input is an information model and supplementary data, and the output is the virtual reality environment.
[0130] Step 4:
[0131] The user immerses themselves in a virtual environment using a head-mounted display and begins learning through an information model. Voice commands and gesture inputs from the user are acquired by the system and sent to the server.
[0132] Step 5:
[0133] The server uses the Google Cloud Natural Language API to analyze speech input. Here, the speech input is converted into text data, and natural language processing operations are performed. The output is the analysis result corresponding to the user's question.
[0134] Step 6:
[0135] The server uses a generative AI model to generate relevant information and answers based on the analysis results. At this stage, the optimal answer to the question is generated and presented as information. The output is response data for the user.
[0136] Step 7:
[0137] The generated information and responses are immediately transmitted to the user's terminal and presented through the user's head-mounted display. Here, the terminal displays the received data and provides information to the user. The output is the response information that the user visually perceives.
[0138] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0139] This invention provides a system that incorporates an emotion engine to recognize user emotions and dynamically adjust learning content in order to enhance the user's learning experience. This system provides emotion-responsive feedback and adjusts the learning environment to promote more effective learning.
[0140] Specifically, the user first logs into the system and selects a learning topic of interest. The device sends this information to the server, which selects 3D models and supplementary materials corresponding to the topic. The selected content is sent to the device, and the user begins learning through the virtual reality environment.
[0141] The key here is that the emotion engine recognizes emotions from the user's voice and gestures. For example, if a user shows a confused expression during learning, the emotion engine recognizes this as "confusion." Based on this information, the server helps the user understand by providing a simple explanation of the learning content or additional supporting materials. Also, when the emotion engine recognizes a positive emotion from the user, it presents more advanced content and supports the user in exploring their interests more deeply.
[0142] As a concrete example, suppose a user chooses the topic of "the expansion of the universe." If the user appears confused while observing a model of the universe in the VR environment, the emotion engine will detect this, and the server will provide an additional animation visualizing the rate of expansion, along with a clear explanation. Through this process, the user can deepen their understanding and maintain their interest.
[0143] Furthermore, the server records changes in emotions and incorporates them into the evaluation of learning progress. This data is used for user feedback and optimizing future learning plans. By integrating with the emotion engine, the system provides a personalized learning experience for each learner, contributing to improved quality of education. In this way, it creates a flexible and dynamic learning environment that is attentive to the user's feelings.
[0144] The following describes the processing flow.
[0145] Step 1:
[0146] The user logs into the system and selects a learning topic of interest. The device sends the user's selection information to the server.
[0147] Step 2:
[0148] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected. This includes extracting information from the database and customizing it according to the user's interests and knowledge level.
[0149] Step 3:
[0150] The server sends the selected content to the terminal. The terminal uses the received content to prepare the virtual reality environment and configure it to be accessible to the user.
[0151] Step 4:
[0152] The user begins learning in a prepared virtual reality environment using a VR headset. The user explores learning topics by manipulating and observing 3D models.
[0153] Step 5:
[0154] The device detects the user's voice and gestures, and the emotion engine analyzes them. Positive or negative emotions are then detected.
[0155] Step 6:
[0156] The server dynamically adjusts learning content based on the user's emotions, as identified by the emotion engine. For example, if the user is confused, the server generates additional explanations or more concise materials and sends them to the device.
[0157] Step 7:
[0158] The device presents the user with customized content. This allows the user to receive a learning experience optimized for their own emotional state.
[0159] Step 8:
[0160] The server records the user's learning behavior, including changes in their emotions. The collected data is used to analyze the user's progress and generate feedback, and is also reflected in future learning plans.
[0161] (Example 2)
[0162] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0163] A challenge in modern education systems is the insufficient customization of learning experiences to suit individual users' emotions and learning progress. Many systems are limited to static content delivery, making it difficult to dynamically adjust content based on users' understanding and interests. As a result, learning effectiveness may decline and motivation may be lost.
[0164] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0165] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting three-dimensional representations and supplementary information based on the learning theme, and means for analyzing the user's feelings and dynamically adjusting the learning content. This makes it possible to provide a flexible and effective learning experience tailored to each individual user.
[0166] A "user" refers to a person who uses the system to learn.
[0167] "Interest" refers to learning themes or topics that users are interested in.
[0168] "Existing knowledge" refers to information and concepts that the user has already acquired or understands.
[0169] "Means" refers to the methods or processes that a system uses to perform a particular function.
[0170] "Three-dimensional representation" refers to visual materials such as 3D models used to visually illustrate learning themes.
[0171] "Supplementary information" refers to additional materials and data that complement the three-dimensional representation and deepen the user's understanding.
[0172] "User equipment" refers to the hardware that a user uses to receive learning content and experience a virtual environment.
[0173] A "virtual environment" refers to a simulated space created through digital technology that allows users to have an immersive experience.
[0174] "Situation analysis" refers to the process of interpreting the user's voice commands and actions to generate information for providing appropriate learning content.
[0175] "Generated information" refers to new data and content created through situational analysis and presented to the user.
[0176] "User sentiment" refers to the emotional state interpreted through the user's voice and facial expressions.
[0177] "Dynamic adjustment" refers to changing learning content in real time to the optimal format based on the user's mood and progress.
[0178] "Recording emotional changes" refers to accumulating data on the evolution of a user's emotional state and using it for subsequent analysis.
[0179] This invention provides a learning system for personalizing and efficiently and dynamically adjusting the user's learning experience. The system aims to detect the user's emotional state in real time and optimize learning content based on those emotions.
[0180] Users access and log in to the system's dedicated application using their own devices. These devices are connected to VR goggles or other VR-compatible devices, allowing users to experience a virtual environment. This environment provides advanced three-dimensional representations and supplementary information, enabling users to explore learning themes based on their interests.
[0181] The server selects an appropriate three-dimensional representation from its database based on the learning theme chosen by the user and sends it to the user's terminal. The server also drives an emotion engine to detect the user's emotions and analyzes the user's voice and gestures using technologies such as voice analysis and image recognition.
[0182] This system dynamically changes the learning content based on the user's emotions during the learning process. For example, if the system determines that the user is confused, the server generates simplified supplementary materials and sends them to the user's device. In this way, it supports the user's understanding and enhances the learning effect.
[0183] As a concrete example, consider using this system when a user is learning about "the expansion of the universe." When the user becomes confused while observing the universe model in a VR environment, the emotion engine can detect this emotion, and the server can provide an additional "animation that visualizes the rate of expansion."
[0184] By utilizing generative AI models, users can further customize their learning experience. For example, prompts could include instructions such as, "Please provide specific instructions on how to respond if the user shows a confused expression while learning about the expansion of the universe in a VR environment."
[0185] These features enable the present invention to provide a dynamic learning experience optimized for individual users, thereby improving the quality of education.
[0186] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0187] Step 1:
[0188] The user logs into the system and selects a learning topic. The user enters their topic of interest via a terminal and presses a select button. Based on this input, the user's interest information is sent from the terminal to the server. The output is information about the selected topic.
[0189] Step 2:
[0190] The device sends information about the selected learning theme to the server. Based on the received information, the server selects relevant three-dimensional representations and supplementary information from its database. This data processing prepares the most suitable learning content for the user. The output is the selected learning content.
[0191] Step 3:
[0192] The server sends the selected learning content to the user's terminal. The terminal receives this content and prepares a virtual environment. The virtual environment is experienced by the user using VR goggles and includes three-dimensional visuals and sound effects related to the learning theme. The output is the learning content represented within the virtual environment.
[0193] Step 4:
[0194] The user begins learning in a virtual environment. The user's voice commands and gestures are collected by the device and sent to the server as input data. During this process, the user's interactions are recorded in real time. The output is behavioral data for obtaining user feedback.
[0195] Step 5:
[0196] The server analyzes the received voice commands and gesture data. The emotion engine uses this data to recognize the user's feelings and intentions. Based on the analysis of the emotional state, dynamic adjustments are made to the learning content. The output is the user's emotional state and the adjusted content.
[0197] Step 6:
[0198] The adjusted content is resent from the server to the device. This allows the device to provide immediate feedback to the user. This process personalizes the user's learning experience. The output is an improved learning experience.
[0199] Step 7:
[0200] The server records changes in the user's emotions and learning progress, and uses this data to optimize future learning plans. This data may be input as prompts to the generating AI model, which then makes suggestions for optimizing learning. The output consists of the accumulated progress data and its analysis results.
[0201] (Application Example 2)
[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0203] Traditional training systems and learning environments have struggled to adapt to the individual emotions and levels of understanding of each user, making it difficult to achieve efficient learning, especially in situations where personalized support is required. This can lead to learners becoming confused or losing motivation, making it difficult to achieve optimal learning outcomes.
[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0205] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting visual models and supplementary materials based on the learning theme, and means for detecting the user's emotions and dynamically adjusting the content based on those emotions. This makes it possible to provide a dynamic learning environment that responds to the emotional state of each individual user and realize an effective learning process tailored to individual needs.
[0206] "Means for inputting user interests and existing knowledge" refers to methods for identifying the user's current areas of interest and existing knowledge, and registering that information in the system.
[0207] "Means for selecting visual models and supplementary materials based on learning themes" refers to methods for identifying and providing the most suitable visual representations and related materials for the learning theme selected by the user.
[0208] "Means for transmitting selected information to a terminal device and preparing a virtual reality environment" refers to a method of transmitting selected visual data and materials to the user's terminal and setting up a virtual environment.
[0209] "Means for detecting user emotions and dynamically adjusting content based on those emotions" refers to a method of understanding the user's emotional state in real time and appropriately modifying the learning content based on that information.
[0210] "Means for presenting analysis results and generated information to users in real time" refers to methods for immediately providing users with the analyzed results and generated information.
[0211] "Means for recording user activity and analyzing progress" refers to methods for tracking a user's learning progress and analyzing how far they have progressed.
[0212] "Means of providing feedback based on progress" refers to methods for providing appropriate feedback according to the user's learning progress.
[0213] "A means of analyzing individual trends and proposing future strategies" refers to a method of understanding a user's learning tendencies and proposing future learning strategies based on those tendencies.
[0214] The system that realizes this invention consists of a user, a terminal device, and a server. The user first logs into the terminal device and selects a learning topic of interest. This information is sent from the terminal to the server. The server selects a suitable visual model and supplementary materials and sends them to the terminal device, preparing to provide the user with a virtual reality environment.
[0215] The terminal device is equipped with sensors to detect the user's facial expressions and voice, which detect the user's emotions in real time. The hardware used includes a camera and microphone, and the software employs emotion recognition algorithms (e.g., DeepFace).
[0216] When a user expresses confusion or lack of understanding during the learning process, that emotional data is transferred to the server. Based on this emotional data, the server dynamically adjusts the learning content. For example, if a user does not understand a task, an additional explanatory video may be displayed on their device.
[0217] Furthermore, the server provides feedback based on the analyzed user's learning progress. This allows for the analysis of the user's individual learning tendencies and the suggestion of a more effective learning plan. This process enables users with questions about specific learning topics to receive immediate support.
[0218] By using a generative AI model, it is possible to generate diverse learning content in response to prompt input. An example of a prompt is, "Please run a program that generates specific supplementary information based on the user's learning content."
[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0220] Step 1:
[0221] The user logs into the terminal device. The user enters a learning topic of interest, and this data is sent from the terminal to the server. The input is the learning topic, and the output is the result of sending the topic information to the server. The terminal packages the topic data in an appropriate format and sends it to the server.
[0222] Step 2:
[0223] The server selects appropriate visual models and supplementary materials based on the received learning theme. The input is theme information submitted by the user, and the output is the selected content. The server searches the database and identifies the most relevant visual content related to the theme.
[0224] Step 3:
[0225] The server sends selected content to the terminal device, preparing the virtual reality environment. The input is the selected content, and the output is the content delivery result to the terminal. The server packages the selected materials in digital format and sends them to the terminal. The terminal receives this and sets up the virtual reality environment for the user.
[0226] Step 4:
[0227] The terminal device uses sensors to detect the user's emotions in real time. Input is sensor information from cameras and microphones, and output is emotion data. The terminal uses a machine learning model (e.g., DeepFace) to analyze the user's emotions from their face and voice.
[0228] Step 5:
[0229] The server dynamically adjusts the learning content based on detected sentiment data. The input is sentiment data, and the output is the adjusted learning content. The server analyzes the sentiment data, re-selects appropriate learning materials and supplementary resources, and provides a learning experience tailored to the user's current situation.
[0230] Step 6:
[0231] The server processes data to record the user's learning progress and generate feedback. The input is the user's learning behavior data, and the output is feedback information. The server analyzes the progress data and generates and provides appropriate feedback to the user.
[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0235] [Second Embodiment]
[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0248] This invention provides an educational platform that generates a virtual reality environment according to the user's interests and learning needs, thereby supporting knowledge acquisition. This system provides a individually optimized learning experience through the user's interactive operation.
[0249] Specifically, users first input their interests and goals to access the system. Once the device receives this information, it sends it to the server, and content preparation based on the selected learning theme begins. The server creates the selected 3D models and supplementary materials, sends them to the device, and sets up the virtual reality environment. The user then uses their smartphone to put on a VR headset and experience immersive learning.
[0250] When a user asks a question to the system using voice or gestures, the server receives the information and analyzes the question using natural language processing technology. The result is generated as an appropriate response and presented to the user via the terminal along with relevant information. In this process, the server performs data processing in real time to help the user resolve their questions. Furthermore, the data generated during learning is accumulated by the server, and the user's learning progress is evaluated based on this data.
[0251] As a concrete example, consider a case where a user becomes interested in "plant photosynthesis" and selects this topic through the system. In this case, the server prepares 3D models of plant cells and visualized materials of the photosynthetic process and sends them to the terminal. The user can observe these in a VR environment and learn about the mechanism of photosynthesis through experience. If the user asks aloud, "What are the final products of photosynthesis?", the server processes the question immediately and provides an answer such as, "Glucose and oxygen."
[0252] In this way, the invention provides an environment in which users can proactively advance their learning, and further contributes to learners' future planning by providing career advice based on information collected by the server. This invention aims to improve the quality of education by providing an individually optimized learning environment.
[0253] The following describes the processing flow.
[0254] Step 1:
[0255] The user accesses the system and logs in. Once logged in, the user proceeds to a screen where they can select learning topics of interest.
[0256] Step 2:
[0257] The device retrieves the user's selected learning theme information and sends it to the server. The theme includes specific fields or topics that the user wants to learn about.
[0258] Step 3:
[0259] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected from the database. The user's interests and existing knowledge level are also considered during the selection process.
[0260] Step 4:
[0261] The server sends selected learning content to the device. The sent content includes 3D models and text materials related to the theme.
[0262] Step 5:
[0263] The device uses the received content to prepare the virtual reality environment, enabling the user to begin the experience using a VR headset. The user becomes immersed in the virtual reality environment and learns interactively.
[0264] Step 6:
[0265] Users can ask the system questions using voice and gestures during the learning process. For example, questions like, "What is the theory behind this phenomenon?"
[0266] Step 7:
[0267] The server receives voice input from the user and analyzes the input using natural language processing techniques. It then generates an appropriate response and prepares relevant 3D objects and information.
[0268] Step 8:
[0269] The server sends the generated responses and information to the terminal and presents them to the user. This allows the user to receive answers to their questions in real time.
[0270] Step 9:
[0271] The server records the user's learning behavior. This data includes information such as what content was spent on and what questions were asked.
[0272] Step 10:
[0273] The server analyzes the collected data to evaluate the user's understanding and progress. Based on these results, it generates and provides individually optimized feedback to the user.
[0274] (Example 1)
[0275] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0276] Traditional education systems have struggled to provide an optimal learning experience tailored to each user's interests and learning style. Furthermore, they lacked immediate answers to user questions during learning and suggestions for future directions based on learning progress. This has resulted in a challenge in providing effective learning support.
[0277] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0278] In this invention, the server includes means for inputting user information and purposes, means for selecting a three-dimensional model and additional materials according to the learning target, and means for analyzing user characteristics and proposing future selections. This enables the provision of individually optimized real-time learning support and an environment that promotes the user's active learning.
[0279] The "user" refers to a person who attempts to obtain a learning experience using the system.
[0280] The "information" refers to data input by the user to record their interests and learning purposes.
[0281] The "three-dimensional model" refers to a digital model generated and provided by the system to visually represent an object or process three-dimensionally.
[0282] The "additional materials" refer to documents and drawings that provide supplementary information related to the learning theme selected by the user in digital form.
[0283] The "voice instruction" refers to instructions and questions input by the user to the system using voice.
[0284] The "action" refers to an operation method input to the system using body movements such as gestures.
[0285] The "related information" refers to information related to the learning content generated by the system in response to the user's instructions and questions.
[0286] The "immediate" refers to the timing of the response that provides information without delay to the user's operations and inquiries.
[0287] The "learning activity" refers to a series of actions by which the user attempts to acquire knowledge and skills using the system.
[0288] The "progress" refers to the state indicating how far the user's learning has advanced.
[0289] A "generative artificial intelligence model" refers to artificial intelligence technology used to analyze a user's learning patterns and generate a learning experience tailored to their individual needs.
[0290] "Characteristics" refer to elements that shape an individual's learning style, such as a user's interests, learning ability, and habits.
[0291] "Choice" refers to future activities or directions that the system suggests to the user based on its analysis results.
[0292] This invention provides an educational platform that generates a virtual reality environment based on the user's interests and learning needs, and assists in knowledge acquisition. This system enables a individually optimized learning experience through user interaction.
[0293] The user first inputs their interests and learning objectives through a terminal. The terminal sends this information to a server, which generates appropriate 3D models and supplementary materials based on the input. This generation process utilizes a generative AI model, efficiently generating data using prompts.
[0294] As a concrete example, consider a case where a user expresses interest in "plant photosynthesis." Based on this information, the server generates a three-dimensional model of plant cells and materials explaining the photosynthetic process, and sends this data to the user's device. The user then uses their smartphone to wear a VR headset and experience an immersive learning environment. Within this virtual reality environment, the user can understand the mechanism of photosynthesis through visual information.
[0295] Furthermore, if a user asks a question via voice, such as "What are the final products of photosynthesis?", the server uses natural language processing technology to analyze the question. Based on the analysis, the server generates an appropriate answer, such as "glucose and oxygen," and immediately presents it to the user via the terminal. This process allows users to resolve their questions in real time and deepen their learning.
[0296] User learning activities are recorded on the server and used for progress evaluation. The accumulated data is analyzed using a generative AI model, and feedback is provided to further improve individual learning experiences. An example of a prompt might be, "Prepare a 3D model and supplementary materials needed if the user says they want to learn about plant photosynthesis."
[0297] In this way, the system provides users with an optimized learning environment, contributing to an improvement in the quality of learning.
[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0299] Step 1:
[0300] The user inputs their interests and learning objectives into the device. For example, by inputting "photosynthesis in plants," the learning theme is transmitted to the server. This information forms the basis for generating learning content based on the user's interests. The device sends this input to the server as digital data.
[0301] Step 2:
[0302] The server generates content using a generative AI model based on the received user information. In this process, three-dimensional models and supplementary materials related to the input theme are selected and generated based on prompt text. This ensures that the most suitable visual materials and explanations are prepared for the learning target. The generated content is processed as digital data on the server and sent to the terminal.
[0303] Step 3:
[0304] The terminal receives data sent from the server and sets up the VR environment. It prepares the application on the terminal to display the 3D model correctly. The user puts on a VR headset and begins the learning experience in virtual reality.
[0305] Step 4:
[0306] The user makes questions or gives instructions using voice or gestures within the VR environment. For example, the user can ask a question verbally such as "What is the final product of photosynthesis?" The terminal converts the voice information into text data and transmits it to the server.
[0307] Step 5:
[0308] The server analyzes the received input data using natural language processing technology. Based on the content of the user's question, the server generates relevant and accurate response data. This response is for helping the user's understanding and is immediately transmitted to the terminal.
[0309] Step 6:
[0310] The terminal receives the response data from the server and presents it to the user visually or aurally. Based on the presented information, the user deepens the learning content. Thereby, interactive and individually optimized learning is realized.
[0311] Step 7:
[0312] The server records the actions and question content of the user during learning and accumulates them in the database. The accumulated data is used for analyzing the user's learning progress and providing feedback. The user can receive continuous learning improvement and progress evaluation through this data.
[0313] (Application Example 1)
[0314] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0315] Modern education demands providing learning environments tailored to the individual needs and interests of each learner. However, general educational content tends to be delivered in a uniform format, making it difficult to customize according to the learner's level of understanding and interests. Furthermore, limited opportunities for learning through real-world experiences mean that knowledge is absorbed only passively and in a fixed manner. It is necessary to improve the quality of education by utilizing virtual environments to provide optimal learning experiences for each individual learner.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0317] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting information models and supplementary data based on the learning theme, and means for transmitting the selected content to the user information processing device and preparing a virtual environment. This enables learning in an individually optimized virtual environment.
[0318] A "user" is a learner who uses the system to access educational content.
[0319] "Interests and existing knowledge" refers to a user's interest in a particular field and the information they have already acquired.
[0320] An "information model" refers to three-dimensional data or virtual objects created based on a learning theme, which aids in visual understanding.
[0321] "Supplemental data" refers to information such as text, images, and videos that complement the information model and deepen the understanding of the learning material.
[0322] A "user information processing device" is a general term for information technology terminals used by users that enable access to a virtual environment.
[0323] A "virtual environment" is a simulation space constructed using digital technology, where users can learn through immersive experiences.
[0324] "Situation analysis" refers to interpreting voice commands and gestures made by the user and processing information accordingly.
[0325] "Relevant information" refers to learning-enhancing answers and additional data provided based on the user's questions and actions.
[0326] "AI technology" refers to technologies that use algorithms and programs related to artificial intelligence to imitate human intelligent behavior.
[0327] "Analyzing voice questions and generating information" refers to the process of receiving a user's voice inquiry, analyzing it, and creating an appropriate response.
[0328] In this invention, users access an interface using a smart device and select a theme based on their interests and learning needs. The server searches for relevant information models and supplementary data based on the data transmitted by the user and transmits the selected content to the user's information processing device. The virtual environment is built using Unity software, and users can obtain an immersive learning experience through it via a head-mounted display or similar device.
[0329] When a user asks a question to the system using voice commands or gestures, the server uses the Google Cloud Natural Language API to analyze the speech. Based on the analysis, AI technology generates relevant information and answers, which are immediately presented to the user. This process is designed to allow users to resolve their questions in real time.
[0330] For example, if the educational theme is "the structure of historical buildings," the server retrieves and transmits an information model of ancient Roman architecture from the system. The user explores this model in a VR environment and asks a voice question such as, "What are the features of this building?", to which the AI responds with appropriate information based on the situation. An example of a prompt is as follows:
[0331] Examples of prompts for a generative AI model:
[0332] "Please describe the structure and features of this ancient building in detail."
[0333] In this way, the present invention provides an individually optimized learning space, enabling users to learn proactively and effectively.
[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0335] Step 1:
[0336] Users access the interface via their smart devices and select learning themes that match their interests and learning needs. The input consists of data about the user's interests and themes, which is sent from the device to the server.
[0337] Step 2:
[0338] The server searches for relevant information models and supplementary data based on the user's theme information received. This process involves matching the database and determining the selected content. The output is the information model and supplementary data that are sent.
[0339] Step 3:
[0340] The selected content is sent from the server to the user's information processing device. The terminal processes the received data and generates a 3D view to prepare the virtual reality environment. The input is an information model and supplementary data, and the output is the virtual reality environment.
[0341] Step 4:
[0342] The user immerses themselves in a virtual environment using a head-mounted display and begins learning through an information model. Voice commands and gesture inputs from the user are acquired by the system and sent to the server.
[0343] Step 5:
[0344] The server uses the Google Cloud Natural Language API to analyze speech input. Here, the speech input is converted into text data, and natural language processing operations are performed. The output is the analysis result corresponding to the user's question.
[0345] Step 6:
[0346] The server uses a generative AI model to generate relevant information and answers based on the analysis results. At this stage, the optimal answer to the question is generated and presented as information. The output is response data for the user.
[0347] Step 7:
[0348] The generated information and responses are immediately transmitted to the user's terminal and presented through the user's head-mounted display. Here, the terminal displays the received data and provides information to the user. The output is the response information that the user visually perceives.
[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0350] This invention provides a system that incorporates an emotion engine to recognize user emotions and dynamically adjust learning content in order to enhance the user's learning experience. This system provides emotion-responsive feedback and adjusts the learning environment to promote more effective learning.
[0351] Specifically, the user first logs into the system and selects a learning topic of interest. The device sends this information to the server, which selects 3D models and supplementary materials corresponding to the topic. The selected content is sent to the device, and the user begins learning through the virtual reality environment.
[0352] The key here is that the emotion engine recognizes emotions from the user's voice and gestures. For example, if a user shows a confused expression during learning, the emotion engine recognizes this as "confusion." Based on this information, the server helps the user understand by providing a simple explanation of the learning content or additional supporting materials. Also, when the emotion engine recognizes a positive emotion from the user, it presents more advanced content and supports the user in exploring their interests more deeply.
[0353] As a concrete example, suppose a user chooses the topic of "the expansion of the universe." If the user appears confused while observing a model of the universe in the VR environment, the emotion engine will detect this, and the server will provide an additional animation visualizing the rate of expansion, along with a clear explanation. Through this process, the user can deepen their understanding and maintain their interest.
[0354] Furthermore, the server records changes in emotions and incorporates them into the evaluation of learning progress. This data is used for user feedback and optimizing future learning plans. By integrating with the emotion engine, the system provides a personalized learning experience for each learner, contributing to improved quality of education. In this way, it creates a flexible and dynamic learning environment that is attentive to the user's feelings.
[0355] The following describes the processing flow.
[0356] Step 1:
[0357] The user logs into the system and selects a learning topic of interest. The device sends the user's selection information to the server.
[0358] Step 2:
[0359] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected. This includes extracting information from the database and customizing it according to the user's interests and knowledge level.
[0360] Step 3:
[0361] The server sends the selected content to the terminal. The terminal uses the received content to prepare the virtual reality environment and configure it to be accessible to the user.
[0362] Step 4:
[0363] The user begins learning in a prepared virtual reality environment using a VR headset. The user explores learning topics by manipulating and observing 3D models.
[0364] Step 5:
[0365] The device detects the user's voice and gestures, and the emotion engine analyzes them. Positive or negative emotions are then detected.
[0366] Step 6:
[0367] The server dynamically adjusts learning content based on the user's emotions, as identified by the emotion engine. For example, if the user is confused, the server generates additional explanations or more concise materials and sends them to the device.
[0368] Step 7:
[0369] The device presents the user with customized content. This allows the user to receive a learning experience optimized for their own emotional state.
[0370] Step 8:
[0371] The server records the user's learning behavior, including changes in their emotions. The collected data is used to analyze the user's progress and generate feedback, and is also reflected in future learning plans.
[0372] (Example 2)
[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0374] A challenge in modern education systems is the insufficient customization of learning experiences to suit individual users' emotions and learning progress. Many systems are limited to static content delivery, making it difficult to dynamically adjust content based on users' understanding and interests. As a result, learning effectiveness may decline and motivation may be lost.
[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0376] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting three-dimensional representations and supplementary information based on the learning theme, and means for analyzing the user's feelings and dynamically adjusting the learning content. This makes it possible to provide a flexible and effective learning experience tailored to each individual user.
[0377] A "user" refers to a person who uses the system to learn.
[0378] "Interest" refers to learning themes or topics that users are interested in.
[0379] "Existing knowledge" refers to information and concepts that the user has already acquired or understands.
[0380] "Means" refers to the methods or processes that a system uses to perform a particular function.
[0381] "Three-dimensional representation" refers to visual materials such as 3D models used to visually illustrate learning themes.
[0382] "Supplementary information" refers to additional materials and data that complement the three-dimensional representation and deepen the user's understanding.
[0383] "User equipment" refers to the hardware that a user uses to receive learning content and experience a virtual environment.
[0384] A "virtual environment" refers to a simulated space created through digital technology that allows users to have an immersive experience.
[0385] "Situation analysis" refers to the process of interpreting the user's voice commands and actions to generate information for providing appropriate learning content.
[0386] "Generated information" refers to new data and content created through situational analysis and presented to the user.
[0387] "User sentiment" refers to the emotional state interpreted through the user's voice and facial expressions.
[0388] "Dynamic adjustment" refers to changing learning content in real time to the optimal format based on the user's mood and progress.
[0389] "Recording emotional changes" refers to accumulating data on the evolution of a user's emotional state and using it for subsequent analysis.
[0390] This invention provides a learning system for personalizing and efficiently and dynamically adjusting the user's learning experience. The system aims to detect the user's emotional state in real time and optimize learning content based on those emotions.
[0391] Users access and log in to the system's dedicated application using their own devices. These devices are connected to VR goggles or other VR-compatible devices, allowing users to experience a virtual environment. This environment provides advanced three-dimensional representations and supplementary information, enabling users to explore learning themes based on their interests.
[0392] The server selects an appropriate three-dimensional representation from its database based on the learning theme chosen by the user and sends it to the user's terminal. The server also drives an emotion engine to detect the user's emotions and analyzes the user's voice and gestures using technologies such as voice analysis and image recognition.
[0393] This system dynamically changes the learning content based on the user's emotions during the learning process. For example, if the system determines that the user is confused, the server generates simplified supplementary materials and sends them to the user's device. In this way, it supports the user's understanding and enhances the learning effect.
[0394] As a concrete example, consider using this system when a user is learning about "the expansion of the universe." When the user becomes confused while observing the universe model in a VR environment, the emotion engine can detect this emotion, and the server can provide an additional "animation that visualizes the rate of expansion."
[0395] By utilizing generative AI models, users can further customize their learning experience. For example, prompts could include instructions such as, "Please provide specific instructions on how to respond if the user shows a confused expression while learning about the expansion of the universe in a VR environment."
[0396] These features enable the present invention to provide a dynamic learning experience optimized for individual users, thereby improving the quality of education.
[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0398] Step 1:
[0399] The user logs into the system and selects a learning topic. The user enters their topic of interest via a terminal and presses a select button. Based on this input, the user's interest information is sent from the terminal to the server. The output is information about the selected topic.
[0400] Step 2:
[0401] The device sends information about the selected learning theme to the server. Based on the received information, the server selects relevant three-dimensional representations and supplementary information from its database. This data processing prepares the most suitable learning content for the user. The output is the selected learning content.
[0402] Step 3:
[0403] The server sends the selected learning content to the user's terminal. The terminal receives this content and prepares a virtual environment. The virtual environment is experienced by the user using VR goggles and includes three-dimensional visuals and sound effects related to the learning theme. The output is the learning content represented within the virtual environment.
[0404] Step 4:
[0405] The user begins learning in a virtual environment. The user's voice commands and gestures are collected by the device and sent to the server as input data. During this process, the user's interactions are recorded in real time. The output is behavioral data for obtaining user feedback.
[0406] Step 5:
[0407] The server analyzes the received voice commands and gesture data. The emotion engine uses this data to recognize the user's feelings and intentions. Based on the analysis of the emotional state, dynamic adjustments are made to the learning content. The output is the user's emotional state and the adjusted content.
[0408] Step 6:
[0409] The adjusted content is resent from the server to the device. This allows the device to provide immediate feedback to the user. This process personalizes the user's learning experience. The output is an improved learning experience.
[0410] Step 7:
[0411] The server records changes in the user's emotions and learning progress, and uses this data to optimize future learning plans. This data may be input as prompts to the generating AI model, which then makes suggestions for optimizing learning. The output consists of the accumulated progress data and its analysis results.
[0412] (Application Example 2)
[0413] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0414] Traditional training systems and learning environments have struggled to adapt to the individual emotions and levels of understanding of each user, making it difficult to achieve efficient learning, especially in situations where personalized support is required. This can lead to learners becoming confused or losing motivation, making it difficult to achieve optimal learning outcomes.
[0415] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0416] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting visual models and supplementary materials based on the learning theme, and means for detecting the user's emotions and dynamically adjusting the content based on those emotions. This makes it possible to provide a dynamic learning environment that responds to the emotional state of each individual user and realize an effective learning process tailored to individual needs.
[0417] "Means for inputting user interests and existing knowledge" refers to methods for identifying the user's current areas of interest and existing knowledge, and registering that information in the system.
[0418] "Means for selecting visual models and supplementary materials based on learning themes" refers to methods for identifying and providing the most suitable visual representations and related materials for the learning theme selected by the user.
[0419] "Means for transmitting selected information to a terminal device and preparing a virtual reality environment" refers to a method of transmitting selected visual data and materials to the user's terminal and setting up a virtual environment.
[0420] "Means for detecting user emotions and dynamically adjusting content based on those emotions" refers to a method of understanding the user's emotional state in real time and appropriately modifying the learning content based on that information.
[0421] "Means for presenting analysis results and generated information to users in real time" refers to methods for immediately providing users with the analyzed results and generated information.
[0422] "Means for recording user activity and analyzing progress" refers to methods for tracking a user's learning progress and analyzing how far they have progressed.
[0423] "Means of providing feedback based on progress" refers to methods for providing appropriate feedback according to the user's learning progress.
[0424] "A means of analyzing individual trends and proposing future strategies" refers to a method of understanding a user's learning tendencies and proposing future learning strategies based on those tendencies.
[0425] The system that realizes this invention consists of a user, a terminal device, and a server. The user first logs into the terminal device and selects a learning topic of interest. This information is sent from the terminal to the server. The server selects a suitable visual model and supplementary materials and sends them to the terminal device, preparing to provide the user with a virtual reality environment.
[0426] The terminal device is equipped with sensors to detect the user's facial expressions and voice, which detect the user's emotions in real time. The hardware used includes a camera and microphone, and the software employs emotion recognition algorithms (e.g., DeepFace).
[0427] When a user expresses confusion or lack of understanding during the learning process, that emotional data is transferred to the server. Based on this emotional data, the server dynamically adjusts the learning content. For example, if a user does not understand a task, an additional explanatory video may be displayed on their device.
[0428] Furthermore, the server provides feedback based on the analyzed user's learning progress. This allows for the analysis of the user's individual learning tendencies and the suggestion of a more effective learning plan. This process enables users with questions about specific learning topics to receive immediate support.
[0429] By using a generative AI model, it is possible to generate diverse learning content in response to prompt input. An example of a prompt is, "Please run a program that generates specific supplementary information based on the user's learning content."
[0430] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0431] Step 1:
[0432] The user logs into the terminal device. The user enters a learning topic of interest, and this data is sent from the terminal to the server. The input is the learning topic, and the output is the result of sending the topic information to the server. The terminal packages the topic data in an appropriate format and sends it to the server.
[0433] Step 2:
[0434] The server selects appropriate visual models and supplementary materials based on the received learning theme. The input is theme information submitted by the user, and the output is the selected content. The server searches the database and identifies the most relevant visual content related to the theme.
[0435] Step 3:
[0436] The server sends selected content to the terminal device, preparing the virtual reality environment. The input is the selected content, and the output is the content delivery result to the terminal. The server packages the selected materials in digital format and sends them to the terminal. The terminal receives this and sets up the virtual reality environment for the user.
[0437] Step 4:
[0438] The terminal device uses sensors to detect the user's emotions in real time. Input is sensor information from cameras and microphones, and output is emotion data. The terminal uses a machine learning model (e.g., DeepFace) to analyze the user's emotions from their face and voice.
[0439] Step 5:
[0440] The server dynamically adjusts the learning content based on detected sentiment data. The input is sentiment data, and the output is the adjusted learning content. The server analyzes the sentiment data, re-selects appropriate learning materials and supplementary resources, and provides a learning experience tailored to the user's current situation.
[0441] Step 6:
[0442] The server processes data to record the user's learning progress and generate feedback. The input is the user's learning behavior data, and the output is feedback information. The server analyzes the progress data and generates and provides appropriate feedback to the user.
[0443] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0444] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0446] [Third Embodiment]
[0447] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0448] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0450] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0454] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0455] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0456] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0457] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0458] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0459] This invention provides an educational platform that generates a virtual reality environment according to the user's interests and learning needs, thereby supporting knowledge acquisition. This system provides a individually optimized learning experience through the user's interactive operation.
[0460] Specifically, users first input their interests and goals to access the system. Once the device receives this information, it sends it to the server, and content preparation based on the selected learning theme begins. The server creates the selected 3D models and supplementary materials, sends them to the device, and sets up the virtual reality environment. The user then uses their smartphone to put on a VR headset and experience immersive learning.
[0461] When a user asks a question to the system using voice or gestures, the server receives the information and analyzes the question using natural language processing technology. The result is generated as an appropriate response and presented to the user via the terminal along with relevant information. In this process, the server performs data processing in real time to help the user resolve their questions. Furthermore, the data generated during learning is accumulated by the server, and the user's learning progress is evaluated based on this data.
[0462] As a concrete example, consider a case where a user becomes interested in "plant photosynthesis" and selects this topic through the system. In this case, the server prepares 3D models of plant cells and visualized materials of the photosynthetic process and sends them to the terminal. The user can observe these in a VR environment and learn about the mechanism of photosynthesis through experience. If the user asks aloud, "What are the final products of photosynthesis?", the server processes the question immediately and provides an answer such as, "Glucose and oxygen."
[0463] In this way, the invention provides an environment in which users can proactively advance their learning, and further contributes to learners' future planning by providing career advice based on information collected by the server. This invention aims to improve the quality of education by providing an individually optimized learning environment.
[0464] The following describes the processing flow.
[0465] Step 1:
[0466] The user accesses the system and logs in. Once logged in, the user proceeds to a screen where they can select learning topics of interest.
[0467] Step 2:
[0468] The device retrieves the user's selected learning theme information and sends it to the server. The theme includes specific fields or topics that the user wants to learn about.
[0469] Step 3:
[0470] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected from the database. The user's interests and existing knowledge level are also considered during the selection process.
[0471] Step 4:
[0472] The server sends selected learning content to the device. The sent content includes 3D models and text materials related to the theme.
[0473] Step 5:
[0474] The device uses the received content to prepare the virtual reality environment, enabling the user to begin the experience using a VR headset. The user becomes immersed in the virtual reality environment and learns interactively.
[0475] Step 6:
[0476] Users can ask the system questions using voice and gestures during the learning process. For example, questions like, "What is the theory behind this phenomenon?"
[0477] Step 7:
[0478] The server receives voice input from the user and analyzes the input using natural language processing techniques. It then generates an appropriate response and prepares relevant 3D objects and information.
[0479] Step 8:
[0480] The server sends the generated responses and information to the terminal and presents them to the user. This allows the user to receive answers to their questions in real time.
[0481] Step 9:
[0482] The server records the user's learning behavior. This data includes information such as what content was spent on and what questions were asked.
[0483] Step 10:
[0484] The server analyzes the collected data to evaluate the user's understanding and progress. Based on these results, it generates and provides individually optimized feedback to the user.
[0485] (Example 1)
[0486] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] Traditional education systems have struggled to provide an optimal learning experience tailored to each user's interests and learning style. Furthermore, they lacked immediate answers to user questions during learning and suggestions for future directions based on learning progress. This has resulted in a challenge in providing effective learning support.
[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0489] In this invention, the server includes means for inputting user information and objectives, means for selecting a three-dimensional model and additional materials according to the learning target, and means for analyzing the user's characteristics and suggesting future choices. This makes it possible to provide individually optimized real-time learning support and an environment that promotes the user's proactive learning.
[0490] A "user" refers to a person who uses a system to gain a learning experience.
[0491] "Information" refers to data that users input to record their interests and learning objectives.
[0492] A "three-dimensional model" refers to a digital model generated and provided by a system to visualize objects or processes in three dimensions.
[0493] "Additional materials" refer to documents and diagrams that provide supplementary information in digital format related to the learning theme selected by the user.
[0494] "Voice instructions" refer to instructions or questions that users input into the system using their voice.
[0495] "Action" refers to a method of inputting information into a system using physical movements such as gestures.
[0496] "Related information" refers to information related to the learning content that the system generates in response to user instructions or questions.
[0497] "Immediate" refers to the timing of responses that provide information without delay in response to user actions or inquiries.
[0498] "Learning activity" refers to a series of actions taken by a user to acquire knowledge and skills using a system.
[0499] "Progress" refers to the state indicating how far the user's learning has progressed.
[0500] A "generative artificial intelligence model" refers to artificial intelligence technology used to analyze a user's learning patterns and generate a learning experience tailored to their individual needs.
[0501] "Characteristics" refer to elements that shape an individual's learning style, such as a user's interests, learning ability, and habits.
[0502] "Choice" refers to future activities or directions that the system suggests to the user based on its analysis results.
[0503] This invention provides an educational platform that generates a virtual reality environment based on the user's interests and learning needs, and assists in knowledge acquisition. This system enables a individually optimized learning experience through user interaction.
[0504] The user first inputs their interests and learning objectives through a terminal. The terminal sends this information to a server, which generates appropriate 3D models and supplementary materials based on the input. This generation process utilizes a generative AI model, efficiently generating data using prompts.
[0505] As a concrete example, consider a case where a user expresses interest in "plant photosynthesis." Based on this information, the server generates a three-dimensional model of plant cells and materials explaining the photosynthetic process, and sends this data to the user's device. The user then uses their smartphone to wear a VR headset and experience an immersive learning environment. Within this virtual reality environment, the user can understand the mechanism of photosynthesis through visual information.
[0506] Furthermore, if a user asks a question via voice, such as "What are the final products of photosynthesis?", the server uses natural language processing technology to analyze the question. Based on the analysis, the server generates an appropriate answer, such as "glucose and oxygen," and immediately presents it to the user via the terminal. This process allows users to resolve their questions in real time and deepen their learning.
[0507] User learning activities are recorded on the server and used for progress evaluation. The accumulated data is analyzed using a generative AI model, and feedback is provided to further improve individual learning experiences. An example of a prompt might be, "Prepare a 3D model and supplementary materials needed if the user says they want to learn about plant photosynthesis."
[0508] In this way, the system provides users with an optimized learning environment, contributing to an improvement in the quality of learning.
[0509] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0510] Step 1:
[0511] The user inputs their interests and learning objectives into the device. For example, by inputting "photosynthesis in plants," the learning theme is transmitted to the server. This information forms the basis for generating learning content based on the user's interests. The device sends this input to the server as digital data.
[0512] Step 2:
[0513] The server generates content using a generative AI model based on the received user information. In this process, three-dimensional models and supplementary materials related to the input theme are selected and generated based on prompt text. This ensures that the most suitable visual materials and explanations are prepared for the learning target. The generated content is processed as digital data on the server and sent to the terminal.
[0514] Step 3:
[0515] The terminal receives data sent from the server and sets up the VR environment. It prepares the application on the terminal to display the 3D model correctly. The user puts on a VR headset and begins the learning experience in virtual reality.
[0516] Step 4:
[0517] Users can ask questions and give instructions using voice and gestures within the VR environment. For example, they can ask a question by voice, such as, "What is the final product of photosynthesis?" The device converts the voice information into text data and sends it to the server.
[0518] Step 5:
[0519] The server analyzes the received input data using natural language processing techniques. Based on the user's question, the server generates relevant and accurate response data. This response is intended to aid the user's understanding and is immediately sent to the terminal.
[0520] Step 6:
[0521] The terminal receives response data from the server and presents it to the user visually or audibly. The user then deepens their understanding of the material based on the presented information. This enables interactive and individually optimized learning.
[0522] Step 7:
[0523] The server records the user's actions and questions during learning and stores them in a database. This stored data is used to analyze the user's learning progress and provide feedback. Through this data, users can receive continuous learning improvement and progress evaluation.
[0524] (Application Example 1)
[0525] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0526] Modern education demands providing learning environments tailored to the individual needs and interests of each learner. However, general educational content tends to be delivered in a uniform format, making it difficult to customize according to the learner's level of understanding and interests. Furthermore, limited opportunities for learning through real-world experiences mean that knowledge is absorbed only passively and in a fixed manner. It is necessary to improve the quality of education by utilizing virtual environments to provide optimal learning experiences for each individual learner.
[0527] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0528] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting information models and supplementary data based on the learning theme, and means for transmitting the selected content to the user information processing device and preparing a virtual environment. This enables learning in an individually optimized virtual environment.
[0529] A "user" is a learner who uses the system to access educational content.
[0530] "Interests and existing knowledge" refers to a user's interest in a particular field and the information they have already acquired.
[0531] An "information model" refers to three-dimensional data or virtual objects created based on a learning theme, which aids in visual understanding.
[0532] "Supplemental data" refers to information such as text, images, and videos that complement the information model and deepen the understanding of the learning material.
[0533] A "user information processing device" is a general term for information technology terminals used by users that enable access to a virtual environment.
[0534] A "virtual environment" is a simulation space constructed using digital technology, where users can learn through immersive experiences.
[0535] "Situation analysis" refers to interpreting voice commands and gestures made by the user and processing information accordingly.
[0536] "Relevant information" refers to learning-enhancing answers and additional data provided based on the user's questions and actions.
[0537] "AI technology" refers to technologies that use algorithms and programs related to artificial intelligence to imitate human intelligent behavior.
[0538] "Analyzing voice questions and generating information" refers to the process of receiving a user's voice inquiry, analyzing it, and creating an appropriate response.
[0539] In this invention, users access an interface using a smart device and select a theme based on their interests and learning needs. The server searches for relevant information models and supplementary data based on the data transmitted by the user and transmits the selected content to the user's information processing device. The virtual environment is built using Unity software, and users can obtain an immersive learning experience through it via a head-mounted display or similar device.
[0540] When a user asks a question to the system using voice commands or gestures, the server uses the Google Cloud Natural Language API to analyze the speech. Based on the analysis, AI technology generates relevant information and answers, which are immediately presented to the user. This process is designed to allow users to resolve their questions in real time.
[0541] For example, if the educational theme is "the structure of historical buildings," the server retrieves and transmits an information model of ancient Roman architecture from the system. The user explores this model in a VR environment and asks a voice question such as, "What are the features of this building?", to which the AI responds with appropriate information based on the situation. An example of a prompt is as follows:
[0542] Examples of prompts for a generative AI model:
[0543] "Please describe the structure and features of this ancient building in detail."
[0544] In this way, the present invention provides an individually optimized learning space, enabling users to learn proactively and effectively.
[0545] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0546] Step 1:
[0547] Users access the interface via their smart devices and select learning themes that match their interests and learning needs. The input consists of data about the user's interests and themes, which is sent from the device to the server.
[0548] Step 2:
[0549] The server searches for relevant information models and supplementary data based on the user's theme information received. This process involves matching the database and determining the selected content. The output is the information model and supplementary data that are sent.
[0550] Step 3:
[0551] The selected content is sent from the server to the user's information processing device. The terminal processes the received data and generates a 3D view to prepare the virtual reality environment. The input is an information model and supplementary data, and the output is the virtual reality environment.
[0552] Step 4:
[0553] The user immerses themselves in a virtual environment using a head-mounted display and begins learning through an information model. Voice commands and gesture inputs from the user are acquired by the system and sent to the server.
[0554] Step 5:
[0555] The server uses the Google Cloud Natural Language API to analyze speech input. Here, the speech input is converted into text data, and natural language processing operations are performed. The output is the analysis result corresponding to the user's question.
[0556] Step 6:
[0557] The server uses a generative AI model to generate relevant information and answers based on the analysis results. At this stage, the optimal answer to the question is generated and presented as information. The output is response data for the user.
[0558] Step 7:
[0559] The generated information and responses are immediately transmitted to the user's terminal and presented through the user's head-mounted display. Here, the terminal displays the received data and provides information to the user. The output is the response information that the user visually perceives.
[0560] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0561] This invention provides a system that incorporates an emotion engine to recognize user emotions and dynamically adjust learning content in order to enhance the user's learning experience. This system provides emotion-responsive feedback and adjusts the learning environment to promote more effective learning.
[0562] Specifically, the user first logs into the system and selects a learning topic of interest. The device sends this information to the server, which selects 3D models and supplementary materials corresponding to the topic. The selected content is sent to the device, and the user begins learning through the virtual reality environment.
[0563] The key here is that the emotion engine recognizes emotions from the user's voice and gestures. For example, if a user shows a confused expression during learning, the emotion engine recognizes this as "confusion." Based on this information, the server helps the user understand by providing a simple explanation of the learning content or additional supporting materials. Also, when the emotion engine recognizes a positive emotion from the user, it presents more advanced content and supports the user in exploring their interests more deeply.
[0564] As a concrete example, suppose a user chooses the topic of "the expansion of the universe." If the user appears confused while observing a model of the universe in the VR environment, the emotion engine will detect this, and the server will provide an additional animation visualizing the rate of expansion, along with a clear explanation. Through this process, the user can deepen their understanding and maintain their interest.
[0565] Furthermore, the server records changes in emotions and incorporates them into the evaluation of learning progress. This data is used for user feedback and optimizing future learning plans. By integrating with the emotion engine, the system provides a personalized learning experience for each learner, contributing to improved quality of education. In this way, it creates a flexible and dynamic learning environment that is attentive to the user's feelings.
[0566] The following describes the processing flow.
[0567] Step 1:
[0568] The user logs into the system and selects a learning topic of interest. The device sends the user's selection information to the server.
[0569] Step 2:
[0570] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected. This includes extracting information from the database and customizing it according to the user's interests and knowledge level.
[0571] Step 3:
[0572] The server sends the selected content to the terminal. The terminal uses the received content to prepare the virtual reality environment and configure it to be accessible to the user.
[0573] Step 4:
[0574] The user begins learning in a prepared virtual reality environment using a VR headset. The user explores learning topics by manipulating and observing 3D models.
[0575] Step 5:
[0576] The device detects the user's voice and gestures, and the emotion engine analyzes them. Positive or negative emotions are then detected.
[0577] Step 6:
[0578] The server dynamically adjusts learning content based on the user's emotions, as identified by the emotion engine. For example, if the user is confused, the server generates additional explanations or more concise materials and sends them to the device.
[0579] Step 7:
[0580] The device presents the user with customized content. This allows the user to receive a learning experience optimized for their own emotional state.
[0581] Step 8:
[0582] The server records the user's learning behavior, including changes in their emotions. The collected data is used to analyze the user's progress and generate feedback, and is also reflected in future learning plans.
[0583] (Example 2)
[0584] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0585] A challenge in modern education systems is the insufficient customization of learning experiences to suit individual users' emotions and learning progress. Many systems are limited to static content delivery, making it difficult to dynamically adjust content based on users' understanding and interests. As a result, learning effectiveness may decline and motivation may be lost.
[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0587] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting three-dimensional representations and supplementary information based on the learning theme, and means for analyzing the user's feelings and dynamically adjusting the learning content. This makes it possible to provide a flexible and effective learning experience tailored to each individual user.
[0588] A "user" refers to a person who uses the system to learn.
[0589] "Interest" refers to learning themes or topics that users are interested in.
[0590] "Existing knowledge" refers to information and concepts that the user has already acquired or understands.
[0591] "Means" refers to the methods or processes that a system uses to perform a particular function.
[0592] "Three-dimensional representation" refers to visual materials such as 3D models used to visually illustrate learning themes.
[0593] "Supplementary information" refers to additional materials and data that complement the three-dimensional representation and deepen the user's understanding.
[0594] "User equipment" refers to the hardware that a user uses to receive learning content and experience a virtual environment.
[0595] A "virtual environment" refers to a simulated space created through digital technology that allows users to have an immersive experience.
[0596] "Situation analysis" refers to the process of interpreting the user's voice commands and actions to generate information for providing appropriate learning content.
[0597] "Generated information" refers to new data and content created through situational analysis and presented to the user.
[0598] "User sentiment" refers to the emotional state interpreted through the user's voice and facial expressions.
[0599] "Dynamic adjustment" refers to changing learning content in real time to the optimal format based on the user's mood and progress.
[0600] "Recording emotional changes" refers to accumulating data on the evolution of a user's emotional state and using it for subsequent analysis.
[0601] This invention provides a learning system for personalizing and efficiently and dynamically adjusting the user's learning experience. The system aims to detect the user's emotional state in real time and optimize learning content based on those emotions.
[0602] Users access and log in to the system's dedicated application using their own devices. These devices are connected to VR goggles or other VR-compatible devices, allowing users to experience a virtual environment. This environment provides advanced three-dimensional representations and supplementary information, enabling users to explore learning themes based on their interests.
[0603] The server selects an appropriate three-dimensional representation from its database based on the learning theme chosen by the user and sends it to the user's terminal. The server also drives an emotion engine to detect the user's emotions and analyzes the user's voice and gestures using technologies such as voice analysis and image recognition.
[0604] This system dynamically changes the learning content based on the user's emotions during the learning process. For example, if the system determines that the user is confused, the server generates simplified supplementary materials and sends them to the user's device. In this way, it supports the user's understanding and enhances the learning effect.
[0605] As a concrete example, consider using this system when a user is learning about "the expansion of the universe." When the user becomes confused while observing the universe model in a VR environment, the emotion engine can detect this emotion, and the server can provide an additional "animation that visualizes the rate of expansion."
[0606] By utilizing generative AI models, users can further customize their learning experience. For example, prompts could include instructions such as, "Please provide specific instructions on how to respond if the user shows a confused expression while learning about the expansion of the universe in a VR environment."
[0607] These features enable the present invention to provide a dynamic learning experience optimized for individual users, thereby improving the quality of education.
[0608] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0609] Step 1:
[0610] The user logs into the system and selects a learning topic. The user enters their topic of interest via a terminal and presses a select button. Based on this input, the user's interest information is sent from the terminal to the server. The output is information about the selected topic.
[0611] Step 2:
[0612] The device sends information about the selected learning theme to the server. Based on the received information, the server selects relevant three-dimensional representations and supplementary information from its database. This data processing prepares the most suitable learning content for the user. The output is the selected learning content.
[0613] Step 3:
[0614] The server sends the selected learning content to the user's terminal. The terminal receives this content and prepares a virtual environment. The virtual environment is experienced by the user using VR goggles and includes three-dimensional visuals and sound effects related to the learning theme. The output is the learning content represented within the virtual environment.
[0615] Step 4:
[0616] The user begins learning in a virtual environment. The user's voice commands and gestures are collected by the device and sent to the server as input data. During this process, the user's interactions are recorded in real time. The output is behavioral data for obtaining user feedback.
[0617] Step 5:
[0618] The server analyzes the received voice commands and gesture data. The emotion engine uses this data to recognize the user's feelings and intentions. Based on the analysis of the emotional state, dynamic adjustments are made to the learning content. The output is the user's emotional state and the adjusted content.
[0619] Step 6:
[0620] The adjusted content is resent from the server to the device. This allows the device to provide immediate feedback to the user. This process personalizes the user's learning experience. The output is an improved learning experience.
[0621] Step 7:
[0622] The server records changes in the user's emotions and learning progress, and uses this data to optimize future learning plans. This data may be input as prompts to the generating AI model, which then makes suggestions for optimizing learning. The output consists of the accumulated progress data and its analysis results.
[0623] (Application Example 2)
[0624] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0625] Traditional training systems and learning environments have struggled to adapt to the individual emotions and levels of understanding of each user, making it difficult to achieve efficient learning, especially in situations where personalized support is required. This can lead to learners becoming confused or losing motivation, making it difficult to achieve optimal learning outcomes.
[0626] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0627] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting visual models and supplementary materials based on the learning theme, and means for detecting the user's emotions and dynamically adjusting the content based on those emotions. This makes it possible to provide a dynamic learning environment that responds to the emotional state of each individual user and realize an effective learning process tailored to individual needs.
[0628] "Means for inputting user interests and existing knowledge" refers to methods for identifying the user's current areas of interest and existing knowledge, and registering that information in the system.
[0629] "Means for selecting visual models and supplementary materials based on learning themes" refers to methods for identifying and providing the most suitable visual representations and related materials for the learning theme selected by the user.
[0630] "Means for transmitting selected information to a terminal device and preparing a virtual reality environment" refers to a method of transmitting selected visual data and materials to the user's terminal and setting up a virtual environment.
[0631] "Means for detecting user emotions and dynamically adjusting content based on those emotions" refers to a method of understanding the user's emotional state in real time and appropriately modifying the learning content based on that information.
[0632] "Means for presenting analysis results and generated information to users in real time" refers to methods for immediately providing users with the analyzed results and generated information.
[0633] "Means for recording user activity and analyzing progress" refers to methods for tracking a user's learning progress and analyzing how far they have progressed.
[0634] "Means of providing feedback based on progress" refers to methods for providing appropriate feedback according to the user's learning progress.
[0635] "A means of analyzing individual trends and proposing future strategies" refers to a method of understanding a user's learning tendencies and proposing future learning strategies based on those tendencies.
[0636] The system that realizes this invention consists of a user, a terminal device, and a server. The user first logs into the terminal device and selects a learning topic of interest. This information is sent from the terminal to the server. The server selects a suitable visual model and supplementary materials and sends them to the terminal device, preparing to provide the user with a virtual reality environment.
[0637] The terminal device is equipped with sensors to detect the user's facial expressions and voice, which detect the user's emotions in real time. The hardware used includes a camera and microphone, and the software employs emotion recognition algorithms (e.g., DeepFace).
[0638] When a user expresses confusion or lack of understanding during the learning process, that emotional data is transferred to the server. Based on this emotional data, the server dynamically adjusts the learning content. For example, if a user does not understand a task, an additional explanatory video may be displayed on their device.
[0639] Furthermore, the server provides feedback based on the analyzed user's learning progress. This allows for the analysis of the user's individual learning tendencies and the suggestion of a more effective learning plan. This process enables users with questions about specific learning topics to receive immediate support.
[0640] By using a generative AI model, it is possible to generate diverse learning content in response to prompt input. An example of a prompt is, "Please run a program that generates specific supplementary information based on the user's learning content."
[0641] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0642] Step 1:
[0643] The user logs into the terminal device. The user enters a learning topic of interest, and this data is sent from the terminal to the server. The input is the learning topic, and the output is the result of sending the topic information to the server. The terminal packages the topic data in an appropriate format and sends it to the server.
[0644] Step 2:
[0645] The server selects appropriate visual models and supplementary materials based on the received learning theme. The input is theme information submitted by the user, and the output is the selected content. The server searches the database and identifies the most relevant visual content related to the theme.
[0646] Step 3:
[0647] The server sends selected content to the terminal device, preparing the virtual reality environment. The input is the selected content, and the output is the content delivery result to the terminal. The server packages the selected materials in digital format and sends them to the terminal. The terminal receives this and sets up the virtual reality environment for the user.
[0648] Step 4:
[0649] The terminal device uses sensors to detect the user's emotions in real time. Input is sensor information from cameras and microphones, and output is emotion data. The terminal uses a machine learning model (e.g., DeepFace) to analyze the user's emotions from their face and voice.
[0650] Step 5:
[0651] The server dynamically adjusts the learning content based on detected sentiment data. The input is sentiment data, and the output is the adjusted learning content. The server analyzes the sentiment data, re-selects appropriate learning materials and supplementary resources, and provides a learning experience tailored to the user's current situation.
[0652] Step 6:
[0653] The server processes data to record the user's learning progress and generate feedback. The input is the user's learning behavior data, and the output is feedback information. The server analyzes the progress data and generates and provides appropriate feedback to the user.
[0654] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0655] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0656] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0657] [Fourth Embodiment]
[0658] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0659] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0660] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0661] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0662] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0663] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0664] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0665] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0666] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0667] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0668] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0669] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0670] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] This invention provides an educational platform that generates a virtual reality environment according to the user's interests and learning needs, thereby supporting knowledge acquisition. This system provides a individually optimized learning experience through the user's interactive operation.
[0672] Specifically, users first input their interests and goals to access the system. Once the device receives this information, it sends it to the server, and content preparation based on the selected learning theme begins. The server creates the selected 3D models and supplementary materials, sends them to the device, and sets up the virtual reality environment. The user then uses their smartphone to put on a VR headset and experience immersive learning.
[0673] When a user asks a question to the system using voice or gestures, the server receives the information and analyzes the question using natural language processing technology. The result is generated as an appropriate response and presented to the user via the terminal along with relevant information. In this process, the server performs data processing in real time to help the user resolve their questions. Furthermore, the data generated during learning is accumulated by the server, and the user's learning progress is evaluated based on this data.
[0674] As a concrete example, consider a case where a user becomes interested in "plant photosynthesis" and selects this topic through the system. In this case, the server prepares 3D models of plant cells and visualized materials of the photosynthetic process and sends them to the terminal. The user can observe these in a VR environment and learn about the mechanism of photosynthesis through experience. If the user asks aloud, "What are the final products of photosynthesis?", the server processes the question immediately and provides an answer such as, "Glucose and oxygen."
[0675] In this way, the invention provides an environment in which users can proactively advance their learning, and further contributes to learners' future planning by providing career advice based on information collected by the server. This invention aims to improve the quality of education by providing an individually optimized learning environment.
[0676] The following describes the processing flow.
[0677] Step 1:
[0678] The user accesses the system and logs in. Once logged in, the user proceeds to a screen where they can select learning topics of interest.
[0679] Step 2:
[0680] The device retrieves the user's selected learning theme information and sends it to the server. The theme includes specific fields or topics that the user wants to learn about.
[0681] Step 3:
[0682] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected from the database. The user's interests and existing knowledge level are also considered during the selection process.
[0683] Step 4:
[0684] The server sends selected learning content to the device. The sent content includes 3D models and text materials related to the theme.
[0685] Step 5:
[0686] The device uses the received content to prepare the virtual reality environment, enabling the user to begin the experience using a VR headset. The user becomes immersed in the virtual reality environment and learns interactively.
[0687] Step 6:
[0688] Users can ask the system questions using voice and gestures during the learning process. For example, questions like, "What is the theory behind this phenomenon?"
[0689] Step 7:
[0690] The server receives voice input from the user and analyzes the input using natural language processing techniques. It then generates an appropriate response and prepares relevant 3D objects and information.
[0691] Step 8:
[0692] The server sends the generated responses and information to the terminal and presents them to the user. This allows the user to receive answers to their questions in real time.
[0693] Step 9:
[0694] The server records the user's learning behavior. This data includes information such as what content was spent on and what questions were asked.
[0695] Step 10:
[0696] The server analyzes the collected data to evaluate the user's understanding and progress. Based on these results, it generates and provides individually optimized feedback to the user.
[0697] (Example 1)
[0698] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0699] Traditional education systems have struggled to provide an optimal learning experience tailored to each user's interests and learning style. Furthermore, they lacked immediate answers to user questions during learning and suggestions for future directions based on learning progress. This has resulted in a challenge in providing effective learning support.
[0700] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0701] In this invention, the server includes means for inputting user information and objectives, means for selecting a three-dimensional model and additional materials according to the learning target, and means for analyzing the user's characteristics and suggesting future choices. This makes it possible to provide individually optimized real-time learning support and an environment that promotes the user's proactive learning.
[0702] A "user" refers to a person who uses a system to gain a learning experience.
[0703] "Information" refers to data that users input to record their interests and learning objectives.
[0704] A "three-dimensional model" refers to a digital model generated and provided by a system to visualize objects or processes in three dimensions.
[0705] "Additional materials" refer to documents and diagrams that provide supplementary information in digital format related to the learning theme selected by the user.
[0706] "Voice instructions" refer to instructions or questions that users input into the system using their voice.
[0707] "Action" refers to a method of inputting information into a system using physical movements such as gestures.
[0708] "Related information" refers to information related to the learning content that the system generates in response to user instructions or questions.
[0709] "Immediate" refers to the timing of responses that provide information without delay in response to user actions or inquiries.
[0710] "Learning activity" refers to a series of actions taken by a user to acquire knowledge and skills using a system.
[0711] "Progress" refers to the state indicating how far the user's learning has progressed.
[0712] A "generative artificial intelligence model" refers to artificial intelligence technology used to analyze a user's learning patterns and generate a learning experience tailored to their individual needs.
[0713] "Characteristics" refer to elements that shape an individual's learning style, such as a user's interests, learning ability, and habits.
[0714] "Choice" refers to future activities or directions that the system suggests to the user based on its analysis results.
[0715] This invention provides an educational platform that generates a virtual reality environment based on the user's interests and learning needs, and assists in knowledge acquisition. This system enables a individually optimized learning experience through user interaction.
[0716] The user first inputs their interests and learning objectives through a terminal. The terminal sends this information to a server, which generates appropriate 3D models and supplementary materials based on the input. This generation process utilizes a generative AI model, efficiently generating data using prompts.
[0717] As a concrete example, consider a case where a user expresses interest in "plant photosynthesis." Based on this information, the server generates a three-dimensional model of plant cells and materials explaining the photosynthetic process, and sends this data to the user's device. The user then uses their smartphone to wear a VR headset and experience an immersive learning environment. Within this virtual reality environment, the user can understand the mechanism of photosynthesis through visual information.
[0718] Furthermore, if a user asks a question via voice, such as "What are the final products of photosynthesis?", the server uses natural language processing technology to analyze the question. Based on the analysis, the server generates an appropriate answer, such as "glucose and oxygen," and immediately presents it to the user via the terminal. This process allows users to resolve their questions in real time and deepen their learning.
[0719] User learning activities are recorded on the server and used for progress evaluation. The accumulated data is analyzed using a generative AI model, and feedback is provided to further improve individual learning experiences. An example of a prompt might be, "Prepare a 3D model and supplementary materials needed if the user says they want to learn about plant photosynthesis."
[0720] In this way, the system provides users with an optimized learning environment, contributing to an improvement in the quality of learning.
[0721] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0722] Step 1:
[0723] The user inputs their interests and learning objectives into the device. For example, by inputting "photosynthesis in plants," the learning theme is transmitted to the server. This information forms the basis for generating learning content based on the user's interests. The device sends this input to the server as digital data.
[0724] Step 2:
[0725] The server generates content using a generative AI model based on the received user information. In this process, three-dimensional models and supplementary materials related to the input theme are selected and generated based on prompt text. This ensures that the most suitable visual materials and explanations are prepared for the learning target. The generated content is processed as digital data on the server and sent to the terminal.
[0726] Step 3:
[0727] The terminal receives data sent from the server and sets up the VR environment. It prepares the application on the terminal to display the 3D model correctly. The user puts on a VR headset and begins the learning experience in virtual reality.
[0728] Step 4:
[0729] Users can ask questions and give instructions using voice and gestures within the VR environment. For example, they can ask a question by voice, such as, "What is the final product of photosynthesis?" The device converts the voice information into text data and sends it to the server.
[0730] Step 5:
[0731] The server analyzes the received input data using natural language processing techniques. Based on the user's question, the server generates relevant and accurate response data. This response is intended to aid the user's understanding and is immediately sent to the terminal.
[0732] Step 6:
[0733] The terminal receives response data from the server and presents it to the user visually or audibly. The user then deepens their understanding of the material based on the presented information. This enables interactive and individually optimized learning.
[0734] Step 7:
[0735] The server records the user's actions and questions during learning and stores them in a database. This stored data is used to analyze the user's learning progress and provide feedback. Through this data, users can receive continuous learning improvement and progress evaluation.
[0736] (Application Example 1)
[0737] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0738] Modern education demands providing learning environments tailored to the individual needs and interests of each learner. However, general educational content tends to be delivered in a uniform format, making it difficult to customize according to the learner's level of understanding and interests. Furthermore, limited opportunities for learning through real-world experiences mean that knowledge is absorbed only passively and in a fixed manner. It is necessary to improve the quality of education by utilizing virtual environments to provide optimal learning experiences for each individual learner.
[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0740] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting information models and supplementary data based on the learning theme, and means for transmitting the selected content to the user information processing device and preparing a virtual environment. This enables learning in an individually optimized virtual environment.
[0741] A "user" is a learner who uses the system to access educational content.
[0742] "Interests and existing knowledge" refers to a user's interest in a particular field and the information they have already acquired.
[0743] An "information model" refers to three-dimensional data or virtual objects created based on a learning theme, which aids in visual understanding.
[0744] "Supplemental data" refers to information such as text, images, and videos that complement the information model and deepen the understanding of the learning material.
[0745] A "user information processing device" is a general term for information technology terminals used by users that enable access to a virtual environment.
[0746] A "virtual environment" is a simulation space constructed using digital technology, where users can learn through immersive experiences.
[0747] "Situation analysis" refers to interpreting voice commands and gestures made by the user and processing information accordingly.
[0748] "Relevant information" refers to learning-enhancing answers and additional data provided based on the user's questions and actions.
[0749] "AI technology" refers to technologies that use algorithms and programs related to artificial intelligence to imitate human intelligent behavior.
[0750] "Analyzing voice questions and generating information" refers to the process of receiving a user's voice inquiry, analyzing it, and creating an appropriate response.
[0751] In this invention, users access an interface using a smart device and select a theme based on their interests and learning needs. The server searches for relevant information models and supplementary data based on the data transmitted by the user and transmits the selected content to the user's information processing device. The virtual environment is built using Unity software, and users can obtain an immersive learning experience through it via a head-mounted display or similar device.
[0752] When a user asks a question to the system using voice commands or gestures, the server uses the Google Cloud Natural Language API to analyze the speech. Based on the analysis, AI technology generates relevant information and answers, which are immediately presented to the user. This process is designed to allow users to resolve their questions in real time.
[0753] For example, if the educational theme is "the structure of historical buildings," the server retrieves and transmits an information model of ancient Roman architecture from the system. The user explores this model in a VR environment and asks a voice question such as, "What are the features of this building?", to which the AI responds with appropriate information based on the situation. An example of a prompt is as follows:
[0754] Examples of prompts for a generative AI model:
[0755] "Please describe the structure and features of this ancient building in detail."
[0756] In this way, the present invention provides an individually optimized learning space, enabling users to learn proactively and effectively.
[0757] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0758] Step 1:
[0759] Users access the interface via their smart devices and select learning themes that match their interests and learning needs. The input consists of data about the user's interests and themes, which is sent from the device to the server.
[0760] Step 2:
[0761] The server searches for relevant information models and supplementary data based on the user's theme information received. This process involves matching the database and determining the selected content. The output is the information model and supplementary data that are sent.
[0762] Step 3:
[0763] The selected content is sent from the server to the user's information processing device. The terminal processes the received data and generates a 3D view to prepare the virtual reality environment. The input is an information model and supplementary data, and the output is the virtual reality environment.
[0764] Step 4:
[0765] The user immerses themselves in a virtual environment using a head-mounted display and begins learning through an information model. Voice commands and gesture inputs from the user are acquired by the system and sent to the server.
[0766] Step 5:
[0767] The server uses the Google Cloud Natural Language API to analyze speech input. Here, the speech input is converted into text data, and natural language processing operations are performed. The output is the analysis result corresponding to the user's question.
[0768] Step 6:
[0769] The server uses a generative AI model to generate relevant information and answers based on the analysis results. At this stage, the optimal answer to the question is generated and presented as information. The output is response data for the user.
[0770] Step 7:
[0771] The generated information and responses are immediately transmitted to the user's terminal and presented through the user's head-mounted display. Here, the terminal displays the received data and provides information to the user. The output is the response information that the user visually perceives.
[0772] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0773] This invention provides a system that incorporates an emotion engine to recognize user emotions and dynamically adjust learning content in order to enhance the user's learning experience. This system provides emotion-responsive feedback and adjusts the learning environment to promote more effective learning.
[0774] Specifically, the user first logs into the system and selects a learning topic of interest. The device sends this information to the server, which selects 3D models and supplementary materials corresponding to the topic. The selected content is sent to the device, and the user begins learning through the virtual reality environment.
[0775] The key here is that the emotion engine recognizes emotions from the user's voice and gestures. For example, if a user shows a confused expression during learning, the emotion engine recognizes this as "confusion." Based on this information, the server helps the user understand by providing a simple explanation of the learning content or additional supporting materials. Also, when the emotion engine recognizes a positive emotion from the user, it presents more advanced content and supports the user in exploring their interests more deeply.
[0776] As a concrete example, suppose a user chooses the topic of "the expansion of the universe." If the user appears confused while observing a model of the universe in the VR environment, the emotion engine will detect this, and the server will provide an additional animation visualizing the rate of expansion, along with a clear explanation. Through this process, the user can deepen their understanding and maintain their interest.
[0777] Furthermore, the server records changes in emotions and incorporates them into the evaluation of learning progress. This data is used for user feedback and optimizing future learning plans. By integrating with the emotion engine, the system provides a personalized learning experience for each learner, contributing to improved quality of education. In this way, it creates a flexible and dynamic learning environment that is attentive to the user's feelings.
[0778] The following describes the processing flow.
[0779] Step 1:
[0780] The user logs into the system and selects a learning topic of interest. The device sends the user's selection information to the server.
[0781] Step 2:
[0782] Based on the theme information received by the server, relevant 3D models and supplementary materials are selected. This includes extracting information from the database and customizing it according to the user's interests and knowledge level.
[0783] Step 3:
[0784] The server sends the selected content to the terminal. The terminal uses the received content to prepare the virtual reality environment and configure it to be accessible to the user.
[0785] Step 4:
[0786] The user begins learning in a prepared virtual reality environment using a VR headset. The user explores learning topics by manipulating and observing 3D models.
[0787] Step 5:
[0788] The device detects the user's voice and gestures, and the emotion engine analyzes them. Positive or negative emotions are then detected.
[0789] Step 6:
[0790] The server dynamically adjusts learning content based on the user's emotions, as identified by the emotion engine. For example, if the user is confused, the server generates additional explanations or more concise materials and sends them to the device.
[0791] Step 7:
[0792] The device presents the user with customized content. This allows the user to receive a learning experience optimized for their own emotional state.
[0793] Step 8:
[0794] The server records the user's learning behavior, including changes in their emotions. The collected data is used to analyze the user's progress and generate feedback, and is also reflected in future learning plans.
[0795] (Example 2)
[0796] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0797] A challenge in modern education systems is the insufficient customization of learning experiences to suit individual users' emotions and learning progress. Many systems are limited to static content delivery, making it difficult to dynamically adjust content based on users' understanding and interests. As a result, learning effectiveness may decline and motivation may be lost.
[0798] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0799] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting three-dimensional representations and supplementary information based on the learning theme, and means for analyzing the user's feelings and dynamically adjusting the learning content. This makes it possible to provide a flexible and effective learning experience tailored to each individual user.
[0800] A "user" refers to a person who uses the system to learn.
[0801] "Interest" refers to learning themes or topics that users are interested in.
[0802] "Existing knowledge" refers to information and concepts that the user has already acquired or understands.
[0803] "Means" refers to the methods or processes that a system uses to perform a particular function.
[0804] "Three-dimensional representation" refers to visual materials such as 3D models used to visually illustrate learning themes.
[0805] "Supplementary information" refers to additional materials and data that complement the three-dimensional representation and deepen the user's understanding.
[0806] "User equipment" refers to the hardware that a user uses to receive learning content and experience a virtual environment.
[0807] A "virtual environment" refers to a simulated space created through digital technology that allows users to have an immersive experience.
[0808] "Situation analysis" refers to the process of interpreting the user's voice commands and actions to generate information for providing appropriate learning content.
[0809] "Generated information" refers to new data and content created through situational analysis and presented to the user.
[0810] "User sentiment" refers to the emotional state interpreted through the user's voice and facial expressions.
[0811] "Dynamic adjustment" refers to changing learning content in real time to the optimal format based on the user's mood and progress.
[0812] "Recording emotional changes" refers to accumulating data on the evolution of a user's emotional state and using it for subsequent analysis.
[0813] This invention provides a learning system for personalizing and efficiently and dynamically adjusting the user's learning experience. The system aims to detect the user's emotional state in real time and optimize learning content based on those emotions.
[0814] Users access and log in to the system's dedicated application using their own devices. These devices are connected to VR goggles or other VR-compatible devices, allowing users to experience a virtual environment. This environment provides advanced three-dimensional representations and supplementary information, enabling users to explore learning themes based on their interests.
[0815] The server selects an appropriate three-dimensional representation from its database based on the learning theme chosen by the user and sends it to the user's terminal. The server also drives an emotion engine to detect the user's emotions and analyzes the user's voice and gestures using technologies such as voice analysis and image recognition.
[0816] This system dynamically changes the learning content based on the user's emotions during the learning process. For example, if the system determines that the user is confused, the server generates simplified supplementary materials and sends them to the user's device. In this way, it supports the user's understanding and enhances the learning effect.
[0817] As a concrete example, consider using this system when a user is learning about "the expansion of the universe." When the user becomes confused while observing the universe model in a VR environment, the emotion engine can detect this emotion, and the server can provide an additional "animation that visualizes the rate of expansion."
[0818] By utilizing generative AI models, users can further customize their learning experience. For example, prompts could include instructions such as, "Please provide specific instructions on how to respond if the user shows a confused expression while learning about the expansion of the universe in a VR environment."
[0819] These features enable the present invention to provide a dynamic learning experience optimized for individual users, thereby improving the quality of education.
[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0821] Step 1:
[0822] The user logs into the system and selects a learning topic. The user enters their topic of interest via a terminal and presses a select button. Based on this input, the user's interest information is sent from the terminal to the server. The output is information about the selected topic.
[0823] Step 2:
[0824] The device sends information about the selected learning theme to the server. Based on the received information, the server selects relevant three-dimensional representations and supplementary information from its database. This data processing prepares the most suitable learning content for the user. The output is the selected learning content.
[0825] Step 3:
[0826] The server sends the selected learning content to the user's terminal. The terminal receives this content and prepares a virtual environment. The virtual environment is experienced by the user using VR goggles and includes three-dimensional visuals and sound effects related to the learning theme. The output is the learning content represented within the virtual environment.
[0827] Step 4:
[0828] The user begins learning in a virtual environment. The user's voice commands and gestures are collected by the device and sent to the server as input data. During this process, the user's interactions are recorded in real time. The output is behavioral data for obtaining user feedback.
[0829] Step 5:
[0830] The server analyzes the received voice commands and gesture data. The emotion engine uses this data to recognize the user's feelings and intentions. Based on the analysis of the emotional state, dynamic adjustments are made to the learning content. The output is the user's emotional state and the adjusted content.
[0831] Step 6:
[0832] The adjusted content is resent from the server to the device. This allows the device to provide immediate feedback to the user. This process personalizes the user's learning experience. The output is an improved learning experience.
[0833] Step 7:
[0834] The server records changes in the user's emotions and learning progress, and uses this data to optimize future learning plans. This data may be input as prompts to the generating AI model, which then makes suggestions for optimizing learning. The output consists of the accumulated progress data and its analysis results.
[0835] (Application Example 2)
[0836] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0837] Traditional training systems and learning environments have struggled to adapt to the individual emotions and levels of understanding of each user, making it difficult to achieve efficient learning, especially in situations where personalized support is required. This can lead to learners becoming confused or losing motivation, making it difficult to achieve optimal learning outcomes.
[0838] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0839] In this invention, the server includes means for inputting the user's interests and existing knowledge, means for selecting visual models and supplementary materials based on the learning theme, and means for detecting the user's emotions and dynamically adjusting the content based on those emotions. This makes it possible to provide a dynamic learning environment that responds to the emotional state of each individual user and realize an effective learning process tailored to individual needs.
[0840] "Means for inputting user interests and existing knowledge" refers to methods for identifying the user's current areas of interest and existing knowledge, and registering that information in the system.
[0841] "Means for selecting visual models and supplementary materials based on learning themes" refers to methods for identifying and providing the most suitable visual representations and related materials for the learning theme selected by the user.
[0842] "Means for transmitting selected information to a terminal device and preparing a virtual reality environment" refers to a method of transmitting selected visual data and materials to the user's terminal and setting up a virtual environment.
[0843] "Means for detecting user emotions and dynamically adjusting content based on those emotions" refers to a method of understanding the user's emotional state in real time and appropriately modifying the learning content based on that information.
[0844] "Means for presenting analysis results and generated information to users in real time" refers to methods for immediately providing users with the analyzed results and generated information.
[0845] "Means for recording user activity and analyzing progress" refers to methods for tracking a user's learning progress and analyzing how far they have progressed.
[0846] "Means of providing feedback based on progress" refers to methods for providing appropriate feedback according to the user's learning progress.
[0847] "A means of analyzing individual trends and proposing future strategies" refers to a method of understanding a user's learning tendencies and proposing future learning strategies based on those tendencies.
[0848] The system that realizes this invention consists of a user, a terminal device, and a server. The user first logs into the terminal device and selects a learning topic of interest. This information is sent from the terminal to the server. The server selects a suitable visual model and supplementary materials and sends them to the terminal device, preparing to provide the user with a virtual reality environment.
[0849] The terminal device is equipped with sensors to detect the user's facial expressions and voice, which detect the user's emotions in real time. The hardware used includes a camera and microphone, and the software employs emotion recognition algorithms (e.g., DeepFace).
[0850] When a user expresses confusion or lack of understanding during the learning process, that emotional data is transferred to the server. Based on this emotional data, the server dynamically adjusts the learning content. For example, if a user does not understand a task, an additional explanatory video may be displayed on their device.
[0851] Furthermore, the server provides feedback based on the analyzed user's learning progress. This allows for the analysis of the user's individual learning tendencies and the suggestion of a more effective learning plan. This process enables users with questions about specific learning topics to receive immediate support.
[0852] By using a generative AI model, it is possible to generate diverse learning content in response to prompt input. An example of a prompt is, "Please run a program that generates specific supplementary information based on the user's learning content."
[0853] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0854] Step 1:
[0855] The user logs into the terminal device. The user enters a learning topic of interest, and this data is sent from the terminal to the server. The input is the learning topic, and the output is the result of sending the topic information to the server. The terminal packages the topic data in an appropriate format and sends it to the server.
[0856] Step 2:
[0857] The server selects appropriate visual models and supplementary materials based on the received learning theme. The input is theme information submitted by the user, and the output is the selected content. The server searches the database and identifies the most relevant visual content related to the theme.
[0858] Step 3:
[0859] The server sends selected content to the terminal device, preparing the virtual reality environment. The input is the selected content, and the output is the content delivery result to the terminal. The server packages the selected materials in digital format and sends them to the terminal. The terminal receives this and sets up the virtual reality environment for the user.
[0860] Step 4:
[0861] The terminal device uses sensors to detect the user's emotions in real time. Input is sensor information from cameras and microphones, and output is emotion data. The terminal uses a machine learning model (e.g., DeepFace) to analyze the user's emotions from their face and voice.
[0862] Step 5:
[0863] The server dynamically adjusts the learning content based on detected sentiment data. The input is sentiment data, and the output is the adjusted learning content. The server analyzes the sentiment data, re-selects appropriate learning materials and supplementary resources, and provides a learning experience tailored to the user's current situation.
[0864] Step 6:
[0865] The server processes data to record the user's learning progress and generate feedback. The input is the user's learning behavior data, and the output is feedback information. The server analyzes the progress data and generates and provides appropriate feedback to the user.
[0866] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0867] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0868] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0869] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0870] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0871] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0872] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0873] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0874] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0875] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0876] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0877] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0878] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0879] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0880] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0881] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0882] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0883] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0884] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0885] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0886] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0887] The following is further disclosed regarding the embodiments described above.
[0888] (Claim 1)
[0889] A means of inputting the user's interests and existing knowledge,
[0890] A means of selecting 3D models and supplementary materials based on the learning theme,
[0891] A means for sending selected content to the user's terminal and preparing a virtual reality environment,
[0892] A means for analyzing context and generating relevant information based on voice commands and gestures from the user,
[0893] A means of presenting analysis results and generated information to the user in real time,
[0894] A means of recording user learning behavior and analyzing progress,
[0895] A means of providing feedback based on progress,
[0896] A means of analyzing user trends and suggesting future career paths,
[0897] A system that includes this.
[0898] (Claim 2)
[0899] The system according to claim 1, which analyzes voice input from a user using natural language processing technology.
[0900] (Claim 3)
[0901] The system according to claim 1, which transmits the generated 3D model and text information to a user terminal.
[0902] "Example 1"
[0903] (Claim 1)
[0904] Means for inputting user information and purpose,
[0905] A means of selecting three-dimensional models and additional materials according to the learning objectives,
[0906] A means of transmitting the selected content to the user's device and constructing a virtual reality environment,
[0907] A means of analyzing content and generating relevant information based on voice instructions and actions from the user,
[0908] A means of immediately displaying the analysis results and generated information to the user,
[0909] A means of recording the user's learning activities and analyzing their progress,
[0910] A means of providing feedback based on the analyzed progress,
[0911] A means of accumulating training data and improving the user's learning content using a generative artificial intelligence model,
[0912] A means of analyzing user characteristics and suggesting future choices,
[0913] A system that includes this.
[0914] (Claim 2)
[0915] The system according to claim 1, which analyzes voice data from a user using natural language processing technology.
[0916] (Claim 3)
[0917] The system according to claim 1, which transmits the generated three-dimensional model and text information to a user device.
[0918] "Application Example 1"
[0919] (Claim 1)
[0920] A means of inputting the user's interests and existing knowledge,
[0921] A means of selecting information models and supplementary data based on the learning theme,
[0922] A means for transmitting selected content to a user information processing device and preparing a virtual environment,
[0923] A means for analyzing the situation and generating relevant information based on voice commands and gestures from the user,
[0924] A means of immediately presenting analysis results and generated information to the user,
[0925] A means of recording user learning behavior and analyzing progress,
[0926] A means of providing feedback based on progress,
[0927] A means of analyzing user trends and suggesting future career paths,
[0928] A means of exploring, purchasing, and experiencing educational content in a virtual environment,
[0929] A means of analyzing voice questions using AI technology and generating information,
[0930] A system that includes this.
[0931] (Claim 2)
[0932] The system according to claim 1, which analyzes voice input from a user using natural language processing technology and enables immediate response.
[0933] (Claim 3)
[0934] The system according to claim 1, which transmits the generated information model and text information to a user information processing device and provides a virtual environment.
[0935] "Example 2 of combining an emotion engine"
[0936] (Claim 1)
[0937] A means of inputting the user's interests and existing knowledge,
[0938] A means of selecting three-dimensional representations and supplementary information based on the learning theme,
[0939] A means for transmitting selected content to the user's device and preparing the virtual environment,
[0940] A means for analyzing the situation and generating relevant information based on voice commands and actions from the user,
[0941] A means of immediately presenting analysis results and generated information to the user,
[0942] A means of recording user learning behavior and analyzing progress,
[0943] A means of providing feedback based on progress,
[0944] A means of analyzing the user's feelings and dynamically adjusting the learning content,
[0945] A means of recording emotional changes and reflecting them in the evaluation of learning progress,
[0946] A system that includes this.
[0947] (Claim 2)
[0948] The system according to claim 1, which processes voice input from a user using natural language processing technology.
[0949] (Claim 3)
[0950] The system according to claim 1, which transmits the generated three-dimensional representation and textual information to a user device.
[0951] "Application example 2 when combining with an emotional engine"
[0952] (Claim 1)
[0953] A means of inputting the user's interests and existing knowledge,
[0954] A means of selecting visual models and supplementary materials based on the learning theme,
[0955] A means for transmitting selected information to a terminal device and preparing a virtual reality environment,
[0956] A means for detecting user emotions and dynamically adjusting content based on those emotions,
[0957] A means of presenting analysis results and generated information to the user in real time,
[0958] A means of recording user activity and analyzing progress,
[0959] A means of providing feedback based on progress,
[0960] A means of analyzing individual trends and proposing future strategies,
[0961] A system that includes this.
[0962] (Claim 2)
[0963] The system according to claim 1, which analyzes voice input from a user using natural language processing technology to detect emotions.
[0964] (Claim 3)
[0965] The system according to claim 1, which transmits the generated visual model and text information to a terminal device. [Explanation of symbols]
[0966] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of inputting the user's interests and existing knowledge, A means of selecting 3D models and supplementary materials based on the learning theme, A means for sending selected content to the user's terminal and preparing a virtual reality environment, A means for analyzing context and generating relevant information based on voice commands and gestures from the user, A means of presenting analysis results and generated information to the user in real time, A means of recording user learning behavior and analyzing progress, A means of providing feedback based on progress, A means of analyzing user trends and suggesting future career paths, A system that includes this.
2. The system according to claim 1, which analyzes voice input from a user using natural language processing technology.
3. The system according to claim 1, which transmits the generated 3D model and text information to a user terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A