system
The system addresses the inefficiencies of conventional learning by using a generative model to analyze comic data, quantify emotions, and collect feedback, resulting in a more engaging and effective educational experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional educational methods lack emotional involvement and visual elements, making learning inefficient and difficult to remember, and there is a need for a mechanism that benefits both creators and learners.
A system that analyzes comic data using a generative model to quantify emotions, dynamically generates stories tailored to learners, and collects user feedback to improve learning content, providing a visual and emotional learning experience.
Enhances learning efficiency and enjoyment by creating memorable and personalized educational experiences through emotional interaction.
Smart Images

Figure 2026068478000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The problems faced by learners in exam preparation are that due to the abstract and monotonous learning content, there is little emotional involvement, making it difficult to form memories and fix information. In conventional educational methods, it is difficult to provide an effective learning experience that includes visual and emotional elements, so there is a need to improve learning efficiency. Furthermore, there is also a problem in that a mechanism that benefits both creators and learners is lacking.
Means for Solving the Problems
[0005] This invention employs a method for analyzing comic data using a generative model and quantifying the emotions of each panel. Based on the quantified emotional information, a story suitable for the learner is dynamically generated and transmitted to the terminal for visual display. Furthermore, user feedback is collected and this data is used to generate the next learning content, providing an effective and enjoyable learning experience for the learner. The invention also provides a user interface that allows the selection of learning themes, enabling flexible content generation based on the selected theme information. By generating numerical values based on emotional categories from the characters' lines and expressions, the emotional flow of the entire story is constructed, creating a learning environment that is easy to remember.
[0006] A "generative model" refers to an algorithm that generates new information from data using a computer program, and specifically refers to one that utilizes artificial intelligence technology.
[0007] "Comic data as source material" refers to pre-prepared visual material in comic book format, including digital information such as story and character depictions.
[0008] "Quantifying emotions" refers to the process of expressing emotions derived from a character's facial expressions and actions as numerical values, thereby enabling the quantitative handling of emotional information.
[0009] "Dynamically generating stories" refers to the process of creating new narrative developments in real time according to the user's needs and circumstances.
[0010] A "terminal" refers to a computer or mobile device that displays information and enables interaction with the user.
[0011] Collecting "user feedback" refers to the act of gathering data on users' experiences and impressions, which is then used to improve the system.
[0012] A "user interface" refers to the screens and operating methods that allow a user to interact with a system, enabling them to input and select information.
[0013] "Emotional category-based numerical values" refer to a method of classifying different emotions into specific categories and expressing them numerically.
[0014] "The emotional flow of the entire story" refers to the emotional changes and developments from the beginning to the end of the narrative, and aims to manage the emotional impact on learners. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention provides a learning system using generative models, realizing an innovative method to make exam preparation more effective and engaging. This system operates through the coordinated interaction of a server, terminals, and users.
[0037] 1. Server-side processing
[0038] The server first receives theme selection information from the user. Based on this information, the server selects appropriate comic data from its database. Using a generative model, the server analyzes the content of each panel of the selected comic data, quantifying emotions from character expressions and dialogue. This organizes the emotional information of each panel and provides data to construct the emotional flow of the story. Based on this emotional data, the server dynamically generates a story tailored to the user's profile and sends it to the terminal.
[0039] 2. Processing on the terminal side
[0040] The user's device receives story data sent from the server and displays it visually as a comic strip. Through the user interface, the user can select their preferred theme and interact with the displayed learning content. For example, if a user studying history selects this theme, the server generates a story based on historical events and displays it in comic strip format on the device, providing a visual and emotional learning experience.
[0041] 3. User-side initiatives
[0042] Users learn through comics displayed on their devices. During the learning process, users can input comments and feedback on points of interest or areas they understand better. The device sends this feedback to the server, which stores it as data to improve the user experience. This feedback allows the server to provide more suitable content to the user when generating future learning stories.
[0043] Based on the above, the present invention provides a system that connects emotions and knowledge, making test preparation more memorable, enjoyable, and efficient. This system adds new value to education and is expected to significantly improve conventional learning methods.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The user operates the terminal's user interface to select the topic they wish to learn about. This topic selection information is then sent from the terminal to the server.
[0047] Step 2:
[0048] Based on the theme information received, the server selects the appropriate comic data from the database. This selection is important for providing stories that are suitable for the learning content.
[0049] Step 3:
[0050] The server uses a generative model to analyze selected comic data, quantifying emotions from the characters' expressions and dialogue in each panel. This organizes the emotional flow of the story as digital data.
[0051] Step 4:
[0052] Based on quantified emotional data, the server dynamically generates the most suitable story for the user. The generated story is customized to fit the user's learning profile.
[0053] Step 5:
[0054] The server sends generated story data to the terminal, which then displays it visually in comic book format. The user views the displayed comic and progresses through the learning process.
[0055] Step 6:
[0056] Users record their learning progress and provide feedback on the content via their devices. This includes entering comments and evaluating their level of understanding.
[0057] Step 7:
[0058] The device sends user feedback data to the server, which then updates its database based on that data and uses it to generate the next story.
[0059] Step 8:
[0060] The server analyzes the collected feedback to gain insights that further improve the user's learning experience. This leads to improvements across the entire system.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] Traditional education systems struggle to provide learning experiences that effectively incorporate learners' interests and emotions, relying instead on learning content that is difficult to remember and easily becomes boring. This has led to the problem of undermining learners' motivation to learn proactively and continuously.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes means for analyzing information data using generation technology and quantifying the emotion of each information unit; means for dynamically generating a story suitable for the learner based on the quantified emotion data; and means for transmitting the generated story to an output device and displaying it visually. This makes it possible to provide an individually optimized learning experience that reflects the learner's emotions and interests, and overcome the limitations of conventional learning methods.
[0066] An "information processing device" is a computing device that has the function of receiving data, processing it, and transmitting the data to other devices as needed.
[0067] "Generative technology" refers to techniques that use machine learning algorithms to create new data and content.
[0068] "Information data" refers to digital information content that is stored in various forms such as text, images, and audio.
[0069] An "information unit" is the smallest analyzable part that makes up information data.
[0070] "Methods for quantifying emotions" refer to the process of using computational techniques to quantify emotions within a unit of information.
[0071] A "learner" is an individual or group that uses a system to learn information.
[0072] A "story" is educational content in story format that is generated based on emotional data.
[0073] An "output device" is hardware used to provide processed information or content to the user visually or audibly.
[0074] A "user interface" is a means of interaction that allows learners to operate a system and input or select information.
[0075] "Characters" refer to the characters or subjects within a story.
[0076] To implement this invention, it is necessary to use a network-connected server and terminals and utilize generation technology that operates on an information processing device. The server analyzes data based on information received from the user and generates appropriate learning content. In this process, the server uses a generation AI model to quantify the emotions of the information data and generates personalized stories that match the user's interests and level of understanding.
[0077] The server connects to a database and retrieves material data tailored to the user's selected learning theme. A commonly available natural language processing technique is used as the specific generative model. This allows for the quantification of the emotions associated with each piece of information, and a narrative is constructed based on this quantified information.
[0078] The device receives stories sent from the server and displays them in a visually viewable format. The device's user interface is used for users to select learning themes and input feedback. This interface allows users to freely record their feelings and understandings as they progress through the learning process. This feedback becomes important input information for future story generation, leading to further improvements in educational effectiveness.
[0079] As a concrete example, when a user studying history selects a theme from a particular era, the server retrieves source data related to that era and analyzes the quantified emotional data using a generative AI model. This generates a story with an emotional flow optimized for the user. On the device, the user can visually experience this story and provide feedback on the parts that interest them.
[0080] An example of a prompt might be, "Generate a story about medieval Europe and describe it with emotion." Using this prompt, the generative AI model creates an appropriate story. In this way, the entire system works together to provide the user with an optimal learning environment.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The server receives learning theme selection information from the user. This input information indicates what the user wants to learn based on their interests. Based on this information, the server selects appropriate source data from the database. The selected source data will form the basis for future story generation.
[0084] Step 2:
[0085] The server applies a generative AI model to the source data retrieved from the database, analyzing and quantifying the emotions of each information unit (e.g., comic panel). In this process, information data related to the selected theme is input, and the generative AI model performs data processing by analyzing facial expressions and dialogue to quantify emotional data. The resulting emotional data is used to construct the story.
[0086] Step 3:
[0087] The server dynamically generates user-specific stories using quantified sentiment data. Here, it combines user profile information with sentiment data to create personalized stories. The output is this customized story data, which is then used in the next processing step.
[0088] Step 4:
[0089] The server sends the generated story data to the device. The transmitted data needs to be converted into a format for display on the device. The converted data becomes a comic book format to provide the user with a visual and emotional learning experience.
[0090] Step 5:
[0091] The device visually displays story data received from the server. This allows users to view the generated comics through the interface. The displayed stories can include interactive elements to help users continue learning.
[0092] Step 6:
[0093] Users learn through the displayed stories and input their thoughts and additional information as feedback on their devices. This feedback reflects the user's learning experience and deepening understanding, and serves as important data for generating future learning content.
[0094] Step 7:
[0095] The device sends feedback received from the user to the server. This feedback data is stored on the server and used in subsequent content creation processes. This feedback is essential for providing a customized learning experience that meets the user's learning needs.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] Traditional learning methods rely on text and static materials, making it difficult to engage learners. Furthermore, providing content tailored to individual learners is challenging, leading to decreased learning efficiency. Therefore, there is a need to develop learning systems that are more memorable and enjoyable.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each fragment, means for dynamically generating a story suitable for the user based on the quantified emotion information, and means for transmitting the generated story to a display device and presenting it visually. This provides learning content that is optimal for each individual learner and enables an interesting and efficient learning experience through emotional interaction.
[0101] A "generative model" is an algorithm that can analyze data and generate new information or content.
[0102] "Information analysis" is the process of analyzing data and extracting useful patterns and features from it.
[0103] A "fragment" refers to an element or part that makes up a whole, and in this context, it refers to individual data units such as audio or images.
[0104] "Quantifying emotions" involves analyzing a character's facial expressions and dialogue to express their emotional state using numerical values.
[0105] A "story tailored to the user" refers to a story customized to a specific user's learning pace and interests.
[0106] "Dynamic generation" refers to the process of creating new content using data in real time.
[0107] A "display device" is a device used to visually present electronically generated information, such as a monitor or display.
[0108] "Visual presentation" means displaying information in an easily understandable way using images, videos, and other visual means.
[0109] "Interaction" refers to a two-way exchange between a person and a computer or content, a process in which the user inputs information and the system responds accordingly.
[0110] The system for realizing this invention is designed so that the server, terminal, and user each fulfill their respective roles. Details of each component are described below.
[0111] First, the server is equipped with a generative AI model, which is used to analyze various data. Specifically, it receives fragments of comic data and quantifies emotions from the characters' facial expressions and dialogue within them. Based on this quantified emotional information, a story optimized for the user is dynamically generated. This process utilizes cloud-based computing resources to perform advanced data processing.
[0112] Next, the device receives the generated story sent from the server and displays it visually. The device functions as a smartphone, tablet, or other display device, providing the user with an interactive learning experience. A high-resolution graphics engine is used for this visual display, generating a variety of visual content.
[0113] The user selects a specific topic they want to learn about and sends it to the server via their terminal as a prompt. The server then generates a relevant story based on this input. For example, a user interested in history might send a prompt such as, "Please generate a comic strip that tells the story of Napoleon's tactics." Upon receiving this prompt, the server generates a story according to the instructions and sends it back to the user's terminal.
[0114] This system allows learners to gain emotionally rich and engaging learning experiences that cannot be obtained from textbooks or familiar materials alone. By emphasizing emotional interaction, it is expected that the learned content will be more easily remembered.
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The user inputs a specific subject they want to learn about using an interface on their device. This input is accompanied by prompts such as, "Generate a comic strip that tells the story of Napoleon's tactics." These prompts trigger the server to generate an appropriate story.
[0118] Step 2:
[0119] The terminal sends the prompt message received from the user to the server. This allows the server to analyze the user's request and prepare to call the appropriate database or generative AI model. The output of this process is the transmission of the prompt message to the server.
[0120] Step 3:
[0121] The server receives a prompt and uses a generative AI model to select relevant data. Specifically, it extracts Napoleon-related comic data from the database and begins analysis to generate a story based on its content. The input data is the prompt, and the output is the selection of data necessary for the analysis.
[0122] Step 4:
[0123] The server inputs the selected data into a generating AI model, quantifying the emotions of each fragment. In this step, the character's facial expressions and dialogue are analyzed, and the emotions they represent are expressed numerically. The quantified emotional data is output, and the story is dynamically generated based on this data.
[0124] Step 5:
[0125] The server generates a story tailored to the user based on quantified emotional data. This generation process also takes user profile information into account, creating individually optimized stories. The output is the generated story data.
[0126] Step 6:
[0127] The server sends the generated story data to the terminal. The terminal receives this data and displays it visually using a high-resolution graphics engine. This allows the user to experience the story in an interactive comic book format.
[0128] Step 7:
[0129] Users read through the provided story and learn as they go. They record feedback on interesting parts and points they understood, and input it into their device. This input feedback is used in the next learning cycle.
[0130] Step 8:
[0131] The device collects user feedback and sends it to the server. The server stores this feedback and uses it to generate learning content for the next cycle. As an output, valuable data is accumulated to improve learning in the next cycle.
[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0133] This invention combines an emotion engine with a learning system that uses a generative model to provide a more personalized learning experience that responds to the learner's emotions. This system is intended to improve learning efficiency through close cooperation between the server, terminal, and user.
[0134] 1. Server-side processing
[0135] The server first receives theme selection information sent from the terminal and uses an emotion engine to acquire emotional information in real time from the user's facial expressions and voice. Using this emotional information, the server selects appropriate comic data from its database, analyzes the facial expressions and dialogue of the characters in each panel using a generative model, and quantifies the emotions. Based on this quantified data, it is reconstructed to generate a story that corresponds to the user's emotions. The server then sends this adjusted story to the terminal.
[0136] 2. Processing on the terminal side
[0137] The device visually displays the received story data as a comic strip and uses an emotion engine to analyze the user's facial expressions and voice. The analyzed emotion data is sent to a server and reflected in the story's development in real time. The device also presents the emotion recognition results to the user through the user interface and provides feedback for learning based on that. For example, if the user shows surprise at what is displayed, the story is adjusted to add new information and pique interest based on that emotion information.
[0138] 3. User-side initiatives
[0139] Users progress through the learning process via their device, viewing comics displayed on the device. During learning, an emotion engine recognizes the user's facial expressions and voice in real time, and the user's feedback and emotion recognition results are reflected in the learning process. By recording past emotion data, more personalized learning content will be provided in the future.
[0140] This invention aims to enhance knowledge acquisition and emotional engagement by leveraging learners' real-time emotions. It also aims to improve educational effectiveness by making the overall learning experience more individualized and dynamic.
[0141] The following describes the processing flow.
[0142] Step 1:
[0143] The user selects a topic they want to learn about through the device's user interface, and the device sends that information to the server.
[0144] Step 2:
[0145] The device activates its built-in emotion engine, analyzes the user's facial expressions and voice in real time to acquire emotional information, and sends this data to the server.
[0146] Step 3:
[0147] Based on the received theme information and user sentiment information, the server selects the appropriate comic data from the database.
[0148] Step 4:
[0149] The server uses a generative model to analyze selected comic data and quantifies the facial expressions and dialogue of the characters in each panel based on emotion categories.
[0150] Step 5:
[0151] The server dynamically generates the most suitable story for the learner based on quantified emotion data and the user's real-time emotion information, and sends it to the terminal.
[0152] Step 6:
[0153] The device displays story data received from the server in comic book format. Simultaneously, it presents the user with emotion recognition results to encourage feedback on the learning experience.
[0154] Step 7:
[0155] Users learn by reading the displayed comics via their devices. User reactions are continuously monitored by an emotion engine, and the story development is adjusted in real time as needed.
[0156] Step 8:
[0157] The device collects user feedback and emotion recognition data, sends it to a server, and updates the database. This data will be used to optimize future learning content.
[0158] (Example 2)
[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0160] Traditional learning systems have struggled to provide personalized content that responds to learners' emotions, resulting in a failure to sustain learners' interest and motivation. There is a need for a system that enables dynamic storytelling that takes learners' real-time emotions into account.
[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0162] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each frame, means for dynamically generating a story suitable for the learner based on the quantified emotion information, and means for transmitting the generated story to a device and displaying it visually. This makes it possible to provide personalized learning content that responds to the learner's real-time emotions.
[0163] A "generative model" is an artificial intelligence algorithm that learns specific patterns and associations from data and has the ability to make predictions and generate new instances.
[0164] "Information" refers to data used for analysis and processing within the system, such as learners' emotions and topic selection information.
[0165] A "frame" is a unit used to represent a single scene or moment within a comic or story.
[0166] "Methods for quantifying emotions" refer to methods and technologies for quantifying emotions obtained from a user's facial expressions and voice, and converting them into analyzable data.
[0167] A "story" refers to learning content presented to learners in story format, which is adjusted to suit the user's emotions.
[0168] "Device" refers to terminals or devices that users use to view or interact with learning content.
[0169] "Real-time" means that data acquisition, processing, and result presentation are all performed instantly, enabling responses that are tailored to the user's current state.
[0170] This invention is a system that uses a generative AI model to provide personalized learning content based on the user's emotions. Specific embodiments for implementing this system are described below.
[0171] Server-side embodiment
[0172] The server runs multiple software components, including a generative AI model and an emotion engine. First, the server processes the user's theme selection information received from the terminal. Then, the emotion engine analyzes the user's facial expressions and voice data sent from the terminal, quantifying their emotions. Using this emotional information, the server uses the generative AI model to select appropriate materials from the database and analyzes the character's facial expressions and dialogue in each frame. From this information, it generates a story tailored to the user and sends it to the terminal.
[0173] As a concrete example, in response to the user's chosen theme of "natural science," the server uses an emotion engine to analyze the user's emotions indicating their curiosity and adjusts the story accordingly.
[0174] An example of a prompt message is, "Generate a story that includes scientific content to stimulate the user's curiosity."
[0175] Terminal-side embodiment
[0176] The device has the ability to display story data received from the server in comic book format. It also uses its built-in camera and microphone to collect the user's facial expressions and voice in real time. This data is sent back to the server via an emotion engine and used to advance the story. The device can also present the results of emotion recognition to the user through its user interface and collect feedback.
[0177] User-side embodiment
[0178] Users participate in the system via a provided device. During learning, they can grasp the learning content through displayed comics and utilize the feedback displayed as the story progresses. Furthermore, the user's emotions are analyzed through real-time recognition of facial expressions and voice, and this data is used to personalize the next learning session. This provides a more interactive and personalized learning experience.
[0179] In this way, the system can dynamically personalize learning by using real-time sentiment analysis, thereby improving learner engagement and learning effectiveness.
[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0181] Step 1:
[0182] The server receives the user's learning theme selection information sent from the terminal. This information becomes input data for determining what kind of story to generate. Based on this input, the server prepares to retrieve corresponding materials from the database. Specifically, for example, if "Animal Science" is selected, the server generates index information to identify related content.
[0183] Step 2:
[0184] The device captures the user's facial expressions and voice in real time. This data, obtained using the camera and microphone, is input into an emotion engine, which converts the user's emotional state into numerical data. This converted numerical data is sent to a server and used to personalize the story. Specifically, if the user shows an expression of enjoyment, that data is quantified as "enjoyment."
[0185] Step 3:
[0186] The server operates a generative AI model based on the received quantified emotion information. In this process, it generates a story that matches the user's emotions, using materials selected from the database. Prompt statements are input to the generative AI model, which dynamically constructs the story. In a concrete example, the prompt "Generate a fantasy story that reflects the user's enjoyment" is used.
[0187] Step 4:
[0188] The server sends the generated story data to the device. The device receives this data and displays it to the user as a visual comic. The output is visual data of a story that is related to the user's learning theme and emotionally resonant. Specifically, a fun adventure story based on the initially selected theme, "Animal Science," is displayed on the device.
[0189] Step 5:
[0190] Users view stories provided through their devices and progress through the learning process. Any new emotional responses the user exhibits are captured again on the device and sent to the server. This feedback loop provides new learning data that will be reflected in the next story generation. Specifically, if a user expresses interest, that data is used to adjust subsequent learning content.
[0191] This series of processing steps allows the learning system to provide users with a more engaging and personalized learning experience.
[0192] (Application Example 2)
[0193] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0194] In modern work environments, workers' emotions and stress levels significantly impact productivity and safety. However, most work systems fail to consider workers' emotions and conditions, merely providing uniform workloads and instructions. As a result, workers often experience unnecessary stress and accumulate fatigue. This invention aims to provide a system that utilizes workers' emotional information in real time to generate individualized work instructions.
[0195] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0196] In this invention, the server includes means for analyzing visual data of materials using a generative model and quantifying the emotions of each scene; means for dynamically generating work instructions suitable for the worker based on the quantified emotional information; and means for transmitting the generated work instructions to a device and displaying them visually or audibly. This makes it possible to provide individualized work instructions that take into account the emotional information of the worker.
[0197] A "generative model" is an artificial intelligence technology that learns patterns from data and generates new data.
[0198] "Visual data" refers to image and video data collected using cameras and sensors.
[0199] "Methods for quantifying emotions" refers to methods of numerically representing human emotions from visual and auditory data using an emotion engine.
[0200] "Dynamic generation methods" refer to methods that automatically create appropriate outputs in response to real-time changing situations and input data.
[0201] "Equipment" refers to machines or devices that have an interface with workers and are used to provide work instructions and information.
[0202] "Means of visual or auditory display" refers to methods of visually displaying information or playing it as sound via displays, speakers, etc.
[0203] "Worker feedback" refers to the opinions, reactions, and experiences obtained from workers during their work, and is used to improve the system and generate future work instructions.
[0204] In an embodiment of this invention, a system is provided that serves as an interface to support work within a factory. This system operates in cooperation with a server, terminals (factory robots and other display devices), and users (workers).
[0205] The server uses generative models to analyze visual data. Specifically, it quantifies workers' emotions based on image and audio data acquired from cameras and sensors. The emotion engine and generative model used are OpenCV and TENSORFLOW®. Other suitable libraries and frameworks can be selected as needed.
[0206] The server generates work instructions tailored to the worker based on this quantified emotional information, dynamically creating the content of these instructions. This generation model uses the AI framework GPT-3®. The server sends the generated instructions to the terminal. The terminal displays the received instructions visually or audibly using a display or audio output device and provides them to the worker.
[0207] Workers proceed with their tasks according to the instructions provided. Emotional changes and feedback during the work are transmitted from the terminal to the server and used to adjust and improve future work instructions. This enables flexible work support tailored to the individual needs of each worker.
[0208] As a concrete example, in a factory line, a robot could scan an employee's face, and if it determines they are fatigued, it could create an instruction to temporarily slow down their work pace and announce, "Please take a short break," thereby protecting the workers' health.
[0209] Example of a prompt:
[0210] "If a worker is deemed to be experiencing stress, please generate instructions for appropriate breaks or changes in work tasks."
[0211] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0212] Step 1:
[0213] The server receives visual and audio data acquired from cameras and sensors, and analyzes the workers' facial expressions and voices. This uses emotion recognition models based on OpenCV and TensorFlow to process the data in real time. The input is image and audio data, and the output is numerical data representing the workers' emotions.
[0214] Step 2:
[0215] The server analyzes quantified sentiment data and uses the generative AI model GPT-3 to generate work instructions suitable for the worker. The input is quantified sentiment data, and the output is customized work instructions tailored to the worker's situation. This process dynamically generates text or voice instructions based on sentiment information.
[0216] Step 3:
[0217] The server sends the generated work instructions to the terminal. The terminal receives them and presents them to the worker visually or audibly via a display or audio output device. The input is the work instruction data, and the output is the instruction display or audio output to the worker.
[0218] Step 4:
[0219] The user, a worker, performs tasks based on instructions received through a terminal. During the work, the terminal collects the worker's feedback and new emotional data. This data is sent from the terminal to the server for use in generating future instructions. The input is the worker's feedback and new emotional data, and the output is the storage of this data in a database used to generate future work instructions.
[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0223] [Second Embodiment]
[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0236] This invention provides a learning system using generative models, realizing an innovative method to make exam preparation more effective and engaging. This system operates through the coordinated interaction of a server, terminals, and users.
[0237] 1. Server-side processing
[0238] The server first receives theme selection information from the user. Based on this information, the server selects appropriate comic data from its database. Using a generative model, the server analyzes the content of each panel of the selected comic data, quantifying emotions from character expressions and dialogue. This organizes the emotional information of each panel and provides data to construct the emotional flow of the story. Based on this emotional data, the server dynamically generates a story tailored to the user's profile and sends it to the terminal.
[0239] 2. Processing on the terminal side
[0240] The user's device receives story data sent from the server and displays it visually as a comic strip. Through the user interface, the user can select their preferred theme and interact with the displayed learning content. For example, if a user studying history selects this theme, the server generates a story based on historical events and displays it in comic strip format on the device, providing a visual and emotional learning experience.
[0241] 3. User-side initiatives
[0242] Users learn through comics displayed on their devices. During the learning process, users can input comments and feedback on points of interest or areas they understand better. The device sends this feedback to the server, which stores it as data to improve the user experience. This feedback allows the server to provide more suitable content to the user when generating future learning stories.
[0243] Based on the above, the present invention provides a system that connects emotions and knowledge, making test preparation more memorable, enjoyable, and efficient. This system adds new value to education and is expected to significantly improve conventional learning methods.
[0244] The following describes the processing flow.
[0245] Step 1:
[0246] The user operates the terminal's user interface to select the topic they wish to learn about. This topic selection information is then sent from the terminal to the server.
[0247] Step 2:
[0248] Based on the theme information received, the server selects the appropriate comic data from the database. This selection is important for providing stories that are suitable for the learning content.
[0249] Step 3:
[0250] The server uses a generative model to analyze selected comic data, quantifying emotions from the characters' expressions and dialogue in each panel. This organizes the emotional flow of the story as digital data.
[0251] Step 4:
[0252] Based on quantified emotional data, the server dynamically generates the most suitable story for the user. The generated story is customized to fit the user's learning profile.
[0253] Step 5:
[0254] The server sends generated story data to the terminal, which then displays it visually in comic book format. The user views the displayed comic and progresses through the learning process.
[0255] Step 6:
[0256] Users record their learning progress and provide feedback on the content via their devices. This includes entering comments and evaluating their level of understanding.
[0257] Step 7:
[0258] The device sends user feedback data to the server, which then updates its database based on that data and uses it to generate the next story.
[0259] Step 8:
[0260] The server analyzes the collected feedback to gain insights that further improve the user's learning experience. This leads to improvements across the entire system.
[0261] (Example 1)
[0262] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0263] Traditional education systems struggle to provide learning experiences that effectively incorporate learners' interests and emotions, relying instead on learning content that is difficult to remember and easily becomes boring. This has led to the problem of undermining learners' motivation to learn proactively and continuously.
[0264] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0265] In this invention, the server includes means for analyzing information data using generation technology and quantifying the emotion of each information unit; means for dynamically generating a story suitable for the learner based on the quantified emotion data; and means for transmitting the generated story to an output device and displaying it visually. This makes it possible to provide an individually optimized learning experience that reflects the learner's emotions and interests, and overcome the limitations of conventional learning methods.
[0266] An "information processing device" is a computing device that has the function of receiving data, processing it, and transmitting the data to other devices as needed.
[0267] "Generative technology" refers to techniques that use machine learning algorithms to create new data and content.
[0268] "Information data" refers to digital information content that is stored in various forms such as text, images, and audio.
[0269] An "information unit" is the smallest analyzable part that makes up information data.
[0270] "Methods for quantifying emotions" refer to the process of using computational techniques to quantify emotions within a unit of information.
[0271] A "learner" is an individual or group that uses a system to learn information.
[0272] A "story" is educational content in story format that is generated based on emotional data.
[0273] An "output device" is hardware used to provide processed information or content to the user visually or audibly.
[0274] A "user interface" is a means of interaction that allows learners to operate a system and input or select information.
[0275] "Characters" refer to the characters or subjects within a story.
[0276] To implement this invention, it is necessary to use a network-connected server and terminals and utilize generation technology that operates on an information processing device. The server analyzes data based on information received from the user and generates appropriate learning content. In this process, the server uses a generation AI model to quantify the emotions of the information data and generates personalized stories that match the user's interests and level of understanding.
[0277] The server connects to a database and retrieves material data tailored to the user's selected learning theme. A commonly available natural language processing technique is used as the specific generative model. This allows for the quantification of the emotions associated with each piece of information, and a narrative is constructed based on this quantified information.
[0278] The device receives stories sent from the server and displays them in a visually viewable format. The device's user interface is used for users to select learning themes and input feedback. This interface allows users to freely record their feelings and understandings as they progress through the learning process. This feedback becomes important input information for future story generation, leading to further improvements in educational effectiveness.
[0279] As a concrete example, when a user studying history selects a theme from a particular era, the server retrieves source data related to that era and analyzes the quantified emotional data using a generative AI model. This generates a story with an emotional flow optimized for the user. On the device, the user can visually experience this story and provide feedback on the parts that interest them.
[0280] An example of a prompt might be, "Generate a story about medieval Europe and describe it with emotion." Using this prompt, the generative AI model creates an appropriate story. In this way, the entire system works together to provide the user with an optimal learning environment.
[0281] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0282] Step 1:
[0283] The server receives selection information of a learning theme from the user. This input information indicates what the user wants to learn based on their interests. Based on this information, the server selects appropriate material data from the database. The selected material data will form the basis for future story generation.
[0284] Step 2:
[0285] The server applies a generation AI model to the material data obtained from the database to analyze and quantify the emotions of each information unit (e.g., a frame of a comic). In this process, information data related to the selected theme is input, and data processing is performed where the generation AI model analyzes expressions and lines to quantify emotion data. The obtained emotion data is used for constructing the story.
[0286] Step 3:
[0287] The server dynamically generates a story suitable for the user using the quantified emotion data. Here, an individualized story is created by combining the user's profile information and the emotion data. What is output is this customized story data, which is used in the next processing step.
[0288] Step 4:
[0289] The server sends the generated story data to the terminal. The sent data needs to be converted into a format for display on the terminal. The converted data becomes a comic format to provide the user with a visual and emotional learning experience.
[0290] Step 5:
[0291] The terminal visually displays the story data received from the server. As a result, the user can view the generated comic through the interface. The displayed story can include interactive elements for continuous learning.
[0292] Step 6:
[0293] Users learn through the displayed stories and input their thoughts and additional information as feedback on their devices. This feedback reflects the user's learning experience and deepening understanding, and serves as important data for generating future learning content.
[0294] Step 7:
[0295] The device sends feedback received from the user to the server. This feedback data is stored on the server and used in subsequent content creation processes. This feedback is essential for providing a customized learning experience that meets the user's learning needs.
[0296] (Application Example 1)
[0297] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0298] Traditional learning methods rely on text and static materials, making it difficult to engage learners. Furthermore, providing content tailored to individual learners is challenging, leading to decreased learning efficiency. Therefore, there is a need to develop learning systems that are more memorable and enjoyable.
[0299] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0300] In this invention, the server includes means for analyzing information using a generation model and quantifying the emotions of each fragment, means for dynamically generating a story suitable for the user based on the quantified emotion information, and means for transmitting the generated story to a display device and visually presenting it. As a result, optimal learning content can be provided to individual learners, and an interesting and efficient learning experience can be achieved through emotional interaction.
[0301] The "generation model" is an algorithm that can analyze data and generate new information and content.
[0302] "Analysis of information" is a process of analyzing data and extracting useful patterns and features from it.
[0303] A "fragment" refers to an element or part that constitutes the whole, and here it refers to individual data units such as voice and images.
[0304] "Quantification of emotions" is to analyze the expressions and lines of characters and represent the emotional state by numerical values.
[0305] A "story suitable for the user" refers to a story customized according to the learning speed and interests of a specific user.
[0306] "Dynamically generate" is a process of creating new content using data in real time.
[0307] A "display device" is a device for visually presenting electronically generated information, such as a monitor or a display.
[0308] "Visually present" is to display information visually in an easy-to-see manner by videos, images, etc.
[0309] "Interaction" refers to the two-way communication between people and computers or content, which is a process in which the user inputs and the system returns a corresponding reaction.
[0310] The system for realizing this invention is designed so that the server, terminal, and user each fulfill their respective roles. Details of each component are described below.
[0311] First, the server is equipped with a generative AI model, which is used to analyze various data. Specifically, it receives fragments of comic data and quantifies emotions from the characters' facial expressions and dialogue within them. Based on this quantified emotional information, a story optimized for the user is dynamically generated. This process utilizes cloud-based computing resources to perform advanced data processing.
[0312] Next, the device receives the generated story sent from the server and displays it visually. The device functions as a smartphone, tablet, or other display device, providing the user with an interactive learning experience. A high-resolution graphics engine is used for this visual display, generating a variety of visual content.
[0313] The user selects a specific topic they want to learn about and sends it to the server via their terminal as a prompt. The server then generates a relevant story based on this input. For example, a user interested in history might send a prompt such as, "Please generate a comic strip that tells the story of Napoleon's tactics." Upon receiving this prompt, the server generates a story according to the instructions and sends it back to the user's terminal.
[0314] This system allows learners to gain emotionally rich and engaging learning experiences that cannot be obtained from textbooks or familiar materials alone. By emphasizing emotional interaction, it is expected that the learned content will be more easily remembered.
[0315] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0316] Step 1:
[0317] The user inputs a specific subject they want to learn about using an interface on their device. This input is accompanied by prompts such as, "Generate a comic strip that tells the story of Napoleon's tactics." These prompts trigger the server to generate an appropriate story.
[0318] Step 2:
[0319] The terminal sends the prompt message received from the user to the server. This allows the server to analyze the user's request and prepare to call the appropriate database or generative AI model. The output of this process is the transmission of the prompt message to the server.
[0320] Step 3:
[0321] The server receives a prompt and uses a generative AI model to select relevant data. Specifically, it extracts Napoleon-related comic data from the database and begins analysis to generate a story based on its content. The input data is the prompt, and the output is the selection of data necessary for the analysis.
[0322] Step 4:
[0323] The server inputs the selected data into a generating AI model, quantifying the emotions of each fragment. In this step, the character's facial expressions and dialogue are analyzed, and the emotions they represent are expressed numerically. The quantified emotional data is output, and the story is dynamically generated based on this data.
[0324] Step 5:
[0325] The server generates a story tailored to the user based on quantified emotional data. This generation process also takes user profile information into account, creating individually optimized stories. The output is the generated story data.
[0326] Step 6:
[0327] The server sends the generated story data to the terminal. The terminal receives this data and displays it visually using a high-resolution graphics engine. This allows the user to experience the story in an interactive comic book format.
[0328] Step 7:
[0329] Users read through the provided story and learn as they go. They record feedback on interesting parts and points they understood, and input it into their device. This input feedback is used in the next learning cycle.
[0330] Step 8:
[0331] The device collects user feedback and sends it to the server. The server stores this feedback and uses it to generate learning content for the next cycle. As an output, valuable data is accumulated to improve learning in the next cycle.
[0332] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0333] This invention combines an emotion engine with a learning system that uses a generative model to provide a more personalized learning experience that responds to the learner's emotions. This system is intended to improve learning efficiency through close cooperation between the server, terminal, and user.
[0334] 1. Server-side processing
[0335] The server first receives theme selection information sent from the terminal and uses an emotion engine to acquire emotional information in real time from the user's facial expressions and voice. Using this emotional information, the server selects appropriate comic data from its database, analyzes the facial expressions and dialogue of the characters in each panel using a generative model, and quantifies the emotions. Based on this quantified data, it is reconstructed to generate a story that corresponds to the user's emotions. The server then sends this adjusted story to the terminal.
[0336] 2. Processing on the terminal side
[0337] The device visually displays the received story data as a comic strip and uses an emotion engine to analyze the user's facial expressions and voice. The analyzed emotion data is sent to a server and reflected in the story's development in real time. The device also presents the emotion recognition results to the user through the user interface and provides feedback for learning based on that. For example, if the user shows surprise at what is displayed, the story is adjusted to add new information and pique interest based on that emotion information.
[0338] 3. User-side initiatives
[0339] Users progress through the learning process via their device, viewing comics displayed on the device. During learning, an emotion engine recognizes the user's facial expressions and voice in real time, and the user's feedback and emotion recognition results are reflected in the learning process. By recording past emotion data, more personalized learning content will be provided in the future.
[0340] This invention aims to enhance knowledge acquisition and emotional engagement by leveraging learners' real-time emotions. It also aims to improve educational effectiveness by making the overall learning experience more individualized and dynamic.
[0341] The following describes the processing flow.
[0342] Step 1:
[0343] The user selects a topic they want to learn about through the device's user interface, and the device sends that information to the server.
[0344] Step 2:
[0345] The device activates its built-in emotion engine, analyzes the user's facial expressions and voice in real time to acquire emotional information, and sends this data to the server.
[0346] Step 3:
[0347] Based on the received theme information and user sentiment information, the server selects the appropriate comic data from the database.
[0348] Step 4:
[0349] The server uses a generative model to analyze selected comic data and quantifies the facial expressions and dialogue of the characters in each panel based on emotion categories.
[0350] Step 5:
[0351] The server dynamically generates the most suitable story for the learner based on quantified emotion data and the user's real-time emotion information, and sends it to the terminal.
[0352] Step 6:
[0353] The device displays story data received from the server in comic book format. Simultaneously, it presents the user with emotion recognition results to encourage feedback on the learning experience.
[0354] Step 7:
[0355] Users learn by reading the displayed comics via their devices. User reactions are continuously monitored by an emotion engine, and the story development is adjusted in real time as needed.
[0356] Step 8:
[0357] The device collects user feedback and emotion recognition data, sends it to a server, and updates the database. This data will be used to optimize future learning content.
[0358] (Example 2)
[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0360] Traditional learning systems have struggled to provide personalized content that responds to learners' emotions, resulting in a failure to sustain learners' interest and motivation. There is a need for a system that enables dynamic storytelling that takes learners' real-time emotions into account.
[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0362] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each frame, means for dynamically generating a story suitable for the learner based on the quantified emotion information, and means for transmitting the generated story to a device and displaying it visually. This makes it possible to provide personalized learning content that responds to the learner's real-time emotions.
[0363] A "generative model" is an artificial intelligence algorithm that learns specific patterns and associations from data and has the ability to make predictions and generate new instances.
[0364] "Information" refers to data used for analysis and processing within the system, such as learners' emotions and topic selection information.
[0365] A "frame" is a unit used to represent a single scene or moment within a comic or story.
[0366] "Methods for quantifying emotions" refer to methods and technologies for quantifying emotions obtained from a user's facial expressions and voice, and converting them into analyzable data.
[0367] A "story" refers to learning content presented to learners in story format, which is adjusted to suit the user's emotions.
[0368] "Device" refers to terminals or devices that users use to view or interact with learning content.
[0369] "Real-time" means that data acquisition, processing, and result presentation are all performed instantly, enabling responses that are tailored to the user's current state.
[0370] This invention is a system that uses a generative AI model to provide personalized learning content based on the user's emotions. Specific embodiments for implementing this system are described below.
[0371] Server-side embodiment
[0372] The server runs multiple software components, including a generative AI model and an emotion engine. First, the server processes the user's theme selection information received from the terminal. Then, the emotion engine analyzes the user's facial expressions and voice data sent from the terminal, quantifying their emotions. Using this emotional information, the server uses the generative AI model to select appropriate materials from the database and analyzes the character's facial expressions and dialogue in each frame. From this information, it generates a story tailored to the user and sends it to the terminal.
[0373] As a concrete example, in response to the user's chosen theme of "natural science," the server uses an emotion engine to analyze the user's emotions indicating their curiosity and adjusts the story accordingly.
[0374] An example of a prompt message is, "Generate a story that includes scientific content to stimulate the user's curiosity."
[0375] Terminal-side embodiment
[0376] The device has the ability to display story data received from the server in comic book format. It also uses its built-in camera and microphone to collect the user's facial expressions and voice in real time. This data is sent back to the server via an emotion engine and used to advance the story. The device can also present the results of emotion recognition to the user through its user interface and collect feedback.
[0377] User-side embodiment
[0378] Users participate in the system via a provided device. During learning, they can grasp the learning content through displayed comics and utilize the feedback displayed as the story progresses. Furthermore, the user's emotions are analyzed through real-time recognition of facial expressions and voice, and this data is used to personalize the next learning session. This provides a more interactive and personalized learning experience.
[0379] In this way, the system can dynamically personalize learning by using real-time sentiment analysis, thereby improving learner engagement and learning effectiveness.
[0380] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0381] Step 1:
[0382] The server receives the user's learning theme selection information sent from the terminal. This information becomes input data for determining what kind of story to generate. Based on this input, the server prepares to retrieve corresponding materials from the database. Specifically, for example, if "Animal Science" is selected, the server generates index information to identify related content.
[0383] Step 2:
[0384] The device captures the user's facial expressions and voice in real time. This data, obtained using the camera and microphone, is input into an emotion engine, which converts the user's emotional state into numerical data. This converted numerical data is sent to a server and used to personalize the story. Specifically, if the user shows an expression of enjoyment, that data is quantified as "enjoyment."
[0385] Step 3:
[0386] The server operates a generative AI model based on the received quantified emotion information. In this process, it generates a story that matches the user's emotions, using materials selected from the database. Prompt statements are input to the generative AI model, which dynamically constructs the story. In a concrete example, the prompt "Generate a fantasy story that reflects the user's enjoyment" is used.
[0387] Step 4:
[0388] The server sends the generated story data to the device. The device receives this data and displays it to the user as a visual comic. The output is visual data of a story that is related to the user's learning theme and emotionally resonant. Specifically, a fun adventure story based on the initially selected theme, "Animal Science," is displayed on the device.
[0389] Step 5:
[0390] Users view stories provided through their devices and progress through the learning process. Any new emotional responses the user exhibits are captured again on the device and sent to the server. This feedback loop provides new learning data that will be reflected in the next story generation. Specifically, if a user expresses interest, that data is used to adjust subsequent learning content.
[0391] This series of processing steps allows the learning system to provide users with a more engaging and personalized learning experience.
[0392] (Application Example 2)
[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0394] In modern work environments, workers' emotions and stress levels significantly impact productivity and safety. However, most work systems fail to consider workers' emotions and conditions, merely providing uniform workloads and instructions. As a result, workers often experience unnecessary stress and accumulate fatigue. This invention aims to provide a system that utilizes workers' emotional information in real time to generate individualized work instructions.
[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0396] In this invention, the server includes means for analyzing visual data of materials using a generative model and quantifying the emotions of each scene; means for dynamically generating work instructions suitable for the worker based on the quantified emotional information; and means for transmitting the generated work instructions to a device and displaying them visually or audibly. This makes it possible to provide individualized work instructions that take into account the emotional information of the worker.
[0397] A "generative model" is an artificial intelligence technology that learns patterns from data and generates new data.
[0398] "Visual data" refers to image and video data collected using cameras and sensors.
[0399] "Methods for quantifying emotions" refers to methods of numerically representing human emotions from visual and auditory data using an emotion engine.
[0400] "Dynamic generation methods" refer to methods that automatically create appropriate outputs in response to real-time changing situations and input data.
[0401] "Equipment" refers to machines or devices that have an interface with workers and are used to provide work instructions and information.
[0402] "Means of visual or auditory display" refers to methods of visually displaying information or playing it as sound via displays, speakers, etc.
[0403] "Worker feedback" refers to the opinions, reactions, and experiences obtained from workers during their work, and is used to improve the system and generate future work instructions.
[0404] In an embodiment of this invention, a system is provided that serves as an interface to support work within a factory. This system operates in cooperation with a server, terminals (factory robots and other display devices), and users (workers).
[0405] The server uses generative models to analyze visual data. Specifically, it quantifies workers' emotions based on image and audio data acquired from cameras and sensors. OpenCV and TensorFlow are used as the emotion engine and generative model. Other suitable libraries and frameworks can be selected as needed.
[0406] The server generates work instructions tailored to the worker based on this quantified emotional information, dynamically creating the content of these instructions. This generation model utilizes the AI framework GPT-3. The server sends the generated instructions to a terminal. The terminal displays the received instructions visually or audibly using a display or audio output device and provides them to the worker.
[0407] Workers proceed with their tasks according to the instructions provided. Emotional changes and feedback during the work are transmitted from the terminal to the server and used to adjust and improve future work instructions. This enables flexible work support tailored to the individual needs of each worker.
[0408] As a concrete example, in a factory line, a robot could scan an employee's face, and if it determines they are fatigued, it could create an instruction to temporarily slow down their work pace and announce, "Please take a short break," thereby protecting the workers' health.
[0409] Example of a prompt:
[0410] "If a worker is deemed to be experiencing stress, please generate instructions for appropriate breaks or changes in work tasks."
[0411] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0412] Step 1:
[0413] The server receives visual and audio data acquired from cameras and sensors, and analyzes the workers' facial expressions and voices. This uses emotion recognition models based on OpenCV and TensorFlow to process the data in real time. The input is image and audio data, and the output is numerical data representing the workers' emotions.
[0414] Step 2:
[0415] The server analyzes quantified sentiment data and uses the generative AI model GPT-3 to generate work instructions suitable for the worker. The input is quantified sentiment data, and the output is customized work instructions tailored to the worker's situation. This process dynamically generates text or voice instructions based on sentiment information.
[0416] Step 3:
[0417] The server sends the generated work instructions to the terminal. The terminal receives them and presents them to the worker visually or audibly via a display or audio output device. The input is the work instruction data, and the output is the instruction display or audio output to the worker.
[0418] Step 4:
[0419] The user, a worker, performs tasks based on instructions received through a terminal. During the work, the terminal collects the worker's feedback and new emotional data. This data is sent from the terminal to the server for use in generating future instructions. The input is the worker's feedback and new emotional data, and the output is the storage of this data in a database used to generate future work instructions.
[0420] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0421] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0422] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0423] [Third Embodiment]
[0424] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0425] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0426] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0427] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0428] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0430] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0431] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0432] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0433] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0434] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0435] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0436] This invention provides a learning system using generative models, realizing an innovative method to make exam preparation more effective and engaging. This system operates through the coordinated interaction of a server, terminals, and users.
[0437] 1. Server-side processing
[0438] The server first receives theme selection information from the user. Based on this information, the server selects appropriate comic data from its database. Using a generative model, the server analyzes the content of each panel of the selected comic data, quantifying emotions from character expressions and dialogue. This organizes the emotional information of each panel and provides data to construct the emotional flow of the story. Based on this emotional data, the server dynamically generates a story tailored to the user's profile and sends it to the terminal.
[0439] 2. Processing on the terminal side
[0440] The user's device receives story data sent from the server and displays it visually as a comic strip. Through the user interface, the user can select their preferred theme and interact with the displayed learning content. For example, if a user studying history selects this theme, the server generates a story based on historical events and displays it in comic strip format on the device, providing a visual and emotional learning experience.
[0441] 3. User-side initiatives
[0442] Users learn through comics displayed on their devices. During the learning process, users can input comments and feedback on points of interest or areas they understand better. The device sends this feedback to the server, which stores it as data to improve the user experience. This feedback allows the server to provide more suitable content to the user when generating future learning stories.
[0443] Based on the above, the present invention provides a system that connects emotions and knowledge, making test preparation more memorable, enjoyable, and efficient. This system adds new value to education and is expected to significantly improve conventional learning methods.
[0444] The following describes the processing flow.
[0445] Step 1:
[0446] The user operates the terminal's user interface to select the topic they wish to learn about. This topic selection information is then sent from the terminal to the server.
[0447] Step 2:
[0448] Based on the theme information received, the server selects the appropriate comic data from the database. This selection is important for providing stories that are suitable for the learning content.
[0449] Step 3:
[0450] The server uses a generative model to analyze selected comic data, quantifying emotions from the characters' expressions and dialogue in each panel. This organizes the emotional flow of the story as digital data.
[0451] Step 4:
[0452] Based on quantified emotional data, the server dynamically generates the most suitable story for the user. The generated story is customized to fit the user's learning profile.
[0453] Step 5:
[0454] The server sends generated story data to the terminal, which then displays it visually in comic book format. The user views the displayed comic and progresses through the learning process.
[0455] Step 6:
[0456] Users record their learning progress and provide feedback on the content via their devices. This includes entering comments and evaluating their level of understanding.
[0457] Step 7:
[0458] The device sends user feedback data to the server, which then updates its database based on that data and uses it to generate the next story.
[0459] Step 8:
[0460] The server analyzes the collected feedback to gain insights that further improve the user's learning experience. This leads to improvements across the entire system.
[0461] (Example 1)
[0462] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0463] Traditional education systems struggle to provide learning experiences that effectively incorporate learners' interests and emotions, relying instead on learning content that is difficult to remember and easily becomes boring. This has led to the problem of undermining learners' motivation to learn proactively and continuously.
[0464] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0465] In this invention, the server includes means for analyzing information data using generation technology and quantifying the emotion of each information unit; means for dynamically generating a story suitable for the learner based on the quantified emotion data; and means for transmitting the generated story to an output device and displaying it visually. This makes it possible to provide an individually optimized learning experience that reflects the learner's emotions and interests, and overcome the limitations of conventional learning methods.
[0466] An "information processing device" is a computing device that has the function of receiving data, processing it, and transmitting the data to other devices as needed.
[0467] "Generative technology" refers to techniques that use machine learning algorithms to create new data and content.
[0468] "Information data" refers to digital information content that is stored in various forms such as text, images, and audio.
[0469] An "information unit" is the smallest analyzable part that makes up information data.
[0470] "Methods for quantifying emotions" refer to the process of using computational techniques to quantify emotions within a unit of information.
[0471] A "learner" is an individual or group that uses a system to learn information.
[0472] A "story" is educational content in story format that is generated based on emotional data.
[0473] An "output device" is hardware used to provide processed information or content to the user visually or audibly.
[0474] A "user interface" is a means of interaction that allows learners to operate a system and input or select information.
[0475] "Characters" refer to the characters or subjects within a story.
[0476] To implement this invention, it is necessary to use a network-connected server and terminals and utilize generation technology that operates on an information processing device. The server analyzes data based on information received from the user and generates appropriate learning content. In this process, the server uses a generation AI model to quantify the emotions of the information data and generates personalized stories that match the user's interests and level of understanding.
[0477] The server connects to a database and retrieves material data tailored to the user's selected learning theme. A commonly available natural language processing technique is used as the specific generative model. This allows for the quantification of the emotions associated with each piece of information, and a narrative is constructed based on this quantified information.
[0478] The device receives stories sent from the server and displays them in a visually viewable format. The device's user interface is used for users to select learning themes and input feedback. This interface allows users to freely record their feelings and understandings as they progress through the learning process. This feedback becomes important input information for future story generation, leading to further improvements in educational effectiveness.
[0479] As a concrete example, when a user studying history selects a theme from a particular era, the server retrieves source data related to that era and analyzes the quantified emotional data using a generative AI model. This generates a story with an emotional flow optimized for the user. On the device, the user can visually experience this story and provide feedback on the parts that interest them.
[0480] An example of a prompt might be, "Generate a story about medieval Europe and describe it with emotion." Using this prompt, the generative AI model creates an appropriate story. In this way, the entire system works together to provide the user with an optimal learning environment.
[0481] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0482] Step 1:
[0483] The server receives learning theme selection information from the user. This input information indicates what the user wants to learn based on their interests. Based on this information, the server selects appropriate source data from the database. The selected source data will form the basis for future story generation.
[0484] Step 2:
[0485] The server applies a generative AI model to the source data retrieved from the database, analyzing and quantifying the emotions of each information unit (e.g., comic panel). In this process, information data related to the selected theme is input, and the generative AI model performs data processing by analyzing facial expressions and dialogue to quantify emotional data. The resulting emotional data is used to construct the story.
[0486] Step 3:
[0487] The server dynamically generates user-specific stories using quantified sentiment data. Here, it combines user profile information with sentiment data to create personalized stories. The output is this customized story data, which is then used in the next processing step.
[0488] Step 4:
[0489] The server sends the generated story data to the device. The transmitted data needs to be converted into a format for display on the device. The converted data becomes a comic book format to provide the user with a visual and emotional learning experience.
[0490] Step 5:
[0491] The device visually displays story data received from the server. This allows users to view the generated comics through the interface. The displayed stories can include interactive elements to help users continue learning.
[0492] Step 6:
[0493] Users learn through the displayed stories and input their thoughts and additional information as feedback on their devices. This feedback reflects the user's learning experience and deepening understanding, and serves as important data for generating future learning content.
[0494] Step 7:
[0495] The device sends feedback received from the user to the server. This feedback data is stored on the server and used in subsequent content creation processes. This feedback is essential for providing a customized learning experience that meets the user's learning needs.
[0496] (Application Example 1)
[0497] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0498] Traditional learning methods rely on text and static materials, making it difficult to engage learners. Furthermore, providing content tailored to individual learners is challenging, leading to decreased learning efficiency. Therefore, there is a need to develop learning systems that are more memorable and enjoyable.
[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0500] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each fragment, means for dynamically generating a story suitable for the user based on the quantified emotion information, and means for transmitting the generated story to a display device and presenting it visually. This provides learning content that is optimal for each individual learner and enables an interesting and efficient learning experience through emotional interaction.
[0501] A "generative model" is an algorithm that can analyze data and generate new information or content.
[0502] "Information analysis" is the process of analyzing data and extracting useful patterns and features from it.
[0503] A "fragment" refers to an element or part that makes up a whole, and in this context, it refers to individual data units such as audio or images.
[0504] "Quantifying emotions" involves analyzing a character's facial expressions and dialogue to express their emotional state using numerical values.
[0505] A "story tailored to the user" refers to a story customized to a specific user's learning pace and interests.
[0506] "Dynamic generation" refers to the process of creating new content using data in real time.
[0507] A "display device" is a device used to visually present electronically generated information, such as a monitor or display.
[0508] "Visual presentation" means displaying information in an easily understandable way using images, videos, and other visual means.
[0509] "Interaction" refers to a two-way exchange between a person and a computer or content, a process in which the user inputs information and the system responds accordingly.
[0510] The system for realizing this invention is designed so that the server, terminal, and user each fulfill their respective roles. Details of each component are described below.
[0511] First, the server is equipped with a generative AI model, which is used to analyze various data. Specifically, it receives fragments of comic data and quantifies emotions from the characters' facial expressions and dialogue within them. Based on this quantified emotional information, a story optimized for the user is dynamically generated. This process utilizes cloud-based computing resources to perform advanced data processing.
[0512] Next, the device receives the generated story sent from the server and displays it visually. The device functions as a smartphone, tablet, or other display device, providing the user with an interactive learning experience. A high-resolution graphics engine is used for this visual display, generating a variety of visual content.
[0513] The user selects a specific topic they want to learn about and sends it to the server via their terminal as a prompt. The server then generates a relevant story based on this input. For example, a user interested in history might send a prompt such as, "Please generate a comic strip that tells the story of Napoleon's tactics." Upon receiving this prompt, the server generates a story according to the instructions and sends it back to the user's terminal.
[0514] This system allows learners to gain emotionally rich and engaging learning experiences that cannot be obtained from textbooks or familiar materials alone. By emphasizing emotional interaction, it is expected that the learned content will be more easily remembered.
[0515] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0516] Step 1:
[0517] The user inputs a specific subject they want to learn about using an interface on their device. This input is accompanied by prompts such as, "Generate a comic strip that tells the story of Napoleon's tactics." These prompts trigger the server to generate an appropriate story.
[0518] Step 2:
[0519] The terminal sends the prompt message received from the user to the server. This allows the server to analyze the user's request and prepare to call the appropriate database or generative AI model. The output of this process is the transmission of the prompt message to the server.
[0520] Step 3:
[0521] The server receives a prompt and uses a generative AI model to select relevant data. Specifically, it extracts Napoleon-related comic data from the database and begins analysis to generate a story based on its content. The input data is the prompt, and the output is the selection of data necessary for the analysis.
[0522] Step 4:
[0523] The server inputs the selected data into a generating AI model, quantifying the emotions of each fragment. In this step, the character's facial expressions and dialogue are analyzed, and the emotions they represent are expressed numerically. The quantified emotional data is output, and the story is dynamically generated based on this data.
[0524] Step 5:
[0525] The server generates a story tailored to the user based on quantified emotional data. This generation process also takes user profile information into account, creating individually optimized stories. The output is the generated story data.
[0526] Step 6:
[0527] The server sends the generated story data to the terminal. The terminal receives this data and displays it visually using a high-resolution graphics engine. This allows the user to experience the story in an interactive comic book format.
[0528] Step 7:
[0529] Users read through the provided story and learn as they go. They record feedback on interesting parts and points they understood, and input it into their device. This input feedback is used in the next learning cycle.
[0530] Step 8:
[0531] The device collects user feedback and sends it to the server. The server stores this feedback and uses it to generate learning content for the next cycle. As an output, valuable data is accumulated to improve learning in the next cycle.
[0532] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0533] This invention combines an emotion engine with a learning system that uses a generative model to provide a more personalized learning experience that responds to the learner's emotions. This system is intended to improve learning efficiency through close cooperation between the server, terminal, and user.
[0534] 1. Server-side processing
[0535] The server first receives theme selection information sent from the terminal and uses an emotion engine to acquire emotional information in real time from the user's facial expressions and voice. Using this emotional information, the server selects appropriate comic data from its database, analyzes the facial expressions and dialogue of the characters in each panel using a generative model, and quantifies the emotions. Based on this quantified data, it is reconstructed to generate a story that corresponds to the user's emotions. The server then sends this adjusted story to the terminal.
[0536] 2. Processing on the terminal side
[0537] The device visually displays the received story data as a comic strip and uses an emotion engine to analyze the user's facial expressions and voice. The analyzed emotion data is sent to a server and reflected in the story's development in real time. The device also presents the emotion recognition results to the user through the user interface and provides feedback for learning based on that. For example, if the user shows surprise at what is displayed, the story is adjusted to add new information and pique interest based on that emotion information.
[0538] 3. User-side initiatives
[0539] Users progress through the learning process via their device, viewing comics displayed on the device. During learning, an emotion engine recognizes the user's facial expressions and voice in real time, and the user's feedback and emotion recognition results are reflected in the learning process. By recording past emotion data, more personalized learning content will be provided in the future.
[0540] This invention aims to enhance knowledge acquisition and emotional engagement by leveraging learners' real-time emotions. It also aims to improve educational effectiveness by making the overall learning experience more individualized and dynamic.
[0541] The following describes the processing flow.
[0542] Step 1:
[0543] The user selects a topic they want to learn about through the device's user interface, and the device sends that information to the server.
[0544] Step 2:
[0545] The device activates its built-in emotion engine, analyzes the user's facial expressions and voice in real time to acquire emotional information, and sends this data to the server.
[0546] Step 3:
[0547] Based on the received theme information and user sentiment information, the server selects the appropriate comic data from the database.
[0548] Step 4:
[0549] The server uses a generative model to analyze selected comic data and quantifies the facial expressions and dialogue of the characters in each panel based on emotion categories.
[0550] Step 5:
[0551] The server dynamically generates the most suitable story for the learner based on quantified emotion data and the user's real-time emotion information, and sends it to the terminal.
[0552] Step 6:
[0553] The device displays story data received from the server in comic book format. Simultaneously, it presents the user with emotion recognition results to encourage feedback on the learning experience.
[0554] Step 7:
[0555] Users learn by reading the displayed comics via their devices. User reactions are continuously monitored by an emotion engine, and the story development is adjusted in real time as needed.
[0556] Step 8:
[0557] The device collects user feedback and emotion recognition data, sends it to a server, and updates the database. This data will be used to optimize future learning content.
[0558] (Example 2)
[0559] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0560] Traditional learning systems have struggled to provide personalized content that responds to learners' emotions, resulting in a failure to sustain learners' interest and motivation. There is a need for a system that enables dynamic storytelling that takes learners' real-time emotions into account.
[0561] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0562] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each frame, means for dynamically generating a story suitable for the learner based on the quantified emotion information, and means for transmitting the generated story to a device and displaying it visually. This makes it possible to provide personalized learning content that responds to the learner's real-time emotions.
[0563] A "generative model" is an artificial intelligence algorithm that learns specific patterns and associations from data and has the ability to make predictions and generate new instances.
[0564] "Information" refers to data used for analysis and processing within the system, such as learners' emotions and topic selection information.
[0565] A "frame" is a unit used to represent a single scene or moment within a comic or story.
[0566] "Methods for quantifying emotions" refer to methods and technologies for quantifying emotions obtained from a user's facial expressions and voice, and converting them into analyzable data.
[0567] A "story" refers to learning content presented to learners in story format, which is adjusted to suit the user's emotions.
[0568] "Device" refers to terminals or devices that users use to view or interact with learning content.
[0569] "Real-time" means that data acquisition, processing, and result presentation are all performed instantly, enabling responses that are tailored to the user's current state.
[0570] This invention is a system that uses a generative AI model to provide personalized learning content based on the user's emotions. Specific embodiments for implementing this system are described below.
[0571] Server-side embodiment
[0572] The server runs multiple software components, including a generative AI model and an emotion engine. First, the server processes the user's theme selection information received from the terminal. Then, the emotion engine analyzes the user's facial expressions and voice data sent from the terminal, quantifying their emotions. Using this emotional information, the server uses the generative AI model to select appropriate materials from the database and analyzes the character's facial expressions and dialogue in each frame. From this information, it generates a story tailored to the user and sends it to the terminal.
[0573] As a concrete example, in response to the user's chosen theme of "natural science," the server uses an emotion engine to analyze the user's emotions indicating their curiosity and adjusts the story accordingly.
[0574] An example of a prompt message is, "Generate a story that includes scientific content to stimulate the user's curiosity."
[0575] Terminal-side embodiment
[0576] The device has the ability to display story data received from the server in comic book format. It also uses its built-in camera and microphone to collect the user's facial expressions and voice in real time. This data is sent back to the server via an emotion engine and used to advance the story. The device can also present the results of emotion recognition to the user through its user interface and collect feedback.
[0577] User-side embodiment
[0578] Users participate in the system via a provided device. During learning, they can grasp the learning content through displayed comics and utilize the feedback displayed as the story progresses. Furthermore, the user's emotions are analyzed through real-time recognition of facial expressions and voice, and this data is used to personalize the next learning session. This provides a more interactive and personalized learning experience.
[0579] In this way, the system can dynamically personalize learning by using real-time sentiment analysis, thereby improving learner engagement and learning effectiveness.
[0580] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0581] Step 1:
[0582] The server receives the user's learning theme selection information sent from the terminal. This information becomes input data for determining what kind of story to generate. Based on this input, the server prepares to retrieve corresponding materials from the database. Specifically, for example, if "Animal Science" is selected, the server generates index information to identify related content.
[0583] Step 2:
[0584] The device captures the user's facial expressions and voice in real time. This data, obtained using the camera and microphone, is input into an emotion engine, which converts the user's emotional state into numerical data. This converted numerical data is sent to a server and used to personalize the story. Specifically, if the user shows an expression of enjoyment, that data is quantified as "enjoyment."
[0585] Step 3:
[0586] The server operates a generative AI model based on the received quantified emotion information. In this process, it generates a story that matches the user's emotions, using materials selected from the database. Prompt statements are input to the generative AI model, which dynamically constructs the story. In a concrete example, the prompt "Generate a fantasy story that reflects the user's enjoyment" is used.
[0587] Step 4:
[0588] The server sends the generated story data to the device. The device receives this data and displays it to the user as a visual comic. The output is visual data of a story that is related to the user's learning theme and emotionally resonant. Specifically, a fun adventure story based on the initially selected theme, "Animal Science," is displayed on the device.
[0589] Step 5:
[0590] Users view stories provided through their devices and progress through the learning process. Any new emotional responses the user exhibits are captured again on the device and sent to the server. This feedback loop provides new learning data that will be reflected in the next story generation. Specifically, if a user expresses interest, that data is used to adjust subsequent learning content.
[0591] This series of processing steps allows the learning system to provide users with a more engaging and personalized learning experience.
[0592] (Application Example 2)
[0593] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0594] In modern work environments, workers' emotions and stress levels significantly impact productivity and safety. However, most work systems fail to consider workers' emotions and conditions, merely providing uniform workloads and instructions. As a result, workers often experience unnecessary stress and accumulate fatigue. This invention aims to provide a system that utilizes workers' emotional information in real time to generate individualized work instructions.
[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0596] In this invention, the server includes means for analyzing visual data of materials using a generative model and quantifying the emotions of each scene; means for dynamically generating work instructions suitable for the worker based on the quantified emotional information; and means for transmitting the generated work instructions to a device and displaying them visually or audibly. This makes it possible to provide individualized work instructions that take into account the emotional information of the worker.
[0597] A "generative model" is an artificial intelligence technology that learns patterns from data and generates new data.
[0598] "Visual data" refers to image and video data collected using cameras and sensors.
[0599] "Methods for quantifying emotions" refers to methods of numerically representing human emotions from visual and auditory data using an emotion engine.
[0600] "Dynamic generation methods" refer to methods that automatically create appropriate outputs in response to real-time changing situations and input data.
[0601] "Equipment" refers to machines or devices that have an interface with workers and are used to provide work instructions and information.
[0602] "Means of visual or auditory display" refers to methods of visually displaying information or playing it as sound via displays, speakers, etc.
[0603] "Worker feedback" refers to the opinions, reactions, and experiences obtained from workers during their work, and is used to improve the system and generate future work instructions.
[0604] In an embodiment of this invention, a system is provided that serves as an interface to support work within a factory. This system operates in cooperation with a server, terminals (factory robots and other display devices), and users (workers).
[0605] The server uses generative models to analyze visual data. Specifically, it quantifies workers' emotions based on image and audio data acquired from cameras and sensors. OpenCV and TensorFlow are used as the emotion engine and generative model. Other suitable libraries and frameworks can be selected as needed.
[0606] The server generates work instructions tailored to the worker based on this quantified emotional information, dynamically creating the content of these instructions. This generation model utilizes the AI framework GPT-3. The server sends the generated instructions to a terminal. The terminal displays the received instructions visually or audibly using a display or audio output device and provides them to the worker.
[0607] Workers proceed with their tasks according to the instructions provided. Emotional changes and feedback during the work are transmitted from the terminal to the server and used to adjust and improve future work instructions. This enables flexible work support tailored to the individual needs of each worker.
[0608] As a concrete example, in a factory line, a robot could scan an employee's face, and if it determines they are fatigued, it could create an instruction to temporarily slow down their work pace and announce, "Please take a short break," thereby protecting the workers' health.
[0609] Example of a prompt:
[0610] "If a worker is deemed to be experiencing stress, please generate instructions for appropriate breaks or changes in work tasks."
[0611] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0612] Step 1:
[0613] The server receives visual and audio data acquired from cameras and sensors, and analyzes the workers' facial expressions and voices. This uses emotion recognition models based on OpenCV and TensorFlow to process the data in real time. The input is image and audio data, and the output is numerical data representing the workers' emotions.
[0614] Step 2:
[0615] The server analyzes quantified sentiment data and uses the generative AI model GPT-3 to generate work instructions suitable for the worker. The input is quantified sentiment data, and the output is customized work instructions tailored to the worker's situation. This process dynamically generates text or voice instructions based on sentiment information.
[0616] Step 3:
[0617] The server sends the generated work instructions to the terminal. The terminal receives them and presents them to the worker visually or audibly via a display or audio output device. The input is the work instruction data, and the output is the instruction display or audio output to the worker.
[0618] Step 4:
[0619] The user, a worker, performs tasks based on instructions received through a terminal. During the work, the terminal collects the worker's feedback and new emotional data. This data is sent from the terminal to the server for use in generating future instructions. The input is the worker's feedback and new emotional data, and the output is the storage of this data in a database used to generate future work instructions.
[0620] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0621] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0622] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0623] [Fourth Embodiment]
[0624] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0625] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0626] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0627] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0628] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0629] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0630] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0631] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0632] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0633] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0634] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0635] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0636] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0637] This invention provides a learning system using generative models, realizing an innovative method to make exam preparation more effective and engaging. This system operates through the coordinated interaction of a server, terminals, and users.
[0638] 1. Server-side processing
[0639] The server first receives theme selection information from the user. Based on this information, the server selects appropriate comic data from its database. Using a generative model, the server analyzes the content of each panel of the selected comic data, quantifying emotions from character expressions and dialogue. This organizes the emotional information of each panel and provides data to construct the emotional flow of the story. Based on this emotional data, the server dynamically generates a story tailored to the user's profile and sends it to the terminal.
[0640] 2. Processing on the terminal side
[0641] The user's device receives story data sent from the server and displays it visually as a comic strip. Through the user interface, the user can select their preferred theme and interact with the displayed learning content. For example, if a user studying history selects this theme, the server generates a story based on historical events and displays it in comic strip format on the device, providing a visual and emotional learning experience.
[0642] 3. User-side initiatives
[0643] Users learn through comics displayed on their devices. During the learning process, users can input comments and feedback on points of interest or areas they understand better. The device sends this feedback to the server, which stores it as data to improve the user experience. This feedback allows the server to provide more suitable content to the user when generating future learning stories.
[0644] Based on the above, the present invention provides a system that connects emotions and knowledge, making test preparation more memorable, enjoyable, and efficient. This system adds new value to education and is expected to significantly improve conventional learning methods.
[0645] The following describes the processing flow.
[0646] Step 1:
[0647] The user operates the terminal's user interface to select the topic they wish to learn about. This topic selection information is then sent from the terminal to the server.
[0648] Step 2:
[0649] Based on the theme information received, the server selects the appropriate comic data from the database. This selection is important for providing stories that are suitable for the learning content.
[0650] Step 3:
[0651] The server uses a generative model to analyze selected comic data, quantifying emotions from the characters' expressions and dialogue in each panel. This organizes the emotional flow of the story as digital data.
[0652] Step 4:
[0653] Based on quantified emotional data, the server dynamically generates the most suitable story for the user. The generated story is customized to fit the user's learning profile.
[0654] Step 5:
[0655] The server sends generated story data to the terminal, which then displays it visually in comic book format. The user views the displayed comic and progresses through the learning process.
[0656] Step 6:
[0657] Users record their learning progress and provide feedback on the content via their devices. This includes entering comments and evaluating their level of understanding.
[0658] Step 7:
[0659] The device sends user feedback data to the server, which then updates its database based on that data and uses it to generate the next story.
[0660] Step 8:
[0661] The server analyzes the collected feedback to gain insights that further improve the user's learning experience. This leads to improvements across the entire system.
[0662] (Example 1)
[0663] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0664] Traditional education systems struggle to provide learning experiences that effectively incorporate learners' interests and emotions, relying instead on learning content that is difficult to remember and easily becomes boring. This has led to the problem of undermining learners' motivation to learn proactively and continuously.
[0665] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0666] In this invention, the server includes means for analyzing information data using generation technology and quantifying the emotion of each information unit; means for dynamically generating a story suitable for the learner based on the quantified emotion data; and means for transmitting the generated story to an output device and displaying it visually. This makes it possible to provide an individually optimized learning experience that reflects the learner's emotions and interests, and overcome the limitations of conventional learning methods.
[0667] An "information processing device" is a computing device that has the function of receiving data, processing it, and transmitting the data to other devices as needed.
[0668] "Generative technology" refers to techniques that use machine learning algorithms to create new data and content.
[0669] "Information data" refers to digital information content that is stored in various forms such as text, images, and audio.
[0670] An "information unit" is the smallest analyzable part that makes up information data.
[0671] "Methods for quantifying emotions" refer to the process of using computational techniques to quantify emotions within a unit of information.
[0672] A "learner" is an individual or group that uses a system to learn information.
[0673] A "story" is educational content in story format that is generated based on emotional data.
[0674] An "output device" is hardware used to provide processed information or content to the user visually or audibly.
[0675] A "user interface" is a means of interaction that allows learners to operate a system and input or select information.
[0676] "Characters" refer to the characters or subjects within a story.
[0677] To implement this invention, it is necessary to use a network-connected server and terminals and utilize generation technology that operates on an information processing device. The server analyzes data based on information received from the user and generates appropriate learning content. In this process, the server uses a generation AI model to quantify the emotions of the information data and generates personalized stories that match the user's interests and level of understanding.
[0678] The server connects to a database and retrieves material data tailored to the user's selected learning theme. A commonly available natural language processing technique is used as the specific generative model. This allows for the quantification of the emotions associated with each piece of information, and a narrative is constructed based on this quantified information.
[0679] The device receives stories sent from the server and displays them in a visually viewable format. The device's user interface is used for users to select learning themes and input feedback. This interface allows users to freely record their feelings and understandings as they progress through the learning process. This feedback becomes important input information for future story generation, leading to further improvements in educational effectiveness.
[0680] As a concrete example, when a user studying history selects a theme from a particular era, the server retrieves source data related to that era and analyzes the quantified emotional data using a generative AI model. This generates a story with an emotional flow optimized for the user. On the device, the user can visually experience this story and provide feedback on the parts that interest them.
[0681] An example of a prompt might be, "Generate a story about medieval Europe and describe it with emotion." Using this prompt, the generative AI model creates an appropriate story. In this way, the entire system works together to provide the user with an optimal learning environment.
[0682] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0683] Step 1:
[0684] The server receives learning theme selection information from the user. This input information indicates what the user wants to learn based on their interests. Based on this information, the server selects appropriate source data from the database. The selected source data will form the basis for future story generation.
[0685] Step 2:
[0686] The server applies a generative AI model to the source data retrieved from the database, analyzing and quantifying the emotions of each information unit (e.g., comic panel). In this process, information data related to the selected theme is input, and the generative AI model performs data processing by analyzing facial expressions and dialogue to quantify emotional data. The resulting emotional data is used to construct the story.
[0687] Step 3:
[0688] The server dynamically generates user-specific stories using quantified sentiment data. Here, it combines user profile information with sentiment data to create personalized stories. The output is this customized story data, which is then used in the next processing step.
[0689] Step 4:
[0690] The server sends the generated story data to the device. The transmitted data needs to be converted into a format for display on the device. The converted data becomes a comic book format to provide the user with a visual and emotional learning experience.
[0691] Step 5:
[0692] The device visually displays story data received from the server. This allows users to view the generated comics through the interface. The displayed stories can include interactive elements to help users continue learning.
[0693] Step 6:
[0694] Users learn through the displayed stories and input their thoughts and additional information as feedback on their devices. This feedback reflects the user's learning experience and deepening understanding, and serves as important data for generating future learning content.
[0695] Step 7:
[0696] The device sends feedback received from the user to the server. This feedback data is stored on the server and used in subsequent content creation processes. This feedback is essential for providing a customized learning experience that meets the user's learning needs.
[0697] (Application Example 1)
[0698] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0699] Traditional learning methods rely on text and static materials, making it difficult to engage learners. Furthermore, providing content tailored to individual learners is challenging, leading to decreased learning efficiency. Therefore, there is a need to develop learning systems that are more memorable and enjoyable.
[0700] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0701] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each fragment, means for dynamically generating a story suitable for the user based on the quantified emotion information, and means for transmitting the generated story to a display device and presenting it visually. This provides learning content that is optimal for each individual learner and enables an interesting and efficient learning experience through emotional interaction.
[0702] A "generative model" is an algorithm that can analyze data and generate new information or content.
[0703] "Information analysis" is the process of analyzing data and extracting useful patterns and features from it.
[0704] A "fragment" refers to an element or part that makes up a whole, and in this context, it refers to individual data units such as audio or images.
[0705] "Quantifying emotions" involves analyzing a character's facial expressions and dialogue to express their emotional state using numerical values.
[0706] A "story tailored to the user" refers to a story customized to a specific user's learning pace and interests.
[0707] "Dynamic generation" refers to the process of creating new content using data in real time.
[0708] A "display device" is a device used to visually present electronically generated information, such as a monitor or display.
[0709] "Visual presentation" means displaying information in an easily understandable way using images, videos, and other visual means.
[0710] "Interaction" refers to a two-way exchange between a person and a computer or content, a process in which the user inputs information and the system responds accordingly.
[0711] The system for realizing this invention is designed so that the server, terminal, and user each fulfill their respective roles. Details of each component are described below.
[0712] First, the server is equipped with a generative AI model, which is used to analyze various data. Specifically, it receives fragments of comic data and quantifies emotions from the characters' facial expressions and dialogue within them. Based on this quantified emotional information, a story optimized for the user is dynamically generated. This process utilizes cloud-based computing resources to perform advanced data processing.
[0713] Next, the device receives the generated story sent from the server and displays it visually. The device functions as a smartphone, tablet, or other display device, providing the user with an interactive learning experience. A high-resolution graphics engine is used for this visual display, generating a variety of visual content.
[0714] The user selects a specific topic they want to learn about and sends it to the server via their terminal as a prompt. The server then generates a relevant story based on this input. For example, a user interested in history might send a prompt such as, "Please generate a comic strip that tells the story of Napoleon's tactics." Upon receiving this prompt, the server generates a story according to the instructions and sends it back to the user's terminal.
[0715] This system allows learners to gain emotionally rich and engaging learning experiences that cannot be obtained from textbooks or familiar materials alone. By emphasizing emotional interaction, it is expected that the learned content will be more easily remembered.
[0716] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0717] Step 1:
[0718] The user inputs a specific subject they want to learn about using an interface on their device. This input is accompanied by prompts such as, "Generate a comic strip that tells the story of Napoleon's tactics." These prompts trigger the server to generate an appropriate story.
[0719] Step 2:
[0720] The terminal sends the prompt message received from the user to the server. This allows the server to analyze the user's request and prepare to call the appropriate database or generative AI model. The output of this process is the transmission of the prompt message to the server.
[0721] Step 3:
[0722] The server receives a prompt and uses a generative AI model to select relevant data. Specifically, it extracts Napoleon-related comic data from the database and begins analysis to generate a story based on its content. The input data is the prompt, and the output is the selection of data necessary for the analysis.
[0723] Step 4:
[0724] The server inputs the selected data into a generating AI model, quantifying the emotions of each fragment. In this step, the character's facial expressions and dialogue are analyzed, and the emotions they represent are expressed numerically. The quantified emotional data is output, and the story is dynamically generated based on this data.
[0725] Step 5:
[0726] The server generates a story tailored to the user based on quantified emotional data. This generation process also takes user profile information into account, creating individually optimized stories. The output is the generated story data.
[0727] Step 6:
[0728] The server sends the generated story data to the terminal. The terminal receives this data and displays it visually using a high-resolution graphics engine. This allows the user to experience the story in an interactive comic book format.
[0729] Step 7:
[0730] Users read through the provided story and learn as they go. They record feedback on interesting parts and points they understood, and input it into their device. This input feedback is used in the next learning cycle.
[0731] Step 8:
[0732] The device collects user feedback and sends it to the server. The server stores this feedback and uses it to generate learning content for the next cycle. As an output, valuable data is accumulated to improve learning in the next cycle.
[0733] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0734] This invention combines an emotion engine with a learning system that uses a generative model to provide a more personalized learning experience that responds to the learner's emotions. This system is intended to improve learning efficiency through close cooperation between the server, terminal, and user.
[0735] 1. Server-side processing
[0736] The server first receives theme selection information sent from the terminal and uses an emotion engine to acquire emotional information in real time from the user's facial expressions and voice. Using this emotional information, the server selects appropriate comic data from its database, analyzes the facial expressions and dialogue of the characters in each panel using a generative model, and quantifies the emotions. Based on this quantified data, it is reconstructed to generate a story that corresponds to the user's emotions. The server then sends this adjusted story to the terminal.
[0737] 2. Processing on the terminal side
[0738] The device visually displays the received story data as a comic strip and uses an emotion engine to analyze the user's facial expressions and voice. The analyzed emotion data is sent to a server and reflected in the story's development in real time. The device also presents the emotion recognition results to the user through the user interface and provides feedback for learning based on that. For example, if the user shows surprise at what is displayed, the story is adjusted to add new information and pique interest based on that emotion information.
[0739] 3. User-side initiatives
[0740] Users progress through the learning process via their device, viewing comics displayed on the device. During learning, an emotion engine recognizes the user's facial expressions and voice in real time, and the user's feedback and emotion recognition results are reflected in the learning process. By recording past emotion data, more personalized learning content will be provided in the future.
[0741] This invention aims to enhance knowledge acquisition and emotional engagement by leveraging learners' real-time emotions. It also aims to improve educational effectiveness by making the overall learning experience more individualized and dynamic.
[0742] The following describes the processing flow.
[0743] Step 1:
[0744] The user selects a topic they want to learn about through the device's user interface, and the device sends that information to the server.
[0745] Step 2:
[0746] The device activates its built-in emotion engine, analyzes the user's facial expressions and voice in real time to acquire emotional information, and sends this data to the server.
[0747] Step 3:
[0748] Based on the received theme information and user sentiment information, the server selects the appropriate comic data from the database.
[0749] Step 4:
[0750] The server uses a generative model to analyze selected comic data and quantifies the facial expressions and dialogue of the characters in each panel based on emotion categories.
[0751] Step 5:
[0752] The server dynamically generates the most suitable story for the learner based on quantified emotion data and the user's real-time emotion information, and sends it to the terminal.
[0753] Step 6:
[0754] The device displays story data received from the server in comic book format. Simultaneously, it presents the user with emotion recognition results to encourage feedback on the learning experience.
[0755] Step 7:
[0756] Users learn by reading the displayed comics via their devices. User reactions are continuously monitored by an emotion engine, and the story development is adjusted in real time as needed.
[0757] Step 8:
[0758] The device collects user feedback and emotion recognition data, sends it to a server, and updates the database. This data will be used to optimize future learning content.
[0759] (Example 2)
[0760] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] Traditional learning systems have struggled to provide personalized content that responds to learners' emotions, resulting in a failure to sustain learners' interest and motivation. There is a need for a system that enables dynamic storytelling that takes learners' real-time emotions into account.
[0762] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0763] In this invention, the server includes means for analyzing information using a generative model and quantifying the emotion of each frame, means for dynamically generating a story suitable for the learner based on the quantified emotion information, and means for transmitting the generated story to a device and displaying it visually. This makes it possible to provide personalized learning content that responds to the learner's real-time emotions.
[0764] A "generative model" is an artificial intelligence algorithm that learns specific patterns and associations from data and has the ability to make predictions and generate new instances.
[0765] "Information" refers to data used for analysis and processing within the system, such as learners' emotions and topic selection information.
[0766] A "frame" is a unit used to represent a single scene or moment within a comic or story.
[0767] "Methods for quantifying emotions" refer to methods and technologies for quantifying emotions obtained from a user's facial expressions and voice, and converting them into analyzable data.
[0768] A "story" refers to learning content presented to learners in story format, which is adjusted to suit the user's emotions.
[0769] "Device" refers to terminals or devices that users use to view or interact with learning content.
[0770] "Real-time" means that data acquisition, processing, and result presentation are all performed instantly, enabling responses that are tailored to the user's current state.
[0771] This invention is a system that uses a generative AI model to provide personalized learning content based on the user's emotions. Specific embodiments for implementing this system are described below.
[0772] Server-side embodiment
[0773] The server runs multiple software components, including a generative AI model and an emotion engine. First, the server processes the user's theme selection information received from the terminal. Then, the emotion engine analyzes the user's facial expressions and voice data sent from the terminal, quantifying their emotions. Using this emotional information, the server uses the generative AI model to select appropriate materials from the database and analyzes the character's facial expressions and dialogue in each frame. From this information, it generates a story tailored to the user and sends it to the terminal.
[0774] As a concrete example, in response to the user's chosen theme of "natural science," the server uses an emotion engine to analyze the user's emotions indicating their curiosity and adjusts the story accordingly.
[0775] An example of a prompt message is, "Generate a story that includes scientific content to stimulate the user's curiosity."
[0776] Terminal-side embodiment
[0777] The device has the ability to display story data received from the server in comic book format. It also uses its built-in camera and microphone to collect the user's facial expressions and voice in real time. This data is sent back to the server via an emotion engine and used to advance the story. The device can also present the results of emotion recognition to the user through its user interface and collect feedback.
[0778] User-side embodiment
[0779] Users participate in the system via a provided device. During learning, they can grasp the learning content through displayed comics and utilize the feedback displayed as the story progresses. Furthermore, the user's emotions are analyzed through real-time recognition of facial expressions and voice, and this data is used to personalize the next learning session. This provides a more interactive and personalized learning experience.
[0780] In this way, the system can dynamically personalize learning by using real-time sentiment analysis, thereby improving learner engagement and learning effectiveness.
[0781] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0782] Step 1:
[0783] The server receives the user's learning theme selection information sent from the terminal. This information becomes input data for determining what kind of story to generate. Based on this input, the server prepares to retrieve corresponding materials from the database. Specifically, for example, if "Animal Science" is selected, the server generates index information to identify related content.
[0784] Step 2:
[0785] The device captures the user's facial expressions and voice in real time. This data, obtained using the camera and microphone, is input into an emotion engine, which converts the user's emotional state into numerical data. This converted numerical data is sent to a server and used to personalize the story. Specifically, if the user shows an expression of enjoyment, that data is quantified as "enjoyment."
[0786] Step 3:
[0787] The server operates a generative AI model based on the received quantified emotion information. In this process, it generates a story that matches the user's emotions, using materials selected from the database. Prompt statements are input to the generative AI model, which dynamically constructs the story. In a concrete example, the prompt "Generate a fantasy story that reflects the user's enjoyment" is used.
[0788] Step 4:
[0789] The server sends the generated story data to the device. The device receives this data and displays it to the user as a visual comic. The output is visual data of a story that is related to the user's learning theme and emotionally resonant. Specifically, a fun adventure story based on the initially selected theme, "Animal Science," is displayed on the device.
[0790] Step 5:
[0791] Users view stories provided through their devices and progress through the learning process. Any new emotional responses the user exhibits are captured again on the device and sent to the server. This feedback loop provides new learning data that will be reflected in the next story generation. Specifically, if a user expresses interest, that data is used to adjust subsequent learning content.
[0792] This series of processing steps allows the learning system to provide users with a more engaging and personalized learning experience.
[0793] (Application Example 2)
[0794] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0795] In modern work environments, workers' emotions and stress levels significantly impact productivity and safety. However, most work systems fail to consider workers' emotions and conditions, merely providing uniform workloads and instructions. As a result, workers often experience unnecessary stress and accumulate fatigue. This invention aims to provide a system that utilizes workers' emotional information in real time to generate individualized work instructions.
[0796] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0797] In this invention, the server includes means for analyzing visual data of materials using a generative model and quantifying the emotions of each scene; means for dynamically generating work instructions suitable for the worker based on the quantified emotional information; and means for transmitting the generated work instructions to a device and displaying them visually or audibly. This makes it possible to provide individualized work instructions that take into account the emotional information of the worker.
[0798] A "generative model" is an artificial intelligence technology that learns patterns from data and generates new data.
[0799] "Visual data" refers to image and video data collected using cameras and sensors.
[0800] "Methods for quantifying emotions" refers to methods of numerically representing human emotions from visual and auditory data using an emotion engine.
[0801] "Dynamic generation methods" refer to methods that automatically create appropriate outputs in response to real-time changing situations and input data.
[0802] "Equipment" refers to machines or devices that have an interface with workers and are used to provide work instructions and information.
[0803] "Means of visual or auditory display" refers to methods of visually displaying information or playing it as sound via displays, speakers, etc.
[0804] "Worker feedback" refers to the opinions, reactions, and experiences obtained from workers during their work, and is used to improve the system and generate future work instructions.
[0805] In an embodiment of this invention, a system is provided that serves as an interface to support work within a factory. This system operates in cooperation with a server, terminals (factory robots and other display devices), and users (workers).
[0806] The server uses generative models to analyze visual data. Specifically, it quantifies workers' emotions based on image and audio data acquired from cameras and sensors. OpenCV and TensorFlow are used as the emotion engine and generative model. Other suitable libraries and frameworks can be selected as needed.
[0807] The server generates work instructions tailored to the worker based on this quantified emotional information, dynamically creating the content of these instructions. This generation model utilizes the AI framework GPT-3. The server sends the generated instructions to a terminal. The terminal displays the received instructions visually or audibly using a display or audio output device and provides them to the worker.
[0808] Workers proceed with their tasks according to the instructions provided. Emotional changes and feedback during the work are transmitted from the terminal to the server and used to adjust and improve future work instructions. This enables flexible work support tailored to the individual needs of each worker.
[0809] As a concrete example, in a factory line, a robot could scan an employee's face, and if it determines they are fatigued, it could create an instruction to temporarily slow down their work pace and announce, "Please take a short break," thereby protecting the workers' health.
[0810] Example of a prompt:
[0811] "If a worker is deemed to be experiencing stress, please generate instructions for appropriate breaks or changes in work tasks."
[0812] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0813] Step 1:
[0814] The server receives visual and audio data acquired from cameras and sensors, and analyzes the workers' facial expressions and voices. This uses emotion recognition models based on OpenCV and TensorFlow to process the data in real time. The input is image and audio data, and the output is numerical data representing the workers' emotions.
[0815] Step 2:
[0816] The server analyzes quantified sentiment data and uses the generative AI model GPT-3 to generate work instructions suitable for the worker. The input is quantified sentiment data, and the output is customized work instructions tailored to the worker's situation. This process dynamically generates text or voice instructions based on sentiment information.
[0817] Step 3:
[0818] The server sends the generated work instructions to the terminal. The terminal receives them and presents them to the worker visually or audibly via a display or audio output device. The input is the work instruction data, and the output is the instruction display or audio output to the worker.
[0819] Step 4:
[0820] The user, a worker, performs tasks based on instructions received through a terminal. During the work, the terminal collects the worker's feedback and new emotional data. This data is sent from the terminal to the server for use in generating future instructions. The input is the worker's feedback and new emotional data, and the output is the storage of this data in a database used to generate future work instructions.
[0821] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0822] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0823] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0824] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0825] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0826] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0827] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0828] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0829] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0830] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0831] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0832] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0833] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0834] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0835] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0836] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0837] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0838] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0839] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0840] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0841] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0842] The following is further disclosed regarding the embodiments described above.
[0843] (Claim 1)
[0844] A method for analyzing comic data using a generative model and quantifying the emotions of each panel,
[0845] A means for dynamically generating stories suitable for learners based on quantified emotional information,
[0846] A means of sending the generated story to a device and displaying it visually,
[0847] A system that includes means for collecting user feedback and storing that data for use in generating future learning content.
[0848] (Claim 2)
[0849] The system according to claim 1, which provides a user interface for selecting a learning theme and transmits the selected theme information to a server.
[0850] (Claim 3)
[0851] The system according to claim 1, which constructs the emotional flow of the entire story by generating numerical values based on emotion categories from the characters' lines and facial expressions.
[0852] "Example 1"
[0853] (Claim 1)
[0854] An information processing device analyzes information data using generation technology and quantifies the emotion of each information unit,
[0855] A means of dynamically generating a story suitable for learners based on quantified emotional data,
[0856] A means for transmitting the generated story to an output device and displaying it visually,
[0857] A system that includes means for collecting user feedback and storing that information for use in generating learning materials for the next time.
[0858] (Claim 2)
[0859] The system according to claim 1, which provides a user interface for selecting information to be learned and transmits the selected information to an information processing device.
[0860] (Claim 3)
[0861] The system according to claim 1, which constructs the overall emotional flow of a story by generating numerical values based on emotional classification from the statements and expressions of the characters.
[0862] "Application Example 1"
[0863] (Claim 1)
[0864] A method for analyzing information using a generative model and quantifying the emotions of each fragment,
[0865] A means of dynamically generating a story suitable for the user based on quantified emotional information,
[0866] A means of transmitting the generated story to a display device and presenting it visually,
[0867] A means of collecting user feedback and storing that data to use in generating future learning content,
[0868] A means of supporting a deeper understanding of knowledge by using stories generated based on the subject matter that learners want to learn,
[0869] A means of adjusting the context of the story generated using prompt statements,
[0870] A system that includes this.
[0871] (Claim 2)
[0872] The system according to claim 1, which provides a user interface for selecting a learning topic and transmits the selected topic information to a base device.
[0873] (Claim 3)
[0874] The system according to claim 1, which generates numerical values based on emotional categories from the dialogue and facial expressions of the characters, thereby constructing the emotional flow of the entire story and enhancing the emotional learning experience.
[0875] "Example 2 of combining an emotion engine"
[0876] (Claim 1)
[0877] A method for analyzing information using a generative model and quantifying the emotions of each frame,
[0878] A means for dynamically generating a story suitable for learners based on quantified emotional information,
[0879] A means for transmitting the generated story to a device and displaying it visually,
[0880] A means of obtaining user responses and saving that data for use in generating the next learning content,
[0881] A means of analyzing the user's facial expressions and voice in real time to acquire emotional information,
[0882] A means of selecting materials corresponding to emotional information from storage methods,
[0883] A system that includes this.
[0884] (Claim 2)
[0885] The system according to claim 1, which provides a user interface for selecting a learning topic and transmits the selected topic information to a server.
[0886] (Claim 3)
[0887] The system according to claim 1, which constructs the overall emotional flow of a story by generating numerical values based on emotional categories from the characters' statements and facial expressions.
[0888] "Application example 2 when combining with an emotional engine"
[0889] (Claim 1)
[0890] A method for analyzing the visual data of materials using a generative model and quantifying the emotions of each scene,
[0891] A means for dynamically generating work instructions suitable for workers based on quantified emotional information,
[0892] A means for transmitting the generated work instructions to the device and displaying them visually or audibly,
[0893] A system that includes means for collecting worker feedback and storing that data for use in generating future work instructions.
[0894] (Claim 2)
[0895] The system according to claim 1, which provides an operation screen for selecting a work theme and transmits the selected work information to a data processing device.
[0896] (Claim 3)
[0897] The system according to claim 1, which constructs the emotional flow of the entire task by generating numerical values based on emotional categories from the dialogue and facial expressions of the person performing the action. [Explanation of Symbols]
[0898] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A method for analyzing comic data using a generative model and quantifying the emotions of each panel, A means for dynamically generating stories suitable for learners based on quantified emotional information, A means of sending the generated story to a device and displaying it visually, A system that includes means for collecting user feedback and storing that data for use in generating future learning content.
2. The system according to claim 1, which provides a user interface for selecting a learning theme and transmits the selected theme information to a server.
3. The system according to claim 1, which constructs the emotional flow of the entire story by generating numerical values based on emotional categories from the characters' lines and facial expressions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A