system

The system addresses language and cultural barriers by integrating multimodal technology and emotion recognition to provide personalized, interactive learning experiences, enabling access to diverse educational content and real-time engagement.

JP2026068329APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional education systems fail to provide learners with diverse educational content accessible across languages and cultures, lack personalization based on learning progress, and do not support real-time interactive learning experiences.

Method used

A system that integrates multimodal technology to collect, translate, and personalize educational content, providing real-time interactive learning through user interfaces and emotion recognition, allowing learners to access diverse content in their native language and engage in live sessions.

Benefits of technology

Enables learners to access diverse educational resources globally, overcome language barriers, receive personalized recommendations, and participate in interactive learning experiences, enhancing the educational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068329000001_ABST
    Figure 2026068329000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for acquiring educational content in various formats, A means for integrating and transforming the aforementioned educational content using multimodal technology, A means for translating the aforementioned educational content into different languages ​​and generating subtitles, A method for analyzing a user's learning history and recommending the next educational content they should learn, A means of providing educational events that users can participate in in real time, A means for delivering educational content to the user's terminal and providing a user interface, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional education system, the mechanism for learners to access diverse educational contents from all over the world according to their own interests and needs and for learning to be individualized was insufficient. Also, due to the barriers of different languages and cultures, it was difficult to effectively learn educational contents of a specific region or country. Furthermore, it was also difficult to recommend flexible contents according to the progress of learning and to conduct real - time interactive learning.

Means for Solving the Problems

[0005] This invention provides a system that enables learners to acquire diverse forms of educational content as integrated content using multimodal technology, and to learn smoothly by translating it into different languages. Furthermore, it promotes a personalized learning experience by recommending the next learning content through analysis of learning history and providing real-time, interactive educational events. In this way, it realizes comprehensive educational support that meets the specific needs of learners.

[0006] "Diverse forms of educational content" refers to the collective term for educational information provided in different media formats, such as videos, audio, text, and images.

[0007] "Multimodal technology" refers to technologies that integrate data of different formats and types (e.g., audio, text, images) to enable comprehensive analysis and understanding.

[0008] "Integrating and converting" is the process of combining multiple different forms of information into a single, consistent format, making it usable.

[0009] Translation is the act of converting content expressed in one language into another language, thereby accurately conveying the original information.

[0010] "Generating subtitles" means creating and displaying supplementary text information to accompany audio or visual content.

[0011] "Learning history" refers to the records and data of what educational content a user has studied in the past.

[0012] An "educational event" is a general term for activities held in real time for educational purposes, such as lectures, seminars, and discussions.

[0013] "User interface" refers to the design and elements of the screens and operations that users use to interact with a system or application. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] The educational support system according to the present invention is a platform that allows learners to easily access diverse educational content from around the world and provides an individualized learning experience. This system consists primarily of a server, terminals, and a user interface.

[0036] The server collects content from various educational providers. The collected content comes in diverse formats, including video, audio, text, and images, and is integrated using multimodal technology. For example, the server receives a video lecture by an Indian mathematician and converts it from audio to text as needed, making the video content available as text.

[0037] Furthermore, the server uses AI technology to translate educational content into different languages. This includes generating subtitles using speech recognition. For example, when an educational lecture provided in English is translated into Japanese, Japanese subtitles are automatically generated. This allows users to learn without language barriers.

[0038] The device provides learners with access to content through a user interface. Users can search for topics they wish to learn about via the device and access appropriate educational content provided by the server. For example, if a user is interested in "European art history," the device will display a list of related lectures and materials and play the selected one.

[0039] The server also records the learner's learning history and uses this to recommend the next content they should study. The learning history is generated based on the user's past viewing history and interests, and supplementary materials are also provided. For example, if a user has watched lectures related to a specific historical event, related continuing learning topics will be automatically recommended.

[0040] Furthermore, this system offers live sessions, allowing users to participate in real time. During live sessions, users can join streaming lectures via their devices and ask questions directly to the instructor.

[0041] Thus, the embodiment of the present invention aims to address the internationalization of education and individual needs by centrally managing different forms of educational content and creating individualized learning experiences.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The server collects educational content provided by educational providers. During collection, it uses APIs and feeds to register metadata (title, category, language, etc.) for each piece of content in a database.

[0045] Step 2:

[0046] The server applies multimodal technology to the collected content, integrating different data formats (video, audio, text, images, etc.) into a consistent format. This makes the content available in a variety of formats.

[0047] Step 3:

[0048] The server uses an AI translation engine to translate collected content into the learner's native language. It also uses speech recognition technology to generate subtitles for video and audio content.

[0049] Step 4:

[0050] Users log in to the platform via their device and search for educational content that interests them. The device then displays a list of relevant content to the user through its user interface.

[0051] Step 5:

[0052] Based on a request from the device, the server sends the selected educational content to the device and makes it playable. This allows the user to view the content in real time.

[0053] Step 6:

[0054] The server recommends the next content a user should learn based on their learning history and current viewing habits. This involves analyzing past history and extracting highly relevant topics.

[0055] Step 7:

[0056] Users participate in live classes through their devices. The server delivers classes in real time using streaming technology and provides an interactive platform for users to ask questions and participate in discussions.

[0057] Step 8:

[0058] After a learning session ends, users enter feedback using a terminal. The server collects this feedback and analyzes it to help improve the system in the future.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In today's educational environment, challenges include access to diverse forms of educational content, learning across language barriers, and providing personalized learning experiences. In particular, there is a need for equitable use of content from around the world, multilingual support, and the suggestion of optimal learning content tailored to the user's learning progress. Furthermore, supporting participation in real-time educational events and creating an interactive learning environment are also essential.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for collecting information, means for integrating the information using multimodal technology, and means for converting speech to text. This makes it possible to integrate diverse forms of content, convert them to different languages, and provide users with a personalized learning experience. It also supports real-time participatory events, enabling an interactive educational environment.

[0064] "Information" refers to a collection of data and knowledge expressed in various forms, including educational content.

[0065] "Means of collection" refers to methods or techniques for efficiently obtaining information from diverse sources.

[0066] "Multimodal technology" is a technology that integrates and processes information in different formats, making data such as video, audio, and text available on a single platform.

[0067] "Means for converting speech to text" refers to technologies that automatically convert speech data into text and express it in text format.

[0068] "Means of translation and subtitle generation" refers to the technology of translating information into another language and creating subtitles to display that translation.

[0069] "Means for recording and analyzing user history" refers to technologies that record past operations and selections and analyze them to enable personalized content recommendations.

[0070] "Means of providing events that users can participate in in real time" refers to the technology and infrastructure that allows participants to instantly access and interactively engage with live events and lectures.

[0071] "Means for distributing information to a user's device and providing an interface" refers to technologies that transmit acquired information to a user's device and provide screens and functions that the user can operate on that device.

[0072] This invention aims to realize an educational support system that provides learners with diverse educational content. This system mainly consists of three components: a server, a terminal, and a user interface.

[0073] The server plays a central role in collecting educational content from various sources. Specifically, it retrieves data in formats such as video, audio, text, and images through APIs and data feeds. For example, it might download a physics lecture video from an online education platform. Multimodal technology is used in this process to integrate data in different formats.

[0074] Furthermore, the server uses speech recognition software to convert the audio of the lecture videos into text in real time. This feature allows users to understand the lecture content not only through audio but also in text format. Simultaneously, the server uses a natural language processing (NLP) engine to perform multilingual translation and automatically generate subtitles. For example, it is possible to translate a science lecture delivered in English into Japanese and provide subtitles.

[0075] The terminal provides a user interface, helping learners easily access content. Users can use the terminal to search for topics of interest and select relevant educational content provided by the server. For example, if a user is interested in "European art history," the search results will display a list of related lectures and materials. From these, the user can choose any content they wish to learn.

[0076] Furthermore, the server records and analyzes the user's learning history and recommends the most suitable next learning content for each individual. This allows learners to deepen their knowledge efficiently and effectively. For example, a user who has watched past lectures related to historical events will be recommended related topics to continue learning.

[0077] The system also provides live streaming lectures in real time as live sessions. By participating using a device, users can send questions to the instructor during the lecture, creating an interactive learning experience.

[0078] A concrete example of a prompt would be, "I would like to translate a new English science lecture into Japanese and provide it with subtitles." This invention will enable learners to access diverse educational resources from around the world and learn across language and format barriers.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The server collects educational content from information sources. Inputs include data in various formats such as video, audio, text, and images, received via APIs and data feeds. These data are integrated using multimodal technology and combined into a single dataset. This process results in an output where data of different formats can be managed in a consistent manner.

[0082] Step 2:

[0083] The server extracts audio data from an integrated dataset and converts it to text using speech recognition software. The input is the audio data from lecture videos. Through this process, the audio information is converted into text output, generating a visually verifiable lecture content.

[0084] Step 3:

[0085] The server uses a generative AI model to translate text into different languages ​​and generate subtitles. The input is lecture content in text format. This translation process outputs educational content with subtitles that enables learning in multiple languages.

[0086] Step 4:

[0087] The device searches for and displays collected and processed content through its user interface. Input consists of the user's interests and search keywords. Based on this, relevant content is displayed, and the user can select from it to proceed with their learning.

[0088] Step 5:

[0089] The server analyzes the user's learning history and recommends the next content they should learn. The input consists of past viewing history and learning content. Based on this, an AI algorithm generates personalized recommendations, outputting content suitable for the next learning session.

[0090] Step 6:

[0091] The terminal provides live sessions and creates an environment for users to participate in real time. The input is information about scheduled live lectures. The output allows users to access the streamed lectures and interact in real time.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In today's educational environment, it is difficult for learners to access diverse information from around the world and obtain personalized learning experiences. Furthermore, support for providing appropriate information tailored to learners' interests and for participating in immediate, interactive learning events is limited. In this context, there is a need to realize educational experiences and information provision optimized for learners.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for integrating and transforming information using multi-format technology, means for distributing information to the user's portable terminal and providing an operation screen, and means for providing relevant information based on topics of interest to the user. This enables learner-specific information access and personalized educational experiences, and facilitates immediate, participatory learning.

[0097] "Information in diverse formats" refers to content provided in multiple different formats, such as video, audio, text, and images.

[0098] "Multi-format technology" refers to technologies that integrate and convert content in different formats, thereby enabling the consistent delivery of information.

[0099] "Means for generating subtitles" refers to methods for converting audio content into text and displaying it as text information on a display device.

[0100] "User" refers to an individual or group that receives information and engages in learning activities through this system.

[0101] "Learning history" refers to a record of content that a user has accessed in the past and topics they have been interested in.

[0102] An "immediately accessible educational event" refers to an interactive learning session that users can participate in in real time.

[0103] A "portable device" refers to an electronic device that can be easily carried by the user, such as a mobile phone or tablet.

[0104] An "operation screen" refers to a screen display that provides an interface for users to search for and access information.

[0105] "Related information" refers to information, materials, or content related to topics that the user is interested in.

[0106] The system according to the present invention consists of a server, a terminal, and a user.

[0107] The server collects diverse forms of information from educational providers and employs multi-format technologies to integrate content in different formats, such as video, audio, text, and images. This is possible, for example, by converting audio accompanying videos into text and generating subtitles using speech recognition and translation technologies so that even foreign language lectures can be easily understood. The content is processed using specific speech recognition and translation technologies, such as the Google® Speech-to-Text API and the Google Cloud Translation API.

[0108] The terminal is a portable device that learners can carry with them and displays an operation screen based on information delivered from the server. Through this operation screen, users can search for topics based on their interests and access related information. In addition, personalized information is provided based on their learning history, enabling them to experience interactive learning.

[0109] Users can participate in real-time educational events using an on-screen interface on their devices. This allows them to access educational resources worldwide and broaden their learning experience. Furthermore, by applying AI technology, content based on their interests can be suggested, effectively supporting their learning progress.

[0110] For example, if a user expresses interest in "Italian Renaissance art," they can use their device to access relevant educational content and real-time sessions, and watch lectures from experts around the world. Furthermore, a generative AI model can be used to recommend relevant materials based on prompts. An example of a prompt would be: "Recommend relevant content based on the topic selected by the user. For example, if 'Italian Renaissance art' is selected, instruct the system to display related videos and materials."

[0111] In this way, the system of the present invention provides learners with an individualized and enriching learning experience and a means of easily accessing educational resources.

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] The server receives information in various formats from educational providers. Specifically, it collects content such as videos, audio, text, and images and stores them in a database. The input includes the entire educational content, and the output is an integrated dataset. To process this data efficiently, the server uses a database management system to organize the storage format and prepare it for subsequent processing.

[0115] Step 2:

[0116] The server integrates and transforms the received information using multi-format technologies. Specifically, it uses the Google Speech-to-Text API to convert speech to text. The input is the corresponding audio data, and the output is the corresponding text data. The server transforms these formats and integrates them into a consistent format, enabling further translation processing for multiple languages.

[0117] Step 3:

[0118] The server translates the integrated information into different languages ​​and generates subtitles. Using the Google Cloud Translation API, the input text is converted into various target languages. The output generated by this process is text and subtitle data in the target languages. The server uses this data to generate multilingual content.

[0119] Step 4:

[0120] The server analyzes the user's learning history and uses an AI model to recommend the next information to learn. By generating prompts and inputting them into the generating AI model, it suggests the most suitable content based on the user's past viewing history and interests. The input is the learning history and its analysis results, and the output is personalized recommendations for the user.

[0121] Step 5:

[0122] The terminal receives information sent from the server and displays an operation screen. Here, users can search for topics based on their interests and access related content and information recommended by the server. Input is data delivered from the server, and output is information dynamically displayed on the terminal. Specifically, the terminal provides an intuitive operation screen based on UI / UX design.

[0123] Step 6:

[0124] Users participate in educational events that are instantly accessible via their devices, engaging in interactive discussions. They can ask questions and join discussions in real time. Input is streaming data received in real time from the server, while output is participant feedback and questions. Through this, users gain a deeper understanding and knowledge.

[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0126] The educational support system according to the present invention is a platform that provides a more personalized learning experience by combining an emotion engine with existing diverse forms of educational content. This system utilizes servers, terminals, and emotion recognition technology to realize dynamic educational delivery that responds to the user's emotions.

[0127] The server continuously collects diverse content formats from educational providers. This includes video, audio, text, and images, which are integrated using multimodal technology. Content in different languages ​​is translated into the user's native language using AI translation technology, and subtitles are generated for audio content.

[0128] This system incorporates an emotion engine that recognizes the user's emotions through the device's camera and microphone during learning. The emotion engine analyzes the user's facial expressions and tone of voice to calculate their emotional state in real time. For example, if the user shows signs of waning concentration while learning a difficult topic, the system will detect this and either lower the difficulty level of the content or display an encouraging message.

[0129] Users access educational content via their devices, and based on the analysis results of the sentiment engine, recommended content is automatically adjusted. The server adds sentiment data to the user's learning history and uses this data to provide feedback on the next learning topic and to boost motivation. For example, if a user shows particular interest in a topic but struggles to understand it, the system analyzes their level of understanding from the sentiment data and presents similar content using different approaches.

[0130] Even during live sessions, the emotion engine remains active, allowing instructors to understand participants' emotional states in real time. This enables them to adjust the pace of the lesson, ask specific questions, and create an interactive environment.

[0131] Thus, the present invention aims to support efficient and effective learning by utilizing emotion recognition technology to provide an optimized educational experience for each user.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] The server collects educational content in video, audio, text, and image formats from registered educational providers and public libraries. The collected content is stored in a database along with metadata.

[0135] Step 2:

[0136] The server uses multimodal technology to integrate the collected educational content and prepare it for delivery in a consistent format. This includes audio-to-text conversion and image analysis and descriptive text generation.

[0137] Step 3:

[0138] The server uses an AI translation engine and speech recognition technology to translate educational content into the user's native language and generate subtitles that correspond to the video and audio.

[0139] Step 4:

[0140] The user logs into the device and searches for a topic they want to learn about. The device displays a list of relevant content through the user interface.

[0141] Step 5:

[0142] When a user starts playing content, the device uses its built-in camera and microphone to record the user's facial expressions and voice, and sends this information to the emotion engine.

[0143] Step 6:

[0144] The emotion engine analyzes user emotions from facial and voice data and calculates parameters such as concentration, interest, and fatigue. These results are sent to the server in real time.

[0145] Step 7:

[0146] Based on data from the emotion engine, the server adjusts content to optimize the user's learning experience. For example, if a user's concentration wanes, it may display a message on the device prompting them to take a break or prioritize recommending relevant, lighter topics.

[0147] Step 8:

[0148] When a user participates in a live session, they join the streaming via their device, and emotion engine data is provided to the instructor in real time. This allows the instructor to understand the emotional trends of the session and adjust the lesson accordingly.

[0149] Step 9:

[0150] After the learning process is complete, the device displays an interface requesting feedback from the user. The collected feedback is sent to a server and used for continuous improvement of the system.

[0151] (Example 2)

[0152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0153] In today's educational environment, it is difficult to efficiently integrate diverse forms of educational content and provide experiences that meet the individual learning needs of users. Furthermore, conventional technologies are insufficient for effectively managing educational events that take into account users' emotions in real time. As a result, there is a problem in providing education that is appropriate for individual learners and improving learning effectiveness.

[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0155] In this invention, the server includes means for acquiring educational content in various formats, means for integrating and transforming educational content using multimodal technology, means for translating educational content into different languages ​​and generating subtitles, means for analyzing the user's learning history and emotional state and recommending the next educational content to be learned, means for providing educational events that can be participated in in real time and for acquiring and analyzing participants' emotional feedback in real time, and means for delivering educational content to the user's terminal, providing a user interface, and personalizing the learning experience based on the user's emotional state. This makes it possible to efficiently utilize educational content in various formats and provide a personalized learning experience.

[0156] "Educational content in diverse formats" refers to educational materials that include different media formats such as videos, audio, text, and images.

[0157] "Multimodal technology" is a technology that integrates data from multiple different media formats and converts it into a common information representation.

[0158] "AI translation technology" is a technology that uses artificial intelligence to automatically translate text and audio between different languages.

[0159] An "emotion engine" is a technology that analyzes a user's emotions from their facial expressions and tone of voice, and calculates their emotional state in real time.

[0160] "Learning history" refers to information about what a user has learned in the past and their progress.

[0161] A "recommendation mechanism" is a system that selects or suggests what a user should learn next based on their learning history and emotional state.

[0162] "Methods for acquiring and analyzing data in real time" refers to technologies that collect data immediately during educational events and perform analysis on the spot.

[0163] A "user interface" is the operating environment through which a user interacts with a system via a terminal and utilizes educational content.

[0164] A "personalized learning experience" refers to the provision of education that is customized according to the individual learning needs and circumstances of each user.

[0165] This educational support system consists of servers, terminals, and emotion recognition technology. The servers continuously collect educational content in various formats provided by educational providers. This collection uses multimodal technology to integrate data from different media formats and convert it into a user-friendly format. For example, it collects English videos, translates them into the user's native language, and generates subtitles.

[0166] The server utilizes AI translation technology to convert content in different languages ​​into the user's native language and adds subtitles to audio content. It also incorporates an emotion engine that analyzes the user's emotions through the device. The resulting emotion data is integrated with the user's learning history and used to determine what to learn next and to provide personalized feedback.

[0167] Users can access educational content through their devices. The devices use cameras and microphones to collect real-time emotional data from the user's facial expressions and tone of voice, and the system uses this information to personalize the learning experience. For example, if a user is feeling stressed while working on a particularly difficult topic, the system will provide simpler explanations or encouraging messages.

[0168] This system captures emotional feedback even during live sessions, allowing instructors to understand participants' emotions in real time and adjust the lesson accordingly. For example, instructors can slow the pace or provide supplementary explanations if participants lose focus.

[0169] As a concrete example, if a user uses a "generative AI model" and inputs a prompt such as "How do I solve a specific problem in mathematical algebra?", the system will use this information to recommend the most suitable learning materials and explanatory videos to the user. In this way, it is possible to provide an optimal educational experience that meets the individual needs of the user.

[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0171] Step 1:

[0172] The server receives diverse forms of educational content from educational providers as input. This input data includes video, audio, text, and images. These are integrated using multimodal technology and converted into a user-friendly format. The output is the integrated content data. Specifically, the server analyzes data from each media format and compiles it into a consistent information package.

[0173] Step 2:

[0174] The server processes integrated educational content using AI translation technology. The input is educational content written in different languages. Through the translation process, the output is text and audio content translated into the user's native language. Specifically, the server automatically adds subtitles to audio content and provides user-selectable subtitle options for videos.

[0175] Step 3:

[0176] The device activates an emotion engine to monitor the user's learning progress. Inputs include the user's facial expressions and voice tone, captured through the device's camera and microphone. This data is analyzed, and real-time emotion data is output. Specifically, the device applies an emotion analysis algorithm to detect different emotional states.

[0177] Step 4:

[0178] The server receives the user's learning history and sentiment data as input and recommends the next educational content to learn. In this step, data analysis determines the appropriate learning materials and learning paths, and provides the user with a list of recommended content as output. Specifically, the server infers the user's interests and level of understanding from past learning data and sentiment feedback, and selects the most suitable learning content.

[0179] Step 5:

[0180] Users access personalized educational content delivered from the server and progress through their learning. At this stage, recommended content serves as input, and the output includes the user's learning progress and new learning history data. Specifically, the user interacts with the learning environment through their device, selecting content as needed.

[0181] Step 6:

[0182] The device generates feedback that allows the instructor to understand participants' emotional states and adjust the lesson pace based on emotional data acquired in real time during live sessions. This feedback provides the instructor with clues to appropriately adjust the pace of the lesson and contributes to interactive lesson management. The input is real-time emotional data from participants, and the output is instructions and suggestions for improvement regarding the lesson pace. Specifically, this could involve the instructor observing participants' reactions and asking questions or suggesting reviewing slides.

[0183] (Application Example 2)

[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0185] In recent years, improving worker performance and safety in factory settings and manufacturing industries has required understanding workers' emotional states in real time and providing appropriate support. However, conventional methods have struggled to accurately recognize workers' emotions and provide immediate, personalized feedback. To address this challenge, there is a need for dynamic support systems that utilize emotion recognition technology in the workplace.

[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0187] In this invention, the server includes means for acquiring knowledge bases in various forms, means for integrating and transforming the knowledge bases using multimodal technology, and means for detecting the user's emotional state during operation and optimizing individual work support. This enables real-time work support based on the worker's emotions.

[0188] A "diverse knowledge base" refers to information or datasets provided in different media formats, including text, images, audio, and video.

[0189] "Multimodal technology" is a technique that integrates and comprehensively analyzes data in different formats, thereby enabling a richer understanding of the data.

[0190] An "information terminal" is a device used by a user, and this includes smartphones, tablets, computers, and wearable devices.

[0191] "Emotional state" refers to the internal psychological state an individual experiences at a given point in time, and includes emotions such as joy, anger, anxiety, and surprise.

[0192] "Optimizing individual work support" means improving work efficiency and safety by providing the most appropriate work instructions and feedback tailored to each worker's current situation and feelings.

[0193] The system implementing this invention consists of a user-worn information terminal, a cloud server, and emotion recognition technology. The information terminal is a wearable device such as smart glasses, equipped with a camera and microphone. This allows for real-time capture of the worker's facial expressions and voice tone, and the acquisition of data.

[0194] The server receives this data and analyzes changes in facial expressions using face recognition libraries such as OpenCV. Simultaneously, it analyzes the tone of voice from the audio data using Google Cloud Speech-to-Text. These analysis results are input into emotion analysis models using TENSORFLOW® or Keras to calculate the worker's emotional state.

[0195] Based on this emotional state information, the server delivers the next task to be performed and encouraging feedback to the user's information terminal, optimizing individual work support. This allows workers to receive support tailored to their emotions, improving work efficiency and safety.

[0196] As a concrete example, consider a scenario where a factory worker is performing a difficult assembly task while wearing smart glasses. If this worker begins to feel stressed, an emotion analysis model detects this, and the server immediately provides feedback based on that information, such as "Take a deep breath and calm down."

[0197] Examples of input prompts for the generating AI model include phrases such as, "What encouraging message should be displayed if the user looks anxious?" or "How should the work procedure be adjusted if the user's tone of voice is different from usual?" This makes it possible to create a flexible and user-friendly work environment.

[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0199] Step 1:

[0200] The terminal acquires the worker's facial expressions and voice. The terminal uses a camera to capture the worker's facial expressions and a microphone to record their voice. This data is sent as input to the next processing step.

[0201] Step 2:

[0202] The server analyzes the facial expression data. Using the facial image data sent from the terminal, it performs facial recognition through the OpenCV library. As output, feature data extracted from the facial expressions is generated. This data serves as basic information for identifying the emotional state of the worker.

[0203] Step 3:

[0204] The server analyzes the audio data. The audio data provided by the terminal is converted to text using Google Cloud Speech-to-Text, and then the tone of the speech is analyzed. The output provides indicators of speech intensity and emotion.

[0205] Step 4:

[0206] The server analyzes the emotional state. It integrates the previously obtained facial feature data and voice tone data, and uses a TensorFlow-based emotion analysis model to estimate the worker's current emotional state. The output is a set of indicators representing the worker's emotional state.

[0207] Step 5:

[0208] The server generates feedback. Based on the emotional state obtained from sentiment analysis, it uses a generative AI model to devise appropriate feedback and work instructions. It creates messages based on questions such as, "What kind of encouraging message should be displayed if the user has an anxious expression?" as a prompt.

[0209] Step 6:

[0210] The terminal displays feedback. Feedback sent from the server is displayed on the worker's terminal, providing real-time support while preventing work interruptions. This allows workers to receive appropriate advice tailored to their emotional state.

[0211] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0212] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0213] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0214] [Second Embodiment]

[0215] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0216] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0217] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0218] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0219] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0221] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0222] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0223] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0224] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0225] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0226] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0227] The educational support system according to the present invention is a platform that allows learners to easily access diverse educational content from around the world and provides an individualized learning experience. This system consists primarily of a server, terminals, and a user interface.

[0228] The server collects content from various educational providers. The collected content comes in diverse formats, including video, audio, text, and images, and is integrated using multimodal technology. For example, the server receives a video lecture by an Indian mathematician and converts it from audio to text as needed, making the video content available as text.

[0229] Furthermore, the server uses AI technology to translate educational content into different languages. This includes generating subtitles using speech recognition. For example, when an educational lecture provided in English is translated into Japanese, Japanese subtitles are automatically generated. This allows users to learn without language barriers.

[0230] The device provides learners with access to content through a user interface. Users can search for topics they wish to learn about via the device and access appropriate educational content provided by the server. For example, if a user is interested in "European art history," the device will display a list of related lectures and materials and play the selected one.

[0231] The server also records the learner's learning history and uses this to recommend the next content they should study. The learning history is generated based on the user's past viewing history and interests, and supplementary materials are also provided. For example, if a user has watched lectures related to a specific historical event, related continuing learning topics will be automatically recommended.

[0232] Furthermore, this system offers live sessions, allowing users to participate in real time. During live sessions, users can join streaming lectures via their devices and ask questions directly to the instructor.

[0233] Thus, the embodiment of the present invention aims to address the internationalization of education and individual needs by centrally managing different forms of educational content and creating individualized learning experiences.

[0234] The following describes the processing flow.

[0235] Step 1:

[0236] The server collects educational content provided by educational providers. During collection, it uses APIs and feeds to register metadata (title, category, language, etc.) for each piece of content in a database.

[0237] Step 2:

[0238] The server applies multimodal technology to the collected content, integrating different data formats (video, audio, text, images, etc.) into a consistent format. This makes the content available in a variety of formats.

[0239] Step 3:

[0240] The server uses an AI translation engine to translate collected content into the learner's native language. It also uses speech recognition technology to generate subtitles for video and audio content.

[0241] Step 4:

[0242] Users log in to the platform via their device and search for educational content that interests them. The device then displays a list of relevant content to the user through its user interface.

[0243] Step 5:

[0244] Based on a request from the device, the server sends the selected educational content to the device and makes it playable. This allows the user to view the content in real time.

[0245] Step 6:

[0246] The server recommends the next content a user should learn based on their learning history and current viewing habits. This involves analyzing past history and extracting highly relevant topics.

[0247] Step 7:

[0248] Users participate in live classes through their devices. The server delivers classes in real time using streaming technology and provides an interactive platform for users to ask questions and participate in discussions.

[0249] Step 8:

[0250] After a learning session ends, users enter feedback using a terminal. The server collects this feedback and analyzes it to help improve the system in the future.

[0251] (Example 1)

[0252] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0253] In today's educational environment, challenges include access to diverse forms of educational content, learning across language barriers, and providing personalized learning experiences. In particular, there is a need for equitable use of content from around the world, multilingual support, and the suggestion of optimal learning content tailored to the user's learning progress. Furthermore, supporting participation in real-time educational events and creating an interactive learning environment are also essential.

[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0255] In this invention, the server includes means for collecting information, means for integrating the information using multimodal technology, and means for converting speech to text. This makes it possible to integrate diverse forms of content, convert them to different languages, and provide users with a personalized learning experience. It also supports real-time participatory events, enabling an interactive educational environment.

[0256] "Information" refers to a collection of data and knowledge expressed in various forms, including educational content.

[0257] "Means of collection" refers to methods or techniques for efficiently obtaining information from diverse sources.

[0258] "Multimodal technology" is a technology that integrates and processes information in different formats, making data such as video, audio, and text available on a single platform.

[0259] "Means for converting speech to text" refers to technologies that automatically convert speech data into text and express it in text format.

[0260] "Means of translation and subtitle generation" refers to the technology of translating information into another language and creating subtitles to display that translation.

[0261] "Means for recording and analyzing user history" refers to technologies that record past operations and selections and analyze them to enable personalized content recommendations.

[0262] "Means of providing events that users can participate in in real time" refers to the technology and infrastructure that allows participants to instantly access and interactively engage with live events and lectures.

[0263] "Means for distributing information to a user's device and providing an interface" refers to technologies that transmit acquired information to a user's device and provide screens and functions that the user can operate on that device.

[0264] This invention aims to realize an educational support system that provides learners with diverse educational content. This system mainly consists of three components: a server, a terminal, and a user interface.

[0265] The server plays a central role in collecting educational content from various sources. Specifically, it retrieves data in formats such as video, audio, text, and images through APIs and data feeds. For example, it might download a physics lecture video from an online education platform. Multimodal technology is used in this process to integrate data in different formats.

[0266] Furthermore, the server uses speech recognition software to convert the audio of the lecture videos into text in real time. This feature allows users to understand the lecture content not only through audio but also in text format. Simultaneously, the server uses a natural language processing (NLP) engine to perform multilingual translation and automatically generate subtitles. For example, it is possible to translate a science lecture delivered in English into Japanese and provide subtitles.

[0267] The terminal provides a user interface, helping learners easily access content. Users can use the terminal to search for topics of interest and select relevant educational content provided by the server. For example, if a user is interested in "European art history," the search results will display a list of related lectures and materials. From these, the user can choose any content they wish to learn.

[0268] Furthermore, the server records and analyzes the user's learning history and recommends the most suitable next learning content for each individual. This allows learners to deepen their knowledge efficiently and effectively. For example, a user who has watched past lectures related to historical events will be recommended related topics to continue learning.

[0269] The system also provides live streaming lectures in real time as live sessions. By participating using a device, users can send questions to the instructor during the lecture, creating an interactive learning experience.

[0270] A concrete example of a prompt would be, "I would like to translate a new English science lecture into Japanese and provide it with subtitles." This invention will enable learners to access diverse educational resources from around the world and learn across language and format barriers.

[0271] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0272] Step 1:

[0273] The server collects educational content from information sources. Inputs include data in various formats such as video, audio, text, and images, received via APIs and data feeds. These data are integrated using multimodal technology and combined into a single dataset. This process results in an output where data of different formats can be managed in a consistent manner.

[0274] Step 2:

[0275] The server extracts audio data from an integrated dataset and converts it to text using speech recognition software. The input is the audio data from lecture videos. Through this process, the audio information is converted into text output, generating a visually verifiable lecture content.

[0276] Step 3:

[0277] The server uses a generative AI model to translate text into different languages ​​and generate subtitles. The input is lecture content in text format. This translation process outputs educational content with subtitles that enables learning in multiple languages.

[0278] Step 4:

[0279] The terminal searches for and displays the collected and processed content through the user interface. The input is the user's interests and search keywords. Based on this, relevant content is displayed, and the user can select from it to obtain an output that allows learning to proceed.

[0280] Step 5:

[0281] The server analyzes the user's learning history and recommends the content to be learned next. The input is the past viewing history and learning content. Based on this, the AI algorithm generates personalized recommendations, and the content suitable for the next learning is output.

[0282] Step 6:

[0283] The terminal provides a live session and prepares an environment for the user to participate in real time. The input is the information regarding the scheduled live lecture. The user can access the streamed lecture and obtain an output that allows real-time interaction.

[0284] (Application Example 1)

[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0286] In the modern educational environment, it is difficult for learners to access diverse information around the world and obtain individualized learning experiences. Also, the provision of appropriate information according to the learners' interests and the support for immediate interactive learning event participation are limited. In such a situation, it is required to realize an educational experience and information provision optimized for learners.

[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0288] In this invention, the server includes means for integrating and transforming information using multi-format technology, means for distributing information to the user's portable terminal and providing an operation screen, and means for providing relevant information based on topics of interest to the user. This enables learner-specific information access and personalized educational experiences, and facilitates immediate, participatory learning.

[0289] "Information in diverse formats" refers to content provided in multiple different formats, such as video, audio, text, and images.

[0290] "Multi-format technology" refers to technologies that integrate and convert content in different formats, thereby enabling the consistent delivery of information.

[0291] "Means for generating subtitles" refers to methods for converting audio content into text and displaying it as text information on a display device.

[0292] "User" refers to an individual or group that receives information and engages in learning activities through this system.

[0293] "Learning history" refers to a record of content that a user has accessed in the past and topics they have been interested in.

[0294] An "immediately accessible educational event" refers to an interactive learning session that users can participate in in real time.

[0295] A "portable device" refers to an electronic device that can be easily carried by the user, such as a mobile phone or tablet.

[0296] An "operation screen" refers to a screen display that provides an interface for users to search for and access information.

[0297] "Related information" refers to information, materials, or content related to topics that the user is interested in.

[0298] The system according to the present invention consists of a server, a terminal, and a user.

[0299] The server collects diverse forms of information from educational providers and employs multi-format technologies to integrate content in different formats, such as video, audio, text, and images. This is possible, for example, by converting audio accompanying videos into text and generating subtitles using speech recognition and translation technologies so that even foreign language lectures can be easily understood. The content is processed using specific speech recognition and translation technologies, such as the Google Speech-to-Text API and the Google Cloud Translation API.

[0300] The terminal is a portable device that learners can carry with them and displays an operation screen based on information delivered from the server. Through this operation screen, users can search for topics based on their interests and access related information. In addition, personalized information is provided based on their learning history, enabling them to experience interactive learning.

[0301] Users can participate in real-time educational events using an on-screen interface on their devices. This allows them to access educational resources worldwide and broaden their learning experience. Furthermore, by applying AI technology, content based on their interests can be suggested, effectively supporting their learning progress.

[0302] As a specific example, when a user shows interest in "Renaissance art in Italy", they can use the terminal to access relevant educational content and real-time sessions, and watch lectures by experts from around the world. Additionally, by using a generative AI model, relevant materials to be learned next can be recommended based on a prompt sentence. Examples of prompt sentences include "Recommend relevant content based on the topic selected by the user. For example, when 'Renaissance art in Italy' is specified, instruct to display related videos and materials."

[0303] In this way, the system of the present invention provides a rich learning experience individualized for learners and realizes a means for easy access to educational resources.

[0304] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0305] Step 1:

[0306] The server receives various forms of information from educational providers. Specifically, it accumulates content such as videos, audio, text, images, etc., and stores them in a database. This input includes the entire educational content, and the output is an integrated dataset. To process this data efficiently, the server uses a database management system to organize the storage format and prepare for subsequent processing.

[0307] Step 2:

[0308] The server integrates and converts the received information using multi-format technology. Specifically, it uses the Google Speech-to-Text API to perform conversion from audio to text. The input is the corresponding audio data, and the output is the corresponding text data. The server converts these formats and integrates them into a consistent format, enabling further translation processing according to multiple languages.

[0309] Step 3:

[0310] The server translates the integrated information into different languages ​​and generates subtitles. Using the Google Cloud Translation API, the input text is converted into various target languages. The output generated by this process is text and subtitle data in the target languages. The server uses this data to generate multilingual content.

[0311] Step 4:

[0312] The server analyzes the user's learning history and uses an AI model to recommend the next information to learn. By generating prompts and inputting them into the generating AI model, it suggests the most suitable content based on the user's past viewing history and interests. The input is the learning history and its analysis results, and the output is personalized recommendations for the user.

[0313] Step 5:

[0314] The terminal receives information sent from the server and displays an operation screen. Here, users can search for topics based on their interests and access related content and information recommended by the server. Input is data delivered from the server, and output is information dynamically displayed on the terminal. Specifically, the terminal provides an intuitive operation screen based on UI / UX design.

[0315] Step 6:

[0316] Users participate in educational events that are instantly accessible via their devices, engaging in interactive discussions. They can ask questions and join discussions in real time. Input is streaming data received in real time from the server, while output is participant feedback and questions. Through this, users gain a deeper understanding and knowledge.

[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0318] The educational support system according to the present invention is a platform that provides a more personalized learning experience by combining an emotion engine with existing diverse forms of educational content. This system utilizes servers, terminals, and emotion recognition technology to realize dynamic educational delivery that responds to the user's emotions.

[0319] The server continuously collects diverse content formats from educational providers. This includes video, audio, text, and images, which are integrated using multimodal technology. Content in different languages ​​is translated into the user's native language using AI translation technology, and subtitles are generated for audio content.

[0320] This system incorporates an emotion engine that recognizes the user's emotions through the device's camera and microphone during learning. The emotion engine analyzes the user's facial expressions and tone of voice to calculate their emotional state in real time. For example, if the user shows signs of waning concentration while learning a difficult topic, the system will detect this and either lower the difficulty level of the content or display an encouraging message.

[0321] Users access educational content via their devices, and based on the analysis results of the sentiment engine, recommended content is automatically adjusted. The server adds sentiment data to the user's learning history and uses this data to provide feedback on the next learning topic and to boost motivation. For example, if a user shows particular interest in a topic but struggles to understand it, the system analyzes their level of understanding from the sentiment data and presents similar content using different approaches.

[0322] Even during live sessions, the emotion engine remains active, allowing instructors to understand participants' emotional states in real time. This enables them to adjust the pace of the lesson, ask specific questions, and create an interactive environment.

[0323] Thus, the present invention aims to support efficient and effective learning by utilizing emotion recognition technology to provide an optimized educational experience for each user.

[0324] The following describes the processing flow.

[0325] Step 1:

[0326] The server collects educational content in video, audio, text, and image formats from registered educational providers and public libraries. The collected content is stored in a database along with metadata.

[0327] Step 2:

[0328] The server uses multimodal technology to integrate the collected educational content and prepare it for delivery in a consistent format. This includes audio-to-text conversion and image analysis and descriptive text generation.

[0329] Step 3:

[0330] The server uses an AI translation engine and speech recognition technology to translate educational content into the user's native language and generate subtitles that correspond to the video and audio.

[0331] Step 4:

[0332] The user logs into the device and searches for a topic they want to learn about. The device displays a list of relevant content through the user interface.

[0333] Step 5:

[0334] When a user starts playing content, the device uses its built-in camera and microphone to record the user's facial expressions and voice, and sends this information to the emotion engine.

[0335] Step 6:

[0336] The emotion engine analyzes user emotions from facial and voice data and calculates parameters such as concentration, interest, and fatigue. These results are sent to the server in real time.

[0337] Step 7:

[0338] Based on data from the emotion engine, the server adjusts content to optimize the user's learning experience. For example, if a user's concentration wanes, it may display a message on the device prompting them to take a break or prioritize recommending relevant, lighter topics.

[0339] Step 8:

[0340] When a user participates in a live session, they join the streaming via their device, and emotion engine data is provided to the instructor in real time. This allows the instructor to understand the emotional trends of the session and adjust the lesson accordingly.

[0341] Step 9:

[0342] After the learning process is complete, the device displays an interface requesting feedback from the user. The collected feedback is sent to a server and used for continuous improvement of the system.

[0343] (Example 2)

[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0345] In today's educational environment, it is difficult to efficiently integrate diverse forms of educational content and provide experiences that meet the individual learning needs of users. Furthermore, conventional technologies are insufficient for effectively managing educational events that take into account users' emotions in real time. As a result, there is a problem in providing education that is appropriate for individual learners and improving learning effectiveness.

[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0347] In this invention, the server includes means for acquiring educational content in various formats, means for integrating and transforming educational content using multimodal technology, means for translating educational content into different languages ​​and generating subtitles, means for analyzing the user's learning history and emotional state and recommending the next educational content to be learned, means for providing educational events that can be participated in in real time and for acquiring and analyzing participants' emotional feedback in real time, and means for delivering educational content to the user's terminal, providing a user interface, and personalizing the learning experience based on the user's emotional state. This makes it possible to efficiently utilize educational content in various formats and provide a personalized learning experience.

[0348] "Educational content in diverse formats" refers to educational materials that include different media formats such as videos, audio, text, and images.

[0349] "Multimodal technology" is a technology that integrates data from multiple different media formats and converts it into a common information representation.

[0350] "AI translation technology" is a technology that uses artificial intelligence to automatically translate text and audio between different languages.

[0351] An "emotion engine" is a technology that analyzes a user's emotions from their facial expressions and tone of voice, and calculates their emotional state in real time.

[0352] "Learning history" refers to information about what a user has learned in the past and their progress.

[0353] A "recommendation mechanism" is a system that selects or suggests what a user should learn next based on their learning history and emotional state.

[0354] "Methods for acquiring and analyzing data in real time" refers to technologies that collect data immediately during educational events and perform analysis on the spot.

[0355] A "user interface" is the operating environment through which a user interacts with a system via a terminal and utilizes educational content.

[0356] A "personalized learning experience" refers to the provision of education that is customized according to the individual learning needs and circumstances of each user.

[0357] This educational support system consists of servers, terminals, and emotion recognition technology. The servers continuously collect educational content in various formats provided by educational providers. This collection uses multimodal technology to integrate data from different media formats and convert it into a user-friendly format. For example, it collects English videos, translates them into the user's native language, and generates subtitles.

[0358] The server utilizes AI translation technology to convert content in different languages ​​into the user's native language and adds subtitles to audio content. It also incorporates an emotion engine that analyzes the user's emotions through the device. The resulting emotion data is integrated with the user's learning history and used to determine what to learn next and to provide personalized feedback.

[0359] Users can access educational content through their devices. The devices use cameras and microphones to collect real-time emotional data from the user's facial expressions and tone of voice, and the system uses this information to personalize the learning experience. For example, if a user is feeling stressed while working on a particularly difficult topic, the system will provide simpler explanations or encouraging messages.

[0360] This system captures emotional feedback even during live sessions, allowing instructors to understand participants' emotions in real time and adjust the lesson accordingly. For example, instructors can slow the pace or provide supplementary explanations if participants lose focus.

[0361] As a concrete example, if a user uses a "generative AI model" and inputs a prompt such as "How do I solve a specific problem in mathematical algebra?", the system will use this information to recommend the most suitable learning materials and explanatory videos to the user. In this way, it is possible to provide an optimal educational experience that meets the individual needs of the user.

[0362] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0363] Step 1:

[0364] The server receives diverse forms of educational content from educational providers as input. This input data includes video, audio, text, and images. These are integrated using multimodal technology and converted into a user-friendly format. The output is the integrated content data. Specifically, the server analyzes data from each media format and compiles it into a consistent information package.

[0365] Step 2:

[0366] The server processes integrated educational content using AI translation technology. The input is educational content written in different languages. Through the translation process, the output is text and audio content translated into the user's native language. Specifically, the server automatically adds subtitles to audio content and provides user-selectable subtitle options for videos.

[0367] Step 3:

[0368] The device activates an emotion engine to monitor the user's learning progress. Inputs include the user's facial expressions and voice tone, captured through the device's camera and microphone. This data is analyzed, and real-time emotion data is output. Specifically, the device applies an emotion analysis algorithm to detect different emotional states.

[0369] Step 4:

[0370] The server receives the user's learning history and sentiment data as input and recommends the next educational content to learn. In this step, data analysis determines the appropriate learning materials and learning paths, and provides the user with a list of recommended content as output. Specifically, the server infers the user's interests and level of understanding from past learning data and sentiment feedback, and selects the most suitable learning content.

[0371] Step 5:

[0372] Users access personalized educational content delivered from the server and progress through their learning. At this stage, recommended content serves as input, and the output includes the user's learning progress and new learning history data. Specifically, the user interacts with the learning environment through their device, selecting content as needed.

[0373] Step 6:

[0374] The device generates feedback that allows the instructor to understand participants' emotional states and adjust the lesson pace based on emotional data acquired in real time during live sessions. This feedback provides the instructor with clues to appropriately adjust the pace of the lesson and contributes to interactive lesson management. The input is real-time emotional data from participants, and the output is instructions and suggestions for improvement regarding the lesson pace. Specifically, this could involve the instructor observing participants' reactions and asking questions or suggesting reviewing slides.

[0375] (Application Example 2)

[0376] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0377] In recent years, improving worker performance and safety in factory settings and manufacturing industries has required understanding workers' emotional states in real time and providing appropriate support. However, conventional methods have struggled to accurately recognize workers' emotions and provide immediate, personalized feedback. To address this challenge, there is a need for dynamic support systems that utilize emotion recognition technology in the workplace.

[0378] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0379] In this invention, the server includes means for acquiring knowledge bases in various forms, means for integrating and transforming the knowledge bases using multimodal technology, and means for detecting the user's emotional state during operation and optimizing individual work support. This enables real-time work support based on the worker's emotions.

[0380] A "diverse knowledge base" refers to information or datasets provided in different media formats, including text, images, audio, and video.

[0381] "Multimodal technology" is a technique that integrates and comprehensively analyzes data in different formats, thereby enabling a richer understanding of the data.

[0382] An "information terminal" is a device used by a user, and this includes smartphones, tablets, computers, and wearable devices.

[0383] "Emotional state" refers to the internal psychological state an individual experiences at a given point in time, and includes emotions such as joy, anger, anxiety, and surprise.

[0384] "Optimizing individual work support" means improving work efficiency and safety by providing the most appropriate work instructions and feedback tailored to each worker's current situation and feelings.

[0385] The system implementing this invention consists of a user-worn information terminal, a cloud server, and emotion recognition technology. The information terminal is a wearable device such as smart glasses, equipped with a camera and microphone. This allows for real-time capture of the worker's facial expressions and voice tone, and the acquisition of data.

[0386] The server receives this data and analyzes changes in facial expressions using face recognition libraries such as OpenCV. Simultaneously, it analyzes the tone of voice from the audio data using Google Cloud Speech-to-Text. These analysis results are input into an emotion analysis model using TensorFlow or Keras to calculate the worker's emotional state.

[0387] Based on this emotional state information, the server delivers the next task to be performed and encouraging feedback to the user's information terminal, optimizing individual work support. This allows workers to receive support tailored to their emotions, improving work efficiency and safety.

[0388] As a concrete example, consider a scenario where a factory worker is performing a difficult assembly task while wearing smart glasses. If this worker begins to feel stressed, an emotion analysis model detects this, and the server immediately provides feedback based on that information, such as "Take a deep breath and calm down."

[0389] Examples of input prompts for the generating AI model include phrases such as, "What encouraging message should be displayed if the user looks anxious?" or "How should the work procedure be adjusted if the user's tone of voice is different from usual?" This makes it possible to create a flexible and user-friendly work environment.

[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0391] Step 1:

[0392] The terminal acquires the worker's facial expressions and voice. The terminal uses a camera to capture the worker's facial expressions and a microphone to record their voice. This data is sent as input to the next processing step.

[0393] Step 2:

[0394] The server analyzes the facial expression data. Using the facial image data sent from the terminal, it performs facial recognition through the OpenCV library. As output, feature data extracted from the facial expressions is generated. This data serves as basic information for identifying the emotional state of the worker.

[0395] Step 3:

[0396] The server analyzes the audio data. The audio data provided by the terminal is converted to text using Google Cloud Speech-to-Text, and then the tone of the speech is analyzed. The output provides indicators of speech intensity and emotion.

[0397] Step 4:

[0398] The server analyzes the emotional state. It integrates the previously obtained facial feature data and voice tone data, and uses a TensorFlow-based emotion analysis model to estimate the worker's current emotional state. The output is a set of indicators representing the worker's emotional state.

[0399] Step 5:

[0400] The server generates feedback. Based on the emotional state obtained from sentiment analysis, it uses a generative AI model to devise appropriate feedback and work instructions. It creates messages based on questions such as, "What kind of encouraging message should be displayed if the user has an anxious expression?" as a prompt.

[0401] Step 6:

[0402] The terminal displays feedback. Feedback sent from the server is displayed on the worker's terminal, providing real-time support while preventing work interruptions. This allows workers to receive appropriate advice tailored to their emotional state.

[0403] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0404] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0405] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0406] [Third Embodiment]

[0407] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0408] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0409] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0410] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0411] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0412] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0413] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0414] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0415] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0416] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0417] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0418] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0419] The educational support system according to the present invention is a platform that allows learners to easily access diverse educational content from around the world and provides an individualized learning experience. This system consists primarily of a server, terminals, and a user interface.

[0420] The server collects content from various educational providers. The collected content comes in diverse formats, including video, audio, text, and images, and is integrated using multimodal technology. For example, the server receives a video lecture by an Indian mathematician and converts it from audio to text as needed, making the video content available as text.

[0421] Furthermore, the server uses AI technology to translate educational content into different languages. This includes generating subtitles using speech recognition. For example, when an educational lecture provided in English is translated into Japanese, Japanese subtitles are automatically generated. This allows users to learn without language barriers.

[0422] The device provides learners with access to content through a user interface. Users can search for topics they wish to learn about via the device and access appropriate educational content provided by the server. For example, if a user is interested in "European art history," the device will display a list of related lectures and materials and play the selected one.

[0423] The server also records the learner's learning history and uses this to recommend the next content they should study. The learning history is generated based on the user's past viewing history and interests, and supplementary materials are also provided. For example, if a user has watched lectures related to a specific historical event, related continuing learning topics will be automatically recommended.

[0424] Furthermore, this system offers live sessions, allowing users to participate in real time. During live sessions, users can join streaming lectures via their devices and ask questions directly to the instructor.

[0425] Thus, the embodiment of the present invention aims to address the internationalization of education and individual needs by centrally managing different forms of educational content and creating individualized learning experiences.

[0426] The following describes the processing flow.

[0427] Step 1:

[0428] The server collects educational content provided by educational providers. During collection, it uses APIs and feeds to register metadata (title, category, language, etc.) for each piece of content in a database.

[0429] Step 2:

[0430] The server applies multimodal technology to the collected content, integrating different data formats (video, audio, text, images, etc.) into a consistent format. This makes the content available in a variety of formats.

[0431] Step 3:

[0432] The server uses an AI translation engine to translate collected content into the learner's native language. It also uses speech recognition technology to generate subtitles for video and audio content.

[0433] Step 4:

[0434] Users log in to the platform via their device and search for educational content that interests them. The device then displays a list of relevant content to the user through its user interface.

[0435] Step 5:

[0436] Based on a request from the device, the server sends the selected educational content to the device and makes it playable. This allows the user to view the content in real time.

[0437] Step 6:

[0438] The server recommends the next content a user should learn based on their learning history and current viewing habits. This involves analyzing past history and extracting highly relevant topics.

[0439] Step 7:

[0440] Users participate in live classes through their devices. The server delivers classes in real time using streaming technology and provides an interactive platform for users to ask questions and participate in discussions.

[0441] Step 8:

[0442] After a learning session ends, users enter feedback using a terminal. The server collects this feedback and analyzes it to help improve the system in the future.

[0443] (Example 1)

[0444] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0445] In today's educational environment, challenges include access to diverse forms of educational content, learning across language barriers, and providing personalized learning experiences. In particular, there is a need for equitable use of content from around the world, multilingual support, and the suggestion of optimal learning content tailored to the user's learning progress. Furthermore, supporting participation in real-time educational events and creating an interactive learning environment are also essential.

[0446] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0447] In this invention, the server includes means for collecting information, means for integrating the information using multimodal technology, and means for converting speech to text. This makes it possible to integrate diverse forms of content, convert them to different languages, and provide users with a personalized learning experience. It also supports real-time participatory events, enabling an interactive educational environment.

[0448] "Information" refers to a collection of data and knowledge expressed in various forms, including educational content.

[0449] "Means of collection" refers to methods or techniques for efficiently obtaining information from diverse sources.

[0450] "Multimodal technology" is a technology that integrates and processes information in different formats, making data such as video, audio, and text available on a single platform.

[0451] "Means for converting speech to text" refers to technologies that automatically convert speech data into text and express it in text format.

[0452] "Means of translation and subtitle generation" refers to the technology of translating information into another language and creating subtitles to display that translation.

[0453] "Means for recording and analyzing user history" refers to technologies that record past operations and selections and analyze them to enable personalized content recommendations.

[0454] "Means of providing events that users can participate in in real time" refers to the technology and infrastructure that allows participants to instantly access and interactively engage with live events and lectures.

[0455] "Means for distributing information to a user's device and providing an interface" refers to technologies that transmit acquired information to a user's device and provide screens and functions that the user can operate on that device.

[0456] This invention aims to realize an educational support system that provides learners with diverse educational content. This system mainly consists of three components: a server, a terminal, and a user interface.

[0457] The server plays a central role in collecting educational content from various sources. Specifically, it retrieves data in formats such as video, audio, text, and images through APIs and data feeds. For example, it might download a physics lecture video from an online education platform. Multimodal technology is used in this process to integrate data in different formats.

[0458] Furthermore, the server uses speech recognition software to convert the audio of the lecture videos into text in real time. This feature allows users to understand the lecture content not only through audio but also in text format. Simultaneously, the server uses a natural language processing (NLP) engine to perform multilingual translation and automatically generate subtitles. For example, it is possible to translate a science lecture delivered in English into Japanese and provide subtitles.

[0459] The terminal provides a user interface, helping learners easily access content. Users can use the terminal to search for topics of interest and select relevant educational content provided by the server. For example, if a user is interested in "European art history," the search results will display a list of related lectures and materials. From these, the user can choose any content they wish to learn.

[0460] Furthermore, the server records and analyzes the user's learning history and recommends the most suitable next learning content for each individual. This allows learners to deepen their knowledge efficiently and effectively. For example, a user who has watched past lectures related to historical events will be recommended related topics to continue learning.

[0461] The system also provides live streaming lectures in real time as live sessions. By participating using a device, users can send questions to the instructor during the lecture, creating an interactive learning experience.

[0462] A concrete example of a prompt would be, "I would like to translate a new English science lecture into Japanese and provide it with subtitles." This invention will enable learners to access diverse educational resources from around the world and learn across language and format barriers.

[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0464] Step 1:

[0465] The server collects educational content from information sources. Inputs include data in various formats such as video, audio, text, and images, received via APIs and data feeds. These data are integrated using multimodal technology and combined into a single dataset. This process results in an output where data of different formats can be managed in a consistent manner.

[0466] Step 2:

[0467] The server extracts audio data from an integrated dataset and converts it to text using speech recognition software. The input is the audio data from lecture videos. Through this process, the audio information is converted into text output, generating a visually verifiable lecture content.

[0468] Step 3:

[0469] The server uses a generative AI model to translate text into different languages ​​and generate subtitles. The input is lecture content in text format. This translation process outputs educational content with subtitles that enables learning in multiple languages.

[0470] Step 4:

[0471] The device searches for and displays collected and processed content through its user interface. Input consists of the user's interests and search keywords. Based on this, relevant content is displayed, and the user can select from it to proceed with their learning.

[0472] Step 5:

[0473] The server analyzes the user's learning history and recommends the next content they should learn. The input consists of past viewing history and learning content. Based on this, an AI algorithm generates personalized recommendations, outputting content suitable for the next learning session.

[0474] Step 6:

[0475] The terminal provides live sessions and creates an environment for users to participate in real time. The input is information about scheduled live lectures. The output allows users to access the streamed lectures and interact in real time.

[0476] (Application Example 1)

[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0478] In today's educational environment, it is difficult for learners to access diverse information from around the world and obtain personalized learning experiences. Furthermore, support for providing appropriate information tailored to learners' interests and for participating in immediate, interactive learning events is limited. In this context, there is a need to realize educational experiences and information provision optimized for learners.

[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0480] In this invention, the server includes means for integrating and transforming information using multi-format technology, means for distributing information to the user's portable terminal and providing an operation screen, and means for providing relevant information based on topics of interest to the user. This enables learner-specific information access and personalized educational experiences, and facilitates immediate, participatory learning.

[0481] "Information in diverse formats" refers to content provided in multiple different formats, such as video, audio, text, and images.

[0482] "Multi-format technology" refers to technologies that integrate and convert content in different formats, thereby enabling the consistent delivery of information.

[0483] "Means for generating subtitles" refers to methods for converting audio content into text and displaying it as text information on a display device.

[0484] "User" refers to an individual or group that receives information and engages in learning activities through this system.

[0485] "Learning history" refers to a record of content that a user has accessed in the past and topics they have been interested in.

[0486] An "immediately accessible educational event" refers to an interactive learning session that users can participate in in real time.

[0487] A "portable device" refers to an electronic device that can be easily carried by the user, such as a mobile phone or tablet.

[0488] An "operation screen" refers to a screen display that provides an interface for users to search for and access information.

[0489] "Related information" refers to information, materials, or content related to topics that the user is interested in.

[0490] The system according to the present invention consists of a server, a terminal, and a user.

[0491] The server collects diverse forms of information from educational providers and employs multi-format technologies to integrate content in different formats, such as video, audio, text, and images. This is possible, for example, by converting audio accompanying videos into text and generating subtitles using speech recognition and translation technologies so that even foreign language lectures can be easily understood. The content is processed using specific speech recognition and translation technologies, such as the Google Speech-to-Text API and the Google Cloud Translation API.

[0492] The terminal is a portable device that learners can carry with them and displays an operation screen based on information delivered from the server. Through this operation screen, users can search for topics based on their interests and access related information. In addition, personalized information is provided based on their learning history, enabling them to experience interactive learning.

[0493] Users can participate in real-time educational events using an on-screen interface on their devices. This allows them to access educational resources worldwide and broaden their learning experience. Furthermore, by applying AI technology, content based on their interests can be suggested, effectively supporting their learning progress.

[0494] For example, if a user expresses interest in "Italian Renaissance art," they can use their device to access relevant educational content and real-time sessions, and watch lectures from experts around the world. Furthermore, a generative AI model can be used to recommend relevant materials based on prompts. An example of a prompt would be: "Recommend relevant content based on the topic selected by the user. For example, if 'Italian Renaissance art' is selected, instruct the system to display related videos and materials."

[0495] In this way, the system of the present invention provides learners with an individualized and enriching learning experience and a means of easily accessing educational resources.

[0496] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0497] Step 1:

[0498] The server receives information in various formats from educational providers. Specifically, it collects content such as videos, audio, text, and images and stores them in a database. The input includes the entire educational content, and the output is an integrated dataset. To process this data efficiently, the server uses a database management system to organize the storage format and prepare it for subsequent processing.

[0499] Step 2:

[0500] The server integrates and transforms the received information using multi-format technologies. Specifically, it uses the Google Speech-to-Text API to convert speech to text. The input is the corresponding audio data, and the output is the corresponding text data. The server transforms these formats and integrates them into a consistent format, enabling further translation processing for multiple languages.

[0501] Step 3:

[0502] The server translates the integrated information into different languages ​​and generates subtitles. Using the Google Cloud Translation API, the input text is converted into various target languages. The output generated by this process is text and subtitle data in the target languages. The server uses this data to generate multilingual content.

[0503] Step 4:

[0504] The server analyzes the user's learning history and uses an AI model to recommend the next information to learn. By generating prompts and inputting them into the generating AI model, it suggests the most suitable content based on the user's past viewing history and interests. The input is the learning history and its analysis results, and the output is personalized recommendations for the user.

[0505] Step 5:

[0506] The terminal receives information sent from the server and displays an operation screen. Here, users can search for topics based on their interests and access related content and information recommended by the server. Input is data delivered from the server, and output is information dynamically displayed on the terminal. Specifically, the terminal provides an intuitive operation screen based on UI / UX design.

[0507] Step 6:

[0508] Users participate in educational events that are instantly accessible via their devices, engaging in interactive discussions. They can ask questions and join discussions in real time. Input is streaming data received in real time from the server, while output is participant feedback and questions. Through this, users gain a deeper understanding and knowledge.

[0509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0510] The educational support system according to the present invention is a platform that provides a more personalized learning experience by combining an emotion engine with existing diverse forms of educational content. This system utilizes servers, terminals, and emotion recognition technology to realize dynamic educational delivery that responds to the user's emotions.

[0511] The server continuously collects diverse content formats from educational providers. This includes video, audio, text, and images, which are integrated using multimodal technology. Content in different languages ​​is translated into the user's native language using AI translation technology, and subtitles are generated for audio content.

[0512] This system incorporates an emotion engine that recognizes the user's emotions through the device's camera and microphone during learning. The emotion engine analyzes the user's facial expressions and tone of voice to calculate their emotional state in real time. For example, if the user shows signs of waning concentration while learning a difficult topic, the system will detect this and either lower the difficulty level of the content or display an encouraging message.

[0513] Users access educational content via their devices, and based on the analysis results of the sentiment engine, recommended content is automatically adjusted. The server adds sentiment data to the user's learning history and uses this data to provide feedback on the next learning topic and to boost motivation. For example, if a user shows particular interest in a topic but struggles to understand it, the system analyzes their level of understanding from the sentiment data and presents similar content using different approaches.

[0514] Even during live sessions, the emotion engine remains active, allowing instructors to understand participants' emotional states in real time. This enables them to adjust the pace of the lesson, ask specific questions, and create an interactive environment.

[0515] Thus, the present invention aims to support efficient and effective learning by utilizing emotion recognition technology to provide an optimized educational experience for each user.

[0516] The following describes the processing flow.

[0517] Step 1:

[0518] The server collects educational content in video, audio, text, and image formats from registered educational providers and public libraries. The collected content is stored in a database along with metadata.

[0519] Step 2:

[0520] The server uses multimodal technology to integrate the collected educational content and prepare it for delivery in a consistent format. This includes audio-to-text conversion and image analysis and descriptive text generation.

[0521] Step 3:

[0522] The server uses an AI translation engine and speech recognition technology to translate educational content into the user's native language and generate subtitles that correspond to the video and audio.

[0523] Step 4:

[0524] The user logs into the device and searches for a topic they want to learn about. The device displays a list of relevant content through the user interface.

[0525] Step 5:

[0526] When a user starts playing content, the device uses its built-in camera and microphone to record the user's facial expressions and voice, and sends this information to the emotion engine.

[0527] Step 6:

[0528] The emotion engine analyzes user emotions from facial and voice data and calculates parameters such as concentration, interest, and fatigue. These results are sent to the server in real time.

[0529] Step 7:

[0530] Based on data from the emotion engine, the server adjusts content to optimize the user's learning experience. For example, if a user's concentration wanes, it may display a message on the device prompting them to take a break or prioritize recommending relevant, lighter topics.

[0531] Step 8:

[0532] When a user participates in a live session, they join the streaming via their device, and emotion engine data is provided to the instructor in real time. This allows the instructor to understand the emotional trends of the session and adjust the lesson accordingly.

[0533] Step 9:

[0534] After the learning process is complete, the device displays an interface requesting feedback from the user. The collected feedback is sent to a server and used for continuous improvement of the system.

[0535] (Example 2)

[0536] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0537] In today's educational environment, it is difficult to efficiently integrate diverse forms of educational content and provide experiences that meet the individual learning needs of users. Furthermore, conventional technologies are insufficient for effectively managing educational events that take into account users' emotions in real time. As a result, there is a problem in providing education that is appropriate for individual learners and improving learning effectiveness.

[0538] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0539] In this invention, the server includes means for acquiring educational content in various formats, means for integrating and transforming educational content using multimodal technology, means for translating educational content into different languages ​​and generating subtitles, means for analyzing the user's learning history and emotional state and recommending the next educational content to be learned, means for providing educational events that can be participated in in real time and for acquiring and analyzing participants' emotional feedback in real time, and means for delivering educational content to the user's terminal, providing a user interface, and personalizing the learning experience based on the user's emotional state. This makes it possible to efficiently utilize educational content in various formats and provide a personalized learning experience.

[0540] "Educational content in diverse formats" refers to educational materials that include different media formats such as videos, audio, text, and images.

[0541] "Multimodal technology" is a technology that integrates data from multiple different media formats and converts it into a common information representation.

[0542] "AI translation technology" is a technology that uses artificial intelligence to automatically translate text and audio between different languages.

[0543] An "emotion engine" is a technology that analyzes a user's emotions from their facial expressions and tone of voice, and calculates their emotional state in real time.

[0544] "Learning history" refers to information about what a user has learned in the past and their progress.

[0545] A "recommendation mechanism" is a system that selects or suggests what a user should learn next based on their learning history and emotional state.

[0546] "Methods for acquiring and analyzing data in real time" refers to technologies that collect data immediately during educational events and perform analysis on the spot.

[0547] A "user interface" is the operating environment through which a user interacts with a system via a terminal and utilizes educational content.

[0548] A "personalized learning experience" refers to the provision of education that is customized according to the individual learning needs and circumstances of each user.

[0549] This educational support system consists of servers, terminals, and emotion recognition technology. The servers continuously collect educational content in various formats provided by educational providers. This collection uses multimodal technology to integrate data from different media formats and convert it into a user-friendly format. For example, it collects English videos, translates them into the user's native language, and generates subtitles.

[0550] The server utilizes AI translation technology to convert content in different languages ​​into the user's native language and adds subtitles to audio content. It also incorporates an emotion engine that analyzes the user's emotions through the device. The resulting emotion data is integrated with the user's learning history and used to determine what to learn next and to provide personalized feedback.

[0551] Users can access educational content through their devices. The devices use cameras and microphones to collect real-time emotional data from the user's facial expressions and tone of voice, and the system uses this information to personalize the learning experience. For example, if a user is feeling stressed while working on a particularly difficult topic, the system will provide simpler explanations or encouraging messages.

[0552] This system captures emotional feedback even during live sessions, allowing instructors to understand participants' emotions in real time and adjust the lesson accordingly. For example, instructors can slow the pace or provide supplementary explanations if participants lose focus.

[0553] As a concrete example, if a user uses a "generative AI model" and inputs a prompt such as "How do I solve a specific problem in mathematical algebra?", the system will use this information to recommend the most suitable learning materials and explanatory videos to the user. In this way, it is possible to provide an optimal educational experience that meets the individual needs of the user.

[0554] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0555] Step 1:

[0556] The server receives diverse forms of educational content from educational providers as input. This input data includes video, audio, text, and images. These are integrated using multimodal technology and converted into a user-friendly format. The output is the integrated content data. Specifically, the server analyzes data from each media format and compiles it into a consistent information package.

[0557] Step 2:

[0558] The server processes integrated educational content using AI translation technology. The input is educational content written in different languages. Through the translation process, the output is text and audio content translated into the user's native language. Specifically, the server automatically adds subtitles to audio content and provides user-selectable subtitle options for videos.

[0559] Step 3:

[0560] The device activates an emotion engine to monitor the user's learning progress. Inputs include the user's facial expressions and voice tone, captured through the device's camera and microphone. This data is analyzed, and real-time emotion data is output. Specifically, the device applies an emotion analysis algorithm to detect different emotional states.

[0561] Step 4:

[0562] The server receives the user's learning history and sentiment data as input and recommends the next educational content to learn. In this step, data analysis determines the appropriate learning materials and learning paths, and provides the user with a list of recommended content as output. Specifically, the server infers the user's interests and level of understanding from past learning data and sentiment feedback, and selects the most suitable learning content.

[0563] Step 5:

[0564] Users access personalized educational content delivered from the server and progress through their learning. At this stage, recommended content serves as input, and the output includes the user's learning progress and new learning history data. Specifically, the user interacts with the learning environment through their device, selecting content as needed.

[0565] Step 6:

[0566] The device generates feedback that allows the instructor to understand participants' emotional states and adjust the lesson pace based on emotional data acquired in real time during live sessions. This feedback provides the instructor with clues to appropriately adjust the pace of the lesson and contributes to interactive lesson management. The input is real-time emotional data from participants, and the output is instructions and suggestions for improvement regarding the lesson pace. Specifically, this could involve the instructor observing participants' reactions and asking questions or suggesting reviewing slides.

[0567] (Application Example 2)

[0568] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0569] In recent years, improving worker performance and safety in factory settings and manufacturing industries has required understanding workers' emotional states in real time and providing appropriate support. However, conventional methods have struggled to accurately recognize workers' emotions and provide immediate, personalized feedback. To address this challenge, there is a need for dynamic support systems that utilize emotion recognition technology in the workplace.

[0570] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0571] In this invention, the server includes means for acquiring knowledge bases in various forms, means for integrating and transforming the knowledge bases using multimodal technology, and means for detecting the user's emotional state during operation and optimizing individual work support. This enables real-time work support based on the worker's emotions.

[0572] A "diverse knowledge base" refers to information or datasets provided in different media formats, including text, images, audio, and video.

[0573] "Multimodal technology" is a technique that integrates and comprehensively analyzes data in different formats, thereby enabling a richer understanding of the data.

[0574] An "information terminal" is a device used by a user, and this includes smartphones, tablets, computers, and wearable devices.

[0575] "Emotional state" refers to the internal psychological state an individual experiences at a given point in time, and includes emotions such as joy, anger, anxiety, and surprise.

[0576] "Optimizing individual work support" means improving work efficiency and safety by providing the most appropriate work instructions and feedback tailored to each worker's current situation and feelings.

[0577] The system implementing this invention consists of a user-worn information terminal, a cloud server, and emotion recognition technology. The information terminal is a wearable device such as smart glasses, equipped with a camera and microphone. This allows for real-time capture of the worker's facial expressions and voice tone, and the acquisition of data.

[0578] The server receives this data and analyzes changes in facial expressions using face recognition libraries such as OpenCV. Simultaneously, it analyzes the tone of voice from the audio data using Google Cloud Speech-to-Text. These analysis results are input into an emotion analysis model using TensorFlow or Keras to calculate the worker's emotional state.

[0579] Based on this emotional state information, the server delivers the next task to be performed and encouraging feedback to the user's information terminal, optimizing individual work support. This allows workers to receive support tailored to their emotions, improving work efficiency and safety.

[0580] As a concrete example, consider a scenario where a factory worker is performing a difficult assembly task while wearing smart glasses. If this worker begins to feel stressed, an emotion analysis model detects this, and the server immediately provides feedback based on that information, such as "Take a deep breath and calm down."

[0581] Examples of input prompts for the generating AI model include phrases such as, "What encouraging message should be displayed if the user looks anxious?" or "How should the work procedure be adjusted if the user's tone of voice is different from usual?" This makes it possible to create a flexible and user-friendly work environment.

[0582] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0583] Step 1:

[0584] The terminal acquires the worker's facial expressions and voice. The terminal uses a camera to capture the worker's facial expressions and a microphone to record their voice. This data is sent as input to the next processing step.

[0585] Step 2:

[0586] The server analyzes the facial expression data. Using the facial image data sent from the terminal, it performs facial recognition through the OpenCV library. As output, feature data extracted from the facial expressions is generated. This data serves as basic information for identifying the emotional state of the worker.

[0587] Step 3:

[0588] The server analyzes the audio data. The audio data provided by the terminal is converted to text using Google Cloud Speech-to-Text, and then the tone of the speech is analyzed. The output provides indicators of speech intensity and emotion.

[0589] Step 4:

[0590] The server analyzes the emotional state. It integrates the previously obtained facial feature data and voice tone data, and uses a TensorFlow-based emotion analysis model to estimate the worker's current emotional state. The output is a set of indicators representing the worker's emotional state.

[0591] Step 5:

[0592] The server generates feedback. Based on the emotional state obtained from sentiment analysis, it uses a generative AI model to devise appropriate feedback and work instructions. It creates messages based on questions such as, "What kind of encouraging message should be displayed if the user has an anxious expression?" as a prompt.

[0593] Step 6:

[0594] The terminal displays feedback. Feedback sent from the server is displayed on the worker's terminal, providing real-time support while preventing work interruptions. This allows workers to receive appropriate advice tailored to their emotional state.

[0595] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0596] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0597] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0598] [Fourth Embodiment]

[0599] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0600] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0601] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0602] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0603] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0604] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0605] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0606] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0607] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0608] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0609] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0610] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0611] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0612] The educational support system according to the present invention is a platform that allows learners to easily access diverse educational content from around the world and provides an individualized learning experience. This system consists primarily of a server, terminals, and a user interface.

[0613] The server collects content from various educational providers. The collected content comes in diverse formats, including video, audio, text, and images, and is integrated using multimodal technology. For example, the server receives a video lecture by an Indian mathematician and converts it from audio to text as needed, making the video content available as text.

[0614] Furthermore, the server uses AI technology to translate educational content into different languages. This includes generating subtitles using speech recognition. For example, when an educational lecture provided in English is translated into Japanese, Japanese subtitles are automatically generated. This allows users to learn without language barriers.

[0615] The device provides learners with access to content through a user interface. Users can search for topics they wish to learn about via the device and access appropriate educational content provided by the server. For example, if a user is interested in "European art history," the device will display a list of related lectures and materials and play the selected one.

[0616] The server also records the learner's learning history and uses this to recommend the next content they should study. The learning history is generated based on the user's past viewing history and interests, and supplementary materials are also provided. For example, if a user has watched lectures related to a specific historical event, related continuing learning topics will be automatically recommended.

[0617] Furthermore, this system offers live sessions, allowing users to participate in real time. During live sessions, users can join streaming lectures via their devices and ask questions directly to the instructor.

[0618] Thus, the embodiment of the present invention aims to address the internationalization of education and individual needs by centrally managing different forms of educational content and creating individualized learning experiences.

[0619] The following describes the processing flow.

[0620] Step 1:

[0621] The server collects educational content provided by educational providers. During collection, it uses APIs and feeds to register metadata (title, category, language, etc.) for each piece of content in a database.

[0622] Step 2:

[0623] The server applies multimodal technology to the collected content, integrating different data formats (video, audio, text, images, etc.) into a consistent format. This makes the content available in a variety of formats.

[0624] Step 3:

[0625] The server uses an AI translation engine to translate collected content into the learner's native language. It also uses speech recognition technology to generate subtitles for video and audio content.

[0626] Step 4:

[0627] Users log in to the platform via their device and search for educational content that interests them. The device then displays a list of relevant content to the user through its user interface.

[0628] Step 5:

[0629] Based on a request from the device, the server sends the selected educational content to the device and makes it playable. This allows the user to view the content in real time.

[0630] Step 6:

[0631] The server recommends the next content a user should learn based on their learning history and current viewing habits. This involves analyzing past history and extracting highly relevant topics.

[0632] Step 7:

[0633] Users participate in live classes through their devices. The server delivers classes in real time using streaming technology and provides an interactive platform for users to ask questions and participate in discussions.

[0634] Step 8:

[0635] After a learning session ends, users enter feedback using a terminal. The server collects this feedback and analyzes it to help improve the system in the future.

[0636] (Example 1)

[0637] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0638] In today's educational environment, challenges include access to diverse forms of educational content, learning across language barriers, and providing personalized learning experiences. In particular, there is a need for equitable use of content from around the world, multilingual support, and the suggestion of optimal learning content tailored to the user's learning progress. Furthermore, supporting participation in real-time educational events and creating an interactive learning environment are also essential.

[0639] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0640] In this invention, the server includes means for collecting information, means for integrating the information using multimodal technology, and means for converting speech to text. This makes it possible to integrate diverse forms of content, convert them to different languages, and provide users with a personalized learning experience. It also supports real-time participatory events, enabling an interactive educational environment.

[0641] "Information" refers to a collection of data and knowledge expressed in various forms, including educational content.

[0642] "Means of collection" refers to methods or techniques for efficiently obtaining information from diverse sources.

[0643] "Multimodal technology" is a technology that integrates and processes information in different formats, making data such as video, audio, and text available on a single platform.

[0644] "Means for converting speech to text" refers to technologies that automatically convert speech data into text and express it in text format.

[0645] "Means of translation and subtitle generation" refers to the technology of translating information into another language and creating subtitles to display that translation.

[0646] "Means for recording and analyzing user history" refers to technologies that record past operations and selections and analyze them to enable personalized content recommendations.

[0647] "Means of providing events that users can participate in in real time" refers to the technology and infrastructure that allows participants to instantly access and interactively engage with live events and lectures.

[0648] "Means for distributing information to a user's device and providing an interface" refers to technologies that transmit acquired information to a user's device and provide screens and functions that the user can operate on that device.

[0649] This invention aims to realize an educational support system that provides learners with diverse educational content. This system mainly consists of three components: a server, a terminal, and a user interface.

[0650] The server plays a central role in collecting educational content from various sources. Specifically, it retrieves data in formats such as video, audio, text, and images through APIs and data feeds. For example, it might download a physics lecture video from an online education platform. Multimodal technology is used in this process to integrate data in different formats.

[0651] Furthermore, the server uses speech recognition software to convert the audio of the lecture videos into text in real time. This feature allows users to understand the lecture content not only through audio but also in text format. Simultaneously, the server uses a natural language processing (NLP) engine to perform multilingual translation and automatically generate subtitles. For example, it is possible to translate a science lecture delivered in English into Japanese and provide subtitles.

[0652] The terminal provides a user interface, helping learners easily access content. Users can use the terminal to search for topics of interest and select relevant educational content provided by the server. For example, if a user is interested in "European art history," the search results will display a list of related lectures and materials. From these, the user can choose any content they wish to learn.

[0653] Furthermore, the server records and analyzes the user's learning history and recommends the most suitable next learning content for each individual. This allows learners to deepen their knowledge efficiently and effectively. For example, a user who has watched past lectures related to historical events will be recommended related topics to continue learning.

[0654] The system also provides live streaming lectures in real time as live sessions. By participating using a device, users can send questions to the instructor during the lecture, creating an interactive learning experience.

[0655] A concrete example of a prompt would be, "I would like to translate a new English science lecture into Japanese and provide it with subtitles." This invention will enable learners to access diverse educational resources from around the world and learn across language and format barriers.

[0656] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0657] Step 1:

[0658] The server collects educational content from information sources. Inputs include data in various formats such as video, audio, text, and images, received via APIs and data feeds. These data are integrated using multimodal technology and combined into a single dataset. This process results in an output where data of different formats can be managed in a consistent manner.

[0659] Step 2:

[0660] The server extracts audio data from an integrated dataset and converts it to text using speech recognition software. The input is the audio data from lecture videos. Through this process, the audio information is converted into text output, generating a visually verifiable lecture content.

[0661] Step 3:

[0662] The server uses a generative AI model to translate text into different languages ​​and generate subtitles. The input is lecture content in text format. This translation process outputs educational content with subtitles that enables learning in multiple languages.

[0663] Step 4:

[0664] The device searches for and displays collected and processed content through its user interface. Input consists of the user's interests and search keywords. Based on this, relevant content is displayed, and the user can select from it to proceed with their learning.

[0665] Step 5:

[0666] The server analyzes the user's learning history and recommends the next content they should learn. The input consists of past viewing history and learning content. Based on this, an AI algorithm generates personalized recommendations, outputting content suitable for the next learning session.

[0667] Step 6:

[0668] The terminal provides live sessions and creates an environment for users to participate in real time. The input is information about scheduled live lectures. The output allows users to access the streamed lectures and interact in real time.

[0669] (Application Example 1)

[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] In today's educational environment, it is difficult for learners to access diverse information from around the world and obtain personalized learning experiences. Furthermore, support for providing appropriate information tailored to learners' interests and for participating in immediate, interactive learning events is limited. In this context, there is a need to realize educational experiences and information provision optimized for learners.

[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0673] In this invention, the server includes means for integrating and transforming information using multi-format technology, means for distributing information to the user's portable terminal and providing an operation screen, and means for providing relevant information based on topics of interest to the user. This enables learner-specific information access and personalized educational experiences, and facilitates immediate, participatory learning.

[0674] "Information in diverse formats" refers to content provided in multiple different formats, such as video, audio, text, and images.

[0675] "Multi-format technology" refers to technologies that integrate and convert content in different formats, thereby enabling the consistent delivery of information.

[0676] "Means for generating subtitles" refers to methods for converting audio content into text and displaying it as text information on a display device.

[0677] "User" refers to an individual or group that receives information and engages in learning activities through this system.

[0678] "Learning history" refers to a record of content that a user has accessed in the past and topics they have been interested in.

[0679] An "immediately accessible educational event" refers to an interactive learning session that users can participate in in real time.

[0680] A "portable device" refers to an electronic device that can be easily carried by the user, such as a mobile phone or tablet.

[0681] An "operation screen" refers to a screen display that provides an interface for users to search for and access information.

[0682] "Related information" refers to information, materials, or content related to topics that the user is interested in.

[0683] The system according to the present invention consists of a server, a terminal, and a user.

[0684] The server collects diverse forms of information from educational providers and employs multi-format technologies to integrate content in different formats, such as video, audio, text, and images. This is possible, for example, by converting audio accompanying videos into text and generating subtitles using speech recognition and translation technologies so that even foreign language lectures can be easily understood. The content is processed using specific speech recognition and translation technologies, such as the Google Speech-to-Text API and the Google Cloud Translation API.

[0685] The terminal is a portable device that learners can carry with them and displays an operation screen based on information delivered from the server. Through this operation screen, users can search for topics based on their interests and access related information. In addition, personalized information is provided based on their learning history, enabling them to experience interactive learning.

[0686] Users can participate in real-time educational events using an on-screen interface on their devices. This allows them to access educational resources worldwide and broaden their learning experience. Furthermore, by applying AI technology, content based on their interests can be suggested, effectively supporting their learning progress.

[0687] For example, if a user expresses interest in "Italian Renaissance art," they can use their device to access relevant educational content and real-time sessions, and watch lectures from experts around the world. Furthermore, a generative AI model can be used to recommend relevant materials based on prompts. An example of a prompt would be: "Recommend relevant content based on the topic selected by the user. For example, if 'Italian Renaissance art' is selected, instruct the system to display related videos and materials."

[0688] In this way, the system of the present invention provides learners with an individualized and enriching learning experience and a means of easily accessing educational resources.

[0689] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0690] Step 1:

[0691] The server receives information in various formats from educational providers. Specifically, it collects content such as videos, audio, text, and images and stores them in a database. The input includes the entire educational content, and the output is an integrated dataset. To process this data efficiently, the server uses a database management system to organize the storage format and prepare it for subsequent processing.

[0692] Step 2:

[0693] The server integrates and transforms the received information using multi-format technologies. Specifically, it uses the Google Speech-to-Text API to convert speech to text. The input is the corresponding audio data, and the output is the corresponding text data. The server transforms these formats and integrates them into a consistent format, enabling further translation processing for multiple languages.

[0694] Step 3:

[0695] The server translates the integrated information into different languages ​​and generates subtitles. Using the Google Cloud Translation API, the input text is converted into various target languages. The output generated by this process is text and subtitle data in the target languages. The server uses this data to generate multilingual content.

[0696] Step 4:

[0697] The server analyzes the user's learning history and uses an AI model to recommend the next information to learn. By generating prompts and inputting them into the generating AI model, it suggests the most suitable content based on the user's past viewing history and interests. The input is the learning history and its analysis results, and the output is personalized recommendations for the user.

[0698] Step 5:

[0699] The terminal receives information sent from the server and displays an operation screen. Here, users can search for topics based on their interests and access related content and information recommended by the server. Input is data delivered from the server, and output is information dynamically displayed on the terminal. Specifically, the terminal provides an intuitive operation screen based on UI / UX design.

[0700] Step 6:

[0701] Users participate in educational events that are instantly accessible via their devices, engaging in interactive discussions. They can ask questions and join discussions in real time. Input is streaming data received in real time from the server, while output is participant feedback and questions. Through this, users gain a deeper understanding and knowledge.

[0702] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0703] The educational support system according to the present invention is a platform that provides a more personalized learning experience by combining an emotion engine with existing diverse forms of educational content. This system utilizes servers, terminals, and emotion recognition technology to realize dynamic educational delivery that responds to the user's emotions.

[0704] The server continuously collects diverse content formats from educational providers. This includes video, audio, text, and images, which are integrated using multimodal technology. Content in different languages ​​is translated into the user's native language using AI translation technology, and subtitles are generated for audio content.

[0705] This system incorporates an emotion engine that recognizes the user's emotions through the device's camera and microphone during learning. The emotion engine analyzes the user's facial expressions and tone of voice to calculate their emotional state in real time. For example, if the user shows signs of waning concentration while learning a difficult topic, the system will detect this and either lower the difficulty level of the content or display an encouraging message.

[0706] Users access educational content via their devices, and based on the analysis results of the sentiment engine, recommended content is automatically adjusted. The server adds sentiment data to the user's learning history and uses this data to provide feedback on the next learning topic and to boost motivation. For example, if a user shows particular interest in a topic but struggles to understand it, the system analyzes their level of understanding from the sentiment data and presents similar content using different approaches.

[0707] Even during live sessions, the emotion engine remains active, allowing instructors to understand participants' emotional states in real time. This enables them to adjust the pace of the lesson, ask specific questions, and create an interactive environment.

[0708] Thus, the present invention aims to support efficient and effective learning by utilizing emotion recognition technology to provide an optimized educational experience for each user.

[0709] The following describes the processing flow.

[0710] Step 1:

[0711] The server collects educational content in video, audio, text, and image formats from registered educational providers and public libraries. The collected content is stored in a database along with metadata.

[0712] Step 2:

[0713] The server uses multimodal technology to integrate the collected educational content and prepare it for delivery in a consistent format. This includes audio-to-text conversion and image analysis and descriptive text generation.

[0714] Step 3:

[0715] The server uses an AI translation engine and speech recognition technology to translate educational content into the user's native language and generate subtitles that correspond to the video and audio.

[0716] Step 4:

[0717] The user logs into the device and searches for a topic they want to learn about. The device displays a list of relevant content through the user interface.

[0718] Step 5:

[0719] When a user starts playing content, the device uses its built-in camera and microphone to record the user's facial expressions and voice, and sends this information to the emotion engine.

[0720] Step 6:

[0721] The emotion engine analyzes user emotions from facial and voice data and calculates parameters such as concentration, interest, and fatigue. These results are sent to the server in real time.

[0722] Step 7:

[0723] Based on data from the emotion engine, the server adjusts content to optimize the user's learning experience. For example, if a user's concentration wanes, it may display a message on the device prompting them to take a break or prioritize recommending relevant, lighter topics.

[0724] Step 8:

[0725] When a user participates in a live session, they join the streaming via their device, and emotion engine data is provided to the instructor in real time. This allows the instructor to understand the emotional trends of the session and adjust the lesson accordingly.

[0726] Step 9:

[0727] After the learning process is complete, the device displays an interface requesting feedback from the user. The collected feedback is sent to a server and used for continuous improvement of the system.

[0728] (Example 2)

[0729] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0730] In today's educational environment, it is difficult to efficiently integrate diverse forms of educational content and provide experiences that meet the individual learning needs of users. Furthermore, conventional technologies are insufficient for effectively managing educational events that take into account users' emotions in real time. As a result, there is a problem in providing education that is appropriate for individual learners and improving learning effectiveness.

[0731] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0732] In this invention, the server includes means for acquiring educational content in various formats, means for integrating and transforming educational content using multimodal technology, means for translating educational content into different languages ​​and generating subtitles, means for analyzing the user's learning history and emotional state and recommending the next educational content to be learned, means for providing educational events that can be participated in in real time and for acquiring and analyzing participants' emotional feedback in real time, and means for delivering educational content to the user's terminal, providing a user interface, and personalizing the learning experience based on the user's emotional state. This makes it possible to efficiently utilize educational content in various formats and provide a personalized learning experience.

[0733] "Educational content in diverse formats" refers to educational materials that include different media formats such as videos, audio, text, and images.

[0734] "Multimodal technology" is a technology that integrates data from multiple different media formats and converts it into a common information representation.

[0735] "AI translation technology" is a technology that uses artificial intelligence to automatically translate text and audio between different languages.

[0736] An "emotion engine" is a technology that analyzes a user's emotions from their facial expressions and tone of voice, and calculates their emotional state in real time.

[0737] "Learning history" refers to information about what a user has learned in the past and their progress.

[0738] A "recommendation mechanism" is a system that selects or suggests what a user should learn next based on their learning history and emotional state.

[0739] "Methods for acquiring and analyzing data in real time" refers to technologies that collect data immediately during educational events and perform analysis on the spot.

[0740] A "user interface" is the operating environment through which a user interacts with a system via a terminal and utilizes educational content.

[0741] A "personalized learning experience" refers to the provision of education that is customized according to the individual learning needs and circumstances of each user.

[0742] This educational support system consists of servers, terminals, and emotion recognition technology. The servers continuously collect educational content in various formats provided by educational providers. This collection uses multimodal technology to integrate data from different media formats and convert it into a user-friendly format. For example, it collects English videos, translates them into the user's native language, and generates subtitles.

[0743] The server utilizes AI translation technology to convert content in different languages ​​into the user's native language and adds subtitles to audio content. It also incorporates an emotion engine that analyzes the user's emotions through the device. The resulting emotion data is integrated with the user's learning history and used to determine what to learn next and to provide personalized feedback.

[0744] Users can access educational content through their devices. The devices use cameras and microphones to collect real-time emotional data from the user's facial expressions and tone of voice, and the system uses this information to personalize the learning experience. For example, if a user is feeling stressed while working on a particularly difficult topic, the system will provide simpler explanations or encouraging messages.

[0745] This system captures emotional feedback even during live sessions, allowing instructors to understand participants' emotions in real time and adjust the lesson accordingly. For example, instructors can slow the pace or provide supplementary explanations if participants lose focus.

[0746] As a concrete example, if a user uses a "generative AI model" and inputs a prompt such as "How do I solve a specific problem in mathematical algebra?", the system will use this information to recommend the most suitable learning materials and explanatory videos to the user. In this way, it is possible to provide an optimal educational experience that meets the individual needs of the user.

[0747] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0748] Step 1:

[0749] The server receives diverse forms of educational content from educational providers as input. This input data includes video, audio, text, and images. These are integrated using multimodal technology and converted into a user-friendly format. The output is the integrated content data. Specifically, the server analyzes data from each media format and compiles it into a consistent information package.

[0750] Step 2:

[0751] The server processes integrated educational content using AI translation technology. The input is educational content written in different languages. Through the translation process, the output is text and audio content translated into the user's native language. Specifically, the server automatically adds subtitles to audio content and provides user-selectable subtitle options for videos.

[0752] Step 3:

[0753] The device activates an emotion engine to monitor the user's learning progress. Inputs include the user's facial expressions and voice tone, captured through the device's camera and microphone. This data is analyzed, and real-time emotion data is output. Specifically, the device applies an emotion analysis algorithm to detect different emotional states.

[0754] Step 4:

[0755] The server receives the user's learning history and sentiment data as input and recommends the next educational content to learn. In this step, data analysis determines the appropriate learning materials and learning paths, and provides the user with a list of recommended content as output. Specifically, the server infers the user's interests and level of understanding from past learning data and sentiment feedback, and selects the most suitable learning content.

[0756] Step 5:

[0757] Users access personalized educational content delivered from the server and progress through their learning. At this stage, recommended content serves as input, and the output includes the user's learning progress and new learning history data. Specifically, the user interacts with the learning environment through their device, selecting content as needed.

[0758] Step 6:

[0759] The device generates feedback that allows the instructor to understand participants' emotional states and adjust the lesson pace based on emotional data acquired in real time during live sessions. This feedback provides the instructor with clues to appropriately adjust the pace of the lesson and contributes to interactive lesson management. The input is real-time emotional data from participants, and the output is instructions and suggestions for improvement regarding the lesson pace. Specifically, this could involve the instructor observing participants' reactions and asking questions or suggesting reviewing slides.

[0760] (Application Example 2)

[0761] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0762] In recent years, improving worker performance and safety in factory settings and manufacturing industries has required understanding workers' emotional states in real time and providing appropriate support. However, conventional methods have struggled to accurately recognize workers' emotions and provide immediate, personalized feedback. To address this challenge, there is a need for dynamic support systems that utilize emotion recognition technology in the workplace.

[0763] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0764] In this invention, the server includes means for acquiring knowledge bases in various forms, means for integrating and transforming the knowledge bases using multimodal technology, and means for detecting the user's emotional state during operation and optimizing individual work support. This enables real-time work support based on the worker's emotions.

[0765] A "diverse knowledge base" refers to information or datasets provided in different media formats, including text, images, audio, and video.

[0766] "Multimodal technology" is a technique that integrates and comprehensively analyzes data in different formats, thereby enabling a richer understanding of the data.

[0767] An "information terminal" is a device used by a user, and this includes smartphones, tablets, computers, and wearable devices.

[0768] "Emotional state" refers to the internal psychological state an individual experiences at a given point in time, and includes emotions such as joy, anger, anxiety, and surprise.

[0769] "Optimizing individual work support" means improving work efficiency and safety by providing the most appropriate work instructions and feedback tailored to each worker's current situation and feelings.

[0770] The system implementing this invention consists of a user-worn information terminal, a cloud server, and emotion recognition technology. The information terminal is a wearable device such as smart glasses, equipped with a camera and microphone. This allows for real-time capture of the worker's facial expressions and voice tone, and the acquisition of data.

[0771] The server receives this data and analyzes changes in facial expressions using face recognition libraries such as OpenCV. Simultaneously, it analyzes the tone of voice from the audio data using Google Cloud Speech-to-Text. These analysis results are input into an emotion analysis model using TensorFlow or Keras to calculate the worker's emotional state.

[0772] Based on this emotional state information, the server delivers the next task to be performed and encouraging feedback to the user's information terminal, optimizing individual work support. This allows workers to receive support tailored to their emotions, improving work efficiency and safety.

[0773] As a concrete example, consider a scenario where a factory worker is performing a difficult assembly task while wearing smart glasses. If this worker begins to feel stressed, an emotion analysis model detects this, and the server immediately provides feedback based on that information, such as "Take a deep breath and calm down."

[0774] Examples of input prompts for the generating AI model include phrases such as, "What encouraging message should be displayed if the user looks anxious?" or "How should the work procedure be adjusted if the user's tone of voice is different from usual?" This makes it possible to create a flexible and user-friendly work environment.

[0775] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0776] Step 1:

[0777] The terminal acquires the worker's facial expressions and voice. The terminal uses a camera to capture the worker's facial expressions and a microphone to record their voice. This data is sent as input to the next processing step.

[0778] Step 2:

[0779] The server analyzes the facial expression data. Using the facial image data sent from the terminal, it performs facial recognition through the OpenCV library. As output, feature data extracted from the facial expressions is generated. This data serves as basic information for identifying the emotional state of the worker.

[0780] Step 3:

[0781] The server analyzes the audio data. The audio data provided by the terminal is converted to text using Google Cloud Speech-to-Text, and then the tone of the speech is analyzed. The output provides indicators of speech intensity and emotion.

[0782] Step 4:

[0783] The server analyzes the emotional state. It integrates the previously obtained facial feature data and voice tone data, and uses a TensorFlow-based emotion analysis model to estimate the worker's current emotional state. The output is a set of indicators representing the worker's emotional state.

[0784] Step 5:

[0785] The server generates feedback. Based on the emotional state obtained from sentiment analysis, it uses a generative AI model to devise appropriate feedback and work instructions. It creates messages based on questions such as, "What kind of encouraging message should be displayed if the user has an anxious expression?" as a prompt.

[0786] Step 6:

[0787] The terminal displays feedback. Feedback sent from the server is displayed on the worker's terminal, providing real-time support while preventing work interruptions. This allows workers to receive appropriate advice tailored to their emotional state.

[0788] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0789] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0790] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0791] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0792] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0793] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0794] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0795] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0796] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0797] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0798] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0799] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0800] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0801] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0802] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0803] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0804] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0805] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0806] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0807] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0808] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0809] The following is further disclosed regarding the embodiments described above.

[0810] (Claim 1)

[0811] Means of acquiring diverse forms of educational content,

[0812] A means for integrating and transforming the aforementioned educational content using multimodal technology,

[0813] A means for translating the aforementioned educational content into different languages ​​and generating subtitles,

[0814] A method for analyzing a user's learning history and recommending the next educational content they should learn,

[0815] A means of providing educational events that users can participate in in real time,

[0816] A means for delivering educational content to the user's terminal and providing a user interface,

[0817] A system that includes this.

[0818] (Claim 2)

[0819] The system according to claim 1, which provides users with an individualized learning experience based on the aforementioned diverse forms of educational content.

[0820] (Claim 3)

[0821] The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned real-time educational event.

[0822] "Example 1"

[0823] (Claim 1)

[0824] Means of collecting information,

[0825] Means for integrating the aforementioned information using multimodal technology,

[0826] A means of converting speech to text,

[0827] A means for translating the aforementioned information into different languages ​​and generating subtitles,

[0828] A means of recording and analyzing user history and recommending the next most suitable information,

[0829] A means of providing events in which users can participate in real time,

[0830] Means for distributing information to the user's device and providing an interface,

[0831] A system that includes this.

[0832] (Claim 2)

[0833] The system according to claim 1, which provides a personalized experience to a user based on the aforementioned diverse forms of information.

[0834] (Claim 3)

[0835] The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned real-time events.

[0836] "Application Example 1"

[0837] (Claim 1)

[0838] Means of acquiring information in various formats,

[0839] A means for integrating and converting the aforementioned information using multi-format technology,

[0840] A means for translating the aforementioned information into different languages ​​and generating subtitles,

[0841] A means of analyzing the user's learning history and recommending the next information they should learn,

[0842] A means of providing educational events that users can participate in immediately,

[0843] A means for distributing information to the user's portable terminal and providing an operation screen,

[0844] A means of providing relevant information based on topics that users have shown interest in,

[0845] A system that includes this.

[0846] (Claim 2)

[0847] The system according to claim 1, which provides a personalized educational experience to a user based on the aforementioned diverse forms of information.

[0848] (Claim 3)

[0849] The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned immediately available educational event.

[0850] "Example 2 of combining an emotion engine"

[0851] (Claim 1)

[0852] Means of acquiring diverse forms of educational content,

[0853] A means for integrating and transforming the aforementioned educational content using multimodal technology,

[0854] A means for translating the aforementioned educational content into different languages ​​and generating subtitles,

[0855] A means of analyzing the user's learning history and emotional state to recommend the next educational content they should learn,

[0856] A means to provide educational events that users can participate in in real time, and to obtain and analyze participants' emotional feedback in real time,

[0857] The means of delivering educational content to the user's terminal, providing a user interface, and personalizing the learning experience based on the user's emotional state,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, which uses emotion recognition technology to dynamically adjust the user's learning experience and provide personalized education.

[0861] (Claim 3)

[0862] The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned real-time educational event, and uses an emotion engine to allow the instructor to understand the emotional state of participants and adjust the progress of the lesson.

[0863] "Application example 2 when combining with an emotional engine"

[0864] (Claim 1)

[0865] Means of acquiring diverse forms of knowledge bases,

[0866] A means for integrating and transforming the aforementioned knowledge base using multimodal technology,

[0867] Means for translating the aforementioned knowledge base into different languages ​​and generating display information,

[0868] A means of analyzing the user's work history and recommending the next task to be performed,

[0869] A means of providing business events that users can participate in in real time,

[0870] A means for detecting the user's emotional state during operation and optimizing individualized work support,

[0871] A means for distributing work details to the user's information terminal and providing a user interface,

[0872] A system that includes this.

[0873] (Claim 2)

[0874] The system according to claim 1, which provides a personalized work support experience to a user based on the aforementioned diverse forms of knowledge bases.

[0875] (Claim 3)

[0876] The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned real-time business event. [Explanation of Symbols]

[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means of acquiring diverse forms of educational content, A means for integrating and transforming the aforementioned educational content using multimodal technology, A means for translating the aforementioned educational content into different languages ​​and generating subtitles, A method for analyzing a user's learning history and recommending the next educational content they should learn, A means of providing educational events that users can participate in in real time, A means for delivering educational content to the user's terminal and providing a user interface, A system that includes this.

2. The system according to claim 1, which provides users with an individualized learning experience based on the aforementioned diverse forms of educational content.

3. The system according to claim 1, which provides question and discussion functions to support participation in the aforementioned real-time educational event.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A