System
The system addresses the challenge of linking diverse content types by using a natural language processing engine to map entrance exam questions to entertainment content, enhancing learning through integrated educational and entertainment experiences.
Patent Information
- Application Number
- JP2024124055
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Students face difficulty in maintaining deep understanding and interest in their studies due to a lack of links between various content types, including textbooks and entertainment media, leading to decreased learning efficiency.
A system utilizing a natural language processing engine to analyze text data, extract keywords and tags, map entrance exam questions to related entertainment content, provide a user interface for searching by age or category, and distribute revenue to content providers, while converting audio and image data into text for dynamic content linking.
Enables students to naturally encounter related entertainment content during their studies, maintaining interest and gaining a deeper understanding through integrated learning and entertainment experiences.
Smart Images

Figure 2026022538000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Today's students are required to acquire knowledge from a variety of content beyond textbooks and reference books, but a lack of links between these contents makes it difficult to maintain deep understanding or interest. Furthermore, learning efficiency can decrease because students are unable to effectively utilize the relationships between a wide variety of entertainment content. For this reason, a system is needed that allows students to naturally discover related entertainment content while solving entrance exam questions and link it to their studies. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means: a system including a means for analyzing text data using a natural language processing engine and extracting keywords and tags, a means for analyzing entrance exam questions and mapping related entertainment content, a means for providing a user interface that allows users to search for content by age group or category, a means for displaying content related to a question when the user solves it, and a means for analyzing the user's learning history and content usage data and distributing revenue to content providers. The system also includes a means for converting audio data and image data into text data and a means for dynamically linking highly relevant content based on the category and keywords of the entrance exam questions. This system allows students to naturally encounter related entertainment content during their studies, maintaining their interest and gaining a deeper understanding.
[0006] A "natural language processing engine" is a computer program that analyzes text data and extracts keywords and tags.
[0007] "Text data" refers to data that contains character information, and includes audio data and image data that have been converted into text.
[0008] A "keyword" is a word or short phrase that has an important meaning in the text data.
[0009] A "tag" is a label that indicates a theme or category related to text data.
[0010] "Entrance examination questions" are questions that are asked when students take the exam, and have academic evaluation criteria.
[0011] "Entertainment content" refers to content provided in a variety of media for entertainment or education purposes, such as novels, movies, music, and manga.
[0012] "User interface" refers to the screen and operation method that users use to operate the system.
[0013] "Learning history" is data that records the content and results of a user's past learning.
[0014] A "content provider" is an individual or organization that produces or distributes entertainment content.
[0015] "Voice recognition technology" is a technology that analyzes voice data and converts it into text data.
[0016] "OCR (Optical Character Recognition) technology" is a technology that extracts character information from image data and converts it into text data.
[0017] A "database" is a system that can efficiently store, manage, and search large amounts of data.
[0018] "Revenue sharing" refers to the distribution of system revenues to stakeholders based on specific criteria. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to a system for linking learning and entertainment content using a natural language processing engine. The following describes an embodiment of the system from the viewpoints of a server, a terminal, and a user.
[0041] Data collection and analysis
[0042] To collect data, the server first acquires content data using APIs from various information providers. This includes audio data, image data, books, and song lyrics. The acquired audio data is converted into text data using voice recognition technology, and image data is converted into text data using OCR technology. This text data is then centrally stored in a database.
[0043] The server then launches a natural language processing engine to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0044] Mapping content to entrance exam questions
[0045] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0046] Providing a user interface
[0047] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific grade level or subject, the server searches for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0048] Learning and Content Links
[0049] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, if a user answers a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from that list, they can gain a deeper understanding of the historical background.
[0050] Data Updates and Revenue Sharing
[0051] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue is distributed. The revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue is distributed to that content provider.
[0052] Specific examples
[0053] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0054] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The server uses the API of the information provider to obtain various content data (books, lyrics, audio data, image data), including audio and image data.
[0058] Step 2:
[0059] The server converts the acquired voice data into text data using voice recognition technology, and converts the image data into text data using OCR (optical character recognition) technology. The converted text data is stored in a database.
[0060] Step 3:
[0061] The server runs a natural language processing (NLP) engine to analyze the text data in the database, extracting important keywords, tags, and content names from the text data.
[0062] Step 4:
[0063] The server associates entrance exam questions with entertainment content based on the keywords and tags extracted through the analysis. For example, a question on Japanese history might be linked to "historical films" or "manga."
[0064] Step 5:
[0065] The server provides a user interface through which users can search for content by age and category.
[0066] Step 6:
[0067] The user terminal sends a search request to the server, for example, if the user wants to search for historical content from a particular era.
[0068] Step 7:
[0069] The server queries the database based on the received search request and returns appropriate results to the user's device, such as a list of movies and novels related to the "Warring States Period" arrow.
[0070] Step 8:
[0071] Users simply launch the learning app, select a specific exam question, and begin answering it. Related entertainment content is automatically displayed while they answer the question.
[0072] Step 9:
[0073] The user device will display related entertainment content based on the category and keywords of the entrance exam questions the user has answered. For example, after answering a question on classical literature, novels and movies related to that topic will be displayed.
[0074] Step 10:
[0075] The server periodically collects and analyzes users' learning history and content usage data, which is then stored in the system as monthly usage statistics.
[0076] Step 11:
[0077] The server calculates revenue based on the collected usage data and distributes it to entertainment content providers, based on the number of views and duration of use.
[0078] Step 12:
[0079] The server notifies both users and content providers of the results of revenue sharing and feedback on learnings, which helps optimize the system.
[0080] The above are the specific processing steps of this system, which effectively link learning content and entertainment content to provide users with a rich learning experience.
[0081] Example 1
[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0083] In conventional learning systems, it is difficult for users to easily access external information or entertainment content directly related to the problem they are solving. Furthermore, they lack a mechanism for effectively managing users' learning data and appropriately distributing revenue to entertainment content providers. This results in low learning effectiveness and makes it difficult to promote deep understanding through entertainment.
[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0085] In this invention, the server includes means for analyzing text data using a natural language processing engine and extracting keywords and tags, means for analyzing entrance exam questions and mapping related entertainment content, means for providing a user interface that allows users to search for content by age group or category, means for converting audio data and image data into text data, means for displaying content related to a question when the user solves it, and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This allows users to easily access related entertainment content while studying, and enables appropriate management of study data and revenue distribution.
[0086] A "natural language processing engine" is a software engine that analyzes text data and extracts important keywords and tags.
[0087] "Keywords" are important words or phrases extracted from text data that are used to categorize and associate content and entrance exam questions.
[0088] "Tags" are metadata attached to text data, and are labels that indicate the characteristics of the content or entrance exam questions.
[0089] "Speech recognition technology" is a technology that converts voice data into text data.
[0090] "Optical character recognition technology" is a technology that analyzes image data and extracts the characters contained therein as text data.
[0091] A "user interface" is a visual interface that allows users to search for content by age group or category.
[0092] "Study history" refers to historical data of entrance exam questions that a user has answered using a learning app.
[0093] "Revenue sharing" refers to the process of appropriately allocating revenue to entertainment content providers based on users' learning history and content usage data.
[0094] "Entrance exam questions" are exam questions designed for users to answer.
[0095] "Entertainment content" refers to content such as books, movies, music, and manga that are provided to allow users to deepen their learning while having fun.
[0096] This invention relates to a system for linking learning and entertainment content using a natural language processing engine, and an embodiment thereof will be described in detail from the viewpoints of a server, a terminal, and a user.
[0097] Data collection and analysis
[0098] To collect data, the server first obtains content data using APIs from various information providers. Specifically, it uses the Google Books API, Spotify API, YouTube Data API, etc. This content data includes audio data, image data, books, and song lyrics. The obtained audio data is converted into text data using Google Cloud Speech-to-Text, and image data is converted into text data using Tesseract OCR technology. This text data is then centrally stored in a MySQL database.
[0099] Next, the server launches a natural language processing engine such as BERT or GPT-3 to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0100] Mapping content to entrance exam questions
[0101] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" could be linked to "historical films" and "manga" that can be obtained from multiple media. In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0102] Providing a user interface
[0103] User devices are provided with an interface developed using React.js, which allows users to search for content by age group and category. When a user selects a specific grade or subject, the server uses Elasticsearch to quickly search for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0104] Learning and Content Links
[0105] When a user launches the learning app and answers entrance exam questions, related entertainment content is displayed. As the user proceeds with the answer, the user's device requests related entertainment content from the server. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, when a user solves a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from the list, the user can gain a deeper understanding of the historical background.
[0106] Data Updates and Revenue Sharing
[0107] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is performed using the Stripe API. Revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue will be distributed to that content provider.
[0108] Specific examples
[0109] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0110] Example prompt: "Please provide entertainment content related to the history of the Heian period. Preferably include novels, movies, and manga."
[0111] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0112] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0113] Step 1:
[0114] To collect data, the server uses APIs from information providers such as Google Books API, Spotify API, and YouTube Data API. The input includes requests to each API. For example, to obtain data such as book information, music information, and video information, an HTTP request is sent to each API. The output is response data from each API. This allows book metadata, music metadata, and video metadata to be obtained. Specifically, the server sends an asynchronous request to each API and waits for a response.
[0115] Step 2:
[0116] The server converts the acquired audio data into text data using Google Cloud Speech-to-Text, and converts the image data into text data using Tesseract OCR. The input includes an audio file and an image file. The audio file is sent to the Google Cloud Speech-to-Text API and text data is received. The image file is also processed with Tesseract OCR to extract characters. The output is text data corresponding to each audio and image data. Specifically, the server sends the audio file to the API and executes the process of receiving text data and extracting characters from the image file.
[0117] Step 3:
[0118] The server centralizes and stores the converted text data in a MySQL database. The input includes text data converted from audio data and image data. SQL commands are executed against the database to store the text data in the corresponding tables. The output is structured text data in the database. Specifically, the server issues SQL queries to insert each piece of text data into a specified column in the table.
[0119] Step 4:
[0120] The server analyzes the stored text data using a natural language processing engine such as BERT or GPT-3. The input includes the text data in the database. This is input into the natural language processing engine, which extracts important keywords and tags. The output is the keywords and tags assigned to each piece of text data. Specifically, the server sends the text data to the engine and executes the process of receiving the analysis results.
[0121] Step 5:
[0122] The server associates entrance exam questions with entertainment content using the extracted keywords and tags. The input includes the entrance exam question data and the extracted keywords and tags. Based on this, related content is selected and links are dynamically generated. The output is a mapping between entrance exam questions and entertainment content. Specifically, the server executes a process of matching the keywords in the entrance exam questions and generating links to the corresponding content.
[0123] Step 6:
[0124] An interface developed using React.js is provided to the user's device. The input includes search criteria selected by the user, such as by era or category. Based on this, a request is sent to the server, and search results for related books, movies, and music are received. As output, a list of content corresponding to the search criteria is displayed on the device. Specifically, the device receives the user's input, sends a request to the server, and receives and displays the results.
[0125] Step 7:
[0126] The user launches the learning app, and related entertainment content is displayed as they answer entrance exam questions. The input includes the entrance exam questions the user answers and their answers. Based on this, a request is sent to the server to retrieve related entertainment content. As an output, the related content is displayed as the user answers. Specifically, the device monitors the user's answering status, sends a request to the server, and retrieves and displays the related content.
[0127] Step 8:
[0128] The server periodically collects user learning history and content usage data and performs revenue sharing using the Stripe API. The input includes learning history data and content usage data. The output is aggregated statistical information and the results of revenue sharing. Specifically, the server periodically aggregates each data and executes a process to calculate and execute revenue sharing.
[0129] (Application example 1)
[0130] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0131] Traditional food delivery systems are limited to simply receiving and delivering orders, and are unable to provide educational or additional entertainment value to users. Users have no access to background knowledge or entertainment content related to the food they order, which limits the food delivery experience.
[0132] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0133] In this invention, the server includes means for analyzing text data using a natural language processing engine to extract keywords and tags, means for analyzing food delivery order details to map related educational and entertainment content, and means for providing a user interface that allows users to search for content by age group or category, thereby enabling users to enjoy educational and entertainment content related to the food they ordered.
[0134] A "natural language processing engine" is software that analyzes the meaning and structure of text data and extracts important keywords and tags.
[0135] "Keywords" are important words that represent the content of text data and are used for searching and classification.
[0136] A "tag" is metadata that represents the attributes or categories of text data, and is used to organize and search content more efficiently.
[0137] "Food delivery" refers to a service that delivers food ordered by users.
[0138] "Entertainment Content" means information or media that provides entertainment to users, including video, audio, and text.
[0139] A "user interface" is a means by which a user interacts with a system, and includes screens and input means for performing operations such as searches and information display.
[0140] "Order details" refers to the food items and their detailed information specified by the user in the food delivery service.
[0141] "Educational content" refers to information and media that provide users with knowledge and information and support learning.
[0142] "Mapping" is the act or process of associating and integrating related information or elements.
[0143] "User's order history" refers to a record of orders a user has placed using the food delivery service in the past.
[0144] "Content provider" means a person or organization that creates and provides entertainment or educational content.
[0145] This invention is a system for providing educational and entertainment content related to food ordered by a user in a food delivery system. This system is composed of a server, a terminal, and a user's perspective.
[0146] Data collection and analysis
[0147] To collect data, the server first acquires content data using APIs from various information providers. This includes text data, video data, cultural information, etc. The acquired video data is converted into text data using OCR and voice recognition technologies. This text data is then centrally stored in a database.
[0148] Next, the server launches a natural language processing engine (e.g., Spacy) to analyze the text data in the database. Through analysis, important keywords and tags are extracted from the text data. For example, from the data on "sushi," tags such as "Japanese cuisine," "history," and "rice" are generated.
[0149] Content and food delivery mapping
[0150] The server uses the acquired keywords and tags to link the user's order with related entertainment content. For example, a user who orders "sushi" might be linked to a documentary about the history of sushi or a video about recipes. In this way, content highly relevant to food delivery is dynamically linked.
[0151] Providing a user interface
[0152] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific food item or category, the server searches for related video and text content and sends the results to the device. This allows users to easily access a variety of content related to the food they ordered.
[0153] Learning and Content Links
[0154] A user launches a food delivery app and places an order. Once the order is confirmed, the user's device displays related educational and entertainment content. For example, when ordering "sushi," a documentary video introducing the history and cultural background of that sushi is displayed, allowing the user to deepen their knowledge by watching it.
[0155] Data Updates and Revenue Sharing
[0156] The server periodically collects user order history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is carried out. Revenue is distributed appropriately to educational and entertainment content providers. Specifically, if a particular video or article is viewed many times, a portion of the revenue will be distributed to that content provider.
[0157] Specific examples
[0158] For example, when a user orders "sushi," the user's device requests related entertainment content from the server based on the analyzed keywords "Japanese cuisine" and "history." The server then searches the database to find documentaries and recipe videos about the history of sushi and provides them to the user. In this way, the user can access a wealth of related content through their order.
[0159] Prompt Sentence Examples
[0160] "Dish: Sushi\nRelated Content: History, Recipes"
[0161] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0162] Step 1:
[0163] The server uses the information provider's API to collect various content data (text data, video data, cultural information, etc.). Specifically, it retrieves the necessary data from the database through API calls and stores it centrally. The input is data from the provider's API, and the collected raw data is stored in the database as output.
[0164] Step 2:
[0165] The server converts the collected video and audio data into text data. Specifically, it uses OCR and voice recognition technology to convert the data into text data and stores it in a database. The input is video data and audio data, and the converted text data is obtained as output.
[0166] Step 3:
[0167] The server runs a natural language processing engine (e.g., Spacy) to analyze the text data in the database. The analysis extracts important keywords and tags from the text data. The converted text data is the input, and the extracted keywords and tags are the output.
[0168] Step 4:
[0169] The server uses the retrieved keywords and tags to map the user's order to relevant educational and entertainment content. For example, a user who orders "sushi" might be mapped to a documentary about the history of sushi and a recipe video. The input is the user's order and the extracted keywords, and the output is the identification of relevant content.
[0170] Step 5:
[0171] The user device provides an interface that allows users to search for content by age or category. Specifically, it includes a content list and a search field. The input is a user selection or search query, and the output is a display of related content.
[0172] Step 6:
[0173] The user places an order, and once the order is confirmed, the user device displays related educational and entertainment content. For example, if you order "sushi," a video about the cultural background of that sushi will be displayed. The input is the user's order and a content list, and the output is the content to be displayed.
[0174] Step 7:
[0175] The server periodically collects user order history and content usage data. Based on this data, it calculates monthly usage statistics and distributes revenue. Specifically, if a particular piece of content is viewed many times, a portion of the revenue is distributed to the content provider. The inputs are usage data and order history, and the output is statistical data and revenue distribution results.
[0176] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0177] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments of the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0178] Data collection and analysis
[0179] To collect data, the server first uses the information provider's API to obtain various content data, including audio data, image data, books, and lyrics. This data is then converted into text data using voice recognition and OCR technology and stored in a database.
[0180] The server runs a natural language processing engine to analyze the text data in the database. Through this analysis, important keywords and tags are extracted from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0181] Mapping content to entrance exam questions
[0182] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." This allows for dynamic mapping of content highly relevant to the entrance exam questions.
[0183] Providing a user interface
[0184] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related books, movies, and music and sends the results to the device.
[0185] Learning and Content Links
[0186] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after solving a proverb question, a novel in which the proverb is used will be displayed, allowing users to deepen their understanding by reading the novel.
[0187] Applying the Emotion Engine
[0188] The server activates an emotion engine that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during the learning process. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0189] Data Updates and Revenue Sharing
[0190] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, which are then stored as monthly usage statistics. Revenues are calculated based on the number of times content is viewed and the duration of use, and distributed to entertainment content providers.
[0191] Specific examples
[0192] For example, suppose a user is taking an English literature entrance exam. The user's device analyzes the keyword "Victorian" and requests related entertainment content from the server. The server searches its database to find the novel "Jane Eyre" and movies set in that era, and provides them to the user.
[0193] The emotion engine also analyzes the user's facial expressions to detect when they are losing focus, and in that case, it will display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0194] The above is a concrete example of how to implement the present invention. This system evolves learning from mere knowledge acquisition to a rich experience that responds to the user's emotional state and interests.
[0195] The processing flow will be explained below.
[0196] A specific embodiment of the present invention will be described below, with the process flow divided into steps from the viewpoints of the server, the terminal, and the user.
[0197] Collecting content data and converting it into text
[0198] Step 1:
[0199] The server uses the information provider's API to obtain content data such as books, music, audio data, and image data.
[0200] Step 2:
[0201] The server converts the acquired voice data into text data using voice recognition technology, and also converts image data into text data using OCR technology. The converted text data is stored in a database.
[0202] Text data analysis and tagging
[0203] Step 3:
[0204] The server runs a natural language processing engine to analyze the text data in the database, extracting important keywords and tags from the text data.
[0205] Step 4:
[0206] The server indexes the text data based on the analysis results and associates relevant content with keywords and tags.
[0207] Mapping to entrance exam questions
[0208] Step 5:
[0209] The server uses the analyzed keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga."
[0210] Providing a learning interface
[0211] Step 6:
[0212] The user terminal provides an interface that allows the user to search for content by age group or category.
[0213] Step 7:
[0214] The user enters search criteria (grade, subject, keywords, etc.) and sends a search request to the server via the terminal.
[0215] Step 8:
[0216] The server queries the database based on the search request to find the appropriate content and returns the results to the user's device.
[0217] Step 9:
[0218] The user terminal displays the search results sent from the server to the user.
[0219] Learning and Content Links
[0220] Step 10:
[0221] The user launches a learning app and answers specific entrance exam questions.
[0222] Step 11:
[0223] The user terminal requests entertainment content related to the entrance exam question being answered from the server.
[0224] Step 12:
[0225] The server searches for highly relevant content and transmits the results to the user terminal.
[0226] Step 13:
[0227] The user device will then display related content to the user, for example, after solving a classical Japanese question, novels or movies related to that topic will be displayed.
[0228] Applying the Emotion Engine
[0229] Step 14:
[0230] The server activates the emotion engine and analyzes the user's facial expressions and voice, and evaluates the user's emotional state based on the analysis results.
[0231] Step 15:
[0232] Based on the evaluation results of the emotion engine, the server suggests appropriate entertainment content to assist learning, for example, displaying relaxing music to a tired user.
[0233] Data Updates and Revenue Sharing
[0234] Step 16:
[0235] The server periodically collects and analyzes the user's learning history, emotional state, and content usage data.
[0236] Step 17:
[0237] The server calculates monthly usage statistics based on the collected data and calculates revenue based on the number of times the content is viewed and the duration of use.
[0238] Step 18:
[0239] The server notifies the content provider of the results of the revenue distribution and distributes the reward.
[0240] These are the specific processing steps of this system, which effectively link learning and entertainment content to provide users with a rich and motivating learning experience.
[0241] Example 2
[0242] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0243] Conventional learning support systems have difficulty dynamically providing learning content according to the user's interests and emotions, which leads to problems such as reduced learning efficiency and motivation.In addition, there are limited ways to link entrance exam questions with entertainment content, making it difficult to attract users' interest.
[0244] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for converting audio data and image data into text data; means for analyzing entrance exam questions and dynamically linking related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for analyzing the user's facial expressions and voice and providing content according to their emotional state during study; and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This makes it possible to provide related content according to the user's emotions and interests, thereby improving study efficiency and motivation.
[0245] A "natural language processing engine" is a program that analyzes text data and extracts keywords and tags.
[0246] "Audio Data" refers to digital files containing speech or audio information.
[0247] "Image data" refers to digital files that contain visual information.
[0248] "Text data" refers to digital files that contain textual information.
[0249] "Keywords" refer to important words or phrases within text data.
[0250] A "tag" refers to a label for classifying text data.
[0251] "Entrance examination questions" refer to questions asked in entrance examinations conducted by educational institutions.
[0252] "Entertainment content" refers to media such as movies, music, and books that are intended to capture the user's interest.
[0253] "User interface" refers to the screen and input devices that users use to operate the system.
[0254] "Learning history" refers to data that records what a user has learned and their progress.
[0255] "Revenue sharing" refers to the process of distributing revenue generated by the system to stakeholders, such as content providers.
[0256] "Facial expression analysis" refers to the technology of analyzing a user's facial expressions to determine their emotional state.
[0257] "Voice analysis" refers to the technology of analyzing a user's voice to determine their emotional state and content.
[0258] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. An embodiment of the present invention will be described in detail below.
[0259] First, the server uses the information provider's API to collect various content data, such as audio data, image data, books, and song lyrics. Specifically, the server converts audio data into text data using Google Cloud Speech-to-Text, and image data into text data using Tesseract OCR software. The converted text data is then stored in the server's database.
[0260] The server then analyzes the stored text data using Hugging Face's Transformers library. This analysis uses generative AI models such as the BERT model to extract important keywords and tags from the text. For example, it generates tags such as "cat," "modern literature," and "author name" from a novel.
[0261] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. For example, a question on "Japanese history" might be associated with "historical films" or "manga." This allows highly relevant content to be mapped to the exam questions.
[0262] The user device provides a user interface that allows users to search for content by age group or category. When a user selects a specific grade or subject, a request is sent to the server. The server searches for related books, movies, and music and sends the results to the user device.
[0263] The user launches the learning app and answers the selected entrance exam questions. As the user answers, the device displays related entertainment content retrieved from the server. For example, after solving a proverb question, the device can display a novel in which the proverb is used, deepening the user's understanding.
[0264] The server also uses Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The emotion engine detects the user's emotional state during training and provides appropriate content. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0265] In addition, the server periodically collects and analyzes users' learning history, content usage data, and emotional data. Based on this, the server stores monthly usage statistics and distributes revenue to content providers. This data is used to calculate remuneration based on which content is viewed how many times and for how long.
[0266] For example, if a user is solving English literature entrance exam questions, the user's device will request related entertainment content from the server based on the analyzed keyword "Victorian era." The server will then search the database to find the novel "Jane Eyre" and movies set in that era and provide them to the user. Furthermore, if the emotion engine analyzes the user's facial expressions and detects that the user is losing concentration, it can display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0267] An example of a prompt might be:
[0268] Suggest content related to the theme "Victorian Era".
[0269] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0270] Step 1: Data collection
[0271] The server uses the APIs of information providers (e.g., book API, music API, image API, etc.) to collect various content data such as audio data, image data, book data, and lyric data. The input is various data obtained from the information providers, and the output is the collected data set. Specifically, it sends a request to the API and stores the obtained data in the storage system within the server.
[0272] Step 2: Text data conversion
[0273] The server uses Google Cloud Speech-to-Text to convert collected voice data to text data, and Tesseract OCR software to convert image data to text data. The input is digital audio or image files, and the output is text data that is converted and stored in a database. Specific operations include reading voice data and sending a conversion request, and reading image data and performing OCR processing.
[0274] Step 3: Natural Language Processing Analysis
[0275] The server uses Hugging Face's Transformers library to analyze the saved text data. The input is the transformed text data, and the output is analyzed data containing keywords and tags extracted from the text data. Specifically, it analyzes the text data using a generative AI model such as the BERT model to extract important keywords and tags. For example, it generates tags such as "cat," "modern literature," and "author name" from novel data.
[0276] Step 4: Mapping content to exam questions
[0277] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. The input is the keywords and tags generated by analysis and the entrance exam questions, and the output is a mapped dataset. Specifically, it matches entrance exam questions with highly relevant keywords with content, for example, associating "historical films" and "manga" with questions on "Japanese history."
[0278] Step 5: Providing a User Interface
[0279] The user device provides an interface that allows users to search for content by age group or category. The input is the user's selected grade level or subject, and the output is search results for related books, movies, and music returned by the server. Specifically, the device receives the user's selection through the interface and sends the request to the server, which searches the database and returns the results.
[0280] Step 6: Learning and Content Linking
[0281] The user launches the learning app and answers the selected entrance exam questions. The input is the user's answer data, and the output is related entertainment content. Specifically, when the user answers a question, books and movies related to that question are displayed. For example, after answering a proverb question, a novel in which that proverb is used is displayed.
[0282] Step 7: Applying the Emotion Engine
[0283] The server invokes Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The input is the user's real-time facial and voice data, and the output is the provision of content based on the analyzed emotional state. Specifically, the system captures the user's facial expressions and voice from the camera and microphone and analyzes them with the emotion engine. For example, if the user is tired, it provides relaxing music or light entertainment content.
[0284] Step 8: Data Updates and Revenue Sharing
[0285] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. The input is various user usage data, and the output is monthly usage statistics and revenue distribution data. Specifically, it analyzes the learning history and usage data stored in the database, tallying up how much of each piece of content was used, and distributes revenue to content providers based on that information.
[0286] (Application example 2)
[0287] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0288] Conventional learning support systems struggle to effectively link the content users are learning with entertainment content, and lack the means to maintain users' motivation and concentration. Furthermore, they rarely provide content that takes into account the user's emotional state, making it difficult to meet individual needs. This results in a limited learning experience and an inability to provide an optimal learning environment for each user.
[0289] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for analyzing entrance exam questions and mapping related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for providing content that corresponds to the user's emotional state using an emotion engine that analyzes the user's emotional state; and means for analyzing the user's learning history and content usage data and distributing revenue to content providers. As a result, users are provided with optimal content that corresponds to their emotional state while studying, enabling them to maintain their motivation and concentration and progress with their studies more effectively.
[0290] A "natural language processing engine" is a technology that analyzes text data and extracts important information such as keywords and tags.
[0291] "Keywords" refer to important words or phrases in text data, and represent the content of a document.
[0292] A "tag" is a label that represents a category or attribute that is assigned to classify text data.
[0293] "Entertainment content" refers to digital content and media such as movies, novels, music, and games that provide users with enjoyment and entertainment.
[0294] "User interface" refers to the screen and operation method that allows a user to interact with a system, making it easier for the user to operate it.
[0295] The "emotion engine" is a technology that analyzes the user's facial expressions and voice to detect their emotional state at that time.
[0296] "Study history" refers to a record of what a user has studied and the questions they have answered, and is data used to understand an individual's learning progress.
[0297] "Revenue sharing" is a method for fairly distributing revenue generated within the system to content providers and other stakeholders.
[0298] "Associating" refers to building semantic connections between different data or content.
[0299] "Dynamic linking" means connecting relevant content and information in real time.
[0300] "Content usage data" refers to records of which content a user has used and to what extent, and is data used to analyze user usage.
[0301] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments for carrying out the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0302] Data collection and analysis
[0303] The server uses the information provider's API to collect various content data, including audio data, image data, and text data. This data is then centrally converted into text data using voice recognition and OCR technology and stored in a database. The server then launches a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags.
[0304] Mapping content to entrance exam questions
[0305] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "literary works." This process is performed using database search and dynamic link generation technology.
[0306] Providing a user interface
[0307] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related movies, books, and music and sends the results to the device. This interface is typically run on devices such as smartphones or head-mounted displays (HMDs).
[0308] Learning and Content Links
[0309] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after answering an English literature question, a related novel will be displayed, allowing the user to read the novel and deepen their understanding.
[0310] Applying the Emotion Engine
[0311] The server launches an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or light entertainment content accordingly.
[0312] Data Updates and Revenue Sharing
[0313] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. This data is stored as monthly usage statistics, and revenue is calculated and distributed to content providers based on the number of times the content is viewed and the duration of use.
[0314] Specific examples
[0315] For example, suppose a user is solving an English literature entrance exam. The user's device requests related entertainment content from the server based on the keyword "Victorian" that was analyzed during the process. The server then searches its database to find the novel "Jane Eyre" and films set in that era, and provides them to the user. The emotion engine also analyzes the user's facial expressions to detect a lapse in concentration. In that case, it displays visually interesting content related to the learning content (such as a documentary film).
[0316] Prompt Sentence Examples
[0317] List entertainment content related to the Victorian era, and provide visually engaging documentaries to help users with short attention spans.
[0318] As described above, this invention evolves learning from mere acquisition of knowledge to a rich experience that responds to the user's emotional state and interests.
[0319] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0320] Step 1:
[0321] The server collects audio data, image data, and text data through the information provider's API. The collected data is stored in raw form on the server (input: audio data, image data, and text data obtained via API; output: raw data stored on the server).
[0322] Step 2:
[0323] The server converts raw data into text data using speech recognition and OCR technology. Voice data is converted into text using speech recognition software (e.g., SpeechRecognition), and image data is converted into text data using OCR software (e.g., Tesseract OCR) (input: raw data; output: converted text data).
[0324] Step 3:
[0325] The server uses a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags, which are then stored in a database in a summarized form (input: text data; output: extracted keywords and tags).
[0326] Step 4:
[0327] The server uses the analyzed keywords and tags to map entrance exam questions to entertainment content. For example, an entrance exam question on "Japanese history" is associated with historical movies and novels (input: entrance exam questions and keywords / tags; output: related entertainment content).
[0328] Step 5:
[0329] It provides an interface on the user's device that allows the user to search for content by age and category. Users can search for a specific grade level or subject and display related movies, books, and music (input: user's search query; output: list of related content).
[0330] Step 6:
[0331] A user uses a learning app to answer entrance exam questions. The user's device displays related entertainment content as the user progresses. For example, if the user answers a question about a proverb, the app displays a literary work in which the proverb is used (input: entrance exam question being answered; output: related entertainment content).
[0332] Step 7:
[0333] The server uses an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) to analyze the user's facial expressions and voice to detect the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or visually interesting content (input: user's facial and voice data; output: user's emotional state and corresponding content).
[0334] Step 8:
[0335] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, generating monthly usage statistics and distributing revenue to content providers based on the number of times content is viewed and the duration of usage (input: learning history, content usage data, emotional data; output: revenue distribution data and statistics on the number of times content is viewed / used).
[0336] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0337] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0338] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0339] [Second embodiment]
[0340] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0341] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0342] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0343] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0344] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0345] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0346] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0347] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0348] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0349] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0350] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0351] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0352] The present invention relates to a system for linking learning and entertainment content using a natural language processing engine. The following describes an embodiment of the system from the viewpoints of a server, a terminal, and a user.
[0353] Data collection and analysis
[0354] To collect data, the server first acquires content data using APIs from various information providers. This includes audio data, image data, books, and song lyrics. The acquired audio data is converted into text data using voice recognition technology, and image data is converted into text data using OCR technology. This text data is then centrally stored in a database.
[0355] The server then launches a natural language processing engine to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0356] Mapping content to entrance exam questions
[0357] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0358] Providing a user interface
[0359] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific grade level or subject, the server searches for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0360] Learning and Content Links
[0361] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, if a user answers a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from that list, they can gain a deeper understanding of the historical background.
[0362] Data Updates and Revenue Sharing
[0363] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue is distributed. The revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue is distributed to that content provider.
[0364] Specific examples
[0365] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0366] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0367] The processing flow will be explained below.
[0368] Step 1:
[0369] The server uses the API of the information provider to obtain various content data (books, lyrics, audio data, image data), including audio and image data.
[0370] Step 2:
[0371] The server converts the acquired voice data into text data using voice recognition technology, and converts the image data into text data using OCR (optical character recognition) technology. The converted text data is stored in a database.
[0372] Step 3:
[0373] The server runs a natural language processing (NLP) engine to analyze the text data in the database, extracting important keywords, tags, and content names from the text data.
[0374] Step 4:
[0375] The server associates entrance exam questions with entertainment content based on the keywords and tags extracted through the analysis. For example, a question on Japanese history might be linked to "historical films" or "manga."
[0376] Step 5:
[0377] The server provides a user interface through which users can search for content by age and category.
[0378] Step 6:
[0379] The user terminal sends a search request to the server, for example, if the user wants to search for historical content from a particular era.
[0380] Step 7:
[0381] The server queries the database based on the received search request and returns appropriate results to the user's device, such as a list of movies and novels related to the "Warring States Period" arrow.
[0382] Step 8:
[0383] Users simply launch the learning app, select a specific exam question, and begin answering it. Related entertainment content is automatically displayed while they answer the question.
[0384] Step 9:
[0385] The user device will display related entertainment content based on the category and keywords of the entrance exam questions the user has answered. For example, after answering a question on classical literature, novels and movies related to that topic will be displayed.
[0386] Step 10:
[0387] The server periodically collects and analyzes users' learning history and content usage data, which is then stored in the system as monthly usage statistics.
[0388] Step 11:
[0389] The server calculates revenue based on the collected usage data and distributes it to entertainment content providers, based on the number of views and duration of use.
[0390] Step 12:
[0391] The server notifies both users and content providers of the results of revenue sharing and feedback on learnings, which helps optimize the system.
[0392] The above are the specific processing steps of this system, which effectively link learning content and entertainment content to provide users with a rich learning experience.
[0393] Example 1
[0394] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0395] In conventional learning systems, it is difficult for users to easily access external information or entertainment content directly related to the problem they are solving. Furthermore, they lack a mechanism for effectively managing users' learning data and appropriately distributing revenue to entertainment content providers. This results in low learning effectiveness and makes it difficult to promote deep understanding through entertainment.
[0396] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0397] In this invention, the server includes means for analyzing text data using a natural language processing engine and extracting keywords and tags, means for analyzing entrance exam questions and mapping related entertainment content, means for providing a user interface that allows users to search for content by age group or category, means for converting audio data and image data into text data, means for displaying content related to a question when the user solves it, and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This allows users to easily access related entertainment content while studying, and enables appropriate management of study data and revenue distribution.
[0398] A "natural language processing engine" is a software engine that analyzes text data and extracts important keywords and tags.
[0399] "Keywords" are important words or phrases extracted from text data that are used to categorize and associate content and entrance exam questions.
[0400] "Tags" are metadata attached to text data, and are labels that indicate the characteristics of the content or entrance exam questions.
[0401] "Speech recognition technology" is a technology that converts voice data into text data.
[0402] "Optical character recognition technology" is a technology that analyzes image data and extracts the characters contained therein as text data.
[0403] A "user interface" is a visual interface that allows users to search for content by age group or category.
[0404] "Study history" refers to historical data of entrance exam questions that a user has answered using a learning app.
[0405] "Revenue sharing" refers to the process of appropriately allocating revenue to entertainment content providers based on users' learning history and content usage data.
[0406] "Entrance exam questions" are exam questions designed for users to answer.
[0407] "Entertainment content" refers to content such as books, movies, music, and manga that are provided to allow users to deepen their learning while having fun.
[0408] This invention relates to a system for linking learning and entertainment content using a natural language processing engine, and an embodiment thereof will be described in detail from the viewpoints of a server, a terminal, and a user.
[0409] Data collection and analysis
[0410] To collect data, the server first obtains content data using APIs from various information providers. Specifically, it uses the Google Books API, Spotify API, YouTube Data API, etc. This content data includes audio data, image data, books, and song lyrics. The obtained audio data is converted into text data using Google Cloud Speech-to-Text, and image data is converted into text data using Tesseract OCR technology. This text data is then centrally stored in a MySQL database.
[0411] Next, the server launches a natural language processing engine such as BERT or GPT-3 to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0412] Mapping content to entrance exam questions
[0413] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" could be linked to "historical films" and "manga" that can be obtained from multiple media. In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0414] Providing a user interface
[0415] User devices are provided with an interface developed using React.js, which allows users to search for content by age group and category. When a user selects a specific grade or subject, the server uses Elasticsearch to quickly search for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0416] Learning and Content Links
[0417] When a user launches the learning app and answers entrance exam questions, related entertainment content is displayed. As the user proceeds with the answer, the user's device requests related entertainment content from the server. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, when a user solves a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from the list, the user can gain a deeper understanding of the historical background.
[0418] Data Updates and Revenue Sharing
[0419] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is performed using the Stripe API. Revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue will be distributed to that content provider.
[0420] Specific examples
[0421] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0422] Example prompt: "Please provide entertainment content related to the history of the Heian period. Preferably include novels, movies, and manga."
[0423] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0424] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0425] Step 1:
[0426] To collect data, the server uses APIs from information providers such as Google Books API, Spotify API, and YouTube Data API. The input includes requests to each API. For example, to obtain data such as book information, music information, and video information, an HTTP request is sent to each API. The output is response data from each API. This allows book metadata, music metadata, and video metadata to be obtained. Specifically, the server sends an asynchronous request to each API and waits for a response.
[0427] Step 2:
[0428] The server converts the acquired audio data into text data using Google Cloud Speech-to-Text, and converts the image data into text data using Tesseract OCR. The input includes an audio file and an image file. The audio file is sent to the Google Cloud Speech-to-Text API and text data is received. The image file is also processed with Tesseract OCR to extract characters. The output is text data corresponding to each audio and image data. Specifically, the server sends the audio file to the API and executes the process of receiving text data and extracting characters from the image file.
[0429] Step 3:
[0430] The server centralizes and stores the converted text data in a MySQL database. The input includes text data converted from audio data and image data. SQL commands are executed against the database to store the text data in the corresponding tables. The output is structured text data in the database. Specifically, the server issues SQL queries to insert each piece of text data into a specified column in the table.
[0431] Step 4:
[0432] The server analyzes the stored text data using a natural language processing engine such as BERT or GPT-3. The input includes the text data in the database. This is input into the natural language processing engine, which extracts important keywords and tags. The output is the keywords and tags assigned to each piece of text data. Specifically, the server sends the text data to the engine and executes the process of receiving the analysis results.
[0433] Step 5:
[0434] The server associates entrance exam questions with entertainment content using the extracted keywords and tags. The input includes the entrance exam question data and the extracted keywords and tags. Based on this, related content is selected and links are dynamically generated. The output is a mapping between entrance exam questions and entertainment content. Specifically, the server executes a process of matching the keywords in the entrance exam questions and generating links to the corresponding content.
[0435] Step 6:
[0436] An interface developed using React.js is provided to the user's device. The input includes search criteria selected by the user, such as by era or category. Based on this, a request is sent to the server, and search results for related books, movies, and music are received. As output, a list of content corresponding to the search criteria is displayed on the device. Specifically, the device receives the user's input, sends a request to the server, and receives and displays the results.
[0437] Step 7:
[0438] The user launches the learning app, and related entertainment content is displayed as they answer entrance exam questions. The input includes the entrance exam questions the user answers and their answers. Based on this, a request is sent to the server to retrieve related entertainment content. As an output, the related content is displayed as the user answers. Specifically, the device monitors the user's answering status, sends a request to the server, and retrieves and displays the related content.
[0439] Step 8:
[0440] The server periodically collects user learning history and content usage data and performs revenue sharing using the Stripe API. The input includes learning history data and content usage data. The output is aggregated statistical information and the results of revenue sharing. Specifically, the server periodically aggregates each data and executes a process to calculate and execute revenue sharing.
[0441] (Application example 1)
[0442] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0443] Traditional food delivery systems are limited to simply receiving and delivering orders, and are unable to provide educational or additional entertainment value to users. Users have no access to background knowledge or entertainment content related to the food they order, which limits the food delivery experience.
[0444] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0445] In this invention, the server includes means for analyzing text data using a natural language processing engine to extract keywords and tags, means for analyzing food delivery order details to map related educational and entertainment content, and means for providing a user interface that allows users to search for content by age group or category, thereby enabling users to enjoy educational and entertainment content related to the food they ordered.
[0446] A "natural language processing engine" is software that analyzes the meaning and structure of text data and extracts important keywords and tags.
[0447] "Keywords" are important words that represent the content of text data and are used for searching and classification.
[0448] A "tag" is metadata that represents the attributes or categories of text data, and is used to organize and search content more efficiently.
[0449] "Food delivery" refers to a service that delivers food ordered by users.
[0450] "Entertainment Content" means information or media that provides entertainment to users, including video, audio, and text.
[0451] A "user interface" is a means by which a user interacts with a system, and includes screens and input means for performing operations such as searches and information display.
[0452] "Order details" refers to the food items and their detailed information specified by the user in the food delivery service.
[0453] "Educational content" refers to information and media that provide users with knowledge and information and support learning.
[0454] "Mapping" is the act or process of associating and integrating related information or elements.
[0455] "User's order history" refers to a record of orders a user has placed using the food delivery service in the past.
[0456] "Content provider" means a person or organization that creates and provides entertainment or educational content.
[0457] This invention is a system for providing educational and entertainment content related to food ordered by a user in a food delivery system. This system is composed of a server, a terminal, and a user's perspective.
[0458] Data collection and analysis
[0459] To collect data, the server first acquires content data using APIs from various information providers. This includes text data, video data, cultural information, etc. The acquired video data is converted into text data using OCR and voice recognition technologies. This text data is then centrally stored in a database.
[0460] Next, the server launches a natural language processing engine (e.g., Spacy) to analyze the text data in the database. Through analysis, important keywords and tags are extracted from the text data. For example, from the data on "sushi," tags such as "Japanese cuisine," "history," and "rice" are generated.
[0461] Content and food delivery mapping
[0462] The server uses the acquired keywords and tags to link the user's order with related entertainment content. For example, a user who orders "sushi" might be linked to a documentary about the history of sushi or a video about recipes. In this way, content highly relevant to food delivery is dynamically linked.
[0463] Providing a user interface
[0464] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific food item or category, the server searches for related video and text content and sends the results to the device. This allows users to easily access a variety of content related to the food they ordered.
[0465] Learning and Content Links
[0466] A user launches a food delivery app and places an order. Once the order is confirmed, the user's device displays related educational and entertainment content. For example, when ordering "sushi," a documentary video introducing the history and cultural background of that sushi is displayed, allowing the user to deepen their knowledge by watching it.
[0467] Data Updates and Revenue Sharing
[0468] The server periodically collects user order history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is carried out. Revenue is distributed appropriately to educational and entertainment content providers. Specifically, if a particular video or article is viewed many times, a portion of the revenue will be distributed to that content provider.
[0469] Specific examples
[0470] For example, when a user orders "sushi," the user's device requests related entertainment content from the server based on the analyzed keywords "Japanese cuisine" and "history." The server then searches the database to find documentaries and recipe videos about the history of sushi and provides them to the user. In this way, the user can access a wealth of related content through their order.
[0471] Prompt Sentence Examples
[0472] "Dish: Sushi\nRelated Content: History, Recipes"
[0473] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0474] Step 1:
[0475] The server uses the information provider's API to collect various content data (text data, video data, cultural information, etc.). Specifically, it retrieves the necessary data from the database through API calls and stores it centrally. The input is data from the provider's API, and the collected raw data is stored in the database as output.
[0476] Step 2:
[0477] The server converts the collected video and audio data into text data. Specifically, it uses OCR and voice recognition technology to convert the data into text data and stores it in a database. The input is video data and audio data, and the converted text data is obtained as output.
[0478] Step 3:
[0479] The server runs a natural language processing engine (e.g., Spacy) to analyze the text data in the database. The analysis extracts important keywords and tags from the text data. The converted text data is the input, and the extracted keywords and tags are the output.
[0480] Step 4:
[0481] The server uses the retrieved keywords and tags to map the user's order to relevant educational and entertainment content. For example, a user who orders "sushi" might be mapped to a documentary about the history of sushi and a recipe video. The input is the user's order and the extracted keywords, and the output is the identification of relevant content.
[0482] Step 5:
[0483] The user device provides an interface that allows users to search for content by age or category. Specifically, it includes a content list and a search field. The input is a user selection or search query, and the output is a display of related content.
[0484] Step 6:
[0485] The user places an order, and once the order is confirmed, the user device displays related educational and entertainment content. For example, if you order "sushi," a video about the cultural background of that sushi will be displayed. The input is the user's order and a content list, and the output is the content to be displayed.
[0486] Step 7:
[0487] The server periodically collects user order history and content usage data. Based on this data, it calculates monthly usage statistics and distributes revenue. Specifically, if a particular piece of content is viewed many times, a portion of the revenue is distributed to the content provider. The inputs are usage data and order history, and the output is statistical data and revenue distribution results.
[0488] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0489] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments of the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0490] Data collection and analysis
[0491] To collect data, the server first uses the information provider's API to obtain various content data, including audio data, image data, books, and lyrics. This data is then converted into text data using voice recognition and OCR technology and stored in a database.
[0492] The server runs a natural language processing engine to analyze the text data in the database. Through this analysis, important keywords and tags are extracted from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0493] Mapping content to entrance exam questions
[0494] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." This allows for dynamic mapping of content highly relevant to the entrance exam questions.
[0495] Providing a user interface
[0496] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related books, movies, and music and sends the results to the device.
[0497] Learning and Content Links
[0498] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after solving a proverb question, a novel in which the proverb is used will be displayed, allowing users to deepen their understanding by reading the novel.
[0499] Applying the Emotion Engine
[0500] The server activates an emotion engine that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during the learning process. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0501] Data Updates and Revenue Sharing
[0502] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, which are then stored as monthly usage statistics. Revenues are calculated based on the number of times content is viewed and the duration of use, and distributed to entertainment content providers.
[0503] Specific examples
[0504] For example, suppose a user is taking an English literature entrance exam. The user's device analyzes the keyword "Victorian" and requests related entertainment content from the server. The server searches its database to find the novel "Jane Eyre" and movies set in that era, and provides them to the user.
[0505] The emotion engine also analyzes the user's facial expressions to detect when they are losing focus, and in that case, it will display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0506] The above is a concrete example of how to implement the present invention. This system evolves learning from mere knowledge acquisition to a rich experience that responds to the user's emotional state and interests.
[0507] The processing flow will be explained below.
[0508] A specific embodiment of the present invention will be described below, with the process flow divided into steps from the viewpoints of the server, the terminal, and the user.
[0509] Collecting content data and converting it into text
[0510] Step 1:
[0511] The server uses the information provider's API to obtain content data such as books, music, audio data, and image data.
[0512] Step 2:
[0513] The server converts the acquired voice data into text data using voice recognition technology, and also converts image data into text data using OCR technology. The converted text data is stored in a database.
[0514] Text data analysis and tagging
[0515] Step 3:
[0516] The server runs a natural language processing engine to analyze the text data in the database, extracting important keywords and tags from the text data.
[0517] Step 4:
[0518] The server indexes the text data based on the analysis results and associates relevant content with keywords and tags.
[0519] Mapping to entrance exam questions
[0520] Step 5:
[0521] The server uses the analyzed keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga."
[0522] Providing a learning interface
[0523] Step 6:
[0524] The user terminal provides an interface that allows the user to search for content by age group or category.
[0525] Step 7:
[0526] The user enters search criteria (grade, subject, keywords, etc.) and sends a search request to the server via the terminal.
[0527] Step 8:
[0528] The server queries the database based on the search request to find the appropriate content and returns the results to the user's device.
[0529] Step 9:
[0530] The user terminal displays the search results sent from the server to the user.
[0531] Learning and Content Links
[0532] Step 10:
[0533] The user launches a learning app and answers specific entrance exam questions.
[0534] Step 11:
[0535] The user terminal requests entertainment content related to the entrance exam question being answered from the server.
[0536] Step 12:
[0537] The server searches for highly relevant content and transmits the results to the user terminal.
[0538] Step 13:
[0539] The user device will then display related content to the user, for example, after solving a classical Japanese question, novels or movies related to that topic will be displayed.
[0540] Applying the Emotion Engine
[0541] Step 14:
[0542] The server activates the emotion engine and analyzes the user's facial expressions and voice, and evaluates the user's emotional state based on the analysis results.
[0543] Step 15:
[0544] Based on the evaluation results of the emotion engine, the server suggests appropriate entertainment content to assist learning, for example, displaying relaxing music to a tired user.
[0545] Data Updates and Revenue Sharing
[0546] Step 16:
[0547] The server periodically collects and analyzes the user's learning history, emotional state, and content usage data.
[0548] Step 17:
[0549] The server calculates monthly usage statistics based on the collected data and calculates revenue based on the number of times the content is viewed and the duration of use.
[0550] Step 18:
[0551] The server notifies the content provider of the results of the revenue distribution and distributes the reward.
[0552] These are the specific processing steps of this system, which effectively link learning and entertainment content to provide users with a rich and motivating learning experience.
[0553] Example 2
[0554] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0555] Conventional learning support systems have difficulty dynamically providing learning content according to the user's interests and emotions, which leads to problems such as reduced learning efficiency and motivation.In addition, there are limited ways to link entrance exam questions with entertainment content, making it difficult to attract users' interest.
[0556] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for converting audio data and image data into text data; means for analyzing entrance exam questions and dynamically linking related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for analyzing the user's facial expressions and voice and providing content according to their emotional state during study; and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This makes it possible to provide related content according to the user's emotions and interests, thereby improving study efficiency and motivation.
[0557] A "natural language processing engine" is a program that analyzes text data and extracts keywords and tags.
[0558] "Audio Data" refers to digital files containing speech or audio information.
[0559] "Image data" refers to digital files that contain visual information.
[0560] "Text data" refers to digital files that contain textual information.
[0561] "Keywords" refer to important words or phrases within text data.
[0562] A "tag" refers to a label for classifying text data.
[0563] "Entrance examination questions" refer to questions asked in entrance examinations conducted by educational institutions.
[0564] "Entertainment content" refers to media such as movies, music, and books that are intended to capture the user's interest.
[0565] "User interface" refers to the screen and input devices that users use to operate the system.
[0566] "Learning history" refers to data that records what a user has learned and their progress.
[0567] "Revenue sharing" refers to the process of distributing revenue generated by the system to stakeholders, such as content providers.
[0568] "Facial expression analysis" refers to the technology of analyzing a user's facial expressions to determine their emotional state.
[0569] "Voice analysis" refers to the technology of analyzing a user's voice to determine their emotional state and content.
[0570] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. An embodiment of the present invention will be described in detail below.
[0571] First, the server uses the information provider's API to collect various content data, such as audio data, image data, books, and song lyrics. Specifically, the server converts audio data into text data using Google Cloud Speech-to-Text, and image data into text data using Tesseract OCR software. The converted text data is then stored in the server's database.
[0572] The server then analyzes the stored text data using Hugging Face's Transformers library. This analysis uses generative AI models such as the BERT model to extract important keywords and tags from the text. For example, it generates tags such as "cat," "modern literature," and "author name" from a novel.
[0573] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. For example, a question on "Japanese history" might be associated with "historical films" or "manga." This allows highly relevant content to be mapped to the exam questions.
[0574] The user device provides a user interface that allows users to search for content by age group or category. When a user selects a specific grade or subject, a request is sent to the server. The server searches for related books, movies, and music and sends the results to the user device.
[0575] The user launches the learning app and answers the selected entrance exam questions. As the user answers, the device displays related entertainment content retrieved from the server. For example, after solving a proverb question, the device can display a novel in which the proverb is used, deepening the user's understanding.
[0576] The server also uses Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The emotion engine detects the user's emotional state during training and provides appropriate content. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0577] In addition, the server periodically collects and analyzes users' learning history, content usage data, and emotional data. Based on this, the server stores monthly usage statistics and distributes revenue to content providers. This data is used to calculate remuneration based on which content is viewed how many times and for how long.
[0578] For example, if a user is solving English literature entrance exam questions, the user's device will request related entertainment content from the server based on the analyzed keyword "Victorian era." The server will then search the database to find the novel "Jane Eyre" and movies set in that era and provide them to the user. Furthermore, if the emotion engine analyzes the user's facial expressions and detects that the user is losing concentration, it can display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0579] An example of a prompt might be:
[0580] Suggest content related to the theme "Victorian Era".
[0581] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0582] Step 1: Data collection
[0583] The server uses the APIs of information providers (e.g., book API, music API, image API, etc.) to collect various content data such as audio data, image data, book data, and lyric data. The input is various data obtained from the information providers, and the output is the collected data set. Specifically, it sends a request to the API and stores the obtained data in the storage system within the server.
[0584] Step 2: Text data conversion
[0585] The server uses Google Cloud Speech-to-Text to convert collected voice data to text data, and Tesseract OCR software to convert image data to text data. The input is digital audio or image files, and the output is text data that is converted and stored in a database. Specific operations include reading voice data and sending a conversion request, and reading image data and performing OCR processing.
[0586] Step 3: Natural Language Processing Analysis
[0587] The server uses Hugging Face's Transformers library to analyze the saved text data. The input is the transformed text data, and the output is analyzed data containing keywords and tags extracted from the text data. Specifically, it analyzes the text data using a generative AI model such as the BERT model to extract important keywords and tags. For example, it generates tags such as "cat," "modern literature," and "author name" from novel data.
[0588] Step 4: Mapping content to exam questions
[0589] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. The input is the keywords and tags generated by analysis and the entrance exam questions, and the output is a mapped dataset. Specifically, it matches entrance exam questions with highly relevant keywords with content, for example, associating "historical films" and "manga" with questions on "Japanese history."
[0590] Step 5: Providing a User Interface
[0591] The user device provides an interface that allows users to search for content by age group or category. The input is the user's selected grade level or subject, and the output is search results for related books, movies, and music returned by the server. Specifically, the device receives the user's selection through the interface and sends the request to the server, which searches the database and returns the results.
[0592] Step 6: Learning and Content Linking
[0593] The user launches the learning app and answers the selected entrance exam questions. The input is the user's answer data, and the output is related entertainment content. Specifically, when the user answers a question, books and movies related to that question are displayed. For example, after answering a proverb question, a novel in which that proverb is used is displayed.
[0594] Step 7: Applying the Emotion Engine
[0595] The server invokes Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The input is the user's real-time facial and voice data, and the output is the provision of content based on the analyzed emotional state. Specifically, the system captures the user's facial expressions and voice from the camera and microphone and analyzes them with the emotion engine. For example, if the user is tired, it provides relaxing music or light entertainment content.
[0596] Step 8: Data Updates and Revenue Sharing
[0597] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. The input is various user usage data, and the output is monthly usage statistics and revenue distribution data. Specifically, it analyzes the learning history and usage data stored in the database, tallying up how much of each piece of content was used, and distributes revenue to content providers based on that information.
[0598] (Application example 2)
[0599] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0600] Conventional learning support systems struggle to effectively link the content users are learning with entertainment content, and lack the means to maintain users' motivation and concentration. Furthermore, they rarely provide content that takes into account the user's emotional state, making it difficult to meet individual needs. This results in a limited learning experience and an inability to provide an optimal learning environment for each user.
[0601] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for analyzing entrance exam questions and mapping related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for providing content that corresponds to the user's emotional state using an emotion engine that analyzes the user's emotional state; and means for analyzing the user's learning history and content usage data and distributing revenue to content providers. As a result, users are provided with optimal content that corresponds to their emotional state while studying, enabling them to maintain their motivation and concentration and progress with their studies more effectively.
[0602] A "natural language processing engine" is a technology that analyzes text data and extracts important information such as keywords and tags.
[0603] "Keywords" refer to important words or phrases in text data, and represent the content of a document.
[0604] A "tag" is a label that represents a category or attribute that is assigned to classify text data.
[0605] "Entertainment content" refers to digital content and media such as movies, novels, music, and games that provide users with enjoyment and entertainment.
[0606] "User interface" refers to the screen and operation method that allows a user to interact with a system, making it easier for the user to operate it.
[0607] The "emotion engine" is a technology that analyzes the user's facial expressions and voice to detect their emotional state at that time.
[0608] "Study history" refers to a record of what a user has studied and the questions they have answered, and is data used to understand an individual's learning progress.
[0609] "Revenue sharing" is a method for fairly distributing revenue generated within the system to content providers and other stakeholders.
[0610] "Associating" refers to building semantic connections between different data or content.
[0611] "Dynamic linking" means connecting relevant content and information in real time.
[0612] "Content usage data" refers to records of which content a user has used and to what extent, and is data used to analyze user usage.
[0613] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments for carrying out the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0614] Data collection and analysis
[0615] The server uses the information provider's API to collect various content data, including audio data, image data, and text data. This data is then centrally converted into text data using voice recognition and OCR technology and stored in a database. The server then launches a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags.
[0616] Mapping content to entrance exam questions
[0617] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "literary works." This process is performed using database search and dynamic link generation technology.
[0618] Providing a user interface
[0619] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related movies, books, and music and sends the results to the device. This interface is typically run on devices such as smartphones or head-mounted displays (HMDs).
[0620] Learning and Content Links
[0621] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after answering an English literature question, a related novel will be displayed, allowing the user to read the novel and deepen their understanding.
[0622] Applying the Emotion Engine
[0623] The server launches an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or light entertainment content accordingly.
[0624] Data Updates and Revenue Sharing
[0625] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. This data is stored as monthly usage statistics, and revenue is calculated and distributed to content providers based on the number of times the content is viewed and the duration of use.
[0626] Specific examples
[0627] For example, suppose a user is solving an English literature entrance exam. The user's device requests related entertainment content from the server based on the keyword "Victorian" that was analyzed during the process. The server then searches its database to find the novel "Jane Eyre" and films set in that era, and provides them to the user. The emotion engine also analyzes the user's facial expressions to detect a lapse in concentration. In that case, it displays visually interesting content related to the learning content (such as a documentary film).
[0628] Prompt Sentence Examples
[0629] List entertainment content related to the Victorian era, and provide visually engaging documentaries to help users with short attention spans.
[0630] As described above, this invention evolves learning from mere acquisition of knowledge to a rich experience that responds to the user's emotional state and interests.
[0631] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0632] Step 1:
[0633] The server collects audio data, image data, and text data through the information provider's API. The collected data is stored in raw form on the server (input: audio data, image data, and text data obtained via API; output: raw data stored on the server).
[0634] Step 2:
[0635] The server converts raw data into text data using speech recognition and OCR technology. Voice data is converted into text using speech recognition software (e.g., SpeechRecognition), and image data is converted into text data using OCR software (e.g., Tesseract OCR) (input: raw data; output: converted text data).
[0636] Step 3:
[0637] The server uses a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags, which are then stored in a database in a summarized form (input: text data; output: extracted keywords and tags).
[0638] Step 4:
[0639] The server uses the analyzed keywords and tags to map entrance exam questions to entertainment content. For example, an entrance exam question on "Japanese history" is associated with historical movies and novels (input: entrance exam questions and keywords / tags; output: related entertainment content).
[0640] Step 5:
[0641] It provides an interface on the user's device that allows the user to search for content by age and category. Users can search for a specific grade level or subject and display related movies, books, and music (input: user's search query; output: list of related content).
[0642] Step 6:
[0643] A user uses a learning app to answer entrance exam questions. The user's device displays related entertainment content as the user progresses. For example, if the user answers a question about a proverb, the app displays a literary work in which the proverb is used (input: entrance exam question being answered; output: related entertainment content).
[0644] Step 7:
[0645] The server uses an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) to analyze the user's facial expressions and voice to detect the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or visually interesting content (input: user's facial and voice data; output: user's emotional state and corresponding content).
[0646] Step 8:
[0647] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, generating monthly usage statistics and distributing revenue to content providers based on the number of times content is viewed and the duration of usage (input: learning history, content usage data, emotional data; output: revenue distribution data and statistics on the number of times content is viewed / used).
[0648] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0649] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0650] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0651] [Third embodiment]
[0652] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0653] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0654] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0655] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0656] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0657] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0658] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0659] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0660] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0661] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0662] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0663] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0664] The present invention relates to a system for linking learning and entertainment content using a natural language processing engine. The following describes an embodiment of the system from the viewpoints of a server, a terminal, and a user.
[0665] Data collection and analysis
[0666] To collect data, the server first acquires content data using APIs from various information providers. This includes audio data, image data, books, and song lyrics. The acquired audio data is converted into text data using voice recognition technology, and image data is converted into text data using OCR technology. This text data is then centrally stored in a database.
[0667] The server then launches a natural language processing engine to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0668] Mapping content to entrance exam questions
[0669] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0670] Providing a user interface
[0671] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific grade level or subject, the server searches for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0672] Learning and Content Links
[0673] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, if a user answers a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from that list, they can gain a deeper understanding of the historical background.
[0674] Data Updates and Revenue Sharing
[0675] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue is distributed. The revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue is distributed to that content provider.
[0676] Specific examples
[0677] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0678] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0679] The processing flow will be explained below.
[0680] Step 1:
[0681] The server uses the API of the information provider to obtain various content data (books, lyrics, audio data, image data), including audio and image data.
[0682] Step 2:
[0683] The server converts the acquired voice data into text data using voice recognition technology, and converts the image data into text data using OCR (optical character recognition) technology. The converted text data is stored in a database.
[0684] Step 3:
[0685] The server runs a natural language processing (NLP) engine to analyze the text data in the database, extracting important keywords, tags, and content names from the text data.
[0686] Step 4:
[0687] The server associates entrance exam questions with entertainment content based on the keywords and tags extracted through the analysis. For example, a question on Japanese history might be linked to "historical films" or "manga."
[0688] Step 5:
[0689] The server provides a user interface through which users can search for content by age and category.
[0690] Step 6:
[0691] The user terminal sends a search request to the server, for example, if the user wants to search for historical content from a particular era.
[0692] Step 7:
[0693] The server queries the database based on the received search request and returns appropriate results to the user's device, such as a list of movies and novels related to the "Warring States Period" arrow.
[0694] Step 8:
[0695] Users simply launch the learning app, select a specific exam question, and begin answering it. Related entertainment content is automatically displayed while they answer the question.
[0696] Step 9:
[0697] The user device will display related entertainment content based on the category and keywords of the entrance exam questions the user has answered. For example, after answering a question on classical literature, novels and movies related to that topic will be displayed.
[0698] Step 10:
[0699] The server periodically collects and analyzes users' learning history and content usage data, which is then stored in the system as monthly usage statistics.
[0700] Step 11:
[0701] The server calculates revenue based on the collected usage data and distributes it to entertainment content providers, based on the number of views and duration of use.
[0702] Step 12:
[0703] The server notifies both users and content providers of the results of revenue sharing and feedback on learnings, which helps optimize the system.
[0704] The above are the specific processing steps of this system, which effectively link learning content and entertainment content to provide users with a rich learning experience.
[0705] Example 1
[0706] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0707] In conventional learning systems, it is difficult for users to easily access external information or entertainment content directly related to the problem they are solving. Furthermore, they lack a mechanism for effectively managing users' learning data and appropriately distributing revenue to entertainment content providers. This results in low learning effectiveness and makes it difficult to promote deep understanding through entertainment.
[0708] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0709] In this invention, the server includes means for analyzing text data using a natural language processing engine and extracting keywords and tags, means for analyzing entrance exam questions and mapping related entertainment content, means for providing a user interface that allows users to search for content by age group or category, means for converting audio data and image data into text data, means for displaying content related to a question when the user solves it, and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This allows users to easily access related entertainment content while studying, and enables appropriate management of study data and revenue distribution.
[0710] A "natural language processing engine" is a software engine that analyzes text data and extracts important keywords and tags.
[0711] "Keywords" are important words or phrases extracted from text data that are used to categorize and associate content and entrance exam questions.
[0712] "Tags" are metadata attached to text data, and are labels that indicate the characteristics of the content or entrance exam questions.
[0713] "Speech recognition technology" is a technology that converts voice data into text data.
[0714] "Optical character recognition technology" is a technology that analyzes image data and extracts the characters contained therein as text data.
[0715] A "user interface" is a visual interface that allows users to search for content by age group or category.
[0716] "Study history" refers to historical data of entrance exam questions that a user has answered using a learning app.
[0717] "Revenue sharing" refers to the process of appropriately allocating revenue to entertainment content providers based on users' learning history and content usage data.
[0718] "Entrance exam questions" are exam questions designed for users to answer.
[0719] "Entertainment content" refers to content such as books, movies, music, and manga that are provided to allow users to deepen their learning while having fun.
[0720] This invention relates to a system for linking learning and entertainment content using a natural language processing engine, and an embodiment thereof will be described in detail from the viewpoints of a server, a terminal, and a user.
[0721] Data collection and analysis
[0722] To collect data, the server first obtains content data using APIs from various information providers. Specifically, it uses the Google Books API, Spotify API, YouTube Data API, etc. This content data includes audio data, image data, books, and song lyrics. The obtained audio data is converted into text data using Google Cloud Speech-to-Text, and image data is converted into text data using Tesseract OCR technology. This text data is then centrally stored in a MySQL database.
[0723] Next, the server launches a natural language processing engine such as BERT or GPT-3 to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0724] Mapping content to entrance exam questions
[0725] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" could be linked to "historical films" and "manga" that can be obtained from multiple media. In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0726] Providing a user interface
[0727] User devices are provided with an interface developed using React.js, which allows users to search for content by age group and category. When a user selects a specific grade or subject, the server uses Elasticsearch to quickly search for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0728] Learning and Content Links
[0729] When a user launches the learning app and answers entrance exam questions, related entertainment content is displayed. As the user proceeds with the answer, the user's device requests related entertainment content from the server. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, when a user solves a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from the list, the user can gain a deeper understanding of the historical background.
[0730] Data Updates and Revenue Sharing
[0731] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is performed using the Stripe API. Revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue will be distributed to that content provider.
[0732] Specific examples
[0733] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0734] Example prompt: "Please provide entertainment content related to the history of the Heian period. Preferably include novels, movies, and manga."
[0735] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0736] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0737] Step 1:
[0738] To collect data, the server uses APIs from information providers such as Google Books API, Spotify API, and YouTube Data API. The input includes requests to each API. For example, to obtain data such as book information, music information, and video information, an HTTP request is sent to each API. The output is response data from each API. This allows book metadata, music metadata, and video metadata to be obtained. Specifically, the server sends an asynchronous request to each API and waits for a response.
[0739] Step 2:
[0740] The server converts the acquired audio data into text data using Google Cloud Speech-to-Text, and converts the image data into text data using Tesseract OCR. The input includes an audio file and an image file. The audio file is sent to the Google Cloud Speech-to-Text API and text data is received. The image file is also processed with Tesseract OCR to extract characters. The output is text data corresponding to each audio and image data. Specifically, the server sends the audio file to the API and executes the process of receiving text data and extracting characters from the image file.
[0741] Step 3:
[0742] The server centralizes and stores the converted text data in a MySQL database. The input includes text data converted from audio data and image data. SQL commands are executed against the database to store the text data in the corresponding tables. The output is structured text data in the database. Specifically, the server issues SQL queries to insert each piece of text data into a specified column in the table.
[0743] Step 4:
[0744] The server analyzes the stored text data using a natural language processing engine such as BERT or GPT-3. The input includes the text data in the database. This is input into the natural language processing engine, which extracts important keywords and tags. The output is the keywords and tags assigned to each piece of text data. Specifically, the server sends the text data to the engine and executes the process of receiving the analysis results.
[0745] Step 5:
[0746] The server associates entrance exam questions with entertainment content using the extracted keywords and tags. The input includes the entrance exam question data and the extracted keywords and tags. Based on this, related content is selected and links are dynamically generated. The output is a mapping between entrance exam questions and entertainment content. Specifically, the server executes a process of matching the keywords in the entrance exam questions and generating links to the corresponding content.
[0747] Step 6:
[0748] An interface developed using React.js is provided to the user's device. The input includes search criteria selected by the user, such as by era or category. Based on this, a request is sent to the server, and search results for related books, movies, and music are received. As output, a list of content corresponding to the search criteria is displayed on the device. Specifically, the device receives the user's input, sends a request to the server, and receives and displays the results.
[0749] Step 7:
[0750] The user launches the learning app, and related entertainment content is displayed as they answer entrance exam questions. The input includes the entrance exam questions the user answers and their answers. Based on this, a request is sent to the server to retrieve related entertainment content. As an output, the related content is displayed as the user answers. Specifically, the device monitors the user's answering status, sends a request to the server, and retrieves and displays the related content.
[0751] Step 8:
[0752] The server periodically collects user learning history and content usage data and performs revenue sharing using the Stripe API. The input includes learning history data and content usage data. The output is aggregated statistical information and the results of revenue sharing. Specifically, the server periodically aggregates each data and executes a process to calculate and execute revenue sharing.
[0753] (Application example 1)
[0754] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0755] Traditional food delivery systems are limited to simply receiving and delivering orders, and are unable to provide educational or additional entertainment value to users. Users have no access to background knowledge or entertainment content related to the food they order, which limits the food delivery experience.
[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0757] In this invention, the server includes means for analyzing text data using a natural language processing engine to extract keywords and tags, means for analyzing food delivery order details to map related educational and entertainment content, and means for providing a user interface that allows users to search for content by age group or category, thereby enabling users to enjoy educational and entertainment content related to the food they ordered.
[0758] A "natural language processing engine" is software that analyzes the meaning and structure of text data and extracts important keywords and tags.
[0759] "Keywords" are important words that represent the content of text data and are used for searching and classification.
[0760] A "tag" is metadata that represents the attributes or categories of text data, and is used to organize and search content more efficiently.
[0761] "Food delivery" refers to a service that delivers food ordered by users.
[0762] "Entertainment Content" means information or media that provides entertainment to users, including video, audio, and text.
[0763] A "user interface" is a means by which a user interacts with a system, and includes screens and input means for performing operations such as searches and information display.
[0764] "Order details" refers to the food items and their detailed information specified by the user in the food delivery service.
[0765] "Educational content" refers to information and media that provide users with knowledge and information and support learning.
[0766] "Mapping" is the act or process of associating and integrating related information or elements.
[0767] "User's order history" refers to a record of orders a user has placed using the food delivery service in the past.
[0768] "Content provider" means a person or organization that creates and provides entertainment or educational content.
[0769] This invention is a system for providing educational and entertainment content related to food ordered by a user in a food delivery system. This system is composed of a server, a terminal, and a user's perspective.
[0770] Data collection and analysis
[0771] To collect data, the server first acquires content data using APIs from various information providers. This includes text data, video data, cultural information, etc. The acquired video data is converted into text data using OCR and voice recognition technologies. This text data is then centrally stored in a database.
[0772] Next, the server launches a natural language processing engine (e.g., Spacy) to analyze the text data in the database. Through analysis, important keywords and tags are extracted from the text data. For example, from the data on "sushi," tags such as "Japanese cuisine," "history," and "rice" are generated.
[0773] Content and food delivery mapping
[0774] The server uses the acquired keywords and tags to link the user's order with related entertainment content. For example, a user who orders "sushi" might be linked to a documentary about the history of sushi or a video about recipes. In this way, content highly relevant to food delivery is dynamically linked.
[0775] Providing a user interface
[0776] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific food item or category, the server searches for related video and text content and sends the results to the device. This allows users to easily access a variety of content related to the food they ordered.
[0777] Learning and Content Links
[0778] A user launches a food delivery app and places an order. Once the order is confirmed, the user's device displays related educational and entertainment content. For example, when ordering "sushi," a documentary video introducing the history and cultural background of that sushi is displayed, allowing the user to deepen their knowledge by watching it.
[0779] Data Updates and Revenue Sharing
[0780] The server periodically collects user order history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is carried out. Revenue is distributed appropriately to educational and entertainment content providers. Specifically, if a particular video or article is viewed many times, a portion of the revenue will be distributed to that content provider.
[0781] Specific examples
[0782] For example, when a user orders "sushi," the user's device requests related entertainment content from the server based on the analyzed keywords "Japanese cuisine" and "history." The server then searches the database to find documentaries and recipe videos about the history of sushi and provides them to the user. In this way, the user can access a wealth of related content through their order.
[0783] Prompt Sentence Examples
[0784] "Dish: Sushi\nRelated Content: History, Recipes"
[0785] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0786] Step 1:
[0787] The server uses the information provider's API to collect various content data (text data, video data, cultural information, etc.). Specifically, it retrieves the necessary data from the database through API calls and stores it centrally. The input is data from the provider's API, and the collected raw data is stored in the database as output.
[0788] Step 2:
[0789] The server converts the collected video and audio data into text data. Specifically, it uses OCR and voice recognition technology to convert the data into text data and stores it in a database. The input is video data and audio data, and the converted text data is obtained as output.
[0790] Step 3:
[0791] The server runs a natural language processing engine (e.g., Spacy) to analyze the text data in the database. The analysis extracts important keywords and tags from the text data. The converted text data is the input, and the extracted keywords and tags are the output.
[0792] Step 4:
[0793] The server uses the retrieved keywords and tags to map the user's order to relevant educational and entertainment content. For example, a user who orders "sushi" might be mapped to a documentary about the history of sushi and a recipe video. The input is the user's order and the extracted keywords, and the output is the identification of relevant content.
[0794] Step 5:
[0795] The user device provides an interface that allows users to search for content by age or category. Specifically, it includes a content list and a search field. The input is a user selection or search query, and the output is a display of related content.
[0796] Step 6:
[0797] The user places an order, and once the order is confirmed, the user device displays related educational and entertainment content. For example, if you order "sushi," a video about the cultural background of that sushi will be displayed. The input is the user's order and a content list, and the output is the content to be displayed.
[0798] Step 7:
[0799] The server periodically collects user order history and content usage data. Based on this data, it calculates monthly usage statistics and distributes revenue. Specifically, if a particular piece of content is viewed many times, a portion of the revenue is distributed to the content provider. The inputs are usage data and order history, and the output is statistical data and revenue distribution results.
[0800] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0801] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments of the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0802] Data collection and analysis
[0803] To collect data, the server first uses the information provider's API to obtain various content data, including audio data, image data, books, and lyrics. This data is then converted into text data using voice recognition and OCR technology and stored in a database.
[0804] The server runs a natural language processing engine to analyze the text data in the database. Through this analysis, important keywords and tags are extracted from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0805] Mapping content to entrance exam questions
[0806] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." This allows for dynamic mapping of content highly relevant to the entrance exam questions.
[0807] Providing a user interface
[0808] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related books, movies, and music and sends the results to the device.
[0809] Learning and Content Links
[0810] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after solving a proverb question, a novel in which the proverb is used will be displayed, allowing users to deepen their understanding by reading the novel.
[0811] Applying the Emotion Engine
[0812] The server activates an emotion engine that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during the learning process. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0813] Data Updates and Revenue Sharing
[0814] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, which are then stored as monthly usage statistics. Revenues are calculated based on the number of times content is viewed and the duration of use, and distributed to entertainment content providers.
[0815] Specific examples
[0816] For example, suppose a user is taking an English literature entrance exam. The user's device analyzes the keyword "Victorian" and requests related entertainment content from the server. The server searches its database to find the novel "Jane Eyre" and movies set in that era, and provides them to the user.
[0817] The emotion engine also analyzes the user's facial expressions to detect when they are losing focus, and in that case, it will display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0818] The above is a concrete example of how to implement the present invention. This system evolves learning from mere knowledge acquisition to a rich experience that responds to the user's emotional state and interests.
[0819] The processing flow will be explained below.
[0820] A specific embodiment of the present invention will be described below, with the process flow divided into steps from the viewpoints of the server, the terminal, and the user.
[0821] Collecting content data and converting it into text
[0822] Step 1:
[0823] The server uses the information provider's API to obtain content data such as books, music, audio data, and image data.
[0824] Step 2:
[0825] The server converts the acquired voice data into text data using voice recognition technology, and also converts image data into text data using OCR technology. The converted text data is stored in a database.
[0826] Text data analysis and tagging
[0827] Step 3:
[0828] The server runs a natural language processing engine to analyze the text data in the database, extracting important keywords and tags from the text data.
[0829] Step 4:
[0830] The server indexes the text data based on the analysis results and associates relevant content with keywords and tags.
[0831] Mapping to entrance exam questions
[0832] Step 5:
[0833] The server uses the analyzed keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga."
[0834] Providing a learning interface
[0835] Step 6:
[0836] The user terminal provides an interface that allows the user to search for content by age group or category.
[0837] Step 7:
[0838] The user enters search criteria (grade, subject, keywords, etc.) and sends a search request to the server via the terminal.
[0839] Step 8:
[0840] The server queries the database based on the search request to find the appropriate content and returns the results to the user's device.
[0841] Step 9:
[0842] The user terminal displays the search results sent from the server to the user.
[0843] Learning and Content Links
[0844] Step 10:
[0845] The user launches a learning app and answers specific entrance exam questions.
[0846] Step 11:
[0847] The user terminal requests entertainment content related to the entrance exam question being answered from the server.
[0848] Step 12:
[0849] The server searches for highly relevant content and transmits the results to the user terminal.
[0850] Step 13:
[0851] The user device will then display related content to the user, for example, after solving a classical Japanese question, novels or movies related to that topic will be displayed.
[0852] Applying the Emotion Engine
[0853] Step 14:
[0854] The server activates the emotion engine and analyzes the user's facial expressions and voice, and evaluates the user's emotional state based on the analysis results.
[0855] Step 15:
[0856] Based on the evaluation results of the emotion engine, the server suggests appropriate entertainment content to assist learning, for example, displaying relaxing music to a tired user.
[0857] Data Updates and Revenue Sharing
[0858] Step 16:
[0859] The server periodically collects and analyzes the user's learning history, emotional state, and content usage data.
[0860] Step 17:
[0861] The server calculates monthly usage statistics based on the collected data and calculates revenue based on the number of times the content is viewed and the duration of use.
[0862] Step 18:
[0863] The server notifies the content provider of the results of the revenue distribution and distributes the reward.
[0864] These are the specific processing steps of this system, which effectively link learning and entertainment content to provide users with a rich and motivating learning experience.
[0865] Example 2
[0866] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0867] Conventional learning support systems have difficulty dynamically providing learning content according to the user's interests and emotions, which leads to problems such as reduced learning efficiency and motivation.In addition, there are limited ways to link entrance exam questions with entertainment content, making it difficult to attract users' interest.
[0868] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for converting audio data and image data into text data; means for analyzing entrance exam questions and dynamically linking related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for analyzing the user's facial expressions and voice and providing content according to their emotional state during study; and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This makes it possible to provide related content according to the user's emotions and interests, thereby improving study efficiency and motivation.
[0869] A "natural language processing engine" is a program that analyzes text data and extracts keywords and tags.
[0870] "Audio Data" refers to digital files containing speech or audio information.
[0871] "Image data" refers to digital files that contain visual information.
[0872] "Text data" refers to digital files that contain textual information.
[0873] "Keywords" refer to important words or phrases within text data.
[0874] A "tag" refers to a label for classifying text data.
[0875] "Entrance examination questions" refer to questions asked in entrance examinations conducted by educational institutions.
[0876] "Entertainment content" refers to media such as movies, music, and books that are intended to capture the user's interest.
[0877] "User interface" refers to the screen and input devices that users use to operate the system.
[0878] "Learning history" refers to data that records what a user has learned and their progress.
[0879] "Revenue sharing" refers to the process of distributing revenue generated by the system to stakeholders, such as content providers.
[0880] "Facial expression analysis" refers to the technology of analyzing a user's facial expressions to determine their emotional state.
[0881] "Voice analysis" refers to the technology of analyzing a user's voice to determine their emotional state and content.
[0882] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. An embodiment of the present invention will be described in detail below.
[0883] First, the server uses the information provider's API to collect various content data, such as audio data, image data, books, and song lyrics. Specifically, the server converts audio data into text data using Google Cloud Speech-to-Text, and image data into text data using Tesseract OCR software. The converted text data is then stored in the server's database.
[0884] The server then analyzes the stored text data using Hugging Face's Transformers library. This analysis uses generative AI models such as the BERT model to extract important keywords and tags from the text. For example, it generates tags such as "cat," "modern literature," and "author name" from a novel.
[0885] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. For example, a question on "Japanese history" might be associated with "historical films" or "manga." This allows highly relevant content to be mapped to the exam questions.
[0886] The user device provides a user interface that allows users to search for content by age group or category. When a user selects a specific grade or subject, a request is sent to the server. The server searches for related books, movies, and music and sends the results to the user device.
[0887] The user launches the learning app and answers the selected entrance exam questions. As the user answers, the device displays related entertainment content retrieved from the server. For example, after solving a proverb question, the device can display a novel in which the proverb is used, deepening the user's understanding.
[0888] The server also uses Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The emotion engine detects the user's emotional state during training and provides appropriate content. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[0889] In addition, the server periodically collects and analyzes users' learning history, content usage data, and emotional data. Based on this, the server stores monthly usage statistics and distributes revenue to content providers. This data is used to calculate remuneration based on which content is viewed how many times and for how long.
[0890] For example, if a user is solving English literature entrance exam questions, the user's device will request related entertainment content from the server based on the analyzed keyword "Victorian era." The server will then search the database to find the novel "Jane Eyre" and movies set in that era and provide them to the user. Furthermore, if the emotion engine analyzes the user's facial expressions and detects that the user is losing concentration, it can display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[0891] An example of a prompt might be:
[0892] Suggest content related to the theme "Victorian Era".
[0893] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0894] Step 1: Data collection
[0895] The server uses the APIs of information providers (e.g., book API, music API, image API, etc.) to collect various content data such as audio data, image data, book data, and lyric data. The input is various data obtained from the information providers, and the output is the collected data set. Specifically, it sends a request to the API and stores the obtained data in the storage system within the server.
[0896] Step 2: Text data conversion
[0897] The server uses Google Cloud Speech-to-Text to convert collected voice data to text data, and Tesseract OCR software to convert image data to text data. The input is digital audio or image files, and the output is text data that is converted and stored in a database. Specific operations include reading voice data and sending a conversion request, and reading image data and performing OCR processing.
[0898] Step 3: Natural Language Processing Analysis
[0899] The server uses Hugging Face's Transformers library to analyze the saved text data. The input is the transformed text data, and the output is analyzed data containing keywords and tags extracted from the text data. Specifically, it analyzes the text data using a generative AI model such as the BERT model to extract important keywords and tags. For example, it generates tags such as "cat," "modern literature," and "author name" from novel data.
[0900] Step 4: Mapping content to exam questions
[0901] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. The input is the keywords and tags generated by analysis and the entrance exam questions, and the output is a mapped dataset. Specifically, it matches entrance exam questions with highly relevant keywords with content, for example, associating "historical films" and "manga" with questions on "Japanese history."
[0902] Step 5: Providing a User Interface
[0903] The user device provides an interface that allows users to search for content by age group or category. The input is the user's selected grade level or subject, and the output is search results for related books, movies, and music returned by the server. Specifically, the device receives the user's selection through the interface and sends the request to the server, which searches the database and returns the results.
[0904] Step 6: Learning and Content Linking
[0905] The user launches the learning app and answers the selected entrance exam questions. The input is the user's answer data, and the output is related entertainment content. Specifically, when the user answers a question, books and movies related to that question are displayed. For example, after answering a proverb question, a novel in which that proverb is used is displayed.
[0906] Step 7: Applying the Emotion Engine
[0907] The server invokes Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The input is the user's real-time facial and voice data, and the output is the provision of content based on the analyzed emotional state. Specifically, the system captures the user's facial expressions and voice from the camera and microphone and analyzes them with the emotion engine. For example, if the user is tired, it provides relaxing music or light entertainment content.
[0908] Step 8: Data Updates and Revenue Sharing
[0909] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. The input is various user usage data, and the output is monthly usage statistics and revenue distribution data. Specifically, it analyzes the learning history and usage data stored in the database, tallying up how much of each piece of content was used, and distributes revenue to content providers based on that information.
[0910] (Application example 2)
[0911] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0912] Conventional learning support systems struggle to effectively link the content users are learning with entertainment content, and lack the means to maintain users' motivation and concentration. Furthermore, they rarely provide content that takes into account the user's emotional state, making it difficult to meet individual needs. This results in a limited learning experience and an inability to provide an optimal learning environment for each user.
[0913] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for analyzing entrance exam questions and mapping related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for providing content that corresponds to the user's emotional state using an emotion engine that analyzes the user's emotional state; and means for analyzing the user's learning history and content usage data and distributing revenue to content providers. As a result, users are provided with optimal content that corresponds to their emotional state while studying, enabling them to maintain their motivation and concentration and progress with their studies more effectively.
[0914] A "natural language processing engine" is a technology that analyzes text data and extracts important information such as keywords and tags.
[0915] "Keywords" refer to important words or phrases in text data, and represent the content of a document.
[0916] A "tag" is a label that represents a category or attribute that is assigned to classify text data.
[0917] "Entertainment content" refers to digital content and media such as movies, novels, music, and games that provide users with enjoyment and entertainment.
[0918] "User interface" refers to the screen and operation method that allows a user to interact with a system, making it easier for the user to operate it.
[0919] The "emotion engine" is a technology that analyzes the user's facial expressions and voice to detect their emotional state at that time.
[0920] "Study history" refers to a record of what a user has studied and the questions they have answered, and is data used to understand an individual's learning progress.
[0921] "Revenue sharing" is a method for fairly distributing revenue generated within the system to content providers and other stakeholders.
[0922] "Associating" refers to building semantic connections between different data or content.
[0923] "Dynamic linking" means connecting relevant content and information in real time.
[0924] "Content usage data" refers to records of which content a user has used and to what extent, and is data used to analyze user usage.
[0925] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments for carrying out the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[0926] Data collection and analysis
[0927] The server uses the information provider's API to collect various content data, including audio data, image data, and text data. This data is then centrally converted into text data using voice recognition and OCR technology and stored in a database. The server then launches a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags.
[0928] Mapping content to entrance exam questions
[0929] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "literary works." This process is performed using database search and dynamic link generation technology.
[0930] Providing a user interface
[0931] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related movies, books, and music and sends the results to the device. This interface is typically run on devices such as smartphones or head-mounted displays (HMDs).
[0932] Learning and Content Links
[0933] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after answering an English literature question, a related novel will be displayed, allowing the user to read the novel and deepen their understanding.
[0934] Applying the Emotion Engine
[0935] The server launches an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or light entertainment content accordingly.
[0936] Data Updates and Revenue Sharing
[0937] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. This data is stored as monthly usage statistics, and revenue is calculated and distributed to content providers based on the number of times the content is viewed and the duration of use.
[0938] Specific examples
[0939] For example, suppose a user is solving an English literature entrance exam. The user's device requests related entertainment content from the server based on the keyword "Victorian" that was analyzed during the process. The server then searches its database to find the novel "Jane Eyre" and films set in that era, and provides them to the user. The emotion engine also analyzes the user's facial expressions to detect a lapse in concentration. In that case, it displays visually interesting content related to the learning content (such as a documentary film).
[0940] Prompt Sentence Examples
[0941] List entertainment content related to the Victorian era, and provide visually engaging documentaries to help users with short attention spans.
[0942] As described above, this invention evolves learning from mere acquisition of knowledge to a rich experience that responds to the user's emotional state and interests.
[0943] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0944] Step 1:
[0945] The server collects audio data, image data, and text data through the information provider's API. The collected data is stored in raw form on the server (input: audio data, image data, and text data obtained via API; output: raw data stored on the server).
[0946] Step 2:
[0947] The server converts raw data into text data using speech recognition and OCR technology. Voice data is converted into text using speech recognition software (e.g., SpeechRecognition), and image data is converted into text data using OCR software (e.g., Tesseract OCR) (input: raw data; output: converted text data).
[0948] Step 3:
[0949] The server uses a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags, which are then stored in a database in a summarized form (input: text data; output: extracted keywords and tags).
[0950] Step 4:
[0951] The server uses the analyzed keywords and tags to map entrance exam questions to entertainment content. For example, an entrance exam question on "Japanese history" is associated with historical movies and novels (input: entrance exam questions and keywords / tags; output: related entertainment content).
[0952] Step 5:
[0953] It provides an interface on the user's device that allows the user to search for content by age and category. Users can search for a specific grade level or subject and display related movies, books, and music (input: user's search query; output: list of related content).
[0954] Step 6:
[0955] A user uses a learning app to answer entrance exam questions. The user's device displays related entertainment content as the user progresses. For example, if the user answers a question about a proverb, the app displays a literary work in which the proverb is used (input: entrance exam question being answered; output: related entertainment content).
[0956] Step 7:
[0957] The server uses an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) to analyze the user's facial expressions and voice to detect the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or visually interesting content (input: user's facial and voice data; output: user's emotional state and corresponding content).
[0958] Step 8:
[0959] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, generating monthly usage statistics and distributing revenue to content providers based on the number of times content is viewed and the duration of usage (input: learning history, content usage data, emotional data; output: revenue distribution data and statistics on the number of times content is viewed / used).
[0960] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0961] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0962] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0963] [Fourth embodiment]
[0964] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0965] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0966] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0967] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0968] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0969] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0970] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0971] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0972] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0973] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0974] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0975] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0976] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0977] The present invention relates to a system for linking learning and entertainment content using a natural language processing engine. The following describes an embodiment of the system from the viewpoints of a server, a terminal, and a user.
[0978] Data collection and analysis
[0979] To collect data, the server first acquires content data using APIs from various information providers. This includes audio data, image data, books, and song lyrics. The acquired audio data is converted into text data using voice recognition technology, and image data is converted into text data using OCR technology. This text data is then centrally stored in a database.
[0980] The server then launches a natural language processing engine to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[0981] Mapping content to entrance exam questions
[0982] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[0983] Providing a user interface
[0984] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific grade level or subject, the server searches for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[0985] Learning and Content Links
[0986] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, if a user answers a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from that list, they can gain a deeper understanding of the historical background.
[0987] Data Updates and Revenue Sharing
[0988] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue is distributed. The revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue is distributed to that content provider.
[0989] Specific examples
[0990] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[0991] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[0992] The processing flow will be explained below.
[0993] Step 1:
[0994] The server uses the API of the information provider to obtain various content data (books, lyrics, audio data, image data), including audio and image data.
[0995] Step 2:
[0996] The server converts the acquired voice data into text data using voice recognition technology, and converts the image data into text data using OCR (optical character recognition) technology. The converted text data is stored in a database.
[0997] Step 3:
[0998] The server runs a natural language processing (NLP) engine to analyze the text data in the database, extracting important keywords, tags, and content names from the text data.
[0999] Step 4:
[1000] The server associates entrance exam questions with entertainment content based on the keywords and tags extracted through the analysis. For example, a question on Japanese history might be linked to "historical films" or "manga."
[1001] Step 5:
[1002] The server provides a user interface through which users can search for content by age and category.
[1003] Step 6:
[1004] The user terminal sends a search request to the server, for example, if the user wants to search for historical content from a particular era.
[1005] Step 7:
[1006] The server queries the database based on the received search request and returns appropriate results to the user's device, such as a list of movies and novels related to the "Warring States Period" arrow.
[1007] Step 8:
[1008] Users simply launch the learning app, select a specific exam question, and begin answering it. Related entertainment content is automatically displayed while they answer the question.
[1009] Step 9:
[1010] The user device will display related entertainment content based on the category and keywords of the entrance exam questions the user has answered. For example, after answering a question on classical literature, novels and movies related to that topic will be displayed.
[1011] Step 10:
[1012] The server periodically collects and analyzes users' learning history and content usage data, which is then stored in the system as monthly usage statistics.
[1013] Step 11:
[1014] The server calculates revenue based on the collected usage data and distributes it to entertainment content providers, based on the number of views and duration of use.
[1015] Step 12:
[1016] The server notifies both users and content providers of the results of revenue sharing and feedback on learnings, which helps optimize the system.
[1017] The above are the specific processing steps of this system, which effectively link learning content and entertainment content to provide users with a rich learning experience.
[1018] Example 1
[1019] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1020] In conventional learning systems, it is difficult for users to easily access external information or entertainment content directly related to the problem they are solving. Furthermore, they lack a mechanism for effectively managing users' learning data and appropriately distributing revenue to entertainment content providers. This results in low learning effectiveness and makes it difficult to promote deep understanding through entertainment.
[1021] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1022] In this invention, the server includes means for analyzing text data using a natural language processing engine and extracting keywords and tags, means for analyzing entrance exam questions and mapping related entertainment content, means for providing a user interface that allows users to search for content by age group or category, means for converting audio data and image data into text data, means for displaying content related to a question when the user solves it, and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This allows users to easily access related entertainment content while studying, and enables appropriate management of study data and revenue distribution.
[1023] A "natural language processing engine" is a software engine that analyzes text data and extracts important keywords and tags.
[1024] "Keywords" are important words or phrases extracted from text data that are used to categorize and associate content and entrance exam questions.
[1025] "Tags" are metadata attached to text data, and are labels that indicate the characteristics of the content or entrance exam questions.
[1026] "Speech recognition technology" is a technology that converts voice data into text data.
[1027] "Optical character recognition technology" is a technology that analyzes image data and extracts the characters contained therein as text data.
[1028] A "user interface" is a visual interface that allows users to search for content by age group or category.
[1029] "Study history" refers to historical data of entrance exam questions that a user has answered using a learning app.
[1030] "Revenue sharing" refers to the process of appropriately allocating revenue to entertainment content providers based on users' learning history and content usage data.
[1031] "Entrance exam questions" are exam questions designed for users to answer.
[1032] "Entertainment content" refers to content such as books, movies, music, and manga that are provided to allow users to deepen their learning while having fun.
[1033] This invention relates to a system for linking learning and entertainment content using a natural language processing engine, and an embodiment thereof will be described in detail from the viewpoints of a server, a terminal, and a user.
[1034] Data collection and analysis
[1035] To collect data, the server first obtains content data using APIs from various information providers. Specifically, it uses the Google Books API, Spotify API, YouTube Data API, etc. This content data includes audio data, image data, books, and song lyrics. The obtained audio data is converted into text data using Google Cloud Speech-to-Text, and image data is converted into text data using Tesseract OCR technology. This text data is then centrally stored in a MySQL database.
[1036] Next, the server launches a natural language processing engine such as BERT or GPT-3 to analyze the text data in the database. This analysis extracts important keywords and tags from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[1037] Mapping content to entrance exam questions
[1038] The server uses the acquired keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" could be linked to "historical films" and "manga" that can be obtained from multiple media. In this way, content highly relevant to the entrance exam questions is dynamically mapped.
[1039] Providing a user interface
[1040] User devices are provided with an interface developed using React.js, which allows users to search for content by age group and category. When a user selects a specific grade or subject, the server uses Elasticsearch to quickly search for related books, movies, and music and sends the results to the device. This allows users to easily access a variety of content relevant to their studies.
[1041] Learning and Content Links
[1042] When a user launches the learning app and answers entrance exam questions, related entertainment content is displayed. As the user proceeds with the answer, the user's device requests related entertainment content from the server. For example, when solving a proverb question, a novel in which the proverb is used is displayed, allowing the user to deepen their understanding by reading the novel. Specifically, when a user solves a history question, a list of movies related to the proverb is displayed, and by watching a "historical documentary film" from the list, the user can gain a deeper understanding of the historical background.
[1043] Data Updates and Revenue Sharing
[1044] The server periodically collects users' learning history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is performed using the Stripe API. Revenue is appropriately distributed to entertainment content providers. Specifically, if a particular novel or movie is viewed many times during learning, a portion of the revenue will be distributed to that content provider.
[1045] Specific examples
[1046] For example, suppose a user is answering entrance exam questions on Japanese literature. During the process, the user's device requests related entertainment content from the server based on the analyzed keyword "Heian period." The server then searches the database, finds the novel "The Tale of Genji" and movies set in that period, and provides them to the user. In this way, the user can access a wealth of related content while answering the entrance exam questions.
[1047] Example prompt: "Please provide entertainment content related to the history of the Heian period. Preferably include novels, movies, and manga."
[1048] The above is a specific embodiment for carrying out the present invention. This system evolves learning from mere knowledge acquisition to deep understanding through entertainment.
[1049] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1050] Step 1:
[1051] To collect data, the server uses APIs from information providers such as Google Books API, Spotify API, and YouTube Data API. The input includes requests to each API. For example, to obtain data such as book information, music information, and video information, an HTTP request is sent to each API. The output is response data from each API. This allows book metadata, music metadata, and video metadata to be obtained. Specifically, the server sends an asynchronous request to each API and waits for a response.
[1052] Step 2:
[1053] The server converts the acquired audio data into text data using Google Cloud Speech-to-Text, and converts the image data into text data using Tesseract OCR. The input includes an audio file and an image file. The audio file is sent to the Google Cloud Speech-to-Text API and text data is received. The image file is also processed with Tesseract OCR to extract characters. The output is text data corresponding to each audio and image data. Specifically, the server sends the audio file to the API and executes the process of receiving text data and extracting characters from the image file.
[1054] Step 3:
[1055] The server centralizes and stores the converted text data in a MySQL database. The input includes text data converted from audio data and image data. SQL commands are executed against the database to store the text data in the corresponding tables. The output is structured text data in the database. Specifically, the server issues SQL queries to insert each piece of text data into a specified column in the table.
[1056] Step 4:
[1057] The server analyzes the stored text data using a natural language processing engine such as BERT or GPT-3. The input includes the text data in the database. This is input into the natural language processing engine, which extracts important keywords and tags. The output is the keywords and tags assigned to each piece of text data. Specifically, the server sends the text data to the engine and executes the process of receiving the analysis results.
[1058] Step 5:
[1059] The server associates entrance exam questions with entertainment content using the extracted keywords and tags. The input includes the entrance exam question data and the extracted keywords and tags. Based on this, related content is selected and links are dynamically generated. The output is a mapping between entrance exam questions and entertainment content. Specifically, the server executes a process of matching the keywords in the entrance exam questions and generating links to the corresponding content.
[1060] Step 6:
[1061] An interface developed using React.js is provided to the user's device. The input includes search criteria selected by the user, such as by era or category. Based on this, a request is sent to the server, and search results for related books, movies, and music are received. As output, a list of content corresponding to the search criteria is displayed on the device. Specifically, the device receives the user's input, sends a request to the server, and receives and displays the results.
[1062] Step 7:
[1063] The user launches the learning app, and related entertainment content is displayed as they answer entrance exam questions. The input includes the entrance exam questions the user answers and their answers. Based on this, a request is sent to the server to retrieve related entertainment content. As an output, the related content is displayed as the user answers. Specifically, the device monitors the user's answering status, sends a request to the server, and retrieves and displays the related content.
[1064] Step 8:
[1065] The server periodically collects user learning history and content usage data and performs revenue sharing using the Stripe API. The input includes learning history data and content usage data. The output is aggregated statistical information and the results of revenue sharing. Specifically, the server periodically aggregates each data and executes a process to calculate and execute revenue sharing.
[1066] (Application example 1)
[1067] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1068] Traditional food delivery systems are limited to simply receiving and delivering orders, and are unable to provide educational or additional entertainment value to users. Users have no access to background knowledge or entertainment content related to the food they order, which limits the food delivery experience.
[1069] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1070] In this invention, the server includes means for analyzing text data using a natural language processing engine to extract keywords and tags, means for analyzing food delivery order details to map related educational and entertainment content, and means for providing a user interface that allows users to search for content by age group or category, thereby enabling users to enjoy educational and entertainment content related to the food they ordered.
[1071] A "natural language processing engine" is software that analyzes the meaning and structure of text data and extracts important keywords and tags.
[1072] "Keywords" are important words that represent the content of text data and are used for searching and classification.
[1073] A "tag" is metadata that represents the attributes or categories of text data, and is used to organize and search content more efficiently.
[1074] "Food delivery" refers to a service that delivers food ordered by users.
[1075] "Entertainment Content" means information or media that provides entertainment to users, including video, audio, and text.
[1076] A "user interface" is a means by which a user interacts with a system, and includes screens and input means for performing operations such as searches and information display.
[1077] "Order details" refers to the food items and their detailed information specified by the user in the food delivery service.
[1078] "Educational content" refers to information and media that provide users with knowledge and information and support learning.
[1079] "Mapping" is the act or process of associating and integrating related information or elements.
[1080] "User's order history" refers to a record of orders a user has placed using the food delivery service in the past.
[1081] "Content provider" means a person or organization that creates and provides entertainment or educational content.
[1082] This invention is a system for providing educational and entertainment content related to food ordered by a user in a food delivery system. This system is composed of a server, a terminal, and a user's perspective.
[1083] Data collection and analysis
[1084] To collect data, the server first acquires content data using APIs from various information providers. This includes text data, video data, cultural information, etc. The acquired video data is converted into text data using OCR and voice recognition technologies. This text data is then centrally stored in a database.
[1085] Next, the server launches a natural language processing engine (e.g., Spacy) to analyze the text data in the database. Through analysis, important keywords and tags are extracted from the text data. For example, from the data on "sushi," tags such as "Japanese cuisine," "history," and "rice" are generated.
[1086] Content and food delivery mapping
[1087] The server uses the acquired keywords and tags to link the user's order with related entertainment content. For example, a user who orders "sushi" might be linked to a documentary about the history of sushi or a video about recipes. In this way, content highly relevant to food delivery is dynamically linked.
[1088] Providing a user interface
[1089] The user device is provided with an interface that allows users to search for content by age group and category. When a user selects a specific food item or category, the server searches for related video and text content and sends the results to the device. This allows users to easily access a variety of content related to the food they ordered.
[1090] Learning and Content Links
[1091] A user launches a food delivery app and places an order. Once the order is confirmed, the user's device displays related educational and entertainment content. For example, when ordering "sushi," a documentary video introducing the history and cultural background of that sushi is displayed, allowing the user to deepen their knowledge by watching it.
[1092] Data Updates and Revenue Sharing
[1093] The server periodically collects user order history and content usage data. Based on this data, monthly usage statistics are calculated and revenue distribution is carried out. Revenue is distributed appropriately to educational and entertainment content providers. Specifically, if a particular video or article is viewed many times, a portion of the revenue will be distributed to that content provider.
[1094] Specific examples
[1095] For example, when a user orders "sushi," the user's device requests related entertainment content from the server based on the analyzed keywords "Japanese cuisine" and "history." The server then searches the database to find documentaries and recipe videos about the history of sushi and provides them to the user. In this way, the user can access a wealth of related content through their order.
[1096] Prompt Sentence Examples
[1097] "Dish: Sushi\nRelated Content: History, Recipes"
[1098] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1099] Step 1:
[1100] The server uses the information provider's API to collect various content data (text data, video data, cultural information, etc.). Specifically, it retrieves the necessary data from the database through API calls and stores it centrally. The input is data from the provider's API, and the collected raw data is stored in the database as output.
[1101] Step 2:
[1102] The server converts the collected video and audio data into text data. Specifically, it uses OCR and voice recognition technology to convert the data into text data and stores it in a database. The input is video data and audio data, and the converted text data is obtained as output.
[1103] Step 3:
[1104] The server runs a natural language processing engine (e.g., Spacy) to analyze the text data in the database. The analysis extracts important keywords and tags from the text data. The converted text data is the input, and the extracted keywords and tags are the output.
[1105] Step 4:
[1106] The server uses the retrieved keywords and tags to map the user's order to relevant educational and entertainment content. For example, a user who orders "sushi" might be mapped to a documentary about the history of sushi and a recipe video. The input is the user's order and the extracted keywords, and the output is the identification of relevant content.
[1107] Step 5:
[1108] The user device provides an interface that allows users to search for content by age or category. Specifically, it includes a content list and a search field. The input is a user selection or search query, and the output is a display of related content.
[1109] Step 6:
[1110] The user places an order, and once the order is confirmed, the user device displays related educational and entertainment content. For example, if you order "sushi," a video about the cultural background of that sushi will be displayed. The input is the user's order and a content list, and the output is the content to be displayed.
[1111] Step 7:
[1112] The server periodically collects user order history and content usage data. Based on this data, it calculates monthly usage statistics and distributes revenue. Specifically, if a particular piece of content is viewed many times, a portion of the revenue is distributed to the content provider. The inputs are usage data and order history, and the output is statistical data and revenue distribution results.
[1113] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1114] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments of the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[1115] Data collection and analysis
[1116] To collect data, the server first uses the information provider's API to obtain various content data, including audio data, image data, books, and lyrics. This data is then converted into text data using voice recognition and OCR technology and stored in a database.
[1117] The server runs a natural language processing engine to analyze the text data in the database. Through this analysis, important keywords and tags are extracted from the text data. For example, from a novel, tags such as "cat," "modern literature," and "author name" are generated.
[1118] Mapping content to entrance exam questions
[1119] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga." This allows for dynamic mapping of content highly relevant to the entrance exam questions.
[1120] Providing a user interface
[1121] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related books, movies, and music and sends the results to the device.
[1122] Learning and Content Links
[1123] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after solving a proverb question, a novel in which the proverb is used will be displayed, allowing users to deepen their understanding by reading the novel.
[1124] Applying the Emotion Engine
[1125] The server activates an emotion engine that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during the learning process. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[1126] Data Updates and Revenue Sharing
[1127] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, which are then stored as monthly usage statistics. Revenues are calculated based on the number of times content is viewed and the duration of use, and distributed to entertainment content providers.
[1128] Specific examples
[1129] For example, suppose a user is taking an English literature entrance exam. The user's device analyzes the keyword "Victorian" and requests related entertainment content from the server. The server searches its database to find the novel "Jane Eyre" and movies set in that era, and provides them to the user.
[1130] The emotion engine also analyzes the user's facial expressions to detect when they are losing focus, and in that case, it will display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[1131] The above is a concrete example of how to implement the present invention. This system evolves learning from mere knowledge acquisition to a rich experience that responds to the user's emotional state and interests.
[1132] The processing flow will be explained below.
[1133] A specific embodiment of the present invention will be described below, with the process flow divided into steps from the viewpoints of the server, the terminal, and the user.
[1134] Collecting content data and converting it into text
[1135] Step 1:
[1136] The server uses the information provider's API to obtain content data such as books, music, audio data, and image data.
[1137] Step 2:
[1138] The server converts the acquired voice data into text data using voice recognition technology, and also converts image data into text data using OCR technology. The converted text data is stored in a database.
[1139] Text data analysis and tagging
[1140] Step 3:
[1141] The server runs a natural language processing engine to analyze the text data in the database, extracting important keywords and tags from the text data.
[1142] Step 4:
[1143] The server indexes the text data based on the analysis results and associates relevant content with keywords and tags.
[1144] Mapping to entrance exam questions
[1145] Step 5:
[1146] The server uses the analyzed keywords and tags to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "manga."
[1147] Providing a learning interface
[1148] Step 6:
[1149] The user terminal provides an interface that allows the user to search for content by age group or category.
[1150] Step 7:
[1151] The user enters search criteria (grade, subject, keywords, etc.) and sends a search request to the server via the terminal.
[1152] Step 8:
[1153] The server queries the database based on the search request to find the appropriate content and returns the results to the user's device.
[1154] Step 9:
[1155] The user terminal displays the search results sent from the server to the user.
[1156] Learning and Content Links
[1157] Step 10:
[1158] The user launches a learning app and answers specific entrance exam questions.
[1159] Step 11:
[1160] The user terminal requests entertainment content related to the entrance exam question being answered from the server.
[1161] Step 12:
[1162] The server searches for highly relevant content and transmits the results to the user terminal.
[1163] Step 13:
[1164] The user device will then display related content to the user, for example, after solving a classical Japanese question, novels or movies related to that topic will be displayed.
[1165] Applying the Emotion Engine
[1166] Step 14:
[1167] The server activates the emotion engine and analyzes the user's facial expressions and voice, and evaluates the user's emotional state based on the analysis results.
[1168] Step 15:
[1169] Based on the evaluation results of the emotion engine, the server suggests appropriate entertainment content to assist learning, for example, displaying relaxing music to a tired user.
[1170] Data Updates and Revenue Sharing
[1171] Step 16:
[1172] The server periodically collects and analyzes the user's learning history, emotional state, and content usage data.
[1173] Step 17:
[1174] The server calculates monthly usage statistics based on the collected data and calculates revenue based on the number of times the content is viewed and the duration of use.
[1175] Step 18:
[1176] The server notifies the content provider of the results of the revenue distribution and distributes the reward.
[1177] These are the specific processing steps of this system, which effectively link learning and entertainment content to provide users with a rich and motivating learning experience.
[1178] Example 2
[1179] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1180] Conventional learning support systems have difficulty dynamically providing learning content according to the user's interests and emotions, which leads to problems such as reduced learning efficiency and motivation.In addition, there are limited ways to link entrance exam questions with entertainment content, making it difficult to attract users' interest.
[1181] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for converting audio data and image data into text data; means for analyzing entrance exam questions and dynamically linking related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for analyzing the user's facial expressions and voice and providing content according to their emotional state during study; and means for analyzing the user's study history and content usage data and distributing revenue to content providers. This makes it possible to provide related content according to the user's emotions and interests, thereby improving study efficiency and motivation.
[1182] A "natural language processing engine" is a program that analyzes text data and extracts keywords and tags.
[1183] "Audio Data" refers to digital files containing speech or audio information.
[1184] "Image data" refers to digital files that contain visual information.
[1185] "Text data" refers to digital files that contain textual information.
[1186] "Keywords" refer to important words or phrases within text data.
[1187] A "tag" refers to a label for classifying text data.
[1188] "Entrance examination questions" refer to questions asked in entrance examinations conducted by educational institutions.
[1189] "Entertainment content" refers to media such as movies, music, and books that are intended to capture the user's interest.
[1190] "User interface" refers to the screen and input devices that users use to operate the system.
[1191] "Learning history" refers to data that records what a user has learned and their progress.
[1192] "Revenue sharing" refers to the process of distributing revenue generated by the system to stakeholders, such as content providers.
[1193] "Facial expression analysis" refers to the technology of analyzing a user's facial expressions to determine their emotional state.
[1194] "Voice analysis" refers to the technology of analyzing a user's voice to determine their emotional state and content.
[1195] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. An embodiment of the present invention will be described in detail below.
[1196] First, the server uses the information provider's API to collect various content data, such as audio data, image data, books, and song lyrics. Specifically, the server converts audio data into text data using Google Cloud Speech-to-Text, and image data into text data using Tesseract OCR software. The converted text data is then stored in the server's database.
[1197] The server then analyzes the stored text data using Hugging Face's Transformers library. This analysis uses generative AI models such as the BERT model to extract important keywords and tags from the text. For example, it generates tags such as "cat," "modern literature," and "author name" from a novel.
[1198] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. For example, a question on "Japanese history" might be associated with "historical films" or "manga." This allows highly relevant content to be mapped to the exam questions.
[1199] The user device provides a user interface that allows users to search for content by age group or category. When a user selects a specific grade or subject, a request is sent to the server. The server searches for related books, movies, and music and sends the results to the user device.
[1200] The user launches the learning app and answers the selected entrance exam questions. As the user answers, the device displays related entertainment content retrieved from the server. For example, after solving a proverb question, the device can display a novel in which the proverb is used, deepening the user's understanding.
[1201] The server also uses Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The emotion engine detects the user's emotional state during training and provides appropriate content. For example, if the user looks tired, it will provide relaxing music or light entertainment content.
[1202] In addition, the server periodically collects and analyzes users' learning history, content usage data, and emotional data. Based on this, the server stores monthly usage statistics and distributes revenue to content providers. This data is used to calculate remuneration based on which content is viewed how many times and for how long.
[1203] For example, if a user is solving English literature entrance exam questions, the user's device will request related entertainment content from the server based on the analyzed keyword "Victorian era." The server will then search the database to find the novel "Jane Eyre" and movies set in that era and provide them to the user. Furthermore, if the emotion engine analyzes the user's facial expressions and detects that the user is losing concentration, it can display visually interesting content (such as a documentary film) that is relevant to the learning content to improve motivation.
[1204] An example of a prompt might be:
[1205] Suggest content related to the theme "Victorian Era".
[1206] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1207] Step 1: Data collection
[1208] The server uses the APIs of information providers (e.g., book API, music API, image API, etc.) to collect various content data such as audio data, image data, book data, and lyric data. The input is various data obtained from the information providers, and the output is the collected data set. Specifically, it sends a request to the API and stores the obtained data in the storage system within the server.
[1209] Step 2: Text data conversion
[1210] The server uses Google Cloud Speech-to-Text to convert collected voice data to text data, and Tesseract OCR software to convert image data to text data. The input is digital audio or image files, and the output is text data that is converted and stored in a database. Specific operations include reading voice data and sending a conversion request, and reading image data and performing OCR processing.
[1211] Step 3: Natural Language Processing Analysis
[1212] The server uses Hugging Face's Transformers library to analyze the saved text data. The input is the transformed text data, and the output is analyzed data containing keywords and tags extracted from the text data. Specifically, it analyzes the text data using a generative AI model such as the BERT model to extract important keywords and tags. For example, it generates tags such as "cat," "modern literature," and "author name" from novel data.
[1213] Step 4: Mapping content to exam questions
[1214] The server dynamically links entrance exam questions to entertainment content based on the extracted keywords and tags. The input is the keywords and tags generated by analysis and the entrance exam questions, and the output is a mapped dataset. Specifically, it matches entrance exam questions with highly relevant keywords with content, for example, associating "historical films" and "manga" with questions on "Japanese history."
[1215] Step 5: Providing a User Interface
[1216] The user device provides an interface that allows users to search for content by age group or category. The input is the user's selected grade level or subject, and the output is search results for related books, movies, and music returned by the server. Specifically, the device receives the user's selection through the interface and sends the request to the server, which searches the database and returns the results.
[1217] Step 6: Learning and Content Linking
[1218] The user launches the learning app and answers the selected entrance exam questions. The input is the user's answer data, and the output is related entertainment content. Specifically, when the user answers a question, books and movies related to that question are displayed. For example, after answering a proverb question, a novel in which that proverb is used is displayed.
[1219] Step 7: Applying the Emotion Engine
[1220] The server invokes Microsoft Azure's Emotion API to analyze the user's facial expressions and voice. The input is the user's real-time facial and voice data, and the output is the provision of content based on the analyzed emotional state. Specifically, the system captures the user's facial expressions and voice from the camera and microphone and analyzes them with the emotion engine. For example, if the user is tired, it provides relaxing music or light entertainment content.
[1221] Step 8: Data Updates and Revenue Sharing
[1222] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. The input is various user usage data, and the output is monthly usage statistics and revenue distribution data. Specifically, it analyzes the learning history and usage data stored in the database, tallying up how much of each piece of content was used, and distributes revenue to content providers based on that information.
[1223] (Application example 2)
[1224] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1225] Conventional learning support systems struggle to effectively link the content users are learning with entertainment content, and lack the means to maintain users' motivation and concentration. Furthermore, they rarely provide content that takes into account the user's emotional state, making it difficult to meet individual needs. This results in a limited learning experience and an inability to provide an optimal learning environment for each user.
[1226] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for analyzing text data using a natural language processing engine and extracting keywords and tags; means for analyzing entrance exam questions and mapping related entertainment content; means for providing a user interface that allows users to search for content by age group or category; means for displaying content related to a question when the user solves it; means for providing content that corresponds to the user's emotional state using an emotion engine that analyzes the user's emotional state; and means for analyzing the user's learning history and content usage data and distributing revenue to content providers. As a result, users are provided with optimal content that corresponds to their emotional state while studying, enabling them to maintain their motivation and concentration and progress with their studies more effectively.
[1227] A "natural language processing engine" is a technology that analyzes text data and extracts important information such as keywords and tags.
[1228] "Keywords" refer to important words or phrases in text data, and represent the content of a document.
[1229] A "tag" is a label that represents a category or attribute that is assigned to classify text data.
[1230] "Entertainment content" refers to digital content and media such as movies, novels, music, and games that provide users with enjoyment and entertainment.
[1231] "User interface" refers to the screen and operation method that allows a user to interact with a system, making it easier for the user to operate it.
[1232] The "emotion engine" is a technology that analyzes the user's facial expressions and voice to detect their emotional state at that time.
[1233] "Study history" refers to a record of what a user has studied and the questions they have answered, and is data used to understand an individual's learning progress.
[1234] "Revenue sharing" is a method for fairly distributing revenue generated within the system to content providers and other stakeholders.
[1235] "Associating" refers to building semantic connections between different data or content.
[1236] "Dynamic linking" means connecting relevant content and information in real time.
[1237] "Content usage data" refers to records of which content a user has used and to what extent, and is data used to analyze user usage.
[1238] The present invention is a learning support system that combines a natural language processing engine and an emotion engine. The embodiments for carrying out the present invention will be specifically described from the viewpoints of a server, a terminal, and a user.
[1239] Data collection and analysis
[1240] The server uses the information provider's API to collect various content data, including audio data, image data, and text data. This data is then centrally converted into text data using voice recognition and OCR technology and stored in a database. The server then launches a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags.
[1241] Mapping content to entrance exam questions
[1242] The server uses the keywords and tags extracted through the analysis to associate entrance exam questions with entertainment content. For example, a question on "Japanese history" might be linked to "historical films" or "literary works." This process is performed using database search and dynamic link generation technology.
[1243] Providing a user interface
[1244] The user device is provided with an interface that allows the user to search for content by age group or category. When the user selects a specific grade or subject, the server searches for related movies, books, and music and sends the results to the device. This interface is typically run on devices such as smartphones or head-mounted displays (HMDs).
[1245] Learning and Content Links
[1246] Users launch the learning app and answer the selected entrance exam questions. As they answer, the user's device displays related entertainment content. For example, after answering an English literature question, a related novel will be displayed, allowing the user to read the novel and deepen their understanding.
[1247] Applying the Emotion Engine
[1248] The server launches an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) that analyzes the user's facial expressions and voice. The emotion engine detects the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or light entertainment content accordingly.
[1249] Data Updates and Revenue Sharing
[1250] The server periodically collects and analyzes users' learning history, content usage data, and emotional data. This data is stored as monthly usage statistics, and revenue is calculated and distributed to content providers based on the number of times the content is viewed and the duration of use.
[1251] Specific examples
[1252] For example, suppose a user is solving an English literature entrance exam. The user's device requests related entertainment content from the server based on the keyword "Victorian" that was analyzed during the process. The server then searches its database to find the novel "Jane Eyre" and films set in that era, and provides them to the user. The emotion engine also analyzes the user's facial expressions to detect a lapse in concentration. In that case, it displays visually interesting content related to the learning content (such as a documentary film).
[1253] Prompt Sentence Examples
[1254] List entertainment content related to the Victorian era, and provide visually engaging documentaries to help users with short attention spans.
[1255] As described above, this invention evolves learning from mere acquisition of knowledge to a rich experience that responds to the user's emotional state and interests.
[1256] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1257] Step 1:
[1258] The server collects audio data, image data, and text data through the information provider's API. The collected data is stored in raw form on the server (input: audio data, image data, and text data obtained via API; output: raw data stored on the server).
[1259] Step 2:
[1260] The server converts raw data into text data using speech recognition and OCR technology. Voice data is converted into text using speech recognition software (e.g., SpeechRecognition), and image data is converted into text data using OCR software (e.g., Tesseract OCR) (input: raw data; output: converted text data).
[1261] Step 3:
[1262] The server uses a natural language processing engine (e.g., Google Cloud Natural Language or Hugging Face Transformers) to analyze the text data and extract important keywords and tags, which are then stored in a database in a summarized form (input: text data; output: extracted keywords and tags).
[1263] Step 4:
[1264] The server uses the analyzed keywords and tags to map entrance exam questions to entertainment content. For example, an entrance exam question on "Japanese history" is associated with historical movies and novels (input: entrance exam questions and keywords / tags; output: related entertainment content).
[1265] Step 5:
[1266] It provides an interface on the user's device that allows the user to search for content by age and category. Users can search for a specific grade level or subject and display related movies, books, and music (input: user's search query; output: list of related content).
[1267] Step 6:
[1268] A user uses a learning app to answer entrance exam questions. The user's device displays related entertainment content as the user progresses. For example, if the user answers a question about a proverb, the app displays a literary work in which the proverb is used (input: entrance exam question being answered; output: related entertainment content).
[1269] Step 7:
[1270] The server uses an emotion engine (e.g., Microsoft Azure Emotion API or OpenCV) to analyze the user's facial expressions and voice to detect the user's emotional state during training. For example, if the user looks tired, it will provide relaxing music or visually interesting content (input: user's facial and voice data; output: user's emotional state and corresponding content).
[1271] Step 8:
[1272] The server periodically collects and analyzes users' learning history, content usage data, and emotional data, generating monthly usage statistics and distributing revenue to content providers based on the number of times content is viewed and the duration of usage (input: learning history, content usage data, emotional data; output: revenue distribution data and statistics on the number of times content is viewed / used).
[1273] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1274] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1275] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1276] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1277] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1278] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1279] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1280] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1281] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1282] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1283] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1284] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1285] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1286] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1287] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1288] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1289] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1290] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1291] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1292] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1293] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1294] The following is further disclosed regarding the above embodiment.
[1295] (Claim 1)
[1296] A means for analyzing text data using a natural language processing engine and extracting keywords and tags;
[1297] A means for analyzing entrance exam questions and mapping related entertainment content;
[1298] A means to provide a user interface that allows users to search for content by age and category;
[1299] a means for displaying content related to a problem when the user solves the problem;
[1300] A means for analyzing the user's learning history and content usage data and distributing revenue to content providers;
[1301] A system including:
[1302] (Claim 2)
[1303] 10. The system of claim 1, further comprising means for converting audio data and image data into text data.
[1304] (Claim 3)
[1305] 10. The system of claim 1, further comprising means for dynamically linking relevant content based on categories and keywords of entrance exam questions.
[1306] "Example 1"
[1307] (Claim 1)
[1308] A means for analyzing text data using a natural language processing engine and extracting keywords and tags;
[1309] A means for analyzing entrance exam questions and mapping related entertainment content;
[1310] A means to provide a user interface that allows users to search for content by age and category;
[1311] means for converting voice data and image data into text data;
[1312] a means for displaying content related to a problem when the user solves the problem;
[1313] A means for analyzing the user's learning history and content usage data and distributing revenue to content providers;
[1314] A system including:
[1315] (Claim 2)
[1316] 10. The system of claim 1, further comprising means for converting the collected voice data into text data using voice recognition techniques and converting the collected image data into text data using optical character recognition techniques.
[1317] (Claim 3)
[1318] 10. The system of claim 1, further comprising means for linking relevant content based on categories and keywords of entrance exam questions.
[1319] "Application Example 1"
[1320] (Claim 1)
[1321] A means for analyzing text data using a natural language processing engine and extracting keywords and tags;
[1322] A means of analyzing food delivery orders to map relevant educational and entertainment content;
[1323] A means to provide a user interface that allows users to search for content by age and category;
[1324] a means for displaying content related to a food item when the user orders the food item;
[1325] A means for analyzing user order history and content usage data and distributing revenue to content providers;
[1326] A system including:
[1327] (Claim 2)
[1328] 10. The system of claim 1, further comprising means for converting audio data and image data into text data.
[1329] (Claim 3)
[1330] 10. The system of claim 1, further comprising means for dynamically linking relevant content based on food categories and keywords.
[1331] "Example 2: Combining Emotion Engines"
[1332] (Claim 1)
[1333] A means for analyzing text data using a natural language processing engine and extracting keywords and tags;
[1334] means for converting voice data and image data into text data;
[1335] A means for analyzing entrance exam questions and dynamically linking related entertainment content;
[1336] A means to provide a user interface that allows users to search for content by age and category;
[1337] a means for displaying content related to a problem when the user solves the problem;
[1338] A means of analyzing the user's facial expressions and voice and providing content that corresponds to the user's emotional state during learning.
[1339] A means for analyzing the user's learning history and content usage data and distributing revenue to content providers;
[1340] A system including:
[1341] (Claim 2)
[1342] 10. The system of claim 1, further comprising means for converting audio data and image data into text data.
[1343] (Claim 3)
[1344] 10. The system of claim 1, further comprising means for analyzing a user's facial expressions and voice and providing content according to the user's emotional state during learning.
[1345] "Application example 2 when combining emotion engines"
[1346] (Claim 1)
[1347] A means for analyzing text data using a natural language processing engine and extracting keywords and tags;
[1348] A means for analyzing entrance exam questions and mapping related entertainment content;
[1349] A means to provide a user interface that allows users to search for content by age and category;
[1350] a means for displaying content related to a problem when the user solves the problem;
[1351] a means for providing content according to the emotional state of the user using an emotion engine that analyzes the emotional state of the user;
[1352] A means for analyzing the user's learning history and content usage data and distributing revenue to content providers;
[1353] A system including:
[1354] (Claim 2)
[1355] 10. The system of claim 1, further comprising means for converting audio data and image data into text data.
[1356] (Claim 3)
[1357] 10. The system of claim 1, further comprising means for dynamically linking relevant content based on categories and keywords of entrance exam questions. [Explanation of symbols]
[1358] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for analyzing text data using a natural language processing engine and extracting keywords and tags; A means for analyzing entrance exam questions and mapping related entertainment content; A means to provide a user interface that allows users to search for content by age and category; a means for displaying content related to a problem when the user solves the problem; A means for analyzing the user's learning history and content usage data and distributing revenue to content providers; A system including:
2. 10. The system of claim 1, further comprising means for converting audio data and image data into text data.
3. The system of claim 1 , further comprising means for dynamically linking relevant content based on categories and keywords of entrance exam questions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A