System
A system that translates and optimizes information in real-time based on user characteristics addresses language barriers, allowing individuals to access education effectively.
Patent Information
- Application Number
- JP2024122825
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Language barriers prevent many individuals, particularly children in impoverished areas, from accessing information and education effectively.
A system that allows users to set their native language, input age and interests, and translates data in real-time using generative AI, optimizing it based on user characteristics for easy understanding.
Enables anyone to access information and education in their native language, tailored to their age and interests, overcoming language barriers.
Smart Images

Figure 2026021143000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, disparities in information and education still exist. In particular, language barriers are a problem that prevent many people from accessing information and education. To ensure that children in impoverished areas have the opportunity to receive an education, a system that removes language barriers and provides easy access is needed. To solve this problem, a system that translates information in real time and optimizes it to suit the characteristics of the user is needed. [Means for solving the problem]
[0005] This invention is a system that includes a means for a user to set their native language, a means for inputting age and interest information, a means for acquiring data from information sources in cooperation with a data collection means, a means for translating the acquired data into the user's native language in real time using a generative AI, a means for optimizing the translated data based on the user's age and interests, and a means for providing the optimized data to the user. This system enables anyone to access information and education in real time, regardless of language barriers.
[0006] "User" refers to the entity that uses the system and is an individual who receives information and services.
[0007] "Mother tongue" refers to the language that a user uses on a daily basis and that is easiest for the user to understand.
[0008] "Age" refers to the number of years based on the user's year of birth, and is a parameter used to optimize information according to the user's life stage.
[0009] "Interest information" refers to information about fields or themes in which a user is interested.
[0010] "Data collection methods" refers to the technologies or processes that obtain data from external sources (e.g., websites, video platforms, news feeds, etc.).
[0011] "Source" refers to the medium that provides the data, such as internet content, television, or video platforms.
[0012] "Generative AI means" refers to technologies that use machine learning and artificial intelligence techniques to generate and translate data in real time.
[0013] "Real-time translation" refers to the process by which information is captured and instantly converted into the required language without delay.
[0014] "Optimization measures" refers to the technology that adjusts the translated data according to the user's basic information (age, interests) and converts it into the most effective and understandable format.
[0015] "Database" refers to a system for efficiently storing and managing information.
[0016] "Means of providing data" refers to the technology or method by which a user receives optimized information. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time and provides it in an optimized format. Specific embodiments of this system are described below.
[0039] Initial System Setup
[0040] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0041] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0042] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0043] Data collection and translation
[0044] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0045] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0046] Information Optimization
[0047] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0048] Providing information
[0049] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0050] Specific examples
[0051] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0052] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0053] The optimized data is sent from the server to the terminal and displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0054] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The user launches the application for the first time.
[0058] Step 2:
[0059] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0060] Step 3:
[0061] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0062] Step 4:
[0063] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0064] Step 5:
[0065] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0066] Step 6:
[0067] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0068] Step 7:
[0069] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0070] Step 8:
[0071] The server periodically sends API requests to collect data from sources (websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0072] Step 9:
[0073] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0074] Step 10:
[0075] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0076] Step 11:
[0077] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0078] Step 12:
[0079] The terminal displays the received data on the screen, allowing users to view information optimized in their native language in real time.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] In today's globalized society, it is becoming increasingly important to access and understand information across language and cultural barriers. A particular challenge is the difficulty for users with different native languages to obtain the most appropriate information in real time, tailored to their age and interests. Furthermore, the lack of integration between the information gathering, translation, and optimization processes makes it difficult to provide users with the information they want quickly and appropriately.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for a user to set their native language, means for a user to input age and interest information, means for collecting data from information sources, means for temporarily storing the collected data, means for translating the acquired data into the user's native language in real time using a generative AI model, means for generating prompt sentences and sending them to the AI model, means for optimizing the translated data based on the user's age and interest information, and means for providing the optimized data to the user, thereby enabling users to obtain optimal information according to their age and interests in real time, regardless of language or cultural barriers.
[0085] The "means for a user to set his / her native language" refers to an interface that allows a user to select his / her native language and input and transmit that information to the system.
[0086] "Means for users to input age and interest information" refers to an interface through which users input information about their age and areas of interest and transmit that information to the system.
[0087] "Means of collecting data from sources" refers to the functionality for obtaining data from external sources such as websites, video platforms, news feeds, etc.
[0088] "Means for temporarily storing collected data" refers to the storage or cache function for temporarily storing collected data.
[0089] "Generative AI model" refers to a generative AI (e.g., OpenAI's GPT-3) used to perform tasks such as natural language processing.
[0090] "Means of translating acquired data into the user's native language in real time" refers to a function that uses a generative AI model to instantly translate acquired data into the user's native language.
[0091] "Means for generating prompt sentences and sending them to an AI model" refers to the function for creating prompt sentences that provide specific translation or simplification instructions to the generative AI model and sending them to the model.
[0092] "Means for optimizing translated data based on user age and interest information" refers to functionality for adjusting and optimizing translated data to make it easier to understand based on the user's age and interests.
[0093] "Means for providing optimized data to a user" refers to an interface or communication means for delivering optimized data to a user and displaying that information.
[0094] This invention is a system that allows users to set their native language and input their age and interest information, and then translates data obtained from various sources in real time and provides it in an optimized format. Specific embodiments of the invention are described below. This system consists of three main elements: a server, a terminal, and a user.
[0095] Initial System Setup
[0096] When a user starts the application, the native language setting interface is displayed first. The user selects their native language on this screen and clicks the "Confirm" button to send the setting information to the server. The server saves the received native language setting information in a database and manages it as a user profile.
[0097] The device then presents the user with a screen to input their age and interest information. The user inputs their age, selects a specific interest area (e.g., "science" or "technology"), and clicks the "Submit" button to send the information to the server. The server stores this information in a database for later data optimization.
[0098] Data collection and translation
[0099] The server connects with sources (e.g. websites, video platforms, news feeds, etc.) and collects the required data. This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server.
[0100] The server then uses a generative AI model (e.g., OpenAI's GPT-3) to translate the collected data into the user's native language in real time. Specifically, the server generates a prompt and sends it to the AI model along with the collected data. Examples of prompts include "Translate this video description into Japanese." or "Please translate a YouTube video about science and technology in a simplified form for a 10-year-old Japanese speaker."
[0101] Information Optimization
[0102] The server optimizes the translated data based on the user's age and interests: for example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and the elements of interest will be highlighted.
[0103] Providing information
[0104] The optimized information is sent from the server to the device. The device receives this information and displays it to the user. The user can easily understand the information that has been translated and optimized into their native language in real time.
[0105] Specific examples
[0106] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0107] The collected video data is translated into Japanese by the server using generative AI and executed using example prompt sentences. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. The optimized data is sent from the server to the device and finally displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0108] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0110] Step 1:
[0111] The user launches an application.
[0112] When a user starts the application, a native language setting interface is displayed. The user selects their native language on this screen and clicks the OK button.
[0113] Input: Select your native language
[0114] Output: Selected native language information is generated
[0115] Step 2:
[0116] The server receives the native language setting information and stores it in a database.
[0117] The server stores the received native language setting information in a database as a user profile.
[0118] Input: Native language setting information
[0119] Output: Native language information stored in a database
[0120] Step 3:
[0121] The device displays a screen where you can enter your age and interests.
[0122] The user enters their age and area of interest (e.g., "science," "technology") and clicks the submit button.
[0123] Input: Enter your age and interests
[0124] Output: Age and interest information entered
[0125] Step 4:
[0126] The server receives the age and interest information and stores it in a database.
[0127] The server stores the received age and interest information as a user profile in a database.
[0128] Input: Age and Interests
[0129] Output: Age and interest information stored in a database
[0130] Step 5:
[0131] The server collects data from the sources.
[0132] The server uses APIs to collect data from websites and video platforms in specified categories, such as "science" or "technology" videos from YouTube.
[0133] Input: Specified data category
[0134] Output: Raw data collected
[0135] Step 6:
[0136] The server temporarily stores the collected data.
[0137] The server stores the collected raw data in temporary storage or cache.
[0138] Input: Raw data collected
[0139] Output: Temporarily saved data
[0140] Step 7:
[0141] The server translates the data using a generative AI model.
[0142] The server generates a prompt and sends it to the generative AI model along with the temporarily saved data. The prompt includes instructions such as "Translate this video description into Japanese."
[0143] Input: Temporarily saved data, prompt text
[0144] Output: Data translated by the generative AI
[0145] Step 8:
[0146] The server optimizes the translation data based on the user's age and interests.
[0147] The server then feeds the translated data back into the generative AI model with specific prompts to tailor the content based on age and interests, such as "Please simplify it for a 10-year-old child."
[0148] Input: Translated data, age and interest information, optimization prompts
[0149] Output: Optimized data
[0150] Step 9:
[0151] The server sends the optimized data to the device.
[0152] The server transmits the optimized data to the user's terminal.
[0153] Input: Optimized data
[0154] Output: Data sent to the terminal
[0155] Step 10:
[0156] The device displays optimized data.
[0157] The device then displays the received data in an appropriate format for the user. For example, in the case of a video platform, translated subtitles and commentary are displayed.
[0158] Input: Optimization data sent from the server
[0159] Output: Displayed optimization data
[0160] Through the above steps, users can obtain information in real time in a format that is easy to understand in their native language, according to their age and interests.
[0161] (Application example 1)
[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0163] Conventional information provision systems often lack the ability to provide content that is individually optimized based on the user's age and interests. This has led to problems, particularly in the field of learning, where information is not provided according to the user's level of understanding, resulting in reduced learning efficiency. Furthermore, few systems have real-time translation capabilities, making it difficult to overcome language barriers. The present invention aims to solve these problems by providing individually optimized learning content to users in real time.
[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0165] In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for providing the optimized data to the user, and means for optimizing and providing learning content based on the information set by the user, thereby making it possible to provide individually optimized learning content in real time according to the user's age and interests.
[0166] The "means for the user to set his / her native language" is an interface or function that allows the user to select his / her native language and set that information.
[0167] The "means for users to input age and interest information" refers to an interface or function that allows users to input their own age and areas of interest.
[0168] "Means for obtaining data from information sources in cooperation with data collection means" refers to the function of collecting data from various information sources (e.g., websites, news feeds, etc.) and incorporating it into the system.
[0169] "Generative AI means for translating acquired data into the user's native language in real time" refers to generative AI technology for instantly translating collected data into the user's native language.
[0170] "Means for optimizing translated data based on the user's age and interests" refers to a function that adjusts translated data based on the user's age and interest information and processes it into an appropriate form.
[0171] The "means for providing optimized data to the user" refers to an interface or function for displaying or providing the data that has been subjected to optimization processing to the user.
[0172] "Means for optimizing and providing learning content based on information set by the user" refers to a function that adjusts and effectively provides learning content based on the native language, age, and interest information set by the user.
[0173] A system for realizing this invention includes a means for a user to set their native language, a means for a user to input age and interest information, a means for acquiring data from information sources in cooperation with a data collection means, a generation AI means for translating the acquired data into the user's native language in real time, a means for optimizing the translated data based on the user's age and interests, a means for providing the optimized data to the user, and a means for optimizing and providing learning content based on the information set by the user.
[0174] Initial System Setup
[0175] When a user launches the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information, and the user enters their age and interests. This information is also sent to the server and used for data optimization.
[0176] Data collection and translation
[0177] The server collects data by interacting with information sources (e.g., websites, news feeds, etc.). The collected data is temporarily stored on the server and then translated in real time using a generative AI method. The target language for translation is the native language selected by the user in the initial settings. The generative AI model used is the Helsinki-NLP / opus-mt-en-jap model from the transformers library.
[0178] Information Optimization
[0179] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0180] Providing information
[0181] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0182] Specific examples
[0183] For example, when a 10-year-old child whose native language is Japanese launches the application for the first time, sets their native language, and enters their age and interests, the server collects data from science-related sources. This collected data is translated into Japanese by the server using generative AI, and the content is simplified based on the user's age and interests before being provided to the user. This allows the user to receive science and technology information in Japanese that is easy to understand.
[0184] Prompt Sentence Examples
[0185] For example, use the following prompt to launch a generative AI model:
[0186] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0187] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0189] Step 1:
[0190] When a user launches the application, a native language setting screen is displayed. Here, the user selects their native language. Input: User's native language. Output: User's native language setting information is sent from the device to the server. The server stores this information in a database and manages it as part of the user profile.
[0191] Step 2:
[0192] The device then displays a screen for the user to enter their age and interest information. The user enters their age and areas of interest. Input: User's age and interest information. Output: This information is sent from the device to the server. The server also stores this in a database for later data optimization.
[0193] Step 3:
[0194] The server collects data from sources (e.g. RSS feeds, websites) periodically or upon user request. Input: URL of source. Output: Collected data. The server temporarily stores the collected data.
[0195] Step 4:
[0196] The server translates the temporarily stored data in real time into the user's native language using the Helsinki-NLP / opus-mt-en-jap model from the transformers library. Input: Collected data, user's native language setting information. Output: Translated data. The server obtains this translated data using a generative AI method.
[0197] Step 5:
[0198] The server optimizes the translated data based on the user's age and interests, for example by replacing technical terms with simpler terms or highlighting information related to their area of interest. Input: Translated data, user's age and interests. Output: Optimized data. Specific actions include filtering and rephrasing the text.
[0199] Step 6:
[0200] The server sends the optimized data to the terminal. Input: Optimized data. Output: Data sent to the terminal. The terminal displays this data to the user. The user can understand this real-time optimized information in their native language.
[0201] These steps allow users to receive personalized, real-time information tailored to their needs. Specifically, the following prompts are used to trigger the generative AI model for translation and optimization:
[0202] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0203] This processing flow is expected to significantly improve information accuracy and learning efficiency.
[0204] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0205] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time, and further recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[0206] Initial System Setup
[0207] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0208] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0209] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0210] Emotion recognition by emotion engine
[0211] The device is equipped with an emotion engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera and microphone, and analyzes the user's emotions from their facial expressions and tone of voice.
[0212] Data collection and translation
[0213] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0214] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0215] Emotional optimization of information
[0216] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine. For example, if the user is excited, it will provide more stimulating content related to their interests, and conversely, if the user is relaxed, it will provide calming content.
[0217] Providing information
[0218] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[0219] Specific examples
[0220] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0221] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0222] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0223] The optimized data is sent from the server to the device and displayed to User A. User A can watch science and technology videos in easy-to-understand Japanese in a format that best suits their emotional state at the time.
[0224] As described above, this system can reduce information and education gaps by optimizing information according to the user's characteristics and emotional state and translating it into their native language in real time.
[0225] The processing flow will be explained below.
[0226] Step 1:
[0227] The user launches the application for the first time.
[0228] Step 2:
[0229] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0230] Step 3:
[0231] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0232] Step 4:
[0233] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0234] Step 5:
[0235] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0236] Step 6:
[0237] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0238] Step 7:
[0239] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0240] Step 8:
[0241] The device uses an emotion engine to recognize the user's emotions in real time, analyzing the user's facial expressions and tone of voice via the camera and microphone to determine their emotional state.
[0242] Step 9:
[0243] The server periodically sends API requests to collect data from sources (e.g., websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0244] Step 10:
[0245] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0246] Step 11:
[0247] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0248] Step 12:
[0249] The server also optimizes the data based on the emotional state recognized by the emotion engine: for example, if the user is excited, it will adjust the content to be more stimulating, and if the user is relaxed, it will adjust the content to be calmer.
[0250] Step 13:
[0251] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0252] Step 14:
[0253] The terminal displays the received data on the screen, allowing users to view optimized information in their native language in real time.
[0254] As a concrete example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0255] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0256] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0257] This allows user A to watch easy-to-understand science and technology videos in Japanese in a way that best suits his or her emotional state at the time.
[0258] Example 2
[0259] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0260] While conventional information provision systems optimize data based on a user's native language, age, and interests, they face the problem of difficulty in providing information that takes into account the user's emotional state. Furthermore, because it is difficult to simultaneously collect and translate information in real time and optimize it based on emotions, it has not been possible to provide truly personalized information to users. The present invention aims to solve this problem.
[0261] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0262] In this invention, the server includes a means for the user to set their native language, a means for the user to input their age and interest information, a means for acquiring data from an information source in cooperation with the data collection means, a generation AI means for translating the acquired data into the user's native language in real time, an emotion recognition means for recognizing the user's emotions using a camera or microphone in the terminal, a means for optimizing the translated data based on the user's age, interests, and emotions, and a means for providing the optimized data to the user. This enables the provision of more highly personalized information based on the user's emotional state as well.
[0263] The "means for the user to set his native language" is a means for the user to specify the language he uses and register that information in the system.
[0264] "Means for users to input age and interest information" refers to means for users to input their own age and interests into the system.
[0265] "Means for acquiring data from information sources in cooperation with data collection means" refers to means for connecting to external information sources and collecting necessary data.
[0266] "Generative AI means" is an artificial intelligence technology that translates collected data into the user's native language in real time.
[0267] The "emotion recognition means" is a means for recognizing the user's emotional state by analyzing the user's facial expressions and tone of voice using a camera or microphone.
[0268] The "means for optimizing translated data based on the age, interests, and emotions of the user" refers to a means for providing translated data in an optimal form according to the age, interests, and emotions of the user.
[0269] The "means for providing optimized data to a user" refers to a means for displaying or distributing the optimized data to a user.
[0270] MODE FOR CARRYING OUT THE INVENTION
[0271] The present invention is a system that provides optimized information in real time, taking into consideration the user's native language, age, interest information, and emotional state. An embodiment of this system will now be described in detail.
[0272] The basic structure of the system consists of a server, a terminal, and a user. Below we explain how each element functions.
[0273] Initial Setup
[0274] When a user starts the application, a native language setting screen is displayed. The user selects their native language using this screen, and the selected information is sent to the server. The server stores the received native language setting information in a database, which is then managed as part of the user profile.
[0275] Next, the device presents the user with a screen for entering their age and interest information. The user enters their age and interest information here, and this information is also sent to the server. The server stores the age and interest information in a database and uses it to provide future information.
[0276] emotion recognition
[0277] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed using the device's built-in camera and microphone. Specifically, the camera analyzes the user's facial expressions and the microphone analyzes the user's tone of voice to recognize the user's emotional state. The results of this analysis are sent to a server in real time, where they are stored and managed.
[0278] Data collection and translation
[0279] The server has the function of collecting data from sources, such as news feeds, video platforms, websites, etc., periodically or upon user request. The collected data is temporarily stored on the server.
[0280] This collected data is then translated in real time into the user's native language using generative AI tools such as the Google Translate API, allowing the user to receive information in their native language of choice.
[0281] Emotional optimization of information
[0282] The translated data is optimized on the server based on the user's age and interests. The user's emotional state, obtained from an emotion recognition engine, is also taken into consideration. Specifically, if the user is excited, more stimulating content is selected, and if the user is relaxed, more calming content is selected. This optimization process is intended to provide information in the most appropriate form for the user.
[0283] Providing information
[0284] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive the information in their native language and in a format that best suits their emotional state at the time.
[0285] Specific examples
[0286] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A launches the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel. The collected video data is translated into Japanese using the Google Translate API. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. In addition, an emotion recognition engine analyzes User A's emotions; if User A is excited, more stimulating science experiment videos are provided, and if User A is calm, educational and calming videos are selected.
[0287] Prompt Sentence Examples
[0288] How can you specifically implement a system that takes into account a user's native language and interests, and optimizes information based on the user's emotional state?
[0289] As described above, this is a system in which the server, terminal, and user work together to provide information in real time that is tailored to the user's characteristics and emotions.
[0290] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0291] Step 1:
[0292] The user launches an application.
[0293] As a specific operation, the application displays a native language setting screen.
[0294] Input: User sets native language.
[0295] Output: Sends the configured native language information to the server.
[0296] The server stores the received native language setting information in a database and manages it as a user profile.
[0297] Step 2:
[0298] The device displays a screen where you can enter your age and interests.
[0299] Input: The user enters age and interest information.
[0300] Output: Send the entered age and interest information to the server.
[0301] The server stores the received age and interest information in a database for future use in providing information.
[0302] Step 3:
[0303] The device's built-in emotion recognition engine uses the camera and microphone to recognize the user's emotions.
[0304] Specifically, the camera analyzes the user's facial expression and the microphone analyzes the user's tone of voice.
[0305] Input: User facial and voice data obtained from camera and microphone.
[0306] Output: The emotion recognition engine analyzes the user's emotional state and sends it to the server.
[0307] The server stores this affective information for use in subsequent data optimization.
[0308] Step 4:
[0309] The server collects data from pre-specified sources.
[0310] Specific sources of information include news feeds, video platforms, websites, etc.
[0311] Input: Various data collected from sources.
[0312] Output: Temporarily stores collected data.
[0313] The server then translates the data in real time using generative AI methods such as the Google Translate API.
[0314] Step 5:
[0315] The server optimizes the translated data based on the user's age, interests and emotional state.
[0316] Input: Translated data, user age, interest, and emotion information.
[0317] Output: Generates optimized data.
[0318] The optimization process involves filtering and categorizing data and emphasizing certain elements, for example, if a user is excited, more stimulating content will be selected, and if they are relaxed, more calming content will be selected.
[0319] Step 6:
[0320] The server sends the optimized data to the device.
[0321] Input: Optimized data.
[0322] Output: Sends data to a terminal and displays it to the user.
[0323] The terminal displays the received optimization data in a form that can be used visually or audibly by the user.
[0324] Through the above steps, the user can receive information optimized for their current emotional state and personal attributes in real time.
[0325] (Application example 2)
[0326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0327] Currently, many systems that provide information to users do not adequately optimize it based on the user's language, age, interests, or even real-time emotions. This can result in information not reaching the user appropriately, leading to reduced receptivity and comprehension. In particular, providing information that does not match the user's emotional state can hinder the user's experience. Therefore, there is a need for a system that optimizes information based on the user's characteristics and real-time emotions and delivers it at the appropriate time.
[0328] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for recognizing the user's emotions and optimizing information, and means for displaying the optimized data to the user. This makes it possible to optimize information based on the user's characteristics and real-time emotional state and provide information in the user's native language at an appropriate time.
[0329] The "means for the user to set his native language" is an interface that allows the user to select the language he or she uses, and to input and save that information.
[0330] The "means for users to input age and interest information" refers to an interface that allows users to input their age and areas of interest and provide that information.
[0331] "Means for acquiring data from information sources in cooperation with data collection means" refers to a mechanism for collecting data from various information sources on a network, and a means for incorporating the data into the system.
[0332] "Generative AI means for translating acquired data into the user's native language in real time" refers to a mechanism that includes a generative AI model for instantly translating collected data into the user's native language.
[0333] "Means for optimizing translated data based on the user's age and interests" refers to a system that processes and organizes translated data in an appropriate form in accordance with the user's age and interest information.
[0334] "Means for recognizing user emotions and optimizing information" refers to a mechanism that analyzes the user's emotions through a camera or microphone and optimizes information in accordance with those emotions.
[0335] The "means for displaying optimized data to the user" refers to an interface for visually or audibly conveying optimized information to the user.
[0336] This invention is a system that translates data acquired from various information sources in real time by allowing a user to input information such as their native language, age, and interests, and also recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[0337] Initial Setup
[0338] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information. The user enters their age and interests and sends the information to the server. This information is stored on the server and used for data optimization.
[0339] emotion recognition
[0340] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera (e.g., Logitech C920) or microphone, and analyzes the user's emotions from their facial expressions and tone of voice. This emotion data is also sent to the server.
[0341] Data collection and translation
[0342] The server collects data by interacting with information sources (e.g., websites, video platforms, news feeds, etc.). This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server and then translated in real time using generative AI means. The target language for translation is the native language selected by the user in the initial settings. This process enables information to be transmitted across language barriers.
[0343] Data Optimization
[0344] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine: if the user is excited, for example, it will provide more stimulating content related to their interests, and conversely, if they are relaxed, it will provide calming content.
[0345] Providing information
[0346] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[0347] Specific examples
[0348] For example, if a user sets "music" and "sports" as their interests and is recognized as having an "excited" emotion, the system will provide the user with the latest sports news and music ranking videos. The system optimizes information based on the user's characteristics and emotional state, and translates it into their native language in real time, thereby reducing information and education gaps.
[0349] Prompt Sentence Examples
[0350] "When users are emotionally excited, generate exciting content related to their interests. Also, translate that content into Japanese. When they are excited, provide content that emphasizes more active elements."
[0351] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0352] Step 1:
[0353] When a user starts an application, the device first displays the native language setting screen. The user selects their native language and enters the information. This input data is sent to the server, which then stores the received native language setting information in a database. Input: Native language information entered by the user. Output: Native language information stored in the database.
[0354] Step 2:
[0355] Next, the terminal displays a screen for the user to enter their age and interest information. The user enters their age and multiple interests and sends this information to the server. The server stores this information in a database. Input: Age and interest information entered by the user. Output: Age and interest information stored in the database.
[0356] Step 3:
[0357] The device uses a camera and microphone to analyze the user's emotions in real time. The emotion recognition engine generates emotion data from the user's facial expressions and tone of voice, and this emotion data is sent to the server. Input: User's facial expressions and voice data captured by the camera and microphone. Output: Emotion data sent to the server.
[0358] Step 4:
[0359] The server periodically collects data from sources (e.g. websites or video platforms). This collected data is temporarily stored on the server. Input: Raw data obtained from sources. Output: Raw data temporarily stored on the server.
[0360] Step 5:
[0361] The server uses generative AI means (e.g., generative AI models) to translate this collected data into the user's native language in real time, thus overcoming language barriers. Input: Temporarily stored raw data. Output: Real-time translated data.
[0362] Step 6:
[0363] The server processes the translated data to optimize it based on the user's age and interest information. It also takes into account the emotional state recognized by the emotion engine to generate optimal information. For example, if the user is excited, the content will be optimized to be more stimulating. Input: Translated data, age and interest information, emotional data. Output: Information optimized for the user.
[0364] Step 7:
[0365] The optimized information is sent from the server to the terminal. The terminal receives this information and displays it to the user in an appropriate interface. This allows the user to receive optimized information in real time. Input: Optimized information. Output: Optimized information displayed on the terminal.
[0366] As described above, by going through each step, a system is realized that allows users to receive information optimized for them in their native language in real time.
[0367] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0368] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0369] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0370] [Second embodiment]
[0371] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0372] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0373] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0374] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0375] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0376] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0377] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0378] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0379] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0380] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0381] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0382] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0383] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time and provides it in an optimized format. Specific embodiments of this system are described below.
[0384] Initial System Setup
[0385] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0386] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0387] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0388] Data collection and translation
[0389] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0390] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0391] Information Optimization
[0392] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0393] Providing information
[0394] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0395] Specific examples
[0396] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0397] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0398] The optimized data is sent from the server to the terminal and displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0399] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0400] The processing flow will be explained below.
[0401] Step 1:
[0402] The user launches the application for the first time.
[0403] Step 2:
[0404] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0405] Step 3:
[0406] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0407] Step 4:
[0408] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0409] Step 5:
[0410] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0411] Step 6:
[0412] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0413] Step 7:
[0414] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0415] Step 8:
[0416] The server periodically sends API requests to collect data from sources (websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0417] Step 9:
[0418] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0419] Step 10:
[0420] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0421] Step 11:
[0422] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0423] Step 12:
[0424] The terminal displays the received data on the screen, allowing users to view information optimized in their native language in real time.
[0425] Example 1
[0426] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0427] In today's globalized society, it is becoming increasingly important to access and understand information across language and cultural barriers. A particular challenge is the difficulty for users with different native languages to obtain the most appropriate information in real time, tailored to their age and interests. Furthermore, the lack of integration between the information gathering, translation, and optimization processes makes it difficult to provide users with the information they want quickly and appropriately.
[0428] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0429] In this invention, the server includes means for a user to set their native language, means for a user to input age and interest information, means for collecting data from information sources, means for temporarily storing the collected data, means for translating the acquired data into the user's native language in real time using a generative AI model, means for generating prompt sentences and sending them to the AI model, means for optimizing the translated data based on the user's age and interest information, and means for providing the optimized data to the user, thereby enabling users to obtain optimal information according to their age and interests in real time, regardless of language or cultural barriers.
[0430] The "means for a user to set his / her native language" refers to an interface that allows a user to select his / her native language and input and transmit that information to the system.
[0431] "Means for users to input age and interest information" refers to an interface through which users input information about their age and areas of interest and transmit that information to the system.
[0432] "Means of collecting data from sources" refers to the functionality for obtaining data from external sources such as websites, video platforms, news feeds, etc.
[0433] "Means for temporarily storing collected data" refers to the storage or cache function for temporarily storing collected data.
[0434] "Generative AI model" refers to a generative AI (e.g., OpenAI's GPT-3) used to perform tasks such as natural language processing.
[0435] "Means of translating acquired data into the user's native language in real time" refers to a function that uses a generative AI model to instantly translate acquired data into the user's native language.
[0436] "Means for generating prompt sentences and sending them to an AI model" refers to the function for creating prompt sentences that provide specific translation or simplification instructions to the generative AI model and sending them to the model.
[0437] "Means for optimizing translated data based on user age and interest information" refers to functionality for adjusting and optimizing translated data to make it easier to understand based on the user's age and interests.
[0438] "Means for providing optimized data to a user" refers to an interface or communication means for delivering optimized data to a user and displaying that information.
[0439] This invention is a system that allows users to set their native language and input their age and interest information, and then translates data obtained from various sources in real time and provides it in an optimized format. Specific embodiments of the invention are described below. This system consists of three main elements: a server, a terminal, and a user.
[0440] Initial System Setup
[0441] When a user starts the application, the native language setting interface is displayed first. The user selects their native language on this screen and clicks the "Confirm" button to send the setting information to the server. The server saves the received native language setting information in a database and manages it as a user profile.
[0442] The device then presents the user with a screen to input their age and interest information. The user inputs their age, selects a specific interest area (e.g., "science" or "technology"), and clicks the "Submit" button to send the information to the server. The server stores this information in a database for later data optimization.
[0443] Data collection and translation
[0444] The server connects with sources (e.g. websites, video platforms, news feeds, etc.) and collects the required data. This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server.
[0445] The server then uses a generative AI model (e.g., OpenAI's GPT-3) to translate the collected data into the user's native language in real time. Specifically, the server generates a prompt and sends it to the AI model along with the collected data. Examples of prompts include "Translate this video description into Japanese." or "Please translate a YouTube video about science and technology in a simplified form for a 10-year-old Japanese speaker."
[0446] Information Optimization
[0447] The server optimizes the translated data based on the user's age and interests: for example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and the elements of interest will be highlighted.
[0448] Providing information
[0449] The optimized information is sent from the server to the device. The device receives this information and displays it to the user. The user can easily understand the information that has been translated and optimized into their native language in real time.
[0450] Specific examples
[0451] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0452] The collected video data is translated into Japanese by the server using generative AI and executed using example prompt sentences. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. The optimized data is sent from the server to the device and finally displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0453] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0454] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0455] Step 1:
[0456] The user launches an application.
[0457] When a user starts the application, a native language setting interface is displayed. The user selects their native language on this screen and clicks the OK button.
[0458] Input: Select your native language
[0459] Output: Selected native language information is generated
[0460] Step 2:
[0461] The server receives the native language setting information and stores it in a database.
[0462] The server stores the received native language setting information in a database as a user profile.
[0463] Input: Native language setting information
[0464] Output: Native language information stored in a database
[0465] Step 3:
[0466] The device displays a screen where you can enter your age and interests.
[0467] The user enters their age and area of interest (e.g., "science," "technology") and clicks the submit button.
[0468] Input: Enter your age and interests
[0469] Output: Age and interest information entered
[0470] Step 4:
[0471] The server receives the age and interest information and stores it in a database.
[0472] The server stores the received age and interest information as a user profile in a database.
[0473] Input: Age and Interests
[0474] Output: Age and interest information stored in a database
[0475] Step 5:
[0476] The server collects data from the sources.
[0477] The server uses APIs to collect data from websites and video platforms in specified categories, such as "science" or "technology" videos from YouTube.
[0478] Input: Specified data category
[0479] Output: Raw data collected
[0480] Step 6:
[0481] The server temporarily stores the collected data.
[0482] The server stores the collected raw data in temporary storage or cache.
[0483] Input: Raw data collected
[0484] Output: Temporarily saved data
[0485] Step 7:
[0486] The server translates the data using a generative AI model.
[0487] The server generates a prompt and sends it to the generative AI model along with the temporarily saved data. The prompt includes instructions such as "Translate this video description into Japanese."
[0488] Input: Temporarily saved data, prompt text
[0489] Output: Data translated by the generative AI
[0490] Step 8:
[0491] The server optimizes the translation data based on the user's age and interests.
[0492] The server then feeds the translated data back into the generative AI model with specific prompts to tailor the content based on age and interests, such as "Please simplify it for a 10-year-old child."
[0493] Input: Translated data, age and interest information, optimization prompts
[0494] Output: Optimized data
[0495] Step 9:
[0496] The server sends the optimized data to the device.
[0497] The server transmits the optimized data to the user's terminal.
[0498] Input: Optimized data
[0499] Output: Data sent to the terminal
[0500] Step 10:
[0501] The device displays optimized data.
[0502] The device then displays the received data in an appropriate format for the user. For example, in the case of a video platform, translated subtitles and commentary are displayed.
[0503] Input: Optimization data sent from the server
[0504] Output: Displayed optimization data
[0505] Through the above steps, users can obtain information in real time in a format that is easy to understand in their native language, according to their age and interests.
[0506] (Application example 1)
[0507] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0508] Conventional information provision systems often lack the ability to provide content that is individually optimized based on the user's age and interests. This has led to problems, particularly in the field of learning, where information is not provided according to the user's level of understanding, resulting in reduced learning efficiency. Furthermore, few systems have real-time translation capabilities, making it difficult to overcome language barriers. The present invention aims to solve these problems by providing individually optimized learning content to users in real time.
[0509] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0510] In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for providing the optimized data to the user, and means for optimizing and providing learning content based on the information set by the user, thereby making it possible to provide individually optimized learning content in real time according to the user's age and interests.
[0511] The "means for the user to set his / her native language" is an interface or function that allows the user to select his / her native language and set that information.
[0512] The "means for users to input age and interest information" refers to an interface or function that allows users to input their own age and areas of interest.
[0513] "Means for obtaining data from information sources in cooperation with data collection means" refers to the function of collecting data from various information sources (e.g., websites, news feeds, etc.) and incorporating it into the system.
[0514] "Generative AI means for translating acquired data into the user's native language in real time" refers to generative AI technology for instantly translating collected data into the user's native language.
[0515] "Means for optimizing translated data based on the user's age and interests" refers to a function that adjusts translated data based on the user's age and interest information and processes it into an appropriate form.
[0516] The "means for providing optimized data to the user" refers to an interface or function for displaying or providing the data that has been subjected to optimization processing to the user.
[0517] "Means for optimizing and providing learning content based on information set by the user" refers to a function that adjusts and effectively provides learning content based on the native language, age, and interest information set by the user.
[0518] A system for realizing this invention includes a means for a user to set their native language, a means for a user to input age and interest information, a means for acquiring data from information sources in cooperation with a data collection means, a generation AI means for translating the acquired data into the user's native language in real time, a means for optimizing the translated data based on the user's age and interests, a means for providing the optimized data to the user, and a means for optimizing and providing learning content based on the information set by the user.
[0519] Initial System Setup
[0520] When a user launches the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information, and the user enters their age and interests. This information is also sent to the server and used for data optimization.
[0521] Data collection and translation
[0522] The server collects data by interacting with information sources (e.g., websites, news feeds, etc.). The collected data is temporarily stored on the server and then translated in real time using a generative AI method. The target language for translation is the native language selected by the user in the initial settings. The generative AI model used is the Helsinki-NLP / opus-mt-en-jap model from the transformers library.
[0523] Information Optimization
[0524] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0525] Providing information
[0526] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0527] Specific examples
[0528] For example, when a 10-year-old child whose native language is Japanese launches the application for the first time, sets their native language, and enters their age and interests, the server collects data from science-related sources. This collected data is translated into Japanese by the server using generative AI, and the content is simplified based on the user's age and interests before being provided to the user. This allows the user to receive science and technology information in Japanese that is easy to understand.
[0529] Prompt Sentence Examples
[0530] For example, use the following prompt to launch a generative AI model:
[0531] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0532] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0533] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0534] Step 1:
[0535] When a user launches the application, a native language setting screen is displayed. Here, the user selects their native language. Input: User's native language. Output: User's native language setting information is sent from the device to the server. The server stores this information in a database and manages it as part of the user profile.
[0536] Step 2:
[0537] The device then displays a screen for the user to enter their age and interest information. The user enters their age and areas of interest. Input: User's age and interest information. Output: This information is sent from the device to the server. The server also stores this in a database for later data optimization.
[0538] Step 3:
[0539] The server collects data from sources (e.g. RSS feeds, websites) periodically or upon user request. Input: URL of source. Output: Collected data. The server temporarily stores the collected data.
[0540] Step 4:
[0541] The server translates the temporarily stored data in real time into the user's native language using the Helsinki-NLP / opus-mt-en-jap model from the transformers library. Input: Collected data, user's native language setting information. Output: Translated data. The server obtains this translated data using a generative AI method.
[0542] Step 5:
[0543] The server optimizes the translated data based on the user's age and interests, for example by replacing technical terms with simpler terms or highlighting information related to their area of interest. Input: Translated data, user's age and interests. Output: Optimized data. Specific actions include filtering and rephrasing the text.
[0544] Step 6:
[0545] The server sends the optimized data to the terminal. Input: Optimized data. Output: Data sent to the terminal. The terminal displays this data to the user. The user can understand this real-time optimized information in their native language.
[0546] These steps allow users to receive personalized, real-time information tailored to their needs. Specifically, the following prompts are used to trigger the generative AI model for translation and optimization:
[0547] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0548] This processing flow is expected to significantly improve information accuracy and learning efficiency.
[0549] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0550] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time, and further recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[0551] Initial System Setup
[0552] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0553] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0554] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0555] Emotion recognition by emotion engine
[0556] The device is equipped with an emotion engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera and microphone, and analyzes the user's emotions from their facial expressions and tone of voice.
[0557] Data collection and translation
[0558] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0559] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0560] Emotional optimization of information
[0561] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine. For example, if the user is excited, it will provide more stimulating content related to their interests, and conversely, if the user is relaxed, it will provide calming content.
[0562] Providing information
[0563] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[0564] Specific examples
[0565] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0566] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0567] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0568] The optimized data is sent from the server to the device and displayed to User A. User A can watch science and technology videos in easy-to-understand Japanese in a format that best suits their emotional state at the time.
[0569] As described above, this system can reduce information and education gaps by optimizing information according to the user's characteristics and emotional state and translating it into their native language in real time.
[0570] The processing flow will be explained below.
[0571] Step 1:
[0572] The user launches the application for the first time.
[0573] Step 2:
[0574] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0575] Step 3:
[0576] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0577] Step 4:
[0578] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0579] Step 5:
[0580] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0581] Step 6:
[0582] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0583] Step 7:
[0584] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0585] Step 8:
[0586] The device uses an emotion engine to recognize the user's emotions in real time, analyzing the user's facial expressions and tone of voice via the camera and microphone to determine their emotional state.
[0587] Step 9:
[0588] The server periodically sends API requests to collect data from sources (e.g., websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0589] Step 10:
[0590] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0591] Step 11:
[0592] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0593] Step 12:
[0594] The server also optimizes the data based on the emotional state recognized by the emotion engine: for example, if the user is excited, it will adjust the content to be more stimulating, and if the user is relaxed, it will adjust the content to be calmer.
[0595] Step 13:
[0596] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0597] Step 14:
[0598] The terminal displays the received data on the screen, allowing users to view optimized information in their native language in real time.
[0599] As a concrete example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0600] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0601] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0602] This allows user A to watch easy-to-understand science and technology videos in Japanese in a way that best suits his or her emotional state at the time.
[0603] Example 2
[0604] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0605] While conventional information provision systems optimize data based on a user's native language, age, and interests, they face the problem of difficulty in providing information that takes into account the user's emotional state. Furthermore, because it is difficult to simultaneously collect and translate information in real time and optimize it based on emotions, it has not been possible to provide truly personalized information to users. The present invention aims to solve this problem.
[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0607] In this invention, the server includes a means for the user to set their native language, a means for the user to input their age and interest information, a means for acquiring data from an information source in cooperation with the data collection means, a generation AI means for translating the acquired data into the user's native language in real time, an emotion recognition means for recognizing the user's emotions using a camera or microphone in the terminal, a means for optimizing the translated data based on the user's age, interests, and emotions, and a means for providing the optimized data to the user. This enables the provision of more highly personalized information based on the user's emotional state as well.
[0608] The "means for the user to set his native language" is a means for the user to specify the language he uses and register that information in the system.
[0609] "Means for users to input age and interest information" refers to means for users to input their own age and interests into the system.
[0610] "Means for acquiring data from information sources in cooperation with data collection means" refers to means for connecting to external information sources and collecting necessary data.
[0611] "Generative AI means" is an artificial intelligence technology that translates collected data into the user's native language in real time.
[0612] The "emotion recognition means" is a means for recognizing the user's emotional state by analyzing the user's facial expressions and tone of voice using a camera or microphone.
[0613] The "means for optimizing translated data based on the age, interests, and emotions of the user" refers to a means for providing translated data in an optimal form according to the age, interests, and emotions of the user.
[0614] The "means for providing optimized data to a user" refers to a means for displaying or distributing the optimized data to a user.
[0615] MODE FOR CARRYING OUT THE INVENTION
[0616] The present invention is a system that provides optimized information in real time, taking into consideration the user's native language, age, interest information, and emotional state. An embodiment of this system will now be described in detail.
[0617] The basic structure of the system consists of a server, a terminal, and a user. Below we explain how each element functions.
[0618] Initial Setup
[0619] When a user starts the application, a native language setting screen is displayed. The user selects their native language using this screen, and the selected information is sent to the server. The server stores the received native language setting information in a database, which is then managed as part of the user profile.
[0620] Next, the device presents the user with a screen for entering their age and interest information. The user enters their age and interest information here, and this information is also sent to the server. The server stores the age and interest information in a database and uses it to provide future information.
[0621] emotion recognition
[0622] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed using the device's built-in camera and microphone. Specifically, the camera analyzes the user's facial expressions and the microphone analyzes the user's tone of voice to recognize the user's emotional state. The results of this analysis are sent to a server in real time, where they are stored and managed.
[0623] Data collection and translation
[0624] The server has the function of collecting data from sources, such as news feeds, video platforms, websites, etc., periodically or upon user request. The collected data is temporarily stored on the server.
[0625] This collected data is then translated in real time into the user's native language using generative AI tools such as the Google Translate API, allowing the user to receive information in their native language of choice.
[0626] Emotional optimization of information
[0627] The translated data is optimized on the server based on the user's age and interests. The user's emotional state, obtained from an emotion recognition engine, is also taken into consideration. Specifically, if the user is excited, more stimulating content is selected, and if the user is relaxed, more calming content is selected. This optimization process is intended to provide information in the most appropriate form for the user.
[0628] Providing information
[0629] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive the information in their native language and in a format that best suits their emotional state at the time.
[0630] Specific examples
[0631] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A launches the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel. The collected video data is translated into Japanese using the Google Translate API. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. In addition, an emotion recognition engine analyzes User A's emotions; if User A is excited, more stimulating science experiment videos are provided, and if User A is calm, educational and calming videos are selected.
[0632] Prompt Sentence Examples
[0633] How can you specifically implement a system that takes into account a user's native language and interests, and optimizes information based on the user's emotional state?
[0634] As described above, this is a system in which the server, terminal, and user work together to provide information in real time that is tailored to the user's characteristics and emotions.
[0635] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0636] Step 1:
[0637] The user launches an application.
[0638] As a specific operation, the application displays a native language setting screen.
[0639] Input: User sets native language.
[0640] Output: Sends the configured native language information to the server.
[0641] The server stores the received native language setting information in a database and manages it as a user profile.
[0642] Step 2:
[0643] The device displays a screen where you can enter your age and interests.
[0644] Input: The user enters age and interest information.
[0645] Output: Send the entered age and interest information to the server.
[0646] The server stores the received age and interest information in a database for future use in providing information.
[0647] Step 3:
[0648] The device's built-in emotion recognition engine uses the camera and microphone to recognize the user's emotions.
[0649] Specifically, the camera analyzes the user's facial expression and the microphone analyzes the user's tone of voice.
[0650] Input: User facial and voice data obtained from camera and microphone.
[0651] Output: The emotion recognition engine analyzes the user's emotional state and sends it to the server.
[0652] The server stores this affective information for use in subsequent data optimization.
[0653] Step 4:
[0654] The server collects data from pre-specified sources.
[0655] Specific sources of information include news feeds, video platforms, websites, etc.
[0656] Input: Various data collected from sources.
[0657] Output: Temporarily stores collected data.
[0658] The server then translates the data in real time using generative AI methods such as the Google Translate API.
[0659] Step 5:
[0660] The server optimizes the translated data based on the user's age, interests and emotional state.
[0661] Input: Translated data, user age, interest, and emotion information.
[0662] Output: Generates optimized data.
[0663] The optimization process involves filtering and categorizing data and emphasizing certain elements, for example, if a user is excited, more stimulating content will be selected, and if they are relaxed, more calming content will be selected.
[0664] Step 6:
[0665] The server sends the optimized data to the device.
[0666] Input: Optimized data.
[0667] Output: Sends data to a terminal and displays it to the user.
[0668] The terminal displays the received optimization data in a form that can be used visually or audibly by the user.
[0669] Through the above steps, the user can receive information optimized for their current emotional state and personal attributes in real time.
[0670] (Application example 2)
[0671] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0672] Currently, many systems that provide information to users do not adequately optimize it based on the user's language, age, interests, or even real-time emotions. This can result in information not reaching the user appropriately, leading to reduced receptivity and comprehension. In particular, providing information that does not match the user's emotional state can hinder the user's experience. Therefore, there is a need for a system that optimizes information based on the user's characteristics and real-time emotions and delivers it at the appropriate time.
[0673] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for recognizing the user's emotions and optimizing information, and means for displaying the optimized data to the user. This makes it possible to optimize information based on the user's characteristics and real-time emotional state and provide information in the user's native language at an appropriate time.
[0674] The "means for the user to set his native language" is an interface that allows the user to select the language he or she uses, and to input and save that information.
[0675] The "means for users to input age and interest information" refers to an interface that allows users to input their age and areas of interest and provide that information.
[0676] "Means for acquiring data from information sources in cooperation with data collection means" refers to a mechanism for collecting data from various information sources on a network, and a means for incorporating the data into the system.
[0677] "Generative AI means for translating acquired data into the user's native language in real time" refers to a mechanism that includes a generative AI model for instantly translating collected data into the user's native language.
[0678] "Means for optimizing translated data based on the user's age and interests" refers to a system that processes and organizes translated data in an appropriate form in accordance with the user's age and interest information.
[0679] "Means for recognizing user emotions and optimizing information" refers to a mechanism that analyzes the user's emotions through a camera or microphone and optimizes information in accordance with those emotions.
[0680] The "means for displaying optimized data to the user" refers to an interface for visually or audibly conveying optimized information to the user.
[0681] This invention is a system that translates data acquired from various information sources in real time by allowing a user to input information such as their native language, age, and interests, and also recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[0682] Initial Setup
[0683] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information. The user enters their age and interests and sends the information to the server. This information is stored on the server and used for data optimization.
[0684] emotion recognition
[0685] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera (e.g., Logitech C920) or microphone, and analyzes the user's emotions from their facial expressions and tone of voice. This emotion data is also sent to the server.
[0686] Data collection and translation
[0687] The server collects data by interacting with information sources (e.g., websites, video platforms, news feeds, etc.). This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server and then translated in real time using generative AI means. The target language for translation is the native language selected by the user in the initial settings. This process enables information to be transmitted across language barriers.
[0688] Data Optimization
[0689] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine: if the user is excited, for example, it will provide more stimulating content related to their interests, and conversely, if they are relaxed, it will provide calming content.
[0690] Providing information
[0691] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[0692] Specific examples
[0693] For example, if a user sets "music" and "sports" as their interests and is recognized as having an "excited" emotion, the system will provide the user with the latest sports news and music ranking videos. The system optimizes information based on the user's characteristics and emotional state, and translates it into their native language in real time, thereby reducing information and education gaps.
[0694] Prompt Sentence Examples
[0695] "When users are emotionally excited, generate exciting content related to their interests. Also, translate that content into Japanese. When they are excited, provide content that emphasizes more active elements."
[0696] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0697] Step 1:
[0698] When a user starts an application, the device first displays the native language setting screen. The user selects their native language and enters the information. This input data is sent to the server, which then stores the received native language setting information in a database. Input: Native language information entered by the user. Output: Native language information stored in the database.
[0699] Step 2:
[0700] Next, the terminal displays a screen for the user to enter their age and interest information. The user enters their age and multiple interests and sends this information to the server. The server stores this information in a database. Input: Age and interest information entered by the user. Output: Age and interest information stored in the database.
[0701] Step 3:
[0702] The device uses a camera and microphone to analyze the user's emotions in real time. The emotion recognition engine generates emotion data from the user's facial expressions and tone of voice, and this emotion data is sent to the server. Input: User's facial expressions and voice data captured by the camera and microphone. Output: Emotion data sent to the server.
[0703] Step 4:
[0704] The server periodically collects data from sources (e.g. websites or video platforms). This collected data is temporarily stored on the server. Input: Raw data obtained from sources. Output: Raw data temporarily stored on the server.
[0705] Step 5:
[0706] The server uses generative AI means (e.g., generative AI models) to translate this collected data into the user's native language in real time, thus overcoming language barriers. Input: Temporarily stored raw data. Output: Real-time translated data.
[0707] Step 6:
[0708] The server processes the translated data to optimize it based on the user's age and interest information. It also takes into account the emotional state recognized by the emotion engine to generate optimal information. For example, if the user is excited, the content will be optimized to be more stimulating. Input: Translated data, age and interest information, emotional data. Output: Information optimized for the user.
[0709] Step 7:
[0710] The optimized information is sent from the server to the terminal. The terminal receives this information and displays it to the user in an appropriate interface. This allows the user to receive optimized information in real time. Input: Optimized information. Output: Optimized information displayed on the terminal.
[0711] As described above, by going through each step, a system is realized that allows users to receive information optimized for them in their native language in real time.
[0712] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0713] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0714] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0715] [Third embodiment]
[0716] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0717] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0718] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0719] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0720] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0721] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0722] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0723] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0724] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0725] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0726] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0727] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0728] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time and provides it in an optimized format. Specific embodiments of this system are described below.
[0729] Initial System Setup
[0730] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0731] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0732] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0733] Data collection and translation
[0734] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0735] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0736] Information Optimization
[0737] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0738] Providing information
[0739] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0740] Specific examples
[0741] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0742] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0743] The optimized data is sent from the server to the terminal and displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0744] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0745] The processing flow will be explained below.
[0746] Step 1:
[0747] The user launches the application for the first time.
[0748] Step 2:
[0749] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0750] Step 3:
[0751] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0752] Step 4:
[0753] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0754] Step 5:
[0755] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0756] Step 6:
[0757] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0758] Step 7:
[0759] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0760] Step 8:
[0761] The server periodically sends API requests to collect data from sources (websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0762] Step 9:
[0763] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0764] Step 10:
[0765] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0766] Step 11:
[0767] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0768] Step 12:
[0769] The terminal displays the received data on the screen, allowing users to view information optimized in their native language in real time.
[0770] Example 1
[0771] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0772] In today's globalized society, it is becoming increasingly important to access and understand information across language and cultural barriers. A particular challenge is the difficulty for users with different native languages to obtain the most appropriate information in real time, tailored to their age and interests. Furthermore, the lack of integration between the information gathering, translation, and optimization processes makes it difficult to provide users with the information they want quickly and appropriately.
[0773] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0774] In this invention, the server includes means for a user to set their native language, means for a user to input age and interest information, means for collecting data from information sources, means for temporarily storing the collected data, means for translating the acquired data into the user's native language in real time using a generative AI model, means for generating prompt sentences and sending them to the AI model, means for optimizing the translated data based on the user's age and interest information, and means for providing the optimized data to the user, thereby enabling users to obtain optimal information according to their age and interests in real time, regardless of language or cultural barriers.
[0775] The "means for a user to set his / her native language" refers to an interface that allows a user to select his / her native language and input and transmit that information to the system.
[0776] "Means for users to input age and interest information" refers to an interface through which users input information about their age and areas of interest and transmit that information to the system.
[0777] "Means of collecting data from sources" refers to the functionality for obtaining data from external sources such as websites, video platforms, news feeds, etc.
[0778] "Means for temporarily storing collected data" refers to the storage or cache function for temporarily storing collected data.
[0779] "Generative AI model" refers to a generative AI (e.g., OpenAI's GPT-3) used to perform tasks such as natural language processing.
[0780] "Means of translating acquired data into the user's native language in real time" refers to a function that uses a generative AI model to instantly translate acquired data into the user's native language.
[0781] "Means for generating prompt sentences and sending them to an AI model" refers to the function for creating prompt sentences that provide specific translation or simplification instructions to the generative AI model and sending them to the model.
[0782] "Means for optimizing translated data based on user age and interest information" refers to functionality for adjusting and optimizing translated data to make it easier to understand based on the user's age and interests.
[0783] "Means for providing optimized data to a user" refers to an interface or communication means for delivering optimized data to a user and displaying that information.
[0784] This invention is a system that allows users to set their native language and input their age and interest information, and then translates data obtained from various sources in real time and provides it in an optimized format. Specific embodiments of the invention are described below. This system consists of three main elements: a server, a terminal, and a user.
[0785] Initial System Setup
[0786] When a user starts the application, the native language setting interface is displayed first. The user selects their native language on this screen and clicks the "Confirm" button to send the setting information to the server. The server saves the received native language setting information in a database and manages it as a user profile.
[0787] The device then presents the user with a screen to input their age and interest information. The user inputs their age, selects a specific interest area (e.g., "science" or "technology"), and clicks the "Submit" button to send the information to the server. The server stores this information in a database for later data optimization.
[0788] Data collection and translation
[0789] The server connects with sources (e.g. websites, video platforms, news feeds, etc.) and collects the required data. This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server.
[0790] The server then uses a generative AI model (e.g., OpenAI's GPT-3) to translate the collected data into the user's native language in real time. Specifically, the server generates a prompt and sends it to the AI model along with the collected data. Examples of prompts include "Translate this video description into Japanese." or "Please translate a YouTube video about science and technology in a simplified form for a 10-year-old Japanese speaker."
[0791] Information Optimization
[0792] The server optimizes the translated data based on the user's age and interests: for example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and the elements of interest will be highlighted.
[0793] Providing information
[0794] The optimized information is sent from the server to the device. The device receives this information and displays it to the user. The user can easily understand the information that has been translated and optimized into their native language in real time.
[0795] Specific examples
[0796] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0797] The collected video data is translated into Japanese by the server using generative AI and executed using example prompt sentences. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. The optimized data is sent from the server to the device and finally displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[0798] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0799] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0800] Step 1:
[0801] The user launches an application.
[0802] When a user starts the application, a native language setting interface is displayed. The user selects their native language on this screen and clicks the OK button.
[0803] Input: Select your native language
[0804] Output: Selected native language information is generated
[0805] Step 2:
[0806] The server receives the native language setting information and stores it in a database.
[0807] The server stores the received native language setting information in a database as a user profile.
[0808] Input: Native language setting information
[0809] Output: Native language information stored in a database
[0810] Step 3:
[0811] The device displays a screen where you can enter your age and interests.
[0812] The user enters their age and area of interest (e.g., "science," "technology") and clicks the submit button.
[0813] Input: Enter your age and interests
[0814] Output: Age and interest information entered
[0815] Step 4:
[0816] The server receives the age and interest information and stores it in a database.
[0817] The server stores the received age and interest information as a user profile in a database.
[0818] Input: Age and Interests
[0819] Output: Age and interest information stored in a database
[0820] Step 5:
[0821] The server collects data from the sources.
[0822] The server uses APIs to collect data from websites and video platforms in specified categories, such as "science" or "technology" videos from YouTube.
[0823] Input: Specified data category
[0824] Output: Raw data collected
[0825] Step 6:
[0826] The server temporarily stores the collected data.
[0827] The server stores the collected raw data in temporary storage or cache.
[0828] Input: Raw data collected
[0829] Output: Temporarily saved data
[0830] Step 7:
[0831] The server translates the data using a generative AI model.
[0832] The server generates a prompt and sends it to the generative AI model along with the temporarily saved data. The prompt includes instructions such as "Translate this video description into Japanese."
[0833] Input: Temporarily saved data, prompt text
[0834] Output: Data translated by the generative AI
[0835] Step 8:
[0836] The server optimizes the translation data based on the user's age and interests.
[0837] The server then feeds the translated data back into the generative AI model with specific prompts to tailor the content based on age and interests, such as "Please simplify it for a 10-year-old child."
[0838] Input: Translated data, age and interest information, optimization prompts
[0839] Output: Optimized data
[0840] Step 9:
[0841] The server sends the optimized data to the device.
[0842] The server transmits the optimized data to the user's terminal.
[0843] Input: Optimized data
[0844] Output: Data sent to the terminal
[0845] Step 10:
[0846] The device displays optimized data.
[0847] The device then displays the received data in an appropriate format for the user. For example, in the case of a video platform, translated subtitles and commentary are displayed.
[0848] Input: Optimization data sent from the server
[0849] Output: Displayed optimization data
[0850] Through the above steps, users can obtain information in real time in a format that is easy to understand in their native language, according to their age and interests.
[0851] (Application example 1)
[0852] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0853] Conventional information provision systems often lack the ability to provide content that is individually optimized based on the user's age and interests. This has led to problems, particularly in the field of learning, where information is not provided according to the user's level of understanding, resulting in reduced learning efficiency. Furthermore, few systems have real-time translation capabilities, making it difficult to overcome language barriers. The present invention aims to solve these problems by providing individually optimized learning content to users in real time.
[0854] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0855] In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for providing the optimized data to the user, and means for optimizing and providing learning content based on the information set by the user, thereby making it possible to provide individually optimized learning content in real time according to the user's age and interests.
[0856] The "means for the user to set his / her native language" is an interface or function that allows the user to select his / her native language and set that information.
[0857] The "means for users to input age and interest information" refers to an interface or function that allows users to input their own age and areas of interest.
[0858] "Means for obtaining data from information sources in cooperation with data collection means" refers to the function of collecting data from various information sources (e.g., websites, news feeds, etc.) and incorporating it into the system.
[0859] "Generative AI means for translating acquired data into the user's native language in real time" refers to generative AI technology for instantly translating collected data into the user's native language.
[0860] "Means for optimizing translated data based on the user's age and interests" refers to a function that adjusts translated data based on the user's age and interest information and processes it into an appropriate form.
[0861] The "means for providing optimized data to the user" refers to an interface or function for displaying or providing the data that has been subjected to optimization processing to the user.
[0862] "Means for optimizing and providing learning content based on information set by the user" refers to a function that adjusts and effectively provides learning content based on the native language, age, and interest information set by the user.
[0863] A system for realizing this invention includes a means for a user to set their native language, a means for a user to input age and interest information, a means for acquiring data from information sources in cooperation with a data collection means, a generation AI means for translating the acquired data into the user's native language in real time, a means for optimizing the translated data based on the user's age and interests, a means for providing the optimized data to the user, and a means for optimizing and providing learning content based on the information set by the user.
[0864] Initial System Setup
[0865] When a user launches the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information, and the user enters their age and interests. This information is also sent to the server and used for data optimization.
[0866] Data collection and translation
[0867] The server collects data by interacting with information sources (e.g., websites, news feeds, etc.). The collected data is temporarily stored on the server and then translated in real time using a generative AI method. The target language for translation is the native language selected by the user in the initial settings. The generative AI model used is the Helsinki-NLP / opus-mt-en-jap model from the transformers library.
[0868] Information Optimization
[0869] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[0870] Providing information
[0871] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[0872] Specific examples
[0873] For example, when a 10-year-old child whose native language is Japanese launches the application for the first time, sets their native language, and enters their age and interests, the server collects data from science-related sources. This collected data is translated into Japanese by the server using generative AI, and the content is simplified based on the user's age and interests before being provided to the user. This allows the user to receive science and technology information in Japanese that is easy to understand.
[0874] Prompt Sentence Examples
[0875] For example, use the following prompt to launch a generative AI model:
[0876] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0877] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[0878] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0879] Step 1:
[0880] When a user launches the application, a native language setting screen is displayed. Here, the user selects their native language. Input: User's native language. Output: User's native language setting information is sent from the device to the server. The server stores this information in a database and manages it as part of the user profile.
[0881] Step 2:
[0882] The device then displays a screen for the user to enter their age and interest information. The user enters their age and areas of interest. Input: User's age and interest information. Output: This information is sent from the device to the server. The server also stores this in a database for later data optimization.
[0883] Step 3:
[0884] The server collects data from sources (e.g. RSS feeds, websites) periodically or upon user request. Input: URL of source. Output: Collected data. The server temporarily stores the collected data.
[0885] Step 4:
[0886] The server translates the temporarily stored data in real time into the user's native language using the Helsinki-NLP / opus-mt-en-jap model from the transformers library. Input: Collected data, user's native language setting information. Output: Translated data. The server obtains this translated data using a generative AI method.
[0887] Step 5:
[0888] The server optimizes the translated data based on the user's age and interests, for example by replacing technical terms with simpler terms or highlighting information related to their area of interest. Input: Translated data, user's age and interests. Output: Optimized data. Specific actions include filtering and rephrasing the text.
[0889] Step 6:
[0890] The server sends the optimized data to the terminal. Input: Optimized data. Output: Data sent to the terminal. The terminal displays this data to the user. The user can understand this real-time optimized information in their native language.
[0891] These steps allow users to receive personalized, real-time information tailored to their needs. Specifically, the following prompts are used to trigger the generative AI model for translation and optimization:
[0892] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[0893] This processing flow is expected to significantly improve information accuracy and learning efficiency.
[0894] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0895] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time, and further recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[0896] Initial System Setup
[0897] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[0898] The server stores the received native language setting information in a database and manages it as part of the user profile.
[0899] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[0900] Emotion recognition by emotion engine
[0901] The device is equipped with an emotion engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera and microphone, and analyzes the user's emotions from their facial expressions and tone of voice.
[0902] Data collection and translation
[0903] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[0904] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[0905] Emotional optimization of information
[0906] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine. For example, if the user is excited, it will provide more stimulating content related to their interests, and conversely, if the user is relaxed, it will provide calming content.
[0907] Providing information
[0908] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[0909] Specific examples
[0910] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0911] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0912] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0913] The optimized data is sent from the server to the device and displayed to User A. User A can watch science and technology videos in easy-to-understand Japanese in a format that best suits their emotional state at the time.
[0914] As described above, this system can reduce information and education gaps by optimizing information according to the user's characteristics and emotional state and translating it into their native language in real time.
[0915] The processing flow will be explained below.
[0916] Step 1:
[0917] The user launches the application for the first time.
[0918] Step 2:
[0919] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[0920] Step 3:
[0921] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[0922] Step 4:
[0923] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[0924] Step 5:
[0925] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[0926] Step 6:
[0927] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[0928] Step 7:
[0929] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[0930] Step 8:
[0931] The device uses an emotion engine to recognize the user's emotions in real time, analyzing the user's facial expressions and tone of voice via the camera and microphone to determine their emotional state.
[0932] Step 9:
[0933] The server periodically sends API requests to collect data from sources (e.g., websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[0934] Step 10:
[0935] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[0936] Step 11:
[0937] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[0938] Step 12:
[0939] The server also optimizes the data based on the emotional state recognized by the emotion engine: for example, if the user is excited, it will adjust the content to be more stimulating, and if the user is relaxed, it will adjust the content to be calmer.
[0940] Step 13:
[0941] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[0942] Step 14:
[0943] The terminal displays the received data on the screen, allowing users to view optimized information in their native language in real time.
[0944] As a concrete example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[0945] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[0946] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[0947] This allows user A to watch easy-to-understand science and technology videos in Japanese in a way that best suits his or her emotional state at the time.
[0948] Example 2
[0949] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0950] While conventional information provision systems optimize data based on a user's native language, age, and interests, they face the problem of difficulty in providing information that takes into account the user's emotional state. Furthermore, because it is difficult to simultaneously collect and translate information in real time and optimize it based on emotions, it has not been possible to provide truly personalized information to users. The present invention aims to solve this problem.
[0951] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0952] In this invention, the server includes a means for the user to set their native language, a means for the user to input their age and interest information, a means for acquiring data from an information source in cooperation with the data collection means, a generation AI means for translating the acquired data into the user's native language in real time, an emotion recognition means for recognizing the user's emotions using a camera or microphone in the terminal, a means for optimizing the translated data based on the user's age, interests, and emotions, and a means for providing the optimized data to the user. This enables the provision of more highly personalized information based on the user's emotional state as well.
[0953] The "means for the user to set his native language" is a means for the user to specify the language he uses and register that information in the system.
[0954] "Means for users to input age and interest information" refers to means for users to input their own age and interests into the system.
[0955] "Means for acquiring data from information sources in cooperation with data collection means" refers to means for connecting to external information sources and collecting necessary data.
[0956] "Generative AI means" is an artificial intelligence technology that translates collected data into the user's native language in real time.
[0957] The "emotion recognition means" is a means for recognizing the user's emotional state by analyzing the user's facial expressions and tone of voice using a camera or microphone.
[0958] The "means for optimizing translated data based on the age, interests, and emotions of the user" refers to a means for providing translated data in an optimal form according to the age, interests, and emotions of the user.
[0959] The "means for providing optimized data to a user" refers to a means for displaying or distributing the optimized data to a user.
[0960] MODE FOR CARRYING OUT THE INVENTION
[0961] The present invention is a system that provides optimized information in real time, taking into consideration the user's native language, age, interest information, and emotional state. An embodiment of this system will now be described in detail.
[0962] The basic structure of the system consists of a server, a terminal, and a user. Below we explain how each element functions.
[0963] Initial Setup
[0964] When a user starts the application, a native language setting screen is displayed. The user selects their native language using this screen, and the selected information is sent to the server. The server stores the received native language setting information in a database, which is then managed as part of the user profile.
[0965] Next, the device presents the user with a screen for entering their age and interest information. The user enters their age and interest information here, and this information is also sent to the server. The server stores the age and interest information in a database and uses it to provide future information.
[0966] emotion recognition
[0967] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed using the device's built-in camera and microphone. Specifically, the camera analyzes the user's facial expressions and the microphone analyzes the user's tone of voice to recognize the user's emotional state. The results of this analysis are sent to a server in real time, where they are stored and managed.
[0968] Data collection and translation
[0969] The server has the function of collecting data from sources, such as news feeds, video platforms, websites, etc., periodically or upon user request. The collected data is temporarily stored on the server.
[0970] This collected data is then translated in real time into the user's native language using generative AI tools such as the Google Translate API, allowing the user to receive information in their native language of choice.
[0971] Emotional optimization of information
[0972] The translated data is optimized on the server based on the user's age and interests. The user's emotional state, obtained from an emotion recognition engine, is also taken into consideration. Specifically, if the user is excited, more stimulating content is selected, and if the user is relaxed, more calming content is selected. This optimization process is intended to provide information in the most appropriate form for the user.
[0973] Providing information
[0974] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive the information in their native language and in a format that best suits their emotional state at the time.
[0975] Specific examples
[0976] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A launches the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel. The collected video data is translated into Japanese using the Google Translate API. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. In addition, an emotion recognition engine analyzes User A's emotions; if User A is excited, more stimulating science experiment videos are provided, and if User A is calm, educational and calming videos are selected.
[0977] Prompt Sentence Examples
[0978] How can you specifically implement a system that takes into account a user's native language and interests, and optimizes information based on the user's emotional state?
[0979] As described above, this is a system in which the server, terminal, and user work together to provide information in real time that is tailored to the user's characteristics and emotions.
[0980] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0981] Step 1:
[0982] The user launches an application.
[0983] As a specific operation, the application displays a native language setting screen.
[0984] Input: User sets native language.
[0985] Output: Sends the configured native language information to the server.
[0986] The server stores the received native language setting information in a database and manages it as a user profile.
[0987] Step 2:
[0988] The device displays a screen where you can enter your age and interests.
[0989] Input: The user enters age and interest information.
[0990] Output: Send the entered age and interest information to the server.
[0991] The server stores the received age and interest information in a database for future use in providing information.
[0992] Step 3:
[0993] The device's built-in emotion recognition engine uses the camera and microphone to recognize the user's emotions.
[0994] Specifically, the camera analyzes the user's facial expression and the microphone analyzes the user's tone of voice.
[0995] Input: User facial and voice data obtained from camera and microphone.
[0996] Output: The emotion recognition engine analyzes the user's emotional state and sends it to the server.
[0997] The server stores this affective information for use in subsequent data optimization.
[0998] Step 4:
[0999] The server collects data from pre-specified sources.
[1000] Specific sources of information include news feeds, video platforms, websites, etc.
[1001] Input: Various data collected from sources.
[1002] Output: Temporarily stores collected data.
[1003] The server then translates the data in real time using generative AI methods such as the Google Translate API.
[1004] Step 5:
[1005] The server optimizes the translated data based on the user's age, interests and emotional state.
[1006] Input: Translated data, user age, interest, and emotion information.
[1007] Output: Generates optimized data.
[1008] The optimization process involves filtering and categorizing data and emphasizing certain elements, for example, if a user is excited, more stimulating content will be selected, and if they are relaxed, more calming content will be selected.
[1009] Step 6:
[1010] The server sends the optimized data to the device.
[1011] Input: Optimized data.
[1012] Output: Sends data to a terminal and displays it to the user.
[1013] The terminal displays the received optimization data in a form that can be used visually or audibly by the user.
[1014] Through the above steps, the user can receive information optimized for their current emotional state and personal attributes in real time.
[1015] (Application example 2)
[1016] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1017] Currently, many systems that provide information to users do not adequately optimize it based on the user's language, age, interests, or even real-time emotions. This can result in information not reaching the user appropriately, leading to reduced receptivity and comprehension. In particular, providing information that does not match the user's emotional state can hinder the user's experience. Therefore, there is a need for a system that optimizes information based on the user's characteristics and real-time emotions and delivers it at the appropriate time.
[1018] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for recognizing the user's emotions and optimizing information, and means for displaying the optimized data to the user. This makes it possible to optimize information based on the user's characteristics and real-time emotional state and provide information in the user's native language at an appropriate time.
[1019] The "means for the user to set his native language" is an interface that allows the user to select the language he or she uses, and to input and save that information.
[1020] The "means for users to input age and interest information" refers to an interface that allows users to input their age and areas of interest and provide that information.
[1021] "Means for acquiring data from information sources in cooperation with data collection means" refers to a mechanism for collecting data from various information sources on a network, and a means for incorporating the data into the system.
[1022] "Generative AI means for translating acquired data into the user's native language in real time" refers to a mechanism that includes a generative AI model for instantly translating collected data into the user's native language.
[1023] "Means for optimizing translated data based on the user's age and interests" refers to a system that processes and organizes translated data in an appropriate form in accordance with the user's age and interest information.
[1024] "Means for recognizing user emotions and optimizing information" refers to a mechanism that analyzes the user's emotions through a camera or microphone and optimizes information in accordance with those emotions.
[1025] The "means for displaying optimized data to the user" refers to an interface for visually or audibly conveying optimized information to the user.
[1026] This invention is a system that translates data acquired from various information sources in real time by allowing a user to input information such as their native language, age, and interests, and also recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[1027] Initial Setup
[1028] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information. The user enters their age and interests and sends the information to the server. This information is stored on the server and used for data optimization.
[1029] emotion recognition
[1030] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera (e.g., Logitech C920) or microphone, and analyzes the user's emotions from their facial expressions and tone of voice. This emotion data is also sent to the server.
[1031] Data collection and translation
[1032] The server collects data by interacting with information sources (e.g., websites, video platforms, news feeds, etc.). This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server and then translated in real time using generative AI means. The target language for translation is the native language selected by the user in the initial settings. This process enables information to be transmitted across language barriers.
[1033] Data Optimization
[1034] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine: if the user is excited, for example, it will provide more stimulating content related to their interests, and conversely, if they are relaxed, it will provide calming content.
[1035] Providing information
[1036] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[1037] Specific examples
[1038] For example, if a user sets "music" and "sports" as their interests and is recognized as having an "excited" emotion, the system will provide the user with the latest sports news and music ranking videos. The system optimizes information based on the user's characteristics and emotional state, and translates it into their native language in real time, thereby reducing information and education gaps.
[1039] Prompt Sentence Examples
[1040] "When users are emotionally excited, generate exciting content related to their interests. Also, translate that content into Japanese. When they are excited, provide content that emphasizes more active elements."
[1041] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1042] Step 1:
[1043] When a user starts an application, the device first displays the native language setting screen. The user selects their native language and enters the information. This input data is sent to the server, which then stores the received native language setting information in a database. Input: Native language information entered by the user. Output: Native language information stored in the database.
[1044] Step 2:
[1045] Next, the terminal displays a screen for the user to enter their age and interest information. The user enters their age and multiple interests and sends this information to the server. The server stores this information in a database. Input: Age and interest information entered by the user. Output: Age and interest information stored in the database.
[1046] Step 3:
[1047] The device uses a camera and microphone to analyze the user's emotions in real time. The emotion recognition engine generates emotion data from the user's facial expressions and tone of voice, and this emotion data is sent to the server. Input: User's facial expressions and voice data captured by the camera and microphone. Output: Emotion data sent to the server.
[1048] Step 4:
[1049] The server periodically collects data from sources (e.g. websites or video platforms). This collected data is temporarily stored on the server. Input: Raw data obtained from sources. Output: Raw data temporarily stored on the server.
[1050] Step 5:
[1051] The server uses generative AI means (e.g., generative AI models) to translate this collected data into the user's native language in real time, thus overcoming language barriers. Input: Temporarily stored raw data. Output: Real-time translated data.
[1052] Step 6:
[1053] The server processes the translated data to optimize it based on the user's age and interest information. It also takes into account the emotional state recognized by the emotion engine to generate optimal information. For example, if the user is excited, the content will be optimized to be more stimulating. Input: Translated data, age and interest information, emotional data. Output: Information optimized for the user.
[1054] Step 7:
[1055] The optimized information is sent from the server to the terminal. The terminal receives this information and displays it to the user in an appropriate interface. This allows the user to receive optimized information in real time. Input: Optimized information. Output: Optimized information displayed on the terminal.
[1056] As described above, by going through each step, a system is realized that allows users to receive information optimized for them in their native language in real time.
[1057] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1058] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1059] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1060] [Fourth embodiment]
[1061] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1062] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1063] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1064] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1065] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1066] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1067] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1068] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1069] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1070] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1071] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1072] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1073] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1074] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time and provides it in an optimized format. Specific embodiments of this system are described below.
[1075] Initial System Setup
[1076] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[1077] The server stores the received native language setting information in a database and manages it as part of the user profile.
[1078] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[1079] Data collection and translation
[1080] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[1081] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[1082] Information Optimization
[1083] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[1084] Providing information
[1085] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[1086] Specific examples
[1087] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[1088] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[1089] The optimized data is sent from the server to the terminal and displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[1090] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[1091] The processing flow will be explained below.
[1092] Step 1:
[1093] The user launches the application for the first time.
[1094] Step 2:
[1095] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[1096] Step 3:
[1097] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[1098] Step 4:
[1099] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[1100] Step 5:
[1101] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[1102] Step 6:
[1103] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[1104] Step 7:
[1105] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[1106] Step 8:
[1107] The server periodically sends API requests to collect data from sources (websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[1108] Step 9:
[1109] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[1110] Step 10:
[1111] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[1112] Step 11:
[1113] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[1114] Step 12:
[1115] The terminal displays the received data on the screen, allowing users to view information optimized in their native language in real time.
[1116] Example 1
[1117] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1118] In today's globalized society, it is becoming increasingly important to access and understand information across language and cultural barriers. A particular challenge is the difficulty for users with different native languages to obtain the most appropriate information in real time, tailored to their age and interests. Furthermore, the lack of integration between the information gathering, translation, and optimization processes makes it difficult to provide users with the information they want quickly and appropriately.
[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1120] In this invention, the server includes means for a user to set their native language, means for a user to input age and interest information, means for collecting data from information sources, means for temporarily storing the collected data, means for translating the acquired data into the user's native language in real time using a generative AI model, means for generating prompt sentences and sending them to the AI model, means for optimizing the translated data based on the user's age and interest information, and means for providing the optimized data to the user, thereby enabling users to obtain optimal information according to their age and interests in real time, regardless of language or cultural barriers.
[1121] The "means for a user to set his / her native language" refers to an interface that allows a user to select his / her native language and input and transmit that information to the system.
[1122] "Means for users to input age and interest information" refers to an interface through which users input information about their age and areas of interest and transmit that information to the system.
[1123] "Means of collecting data from sources" refers to the functionality for obtaining data from external sources such as websites, video platforms, news feeds, etc.
[1124] "Means for temporarily storing collected data" refers to the storage or cache function for temporarily storing collected data.
[1125] "Generative AI model" refers to a generative AI (e.g., OpenAI's GPT-3) used to perform tasks such as natural language processing.
[1126] "Means of translating acquired data into the user's native language in real time" refers to a function that uses a generative AI model to instantly translate acquired data into the user's native language.
[1127] "Means for generating prompt sentences and sending them to an AI model" refers to the function for creating prompt sentences that provide specific translation or simplification instructions to the generative AI model and sending them to the model.
[1128] "Means for optimizing translated data based on user age and interest information" refers to functionality for adjusting and optimizing translated data to make it easier to understand based on the user's age and interests.
[1129] "Means for providing optimized data to a user" refers to an interface or communication means for delivering optimized data to a user and displaying that information.
[1130] This invention is a system that allows users to set their native language and input their age and interest information, and then translates data obtained from various sources in real time and provides it in an optimized format. Specific embodiments of the invention are described below. This system consists of three main elements: a server, a terminal, and a user.
[1131] Initial System Setup
[1132] When a user starts the application, the native language setting interface is displayed first. The user selects their native language on this screen and clicks the "Confirm" button to send the setting information to the server. The server saves the received native language setting information in a database and manages it as a user profile.
[1133] The device then presents the user with a screen to input their age and interest information. The user inputs their age, selects a specific interest area (e.g., "science" or "technology"), and clicks the "Submit" button to send the information to the server. The server stores this information in a database for later data optimization.
[1134] Data collection and translation
[1135] The server connects with sources (e.g. websites, video platforms, news feeds, etc.) and collects the required data. This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server.
[1136] The server then uses a generative AI model (e.g., OpenAI's GPT-3) to translate the collected data into the user's native language in real time. Specifically, the server generates a prompt and sends it to the AI model along with the collected data. Examples of prompts include "Translate this video description into Japanese." or "Please translate a YouTube video about science and technology in a simplified form for a 10-year-old Japanese speaker."
[1137] Information Optimization
[1138] The server optimizes the translated data based on the user's age and interests: for example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and the elements of interest will be highlighted.
[1139] Providing information
[1140] The optimized information is sent from the server to the device. The device receives this information and displays it to the user. The user can easily understand the information that has been translated and optimized into their native language in real time.
[1141] Specific examples
[1142] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[1143] The collected video data is translated into Japanese by the server using generative AI and executed using example prompt sentences. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. The optimized data is sent from the server to the device and finally displayed to User A. This allows User A to watch science and technology videos in Japanese that are easy to understand.
[1144] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[1145] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1146] Step 1:
[1147] The user launches an application.
[1148] When a user starts the application, a native language setting interface is displayed. The user selects their native language on this screen and clicks the OK button.
[1149] Input: Select your native language
[1150] Output: Selected native language information is generated
[1151] Step 2:
[1152] The server receives the native language setting information and stores it in a database.
[1153] The server stores the received native language setting information in a database as a user profile.
[1154] Input: Native language setting information
[1155] Output: Native language information stored in a database
[1156] Step 3:
[1157] The device displays a screen where you can enter your age and interests.
[1158] The user enters their age and area of interest (e.g., "science," "technology") and clicks the submit button.
[1159] Input: Enter your age and interests
[1160] Output: Age and interest information entered
[1161] Step 4:
[1162] The server receives the age and interest information and stores it in a database.
[1163] The server stores the received age and interest information as a user profile in a database.
[1164] Input: Age and Interests
[1165] Output: Age and interest information stored in a database
[1166] Step 5:
[1167] The server collects data from the sources.
[1168] The server uses APIs to collect data from websites and video platforms in specified categories, such as "science" or "technology" videos from YouTube.
[1169] Input: Specified data category
[1170] Output: Raw data collected
[1171] Step 6:
[1172] The server temporarily stores the collected data.
[1173] The server stores the collected raw data in temporary storage or cache.
[1174] Input: Raw data collected
[1175] Output: Temporarily saved data
[1176] Step 7:
[1177] The server translates the data using a generative AI model.
[1178] The server generates a prompt and sends it to the generative AI model along with the temporarily saved data. The prompt includes instructions such as "Translate this video description into Japanese."
[1179] Input: Temporarily saved data, prompt text
[1180] Output: Data translated by the generative AI
[1181] Step 8:
[1182] The server optimizes the translation data based on the user's age and interests.
[1183] The server then feeds the translated data back into the generative AI model with specific prompts to tailor the content based on age and interests, such as "Please simplify it for a 10-year-old child."
[1184] Input: Translated data, age and interest information, optimization prompts
[1185] Output: Optimized data
[1186] Step 9:
[1187] The server sends the optimized data to the device.
[1188] The server transmits the optimized data to the user's terminal.
[1189] Input: Optimized data
[1190] Output: Data sent to the terminal
[1191] Step 10:
[1192] The device displays optimized data.
[1193] The device then displays the received data in an appropriate format for the user. For example, in the case of a video platform, translated subtitles and commentary are displayed.
[1194] Input: Optimization data sent from the server
[1195] Output: Displayed optimization data
[1196] Through the above steps, users can obtain information in real time in a format that is easy to understand in their native language, according to their age and interests.
[1197] (Application example 1)
[1198] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1199] Conventional information provision systems often lack the ability to provide content that is individually optimized based on the user's age and interests. This has led to problems, particularly in the field of learning, where information is not provided according to the user's level of understanding, resulting in reduced learning efficiency. Furthermore, few systems have real-time translation capabilities, making it difficult to overcome language barriers. The present invention aims to solve these problems by providing individually optimized learning content to users in real time.
[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1201] In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for providing the optimized data to the user, and means for optimizing and providing learning content based on the information set by the user, thereby making it possible to provide individually optimized learning content in real time according to the user's age and interests.
[1202] The "means for the user to set his / her native language" is an interface or function that allows the user to select his / her native language and set that information.
[1203] The "means for users to input age and interest information" refers to an interface or function that allows users to input their own age and areas of interest.
[1204] "Means for obtaining data from information sources in cooperation with data collection means" refers to the function of collecting data from various information sources (e.g., websites, news feeds, etc.) and incorporating it into the system.
[1205] "Generative AI means for translating acquired data into the user's native language in real time" refers to generative AI technology for instantly translating collected data into the user's native language.
[1206] "Means for optimizing translated data based on the user's age and interests" refers to a function that adjusts translated data based on the user's age and interest information and processes it into an appropriate form.
[1207] The "means for providing optimized data to the user" refers to an interface or function for displaying or providing the data that has been subjected to optimization processing to the user.
[1208] "Means for optimizing and providing learning content based on information set by the user" refers to a function that adjusts and effectively provides learning content based on the native language, age, and interest information set by the user.
[1209] A system for realizing this invention includes a means for a user to set their native language, a means for a user to input age and interest information, a means for acquiring data from information sources in cooperation with a data collection means, a generation AI means for translating the acquired data into the user's native language in real time, a means for optimizing the translated data based on the user's age and interests, a means for providing the optimized data to the user, and a means for optimizing and providing learning content based on the information set by the user.
[1210] Initial System Setup
[1211] When a user launches the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information, and the user enters their age and interests. This information is also sent to the server and used for data optimization.
[1212] Data collection and translation
[1213] The server collects data by interacting with information sources (e.g., websites, news feeds, etc.). The collected data is temporarily stored on the server and then translated in real time using a generative AI method. The target language for translation is the native language selected by the user in the initial settings. The generative AI model used is the Helsinki-NLP / opus-mt-en-jap model from the transformers library.
[1214] Information Optimization
[1215] The server optimizes the translated data based on the user's age and interests. For example, if the user is a 10-year-old child and is interested in "science" and "technology," the difficulty level will be lowered to make the content easier to understand, and information related to the subject of interest will be emphasized.
[1216] Providing information
[1217] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive information in real time in their native language in an easy-to-understand format.
[1218] Specific examples
[1219] For example, when a 10-year-old child whose native language is Japanese launches the application for the first time, sets their native language, and enters their age and interests, the server collects data from science-related sources. This collected data is translated into Japanese by the server using generative AI, and the content is simplified based on the user's age and interests before being provided to the user. This allows the user to receive science and technology information in Japanese that is easy to understand.
[1220] Prompt Sentence Examples
[1221] For example, use the following prompt to launch a generative AI model:
[1222] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[1223] As described above, this system can reduce information and education disparities by optimizing information according to the user's characteristics and translating it into the user's native language in real time.
[1224] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1225] Step 1:
[1226] When a user launches the application, a native language setting screen is displayed. Here, the user selects their native language. Input: User's native language. Output: User's native language setting information is sent from the device to the server. The server stores this information in a database and manages it as part of the user profile.
[1227] Step 2:
[1228] The device then displays a screen for the user to enter their age and interest information. The user enters their age and areas of interest. Input: User's age and interest information. Output: This information is sent from the device to the server. The server also stores this in a database for later data optimization.
[1229] Step 3:
[1230] The server collects data from sources (e.g. RSS feeds, websites) periodically or upon user request. Input: URL of source. Output: Collected data. The server temporarily stores the collected data.
[1231] Step 4:
[1232] The server translates the temporarily stored data in real time into the user's native language using the Helsinki-NLP / opus-mt-en-jap model from the transformers library. Input: Collected data, user's native language setting information. Output: Translated data. The server obtains this translated data using a generative AI method.
[1233] Step 5:
[1234] The server optimizes the translated data based on the user's age and interests, for example by replacing technical terms with simpler terms or highlighting information related to their area of interest. Input: Translated data, user's age and interests. Output: Optimized data. Specific actions include filtering and rephrasing the text.
[1235] Step 6:
[1236] The server sends the optimized data to the terminal. Input: Optimized data. Output: Data sent to the terminal. The terminal displays this data to the user. The user can understand this real-time optimized information in their native language.
[1237] These steps allow users to receive personalized, real-time information tailored to their needs. Specifically, the following prompts are used to trigger the generative AI model for translation and optimization:
[1238] "Translate the following English text to Japanese considering the user is a 10-year-old child interested in science and technology: [insert English text here]"
[1239] This processing flow is expected to significantly improve information accuracy and learning efficiency.
[1240] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1241] This invention is a system that allows a user to set their native language and input their age and interests, and then translates data acquired from various information sources in real time, and further recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[1242] Initial System Setup
[1243] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on that screen and sends the setting information to the server.
[1244] The server stores the received native language setting information in a database and manages it as part of the user profile.
[1245] Next, the device presents the user with a screen to input their age and interest information. The user inputs their age and interests and sends the information to the server, where it is stored and used for data optimization.
[1246] Emotion recognition by emotion engine
[1247] The device is equipped with an emotion engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera and microphone, and analyzes the user's emotions from their facial expressions and tone of voice.
[1248] Data collection and translation
[1249] The server collects data by interacting with sources (e.g., websites, video platforms, news feeds, etc.) This data collection occurs periodically or upon user request.
[1250] The collected data is temporarily stored on a server and then translated in real time using a generative AI method. The language to be translated is the native language selected by the user in the initial settings. This process makes it possible to communicate information across language barriers.
[1251] Emotional optimization of information
[1252] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine. For example, if the user is excited, it will provide more stimulating content related to their interests, and conversely, if the user is relaxed, it will provide calming content.
[1253] Providing information
[1254] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[1255] Specific examples
[1256] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[1257] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[1258] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[1259] The optimized data is sent from the server to the device and displayed to User A. User A can watch science and technology videos in easy-to-understand Japanese in a format that best suits their emotional state at the time.
[1260] As described above, this system can reduce information and education gaps by optimizing information according to the user's characteristics and emotional state and translating it into their native language in real time.
[1261] The processing flow will be explained below.
[1262] Step 1:
[1263] The user launches the application for the first time.
[1264] Step 2:
[1265] The terminal displays a native language setting screen, where the user selects their native language and presses a button to confirm the setting.
[1266] Step 3:
[1267] The terminal transmits the selected native language to the server. The transmitted data includes the user ID and the selected language code.
[1268] Step 4:
[1269] The server stores the received native language setting information in a database, whereby the native language information is managed as part of the user profile.
[1270] Step 5:
[1271] The terminal then displays a screen for inputting age and interest information. The user inputs the age and interests and presses the send button.
[1272] Step 6:
[1273] The terminal transmits the input age and interest information to the server. The transmitted data includes the user ID, age, and interest information.
[1274] Step 7:
[1275] The server stores the received age and interest information in a database, which is used for subsequent data optimization.
[1276] Step 8:
[1277] The device uses an emotion engine to recognize the user's emotions in real time, analyzing the user's facial expressions and tone of voice via the camera and microphone to determine their emotional state.
[1278] Step 9:
[1279] The server periodically sends API requests to collect data from sources (e.g., websites, video platforms, news feeds, etc.) or retrieves the latest data upon user request.
[1280] Step 10:
[1281] The server temporarily stores the collected data, and at the same time, it uses generative AI to translate this data in real time into the user's native language.
[1282] Step 11:
[1283] The server optimizes the translated data based on the user's age and interests, for example simplifying the content for a 10-year-old and highlighting information related to their areas of interest.
[1284] Step 12:
[1285] The server also optimizes the data based on the emotional state recognized by the emotion engine: for example, if the user is excited, it will adjust the content to be more stimulating, and if the user is relaxed, it will adjust the content to be calmer.
[1286] Step 13:
[1287] The server transmits the optimized data to the terminal, which includes the translation and optimized information corresponding to the user ID.
[1288] Step 14:
[1289] The terminal displays the received data on the screen, allowing users to view optimized information in their native language in real time.
[1290] As a concrete example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." After User A starts the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel.
[1291] The collected video data is translated into Japanese by the server using generative AI. After translation, the content is simplified based on User A's age and interests, and relevant scientific and technological elements are emphasized.
[1292] Furthermore, the emotion engine analyzes the emotions of user A, and if user A is excited, it will provide more stimulating videos of science experiments. On the other hand, if user A is calm, it will select videos with educational and calming content.
[1293] This allows user A to watch easy-to-understand science and technology videos in Japanese in a way that best suits his or her emotional state at the time.
[1294] Example 2
[1295] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1296] While conventional information provision systems optimize data based on a user's native language, age, and interests, they face the problem of difficulty in providing information that takes into account the user's emotional state. Furthermore, because it is difficult to simultaneously collect and translate information in real time and optimize it based on emotions, it has not been possible to provide truly personalized information to users. The present invention aims to solve this problem.
[1297] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1298] In this invention, the server includes a means for the user to set their native language, a means for the user to input their age and interest information, a means for acquiring data from an information source in cooperation with the data collection means, a generation AI means for translating the acquired data into the user's native language in real time, an emotion recognition means for recognizing the user's emotions using a camera or microphone in the terminal, a means for optimizing the translated data based on the user's age, interests, and emotions, and a means for providing the optimized data to the user. This enables the provision of more highly personalized information based on the user's emotional state as well.
[1299] The "means for the user to set his native language" is a means for the user to specify the language he uses and register that information in the system.
[1300] "Means for users to input age and interest information" refers to means for users to input their own age and interests into the system.
[1301] "Means for acquiring data from information sources in cooperation with data collection means" refers to means for connecting to external information sources and collecting necessary data.
[1302] "Generative AI means" is an artificial intelligence technology that translates collected data into the user's native language in real time.
[1303] The "emotion recognition means" is a means for recognizing the user's emotional state by analyzing the user's facial expressions and tone of voice using a camera or microphone.
[1304] The "means for optimizing translated data based on the age, interests, and emotions of the user" refers to a means for providing translated data in an optimal form according to the age, interests, and emotions of the user.
[1305] The "means for providing optimized data to a user" refers to a means for displaying or distributing the optimized data to a user.
[1306] MODE FOR CARRYING OUT THE INVENTION
[1307] The present invention is a system that provides optimized information in real time, taking into consideration the user's native language, age, interest information, and emotional state. An embodiment of this system will now be described in detail.
[1308] The basic structure of the system consists of a server, a terminal, and a user. Below we explain how each element functions.
[1309] Initial Setup
[1310] When a user starts the application, a native language setting screen is displayed. The user selects their native language using this screen, and the selected information is sent to the server. The server stores the received native language setting information in a database, which is then managed as part of the user profile.
[1311] Next, the device presents the user with a screen for entering their age and interest information. The user enters their age and interest information here, and this information is also sent to the server. The server stores the age and interest information in a database and uses it to provide future information.
[1312] emotion recognition
[1313] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed using the device's built-in camera and microphone. Specifically, the camera analyzes the user's facial expressions and the microphone analyzes the user's tone of voice to recognize the user's emotional state. The results of this analysis are sent to a server in real time, where they are stored and managed.
[1314] Data collection and translation
[1315] The server has the function of collecting data from sources, such as news feeds, video platforms, websites, etc., periodically or upon user request. The collected data is temporarily stored on the server.
[1316] This collected data is then translated in real time into the user's native language using generative AI tools such as the Google Translate API, allowing the user to receive information in their native language of choice.
[1317] Emotional optimization of information
[1318] The translated data is optimized on the server based on the user's age and interests. The user's emotional state, obtained from an emotion recognition engine, is also taken into consideration. Specifically, if the user is excited, more stimulating content is selected, and if the user is relaxed, more calming content is selected. This optimization process is intended to provide information in the most appropriate form for the user.
[1319] Providing information
[1320] The optimized information is sent from the server to the device, which receives it and displays it to the user, allowing the user to receive the information in their native language and in a format that best suits their emotional state at the time.
[1321] Specific examples
[1322] For example, User A is a 10-year-old child whose native language is Japanese and whose interests are "science" and "technology." When User A launches the application for the first time, sets his native language, and enters his age and interest information, the server collects video data from YouTube's science channel. The collected video data is translated into Japanese using the Google Translate API. After translation, the content is simplified based on User A's age and interests, and relevant science and technology elements are emphasized. In addition, an emotion recognition engine analyzes User A's emotions; if User A is excited, more stimulating science experiment videos are provided, and if User A is calm, educational and calming videos are selected.
[1323] Prompt Sentence Examples
[1324] How can you specifically implement a system that takes into account a user's native language and interests, and optimizes information based on the user's emotional state?
[1325] As described above, this is a system in which the server, terminal, and user work together to provide information in real time that is tailored to the user's characteristics and emotions.
[1326] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1327] Step 1:
[1328] The user launches an application.
[1329] As a specific operation, the application displays a native language setting screen.
[1330] Input: User sets native language.
[1331] Output: Sends the configured native language information to the server.
[1332] The server stores the received native language setting information in a database and manages it as a user profile.
[1333] Step 2:
[1334] The device displays a screen where you can enter your age and interests.
[1335] Input: The user enters age and interest information.
[1336] Output: Send the entered age and interest information to the server.
[1337] The server stores the received age and interest information in a database for future use in providing information.
[1338] Step 3:
[1339] The device's built-in emotion recognition engine uses the camera and microphone to recognize the user's emotions.
[1340] Specifically, the camera analyzes the user's facial expression and the microphone analyzes the user's tone of voice.
[1341] Input: User facial and voice data obtained from camera and microphone.
[1342] Output: The emotion recognition engine analyzes the user's emotional state and sends it to the server.
[1343] The server stores this affective information for use in subsequent data optimization.
[1344] Step 4:
[1345] The server collects data from pre-specified sources.
[1346] Specific sources of information include news feeds, video platforms, websites, etc.
[1347] Input: Various data collected from sources.
[1348] Output: Temporarily stores collected data.
[1349] The server then translates the data in real time using generative AI methods such as the Google Translate API.
[1350] Step 5:
[1351] The server optimizes the translated data based on the user's age, interests and emotional state.
[1352] Input: Translated data, user age, interest, and emotion information.
[1353] Output: Generates optimized data.
[1354] The optimization process involves filtering and categorizing data and emphasizing certain elements, for example, if a user is excited, more stimulating content will be selected, and if they are relaxed, more calming content will be selected.
[1355] Step 6:
[1356] The server sends the optimized data to the device.
[1357] Input: Optimized data.
[1358] Output: Sends data to a terminal and displays it to the user.
[1359] The terminal displays the received optimization data in a form that can be used visually or audibly by the user.
[1360] Through the above steps, the user can receive information optimized for their current emotional state and personal attributes in real time.
[1361] (Application example 2)
[1362] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1363] Currently, many systems that provide information to users do not adequately optimize it based on the user's language, age, interests, or even real-time emotions. This can result in information not reaching the user appropriately, leading to reduced receptivity and comprehension. In particular, providing information that does not match the user's emotional state can hinder the user's experience. Therefore, there is a need for a system that optimizes information based on the user's characteristics and real-time emotions and delivers it at the appropriate time.
[1364] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to set their native language, means for the user to input age and interest information, means for acquiring data from information sources in cooperation with the data collection means, generation AI means for translating the acquired data into the user's native language in real time, means for optimizing the translated data based on the user's age and interests, means for recognizing the user's emotions and optimizing information, and means for displaying the optimized data to the user. This makes it possible to optimize information based on the user's characteristics and real-time emotional state and provide information in the user's native language at an appropriate time.
[1365] The "means for the user to set his native language" is an interface that allows the user to select the language he or she uses, and to input and save that information.
[1366] The "means for users to input age and interest information" refers to an interface that allows users to input their age and areas of interest and provide that information.
[1367] "Means for acquiring data from information sources in cooperation with data collection means" refers to a mechanism for collecting data from various information sources on a network, and a means for incorporating the data into the system.
[1368] "Generative AI means for translating acquired data into the user's native language in real time" refers to a mechanism that includes a generative AI model for instantly translating collected data into the user's native language.
[1369] "Means for optimizing translated data based on the user's age and interests" refers to a system that processes and organizes translated data in an appropriate form in accordance with the user's age and interest information.
[1370] "Means for recognizing user emotions and optimizing information" refers to a mechanism that analyzes the user's emotions through a camera or microphone and optimizes information in accordance with those emotions.
[1371] The "means for displaying optimized data to the user" refers to an interface for visually or audibly conveying optimized information to the user.
[1372] This invention is a system that translates data acquired from various information sources in real time by allowing a user to input information such as their native language, age, and interests, and also recognizes the user's emotions to provide optimized information. Specific embodiments of this system are described below.
[1373] Initial Setup
[1374] When a user starts the application, a native language setting screen is displayed first. The user selects their native language on this screen and sends the setting information to the server. The server stores the received native language setting information in a database and manages it as part of the user profile. Next, the device presents the user with a screen for entering age and interest information. The user enters their age and interests and sends the information to the server. This information is stored on the server and used for data optimization.
[1375] emotion recognition
[1376] The device is equipped with an emotion recognition engine that recognizes the user's emotions in real time. This emotion recognition is performed through a camera (e.g., Logitech C920) or microphone, and analyzes the user's emotions from their facial expressions and tone of voice. This emotion data is also sent to the server.
[1377] Data collection and translation
[1378] The server collects data by interacting with information sources (e.g., websites, video platforms, news feeds, etc.). This data collection occurs periodically or upon user request. The collected data is temporarily stored on the server and then translated in real time using generative AI means. The target language for translation is the native language selected by the user in the initial settings. This process enables information to be transmitted across language barriers.
[1379] Data Optimization
[1380] The server optimizes the translated data based on the user's age and interests, and also reflects the emotional state recognized by the emotion engine: if the user is excited, for example, it will provide more stimulating content related to their interests, and conversely, if they are relaxed, it will provide calming content.
[1381] Providing information
[1382] The optimized information is sent from the server to the terminal, which receives the information and displays it to the user, allowing the user to receive the optimized information in their native language in real time.
[1383] Specific examples
[1384] For example, if a user sets "music" and "sports" as their interests and is recognized as having an "excited" emotion, the system will provide the user with the latest sports news and music ranking videos. The system optimizes information based on the user's characteristics and emotional state, and translates it into their native language in real time, thereby reducing information and education gaps.
[1385] Prompt Sentence Examples
[1386] "When users are emotionally excited, generate exciting content related to their interests. Also, translate that content into Japanese. When they are excited, provide content that emphasizes more active elements."
[1387] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1388] Step 1:
[1389] When a user starts an application, the device first displays the native language setting screen. The user selects their native language and enters the information. This input data is sent to the server, which then stores the received native language setting information in a database. Input: Native language information entered by the user. Output: Native language information stored in the database.
[1390] Step 2:
[1391] Next, the terminal displays a screen for the user to enter their age and interest information. The user enters their age and multiple interests and sends this information to the server. The server stores this information in a database. Input: Age and interest information entered by the user. Output: Age and interest information stored in the database.
[1392] Step 3:
[1393] The device uses a camera and microphone to analyze the user's emotions in real time. The emotion recognition engine generates emotion data from the user's facial expressions and tone of voice, and this emotion data is sent to the server. Input: User's facial expressions and voice data captured by the camera and microphone. Output: Emotion data sent to the server.
[1394] Step 4:
[1395] The server periodically collects data from sources (e.g. websites or video platforms). This collected data is temporarily stored on the server. Input: Raw data obtained from sources. Output: Raw data temporarily stored on the server.
[1396] Step 5:
[1397] The server uses generative AI means (e.g., generative AI models) to translate this collected data into the user's native language in real time, thus overcoming language barriers. Input: Temporarily stored raw data. Output: Real-time translated data.
[1398] Step 6:
[1399] The server processes the translated data to optimize it based on the user's age and interest information. It also takes into account the emotional state recognized by the emotion engine to generate optimal information. For example, if the user is excited, the content will be optimized to be more stimulating. Input: Translated data, age and interest information, emotional data. Output: Information optimized for the user.
[1400] Step 7:
[1401] The optimized information is sent from the server to the terminal. The terminal receives this information and displays it to the user in an appropriate interface. This allows the user to receive optimized information in real time. Input: Optimized information. Output: Optimized information displayed on the terminal.
[1402] As described above, by going through each step, a system is realized that allows users to receive information optimized for them in their native language in real time.
[1403] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1404] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1405] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1406] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1407] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1408] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1409] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1410] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1411] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1412] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1413] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1414] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1415] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1416] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1417] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1418] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1419] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1420] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1421] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1422] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1423] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1424] The following is further disclosed regarding the above embodiment.
[1425] (Claim 1)
[1426] a means for a user to set their native language;
[1427] a means for a user to input age and interest information;
[1428] means of obtaining data from sources in conjunction with data collection instruments;
[1429] A generative AI means for translating the acquired data into the user's native language in real time;
[1430] means for optimizing the translated data based on the age and interests of the user;
[1431] means for providing optimized data to a user;
[1432] A system including:
[1433] (Claim 2)
[1434] 10. The system of claim 1, further comprising means for storing the acquired data.
[1435] (Claim 3)
[1436] 10. The system of claim 1, further comprising means for storing the user's native language setting in a database.
[1437] "Example 1"
[1438] (Claim 1)
[1439] a means for a user to set their native language;
[1440] a means for a user to input age and interest information;
[1441] a means of collecting data from sources;
[1442] a means for temporarily storing the collected data;
[1443] A means for translating the acquired data into the user's native language in real time using a generative AI model; and
[1444] A means for generating and sending prompt sentences to the AI model;
[1445] means for optimizing the translated data based on the user's age and interest information;
[1446] means for providing optimized data to a user;
[1447] A system including:
[1448] (Claim 2)
[1449] 10. The system of claim 1, further comprising means for storing the translated and optimized data.
[1450] (Claim 3)
[1451] 10. The system of claim 1, further comprising means for storing the user's native language preference, age, and interest information in the database.
[1452] "Application Example 1"
[1453] (Claim 1)
[1454] a means for a user to set their native language;
[1455] a means for a user to input age and interest information;
[1456] means of obtaining data from sources in conjunction with data collection instruments;
[1457] A generative AI means for translating the acquired data into the user's native language in real time;
[1458] means for optimizing the translated data based on the age and interests of the user;
[1459] means for providing optimized data to a user;
[1460] A means for optimizing and providing learning content based on information set by the user;
[1461] A system including:
[1462] (Claim 2)
[1463] 10. The system of claim 1, further comprising means for storing the acquired data.
[1464] (Claim 3)
[1465] 10. The system of claim 1, further comprising means for storing the user's native language setting in a database.
[1466] "Example 2: Combining Emotion Engines"
[1467] (Claim 1)
[1468] a means for a user to set their native language;
[1469] a means for a user to input age and interest information;
[1470] means of obtaining data from sources in conjunction with data collection instruments;
[1471] A generative AI means for translating the acquired data into the user's native language in real time;
[1472] An emotion recognition means for recognizing an emotion of a user using a camera or a microphone of the terminal;
[1473] means for optimizing the translated data based on the user's age, interests, and emotions;
[1474] means for providing optimized data to a user;
[1475] A system including:
[1476] (Claim 2)
[1477] 10. The system of claim 1, further comprising means for storing the acquired data.
[1478] (Claim 3)
[1479] 10. The system of claim 1, further comprising means for storing the user's native language setting in a database.
[1480] "Application example 2 when combining emotion engines"
[1481] (Claim 1)
[1482] a means for a user to set their native language;
[1483] a means for a user to input age and interest information;
[1484] means of obtaining data from sources in conjunction with data collection instruments;
[1485] A generative AI means for translating the acquired data into the user's native language in real time;
[1486] means for optimizing the translated data based on the age and interests of the user;
[1487] A means for recognizing user emotions and optimizing information;
[1488] means for displaying the optimized data to a user;
[1489] A system including:
[1490] (Claim 2)
[1491] 10. The system of claim 1, further comprising means for storing the acquired data.
[1492] (Claim 3)
[1493] 10. The system of claim 1, further comprising means for storing the user's native language setting in a database. [Explanation of symbols]
[1494] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for a user to set their native language; a means for a user to input age and interest information; means of obtaining data from sources in conjunction with data collection instruments; A generative AI means for translating the acquired data into the user's native language in real time; means for optimizing the translated data based on the age and interests of the user; means for providing optimized data to a user; A system including:
2. The system of claim 1 further comprising means for storing the acquired data.
3. 10. The system of claim 1, further comprising means for storing the user's native language setting in a database.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A