System

A system using chat-based and automatic composition AIs generates mnemonics and melodies to facilitate efficient and enjoyable learning by converting user-input information into audio data.

JP2026017442APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118224
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Existing learning tools lack support for efficiently and enjoyably memorizing information using mnemonics and melodies, requiring manual effort for creation.

Method used

A system that includes a chat-based AI to generate mnemonics and an automatic composition AI to create melodies, converting user-input information into audio data for efficient and enjoyable learning.

Benefits of technology

Enables users to easily and effectively memorize information through the combination of mnemonics and melodies, improving learning efficiency and enjoyment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017442000001_ABST
    Figure 2026017442000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving information desired to be memorized from a user and transmitting the information to a chat system artificial intelligence; means for causing the chat system artificial intelligence to generate a bargain based on the information and return the bargain; means for transmitting the received bargain to an automatic composition system artificial intelligence; and means for causing the automatic composition system artificial intelligence to generate a melody suitable for the bargain, generate voice data in which the melody is incorporated in the bargain, and return the voice data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Using mnemonics and appropriate melodies is an effective way for learners to memorize information effectively and efficiently. However, manually creating mnemonics and then coming up with melodies takes time and effort. As a result, there is a lack of support tools to help learners memorize information efficiently. [Means for solving the problem]

[0005] In order to solve the above problems, the present invention provides the following means.

[0006] The system includes: means for receiving information to be memorized from a user and transmitting the information to a chat-based AI; means for the chat-based AI to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic to an automatic composition AI; means for the automatic composition AI to generate a melody suited to the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; and means for providing the received audio data to a user. This allows learners to easily use a learning support tool that combines mnemonics and melodies, enabling them to efficiently memorize information to be memorized.

[0007] A "user" is a person who uses this system to input information they want to remember and receives learning support.

[0008] "Terminal" refers to a device through which a user inputs information and receives generated mnemonics and melodies, and includes smartphones, PCs, etc.

[0009] "Chat AI" is AI that analyzes information entered by users and generates appropriate mnemonics.

[0010] "Automatic composition AI" is an AI that generates a suitable melody for a mnemonic generated by a chat AI and returns it as audio data.

[0011] A mnemonic is a phrase or short sentence that converts information to be remembered into a concise and easy-to-remember form.

[0012] A "melody" is a musical melody generated to match the mnemonic lyrics.

[0013] "Audio data" is a digital audio file that combines mnemonics and melodies generated by an automatic composition AI.

[0014] "Return" is the act of returning the results generated by artificial intelligence to the terminal. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention describes a system that assists learning by inputting information that a user wants to remember and converting it into mnemonics and melodies. This system is composed of a terminal, a chat system AI, and an automatic composition system AI.

[0037] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. Once input is complete, the device sends this information to a chat-based AI located on a server. The chat-based AI analyzes the received information and generates an appropriate mnemonic. The mnemonic is a concise and easy-to-remember version of the original information (for example, "tai, water, gold, earth, fire, wood, earth, heaven, sea").

[0038] The generated mnemonic is sent back to the device. The device then sends the mnemonic to an automatic composition AI on the server. The automatic composition AI creates an appropriate melody based on the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The audio data is generated in a music file format (e.g., MP3 format) and sent back to the device.

[0039] The terminal provides the received audio data to the user. By playing the audio data, the user can listen to the mnemonic set to the melody, allowing the user to memorize the information they want to remember in an enjoyable and efficient way.

[0040] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the automatic composition AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who will play it back and study.

[0041] This system allows users to easily memorize information, improving learning efficiency. Furthermore, learning becomes more enjoyable because information is more easily retained in memory in the form of music.

[0042] The processing flow will be explained below.

[0043] Step 1:

[0044] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0045] Step 2:

[0046] The device receives the input information and sends it to the chat AI, which can use protocols such as APIs.

[0047] Step 3:

[0048] The chat AI on the server analyzes the information it receives. For example, if the input is about the order of the planets in the solar system, it will extract the necessary data.

[0049] Step 4:

[0050] The chat AI generates simple and easy-to-remember mnemonics based on the extracted data, such as "tai (sea), water, gold (metal), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)."

[0051] Step 5:

[0052] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0053] Step 6:

[0054] The device then sends the received mnemonics to the automatic composition AI on the server. The transmission protocol is the same as in step 2.

[0055] Step 7:

[0056] The automatic composition AI analyzes the received mnemonics and generates a suitable melody that incorporates the mnemonics as lyrics.

[0057] Step 8:

[0058] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0059] Step 9:

[0060] The terminal provides the received audio data to the user, who can then play the audio data on the terminal and listen to the mnemonic set to the melody.

[0061] Step 10:

[0062] By repeatedly listening to the audio data, users can efficiently memorize the information they want to remember, and this process improves learning effectiveness.

[0063] Example 1

[0064] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0065] Conventional learning support systems simply present the information users want to remember as text or images, making it difficult for them to memorize the information effectively and in an enjoyable way. Furthermore, they lacked an appealing method to help users solidify their memories. While memorizing information accompanied by music is expected to help users remember it longer, no such system existed. Therefore, a method was needed to enable users to learn the information they want to memorize efficiently and in an enjoyable way.

[0066] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0067] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a natural language processing AI, means for the natural language processing AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to a music generation AI, means for the music generation AI to generate music suitable for the mnemonic, generate audio data incorporating the music into the mnemonic, and return the audio data, and means for providing the received audio data to the user. This allows the user to learn the information they want to memorize in an enjoyable and effective way.

[0068] "User" refers to the entity that inputs the information they want to remember into the system and learns.

[0069] "Terminal" refers to a device that allows a user to input information or receive voice data. Examples include smartphones, tablets, and PCs.

[0070] "Server" refers to the central part of the computer network where the chat AI and automatic composition AI run.

[0071] "Natural language processing AI" refers to artificial intelligence programs designed to understand and analyze human language. Examples include GPT-3.

[0072] "Generative music AI" refers to artificial intelligence programs that generate music based on given text or information. Examples include Jukedeck and Amper Music.

[0073] "Information to be remembered" refers to specific data or knowledge that the user wants to remember. For example, "the order of the planets in the solar system."

[0074] A mnemonic is a phrase that converts information you want to remember into a concise and easy-to-remember form.

[0075] "Audio Data" refers to data formats containing music generated by automatic music composition AI, such as MP3 format.

[0076] This invention relates to a system that assists learning by converting information that a user wants to remember into mnemonics and melodies. This system is composed of a terminal, a natural language processing AI, and a music generation AI.

[0077] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into a terminal. The terminal can be a smartphone, tablet, PC, or the like. Once the user inputs the information, the terminal has a means for sending this information to a server. This means sends the information to the server using a protocol such as an HTTP request.

[0078] A natural language processing AI (e.g., GPT-3) runs on the server, analyzes the received information, and generates an appropriate mnemonic. For example, if a user inputs, "I want to remember the order of the planets in the solar system," the natural language processing AI generates a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is stored in memory and sent back to the device.

[0079] The device then sends the acquired mnemonic back to the server, this time passing it on to a music generation AI (e.g., Jukedeck or Amper Music). The music generation AI then creates an appropriate melody based on the mnemonic. Specifically, it incorporates the generated mnemonic as lyrics and generates the audio in an audio data format (e.g., MP3 format).

[0080] The generated audio data is sent back from the server to the device. The device has the means to store the audio data in local storage and provide it to the user. The user can play this audio data using the device's music playback application. This allows the user to enjoy learning while listening to mnemonics set to the melody.

[0081] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system" into a device, the natural language processing AI will generate a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the music generation AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who can play it back to study.

[0082] An example of a prompt is as follows:

[0083] User: I want to remember the order of the planets in the solar system.

[0084] This system allows users to memorize information in a fun and efficient way, improving learning efficiency.In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0085] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0086] Step 1:

[0087] The user inputs the information they want to remember into the device. For example, they input "I want to remember the order of the planets in the solar system." This input data is saved in an input field on the device.

[0088] Step 2:

[0089] The device sends the information entered to the server. Specifically, it creates an HTTP request and sends the information entered by the user to the server's natural language processing AI. The input is the information entered by the user, and the output is a notification to the server that transmission has been completed.

[0090] Step 3:

[0091] The server's natural language processing AI analyzes the information sent. A chat-based AI (e.g., GPT-3) analyzes the user's input and generates an appropriate mnemonic. For example, an input such as "I want to remember the order of the planets in the solar system" is converted into a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." The input is the information sent by the user, and the output is the generated mnemonic.

[0092] Step 4:

[0093] The server sends the generated mnemonic to the terminal. Specifically, it returns this mnemonic as an HTTP response. The terminal receives this response. The input is the generated mnemonic, and the output is a notification to the terminal that transmission has been completed.

[0094] Step 5:

[0095] The device sends the received mnemonic to the music generation AI on the server. The generated mnemonic is sent again to the server, this time requesting processing from the music generation AI. Specifically, an HTTP request is created and the mnemonic data is sent to the music generation AI endpoint. The input is the received mnemonic, and the output is a notification to the server that transmission has been completed.

[0096] Step 6:

[0097] The server's music generation AI generates a melody that matches the mnemonic. Specifically, it creates a melody based on this mnemonic and generates audio data that incorporates the mnemonic as lyrics. For example, it generates audio data (MP3 format, etc.) with the lyrics "tai, water, gold, earth, fire, wood, earth, heaven, sea" set to a melody. The input is the mnemonic, and the output is the generated audio data.

[0098] Step 7:

[0099] The server sends the generated audio data to the terminal. Specifically, it returns the generated audio data to the terminal as an HTTP response. The input is the generated audio data, and the output is a transmission completion notification to the terminal.

[0100] Step 8:

[0101] The device provides the received audio data to the user. The audio data stored on the device can be played back upon user request. The user can listen to this audio data using the device's music playback application. The input is the received audio data, and the output is the audio data provided in a playable format.

[0102] This system allows users to memorize information in a fun and efficient way, improving learning efficiency. In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0103] (Application example 1)

[0104] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0105] While there are many methods for efficiently memorizing information, there are only a limited number of ways for users to study in an enjoyable way. Furthermore, there are not enough methods for effectively tracking learning progress. Therefore, there is a need for a method to maintain motivation while solidifying information in memory.

[0106] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0107] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a chat-based AI, means for the chat-based AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to an automatic composition-based AI, means for the automatic composition-based AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data, means for providing the received audio data to the user, means for playing the audio data through a smartphone application, and means for tracking the user's learning progress. This allows the user to memorize information in a fun and efficient way, and enables the user's learning progress to be effectively tracked.

[0108] "User" refers to a person who uses the system of the present invention to learn information.

[0109] "Memorable information" refers to facts or data that a user wishes to remember.

[0110] "Chat AI" is AI that conducts dialogue based on input data, analyzes it, and generates information.

[0111] A mnemonic is a conversion of original information into easy-to-remember words or phrases.

[0112] "Automatic composition artificial intelligence" is an artificial intelligence that has the ability to automatically create music.

[0113] A "melody" is a series of notes combined together to form a section of music that is aurally pleasing.

[0114] "Audio data" is data in the form of a digital file that contains audio information.

[0115] A "smartphone application" is a type of software that runs on a smartphone.

[0116] "Study Progress Tracking" means the act of continuously recording and monitoring a user's learning progress.

[0117] The "Melody Learning Assistant" system of the present invention is designed to help users memorize information they want to remember in a fun and effective way. First, the user inputs the information they want to memorize using a smartphone application. This information is called "information to memorize." Once the user inputs the information, it is sent from the device to a server.

[0118] A "chat AI" is placed on the server, and this AI analyzes the input information and generates easy-to-remember "mnemonics." The chat AI creates mnemonics using a generative AI model. These mnemonics convert the original information into a concise and easy-to-remember format. For example, if a user inputs "I want to remember the order of the planets in the solar system," the chat AI generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi."

[0119] The generated mnemonics are stored on the server and then sent to an "automatic composition AI." This automatic composition AI also uses a generative AI model to create a "melody" that suits the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The generated audio data is generated in a music file format such as MP3 and sent back from the server to the device.

[0120] The smartphone application has the function of providing the received audio data to the user and playing it back. This allows the user to listen to the mnemonics set to the melody, making it fun to memorize the information. The application also has the function of tracking the user's learning progress, recording and monitoring how much progress has been made.

[0121] This system not only enables users to memorize information efficiently, but also enables them to effectively manage their learning progress. Below are some examples of specific prompt sentences.

[0122] (Example of a prompt)

[0123] 1. Chat AI prompt:

[0124] Turn the following information into a mnemonic: The order of the planets in the solar system

[0125] 2. Prompts for automatic composition AI:

[0126] Compose a melody based on the following puns: Tai (tai), Mizu (sui), Kin (kin), Chi (chi), Ka (hi), Moku (ki), Tsuchi (do), Ten (ten), Umi (kai)

[0127] As described above, the system of the present invention allows users to memorize information while having fun and effectively manage their learning progress.

[0128] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0129] Step 1:

[0130] The user enters the information they want to remember.

[0131] The user launches the smartphone application and inputs the information they want to remember. The input information is sent to the server as "information to remember." The input here is in text format, such as "the order of the planets in the solar system."

[0132] Step 2:

[0133] The server sends the information to the chat AI.

[0134] The server sends the received information to a chat-based AI, which analyzes the input information and generates easy-to-remember mnemonics. The AI ​​uses a generative AI model to analyze the data and output mnemonics such as "tai (tai), water (mizu), kin (kin), chi (chi), ka (hi), ki (moku), tsuchi (tsuchi), ten (ten), umi (umi)."

[0135] Step 3:

[0136] The generated mnemonics are sent from the server to an automatic composition AI.

[0137] The server receives the mnemonic returned by the chat AI and sends it to the automatic composition AI. The automatic composition AI uses the mnemonic as input data to generate an appropriate melody. The melody is generated using a prompt sentence.

[0138] Step 4:

[0139] An automatic composition AI generates a melody and returns it as audio data.

[0140] The automatic composition AI creates a melody based on the received mnemonic and generates it as audio data. This audio data is sent back to the server in a playable format such as MP3. The output audio data is provided to the user in a format that allows them to learn in an enjoyable way.

[0141] Step 5:

[0142] The server sends the audio data back to the smartphone application.

[0143] The server sends the audio data received from the automatic composition AI back to the smartphone application, at which point the transmission of the audio data is complete.

[0144] Step 6:

[0145] The smartphone application plays the audio data and provides it to the user.

[0146] The smartphone application plays the received audio data and provides it to the user. By listening to the mnemonics set to the melody, the user can memorize information in a fun and efficient way. Data processing and calculations are performed in real time during playback.

[0147] Step 7:

[0148] A smartphone application tracks learning progress.

[0149] The smartphone application tracks the user's learning progress and records their achievements and progress, allowing them to effectively manage their own learning. Data is periodically sent to a server and stored.

[0150] The above are the specific processing steps of the "Melody Learning Assistant" system.

[0151] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0152] This paper describes a system that supports learning by inputting information that a user wants to remember and converting it into mnemonics and melodies, and further combines this with an emotion engine that recognizes the user's emotions, enabling individually optimized learning support. This system is composed of a terminal, chat AI, automatic composition AI, and the emotion engine.

[0153] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. The device then sends the input information to a chat-based AI on the server. A unique feature of the device is that it is equipped with an emotion engine that analyzes emotional information (for example, voice tone, facial expression, input content, etc.) during the user's input process. The results of this analysis are sent to the chat-based AI and the automatic composition AI, and are used to adjust the generated mnemonics and melodies.

[0154] The chat AI analyzes the received information and generates appropriate mnemonics. The mnemonics are a concise and easy-to-remember version of the original information, such as "tai, water, gold, earth, fire, wood, earth, heaven, sea." Based on the analysis results of the emotion engine, the mnemonics can be adjusted according to the user's emotional state.

[0155] The generated mnemonic is sent back to the device. The device then sends the mnemonic and the analysis results of the emotion engine to an automatic composition AI on the server. The automatic composition AI generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. Furthermore, based on the results of the emotion engine, the tempo and key of the melody are adjusted to match the user's emotional state.

[0156] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user. The user can play the audio data on the device and listen to the mnemonics that go with the melody. The user can also download the data and listen to it repeatedly.

[0157] For example, if a user types "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and emotional information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0158] This system allows users to easily memorize information they want to remember in an individually optimized way, improving learning efficiency. The introduction of an emotion engine makes it possible to further customize the system to match the user's emotional state, resulting in more effective learning support.

[0159] The processing flow will be explained below.

[0160] Step 1:

[0161] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0162] Step 2:

[0163] As soon as the device receives the information entered by the user, it activates an emotion engine to analyze the user's emotional state, based on the user's tone of voice, facial expression, and typing speed.

[0164] Step 3:

[0165] The device sends the input information and the emotional state analyzed by the emotion engine to the chat AI on the server, and the data includes emotional information.

[0166] Step 4:

[0167] The chat AI on the server analyzes the received information and generates appropriate mnemonics. For example, it creates mnemonics such as "tai (sea), water, gold (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)." It also adjusts the expressions of the mnemonics based on emotional information.

[0168] Step 5:

[0169] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0170] Step 6:

[0171] The device then sends the received mnemonic and the analysis results of the emotion engine to the automatic composition AI on the server. The transmission protocol is the same as in step 3.

[0172] Step 7:

[0173] The automatic composition AI on the server analyzes the received mnemonics and emotional information. Based on this, it generates a melody that matches the mnemonics. This melody is generated in a format that incorporates the mnemonics as lyrics. It also adjusts the tempo and key of the melody based on the emotional information.

[0174] Step 8:

[0175] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0176] Step 9:

[0177] The device then provides the received audio data to the user. The user can then play the audio data on the device and listen to the mnemonics set to the melody. The audio data is also provided in a downloadable format, so it can be listened to repeatedly.

[0178] Step 10:

[0179] By repeatedly listening to audio data, users can efficiently memorize the information they want to remember. This process improves learning effectiveness and is personalized to their emotional state, resulting in more effective learning.

[0180] Example 2

[0181] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0182] Conventional learning support systems lack the ingenuity to help users effectively memorize the information they want to remember, and in particular, they lack individual optimization based on the user's emotional state, which results in a decrease in learning efficiency.In addition, there is a demand for a method that makes learning fun rather than simply memorizing.

[0183] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be memorized from a user and transmitting the information to a natural language processing system; means for the natural language processing system to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic and the user's emotional information to an automatic composition system; means for the automatic composition system to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for analyzing the user's emotions and transmitting the analysis results to the natural language processing system and the automatic composition system; and means for providing the received audio data to the user. This enables individually optimized learning support tailored to the user's emotional state, resulting in enjoyable and efficient memorization.

[0184] "User" refers to anyone who wants to use the system to remember information.

[0185] "Information to be remembered" refers to specific data or facts that the user wants to remember.

[0186] A "natural language processing system" refers to an artificial intelligence-based processing system that analyzes input text and generates appropriate output.

[0187] A mnemonic is a sentence or phrase designed to make it easier to remember information.

[0188] An "automated composition system" refers to an artificial intelligence-based system that can generate melodies and music based on input linguistic data.

[0189] "Audio data" refers to the data of an acoustic signal that combines the generated melody and mnemonic.

[0190] "Emotional information" is data that indicates the user's emotional state during input, and refers to the analysis results obtained from voice tone, facial expression, input content, etc.

[0191] This system supports learning by allowing users to input information they want to remember and converting it into mnemonics and melodies. By combining this with an emotion engine that recognizes the user's emotions, it enables individually optimized learning support. This system is comprised of a terminal, a natural language processing system, an automatic composition system, and an emotion recognition engine.

[0192] First, the user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." In this case, it is desirable for the device to have an interface such as voice input or touch panel input.

[0193] Next, the device sends the input information to a natural language processing system on a server. For example, GPT-3 or GPT-4 is used for the natural language processing system. During this process, the device's built-in emotion recognition engine recognizes and analyzes the user's emotions from voice tone, facial expressions, etc. The analysis results are sent to the natural language processing system or automatic composition system, making it possible to provide output tailored to the user's emotional state.

[0194] The natural language processing system analyzes the received information and generates an appropriate mnemonic. For example, it converts the information "the order of the planets in the solar system" into "tai, water, gold, earth, fire, wood, earth, sky, sea." This mnemonic may then be adjusted according to the user's emotional state based on the analysis results of an emotion recognition engine.

[0195] The generated mnemonic is sent back to the device, which then sends the mnemonic and the user's emotion analysis results to an automatic composition system on the server. Examples of automatic composition systems include AWS DeepComposer and Magenta. The automatic composition system generates a melody that matches the mnemonic and generates audio data that incorporates this melody as lyrics. The tempo and key of the melody are also adjusted based on the analysis results of the emotion recognition engine.

[0196] The audio data generated by the automatic composition system is usually sent to the device in MP3 or WAV format. The device receives this audio data and provides it to the user. The user can play the audio data on their device and listen to the mnemonics that go with the melody, improving memorization efficiency by integrating visual and auditory information. The audio data can also be downloaded and listened to repeatedly.

[0197] Example prompt sentence:

[0198] "Please explain in detail the process by which the system inputs information that the user wants to remember and converts it into a mnemonic and melody."

[0199] In this way, the present invention provides efficient and enjoyable learning support that incorporates customization based on the user's emotions, and can improve upon the problems of conventional learning support systems.

[0200] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0201] Step 1:

[0202] The user inputs the information they want to remember into the terminal.

[0203] Input: Information you want to remember (e.g., "I want to remember the order of the planets in the solar system")

[0204] Output: User input data

[0205] Specific operation: The user inputs text information using the device's interface, and the device stores the input data in its internal memory.

[0206] Step 2:

[0207] The terminal transmits the input information to a natural language processing system.

[0208] Input: User-entered data

[0209] Output: Data (text data) to be sent to the natural language processing system

[0210] Specific operation: The device acquires user input data stored in its internal memory and transmits it to the server's natural language processing system via the Internet.

[0211] Step 3:

[0212] The device analyzes the user's emotional information and sends it to a natural language processing system.

[0213] Input: User's tone of voice, facial expressions, and input

[0214] Output: Sentiment analysis data

[0215] Specific operation: The device uses the emotion engine to analyze emotional information such as voice tone and facial expressions. The analysis results are sent to the server's natural language processing system as emotion analysis data.

[0216] Step 4:

[0217] The server generates mnemonics using a natural language processing system.

[0218] Input: User input data, sentiment analysis data

[0219] Output: Mnemonic data

[0220] Specific operation: The natural language processing system analyzes user input data and generates easy-to-remember mnemonics. The mnemonics are adjusted based on sentiment analysis data. The generated mnemonics are sent back to the device from the server.

[0221] Step 5:

[0222] The terminal again transmits the mnemonic and emotional information to the automatic composition system on the server.

[0223] Input: mnemonic data, sentiment analysis data

[0224] Output: Data to be sent to the automated composition system

[0225] Specific operation: The device combines the mnemonics received from the server with the emotion analysis data and sends it to the automatic composition system.

[0226] Step 6:

[0227] The server generates a melody using an automatic composition system.

[0228] Input: mnemonic data, sentiment analysis data

[0229] Output: Audio data

[0230] Specific operation: The automatic composition system generates a melody that matches the user's emotional state based on the received mnemonic. Audio data is then generated with the mnemonic incorporated into the melody as lyrics.

[0231] Step 7:

[0232] The server transmits the generated voice data to the terminal.

[0233] Input: Audio data

[0234] Output: Audio data sent to the device (e.g. MP3 or WAV format)

[0235] Specific operation: The server sends the generated audio data back to the device, usually in MP3 or WAV format.

[0236] Step 8:

[0237] The terminal provides the audio data to the user.

[0238] Input: Audio data

[0239] Output: Playable audio data

[0240] Specific operation: The device uses playback software to provide the received audio data to the user. The user plays the audio data and listens to the mnemonic set to the melody. The user can also download the audio data and listen to it repeatedly.

[0241] (Application example 2)

[0242] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0243] Conventional learning support systems do not provide a method for users to efficiently memorize information. Furthermore, they are unable to optimize the content of learning support according to the user's emotional state. As a result, learning efficiency is low and it is difficult for users to easily memorize information. The present invention aims to solve these problems by analyzing the user's emotions and providing optimized learning support based on the analysis.

[0244] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be remembered from a user and transmitting the information to a chat-type AI; means for the chat-type AI to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic to an automatic composition-type AI; means for the automatic composition-type AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for providing the received audio data to the user; and means for incorporating an emotion engine that analyzes the user's emotions and optimizing a series of processes based on the emotion information analyzed by the emotion engine. This enables the user to memorize information in a form optimized according to their emotional state.

[0245] A "user" is a person who uses the system to input information and receive learning support.

[0246] "Information to be remembered" is content or data that the user wants to remember.

[0247] "Chat AI" refers to algorithms or programs that generate mnemonics based on the information they receive.

[0248] A mnemonic is a simple word that uses sound and rhythm to make difficult-to-remember information easier to remember.

[0249] "Automatic composition artificial intelligence" refers to algorithms and programs that generate melodies suitable for mnemonics and incorporate them as audio data.

[0250] A "melody" is a musical melody, a series of sounds generated according to a mnemonic.

[0251] "Audio data" refers to digital sound information including the generated melody.

[0252] An "emotion engine" refers to an algorithm or program that analyzes the user's emotional state and optimizes the system's operation based on that information.

[0253] "Playback" refers to the act of letting the user listen to the generated audio data.

[0254] A "downloadable format" is a digital data format that allows a user to save and later play the content.

[0255] This invention relates to a learning support system that inputs information that a user wants to remember and converts it into mnemonics and melodies. The system is composed of the following elements: a terminal, chat AI, automatic composition AI, and an emotion engine. This system provides personalized and optimized learning support according to the user's emotional state.

[0256] First, the user inputs the information they want to remember into a device. This device could be a smartphone, smart glasses, or a head-mounted display. The input information is sent to a chat-based AI on a server. The chat-based AI generates an appropriate mnemonic based on the received information.

[0257] Next, the device is equipped with an emotion engine that analyzes the emotional information as the user inputs. The analysis results are sent to chat AI and automatic composition AI. The emotion engine determines the user's emotions based on voice tone, facial expressions, input content, etc. Emotion engines that can be used include Affectiva and Microsoft Azure Emotional Analysis API.

[0258] The generated mnemonics and emotional information are sent to an automatic composition AI, which then generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonics as lyrics. Possible AIs that could be used include OpenAI's MuseNet.

[0259] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user and plays it back. A smartphone's Pygame library is an effective way to play the audio. Users can also download the audio data and listen to it repeatedly.

[0260] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and the emotion information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0261] An example prompt might have the following format:

[0262] Chat AI prompt:

[0263] User input: I want to remember the order of the planets in the solar system.

[0264] User emotion: Excited

[0265] Prompt for automatic composition AI:

[0266] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0267] User emotion: Excited

[0268] In the above-described manner, the present invention enables learning support that is optimized for the emotional state of the user.

[0269] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0270] Step 1:

[0271] The user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." The device then sends this input information to a chat-based AI on the server.

[0272] Step 2:

[0273] The chat AI on the server analyzes the received information and generates a mnemonic. For example, in response to the input information "I want to remember the order of the planets in the solar system," it generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is then sent back to the device.

[0274] Step 3:

[0275] At the same time, the device uses an emotion engine to analyze the emotional information obtained during the user's input process. For example, it may analyze the user's tone of voice and facial expression as "excited." This emotional information is then sent to the server.

[0276] Step 4:

[0277] The device sends the received mnemonic and emotional information to the automatic composition AI on the server, using the following format as a prompt:

[0278] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0279] User emotion: Excited

[0280] Step 5:

[0281] The automatic composition AI on the server generates an appropriate melody based on mnemonics and emotional information. For example, it creates a bright, fast-paced melody for the mnemonic "tai (sea), water, gold (sea), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), sea (ocean)." This melody is generated as audio data and sent back to the device.

[0282] Step 6:

[0283] The device provides the received audio data to the user. The audio data is usually in MP3 or WAV format and is played on the device. For example, you can play audio using the Pygame library.

[0284] Step 7:

[0285] The user can listen to the audio data played on the device and confirm that the mnemonic is set to the melody. The user can also download the audio data and play it as many times as they like.

[0286] These steps allow the user to remember information in a way that is optimized for their emotional state.

[0287] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0288] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0289] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0290] [Second embodiment]

[0291] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0292] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0293] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0294] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0295] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0296] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0297] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0298] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0299] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0300] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0301] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0302] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0303] This invention describes a system that assists learning by inputting information that a user wants to remember and converting it into mnemonics and melodies. This system is composed of a terminal, a chat system AI, and an automatic composition system AI.

[0304] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. Once input is complete, the device sends this information to a chat-based AI located on a server. The chat-based AI analyzes the received information and generates an appropriate mnemonic. The mnemonic is a concise and easy-to-remember version of the original information (for example, "tai, water, gold, earth, fire, wood, earth, heaven, sea").

[0305] The generated mnemonic is sent back to the device. The device then sends the mnemonic to an automatic composition AI on the server. The automatic composition AI creates an appropriate melody based on the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The audio data is generated in a music file format (e.g., MP3 format) and sent back to the device.

[0306] The terminal provides the received audio data to the user. By playing the audio data, the user can listen to the mnemonic set to the melody, allowing the user to memorize the information they want to remember in an enjoyable and efficient way.

[0307] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the automatic composition AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who will play it back and study.

[0308] This system allows users to easily memorize information, improving learning efficiency. Furthermore, learning becomes more enjoyable because information is more easily retained in memory in the form of music.

[0309] The processing flow will be explained below.

[0310] Step 1:

[0311] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0312] Step 2:

[0313] The device receives the input information and sends it to the chat AI, which can use protocols such as APIs.

[0314] Step 3:

[0315] The chat AI on the server analyzes the information it receives. For example, if the input is about the order of the planets in the solar system, it will extract the necessary data.

[0316] Step 4:

[0317] The chat AI generates simple and easy-to-remember mnemonics based on the extracted data, such as "tai (sea), water, gold (metal), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)."

[0318] Step 5:

[0319] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0320] Step 6:

[0321] The device then sends the received mnemonics to the automatic composition AI on the server. The transmission protocol is the same as in step 2.

[0322] Step 7:

[0323] The automatic composition AI analyzes the received mnemonics and generates a suitable melody that incorporates the mnemonics as lyrics.

[0324] Step 8:

[0325] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0326] Step 9:

[0327] The terminal provides the received audio data to the user, who can then play the audio data on the terminal and listen to the mnemonic set to the melody.

[0328] Step 10:

[0329] By repeatedly listening to the audio data, users can efficiently memorize the information they want to remember, and this process improves learning effectiveness.

[0330] Example 1

[0331] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0332] Conventional learning support systems simply present the information users want to remember as text or images, making it difficult for them to memorize the information effectively and in an enjoyable way. Furthermore, they lacked an appealing method to help users solidify their memories. While memorizing information accompanied by music is expected to help users remember it longer, no such system existed. Therefore, a method was needed to enable users to learn the information they want to memorize efficiently and in an enjoyable way.

[0333] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0334] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a natural language processing AI, means for the natural language processing AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to a music generation AI, means for the music generation AI to generate music suitable for the mnemonic, generate audio data incorporating the music into the mnemonic, and return the audio data, and means for providing the received audio data to the user. This allows the user to learn the information they want to memorize in an enjoyable and effective way.

[0335] "User" refers to the entity that inputs the information they want to remember into the system and learns.

[0336] "Terminal" refers to a device that allows a user to input information or receive voice data. Examples include smartphones, tablets, and PCs.

[0337] "Server" refers to the central part of the computer network where the chat AI and automatic composition AI run.

[0338] "Natural language processing AI" refers to artificial intelligence programs designed to understand and analyze human language. Examples include GPT-3.

[0339] "Generative music AI" refers to artificial intelligence programs that generate music based on given text or information. Examples include Jukedeck and Amper Music.

[0340] "Information to be remembered" refers to specific data or knowledge that the user wants to remember. For example, "the order of the planets in the solar system."

[0341] A mnemonic is a phrase that converts information you want to remember into a concise and easy-to-remember form.

[0342] "Audio Data" refers to data formats containing music generated by automatic music composition AI, such as MP3 format.

[0343] This invention relates to a system that assists learning by converting information that a user wants to remember into mnemonics and melodies. This system is composed of a terminal, a natural language processing AI, and a music generation AI.

[0344] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into a terminal. The terminal can be a smartphone, tablet, PC, or the like. Once the user inputs the information, the terminal has a means for sending this information to a server. This means sends the information to the server using a protocol such as an HTTP request.

[0345] A natural language processing AI (e.g., GPT-3) runs on the server, analyzes the received information, and generates an appropriate mnemonic. For example, if a user inputs, "I want to remember the order of the planets in the solar system," the natural language processing AI generates a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is stored in memory and sent back to the device.

[0346] The device then sends the acquired mnemonic back to the server, this time passing it on to a music generation AI (e.g., Jukedeck or Amper Music). The music generation AI then creates an appropriate melody based on the mnemonic. Specifically, it incorporates the generated mnemonic as lyrics and generates the audio in an audio data format (e.g., MP3 format).

[0347] The generated audio data is sent back from the server to the device. The device has the means to store the audio data in local storage and provide it to the user. The user can play this audio data using the device's music playback application. This allows the user to enjoy learning while listening to mnemonics set to the melody.

[0348] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system" into a device, the natural language processing AI will generate a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the music generation AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who can play it back to study.

[0349] An example of a prompt is as follows:

[0350] User: I want to remember the order of the planets in the solar system.

[0351] This system allows users to memorize information in a fun and efficient way, improving learning efficiency.In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0352] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0353] Step 1:

[0354] The user inputs the information they want to remember into the device. For example, they input "I want to remember the order of the planets in the solar system." This input data is saved in an input field on the device.

[0355] Step 2:

[0356] The device sends the information entered to the server. Specifically, it creates an HTTP request and sends the information entered by the user to the server's natural language processing AI. The input is the information entered by the user, and the output is a notification to the server that transmission has been completed.

[0357] Step 3:

[0358] The server's natural language processing AI analyzes the information sent. A chat-based AI (e.g., GPT-3) analyzes the user's input and generates an appropriate mnemonic. For example, an input such as "I want to remember the order of the planets in the solar system" is converted into a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." The input is the information sent by the user, and the output is the generated mnemonic.

[0359] Step 4:

[0360] The server sends the generated mnemonic to the terminal. Specifically, it returns this mnemonic as an HTTP response. The terminal receives this response. The input is the generated mnemonic, and the output is a notification to the terminal that transmission has been completed.

[0361] Step 5:

[0362] The device sends the received mnemonic to the music generation AI on the server. The generated mnemonic is sent again to the server, this time requesting processing from the music generation AI. Specifically, an HTTP request is created and the mnemonic data is sent to the music generation AI endpoint. The input is the received mnemonic, and the output is a notification to the server that transmission has been completed.

[0363] Step 6:

[0364] The server's music generation AI generates a melody that matches the mnemonic. Specifically, it creates a melody based on this mnemonic and generates audio data that incorporates the mnemonic as lyrics. For example, it generates audio data (MP3 format, etc.) with the lyrics "tai, water, gold, earth, fire, wood, earth, heaven, sea" set to a melody. The input is the mnemonic, and the output is the generated audio data.

[0365] Step 7:

[0366] The server sends the generated audio data to the terminal. Specifically, it returns the generated audio data to the terminal as an HTTP response. The input is the generated audio data, and the output is a transmission completion notification to the terminal.

[0367] Step 8:

[0368] The device provides the received audio data to the user. The audio data stored on the device can be played back upon user request. The user can listen to this audio data using the device's music playback application. The input is the received audio data, and the output is the audio data provided in a playable format.

[0369] This system allows users to memorize information in a fun and efficient way, improving learning efficiency. In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0370] (Application example 1)

[0371] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0372] While there are many methods for efficiently memorizing information, there are only a limited number of ways for users to study in an enjoyable way. Furthermore, there are not enough methods for effectively tracking learning progress. Therefore, there is a need for a method to maintain motivation while solidifying information in memory.

[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0374] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a chat-based AI, means for the chat-based AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to an automatic composition-based AI, means for the automatic composition-based AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data, means for providing the received audio data to the user, means for playing the audio data through a smartphone application, and means for tracking the user's learning progress. This allows the user to memorize information in a fun and efficient way, and enables the user's learning progress to be effectively tracked.

[0375] "User" refers to a person who uses the system of the present invention to learn information.

[0376] "Memorable information" refers to facts or data that a user wishes to remember.

[0377] "Chat AI" is AI that conducts dialogue based on input data, analyzes it, and generates information.

[0378] A mnemonic is a conversion of original information into easy-to-remember words or phrases.

[0379] "Automatic composition artificial intelligence" is an artificial intelligence that has the ability to automatically create music.

[0380] A "melody" is a series of notes combined together to form a section of music that is aurally pleasing.

[0381] "Audio data" is data in the form of a digital file that contains audio information.

[0382] A "smartphone application" is a type of software that runs on a smartphone.

[0383] "Study Progress Tracking" means the act of continuously recording and monitoring a user's learning progress.

[0384] The "Melody Learning Assistant" system of the present invention is designed to help users memorize information they want to remember in a fun and effective way. First, the user inputs the information they want to memorize using a smartphone application. This information is called "information to memorize." Once the user inputs the information, it is sent from the device to a server.

[0385] A "chat AI" is placed on the server, and this AI analyzes the input information and generates easy-to-remember "mnemonics." The chat AI creates mnemonics using a generative AI model. These mnemonics convert the original information into a concise and easy-to-remember format. For example, if a user inputs "I want to remember the order of the planets in the solar system," the chat AI generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi."

[0386] The generated mnemonics are stored on the server and then sent to an "automatic composition AI." This automatic composition AI also uses a generative AI model to create a "melody" that suits the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The generated audio data is generated in a music file format such as MP3 and sent back from the server to the device.

[0387] The smartphone application has the function of providing the received audio data to the user and playing it back. This allows the user to listen to the mnemonics set to the melody, making it fun to memorize the information. The application also has the function of tracking the user's learning progress, recording and monitoring how much progress has been made.

[0388] This system not only enables users to memorize information efficiently, but also enables them to effectively manage their learning progress. Below are some examples of specific prompt sentences.

[0389] (Example of a prompt)

[0390] 1. Chat AI prompt:

[0391] Turn the following information into a mnemonic: The order of the planets in the solar system

[0392] 2. Prompts for automatic composition AI:

[0393] Compose a melody based on the following puns: Tai (tai), Mizu (sui), Kin (kin), Chi (chi), Ka (hi), Moku (ki), Tsuchi (do), Ten (ten), Umi (kai)

[0394] As described above, the system of the present invention allows users to memorize information while having fun and effectively manage their learning progress.

[0395] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0396] Step 1:

[0397] The user enters the information they want to remember.

[0398] The user launches the smartphone application and inputs the information they want to remember. The input information is sent to the server as "information to remember." The input here is in text format, such as "the order of the planets in the solar system."

[0399] Step 2:

[0400] The server sends the information to the chat AI.

[0401] The server sends the received information to a chat-based AI, which analyzes the input information and generates easy-to-remember mnemonics. The AI ​​uses a generative AI model to analyze the data and output mnemonics such as "tai (tai), water (mizu), kin (kin), chi (chi), ka (hi), ki (moku), tsuchi (tsuchi), ten (ten), umi (umi)."

[0402] Step 3:

[0403] The generated mnemonics are sent from the server to an automatic composition AI.

[0404] The server receives the mnemonic returned by the chat AI and sends it to the automatic composition AI. The automatic composition AI uses the mnemonic as input data to generate an appropriate melody. The melody is generated using a prompt sentence.

[0405] Step 4:

[0406] An automatic composition AI generates a melody and returns it as audio data.

[0407] The automatic composition AI creates a melody based on the received mnemonic and generates it as audio data. This audio data is sent back to the server in a playable format such as MP3. The output audio data is provided to the user in a format that allows them to learn in an enjoyable way.

[0408] Step 5:

[0409] The server sends the audio data back to the smartphone application.

[0410] The server sends the audio data received from the automatic composition AI back to the smartphone application, at which point the transmission of the audio data is complete.

[0411] Step 6:

[0412] The smartphone application plays the audio data and provides it to the user.

[0413] The smartphone application plays the received audio data and provides it to the user. By listening to the mnemonics set to the melody, the user can memorize information in a fun and efficient way. Data processing and calculations are performed in real time during playback.

[0414] Step 7:

[0415] A smartphone application tracks learning progress.

[0416] The smartphone application tracks the user's learning progress and records their achievements and progress, allowing them to effectively manage their own learning. Data is periodically sent to a server and stored.

[0417] The above are the specific processing steps of the "Melody Learning Assistant" system.

[0418] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0419] This paper describes a system that supports learning by inputting information that a user wants to remember and converting it into mnemonics and melodies, and further combines this with an emotion engine that recognizes the user's emotions, enabling individually optimized learning support. This system is composed of a terminal, chat AI, automatic composition AI, and the emotion engine.

[0420] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. The device then sends the input information to a chat-based AI on the server. A unique feature of the device is that it is equipped with an emotion engine that analyzes emotional information (for example, voice tone, facial expression, input content, etc.) during the user's input process. The results of this analysis are sent to the chat-based AI and the automatic composition AI, and are used to adjust the generated mnemonics and melodies.

[0421] The chat AI analyzes the received information and generates appropriate mnemonics. The mnemonics are a concise and easy-to-remember version of the original information, such as "tai, water, gold, earth, fire, wood, earth, heaven, sea." Based on the analysis results of the emotion engine, the mnemonics can be adjusted according to the user's emotional state.

[0422] The generated mnemonic is sent back to the device. The device then sends the mnemonic and the analysis results of the emotion engine to an automatic composition AI on the server. The automatic composition AI generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. Furthermore, based on the results of the emotion engine, the tempo and key of the melody are adjusted to match the user's emotional state.

[0423] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user. The user can play the audio data on the device and listen to the mnemonics that go with the melody. The user can also download the data and listen to it repeatedly.

[0424] For example, if a user types "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and emotional information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0425] This system allows users to easily memorize information they want to remember in an individually optimized way, improving learning efficiency. The introduction of an emotion engine makes it possible to further customize the system to match the user's emotional state, resulting in more effective learning support.

[0426] The processing flow will be explained below.

[0427] Step 1:

[0428] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0429] Step 2:

[0430] As soon as the device receives the information entered by the user, it activates an emotion engine to analyze the user's emotional state, based on the user's tone of voice, facial expression, and typing speed.

[0431] Step 3:

[0432] The device sends the input information and the emotional state analyzed by the emotion engine to the chat AI on the server, and the data includes emotional information.

[0433] Step 4:

[0434] The chat AI on the server analyzes the received information and generates appropriate mnemonics. For example, it creates mnemonics such as "tai (sea), water, gold (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)." It also adjusts the expressions of the mnemonics based on emotional information.

[0435] Step 5:

[0436] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0437] Step 6:

[0438] The device then sends the received mnemonic and the analysis results of the emotion engine to the automatic composition AI on the server. The transmission protocol is the same as in step 3.

[0439] Step 7:

[0440] The automatic composition AI on the server analyzes the received mnemonics and emotional information. Based on this, it generates a melody that matches the mnemonics. This melody is generated in a format that incorporates the mnemonics as lyrics. It also adjusts the tempo and key of the melody based on the emotional information.

[0441] Step 8:

[0442] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0443] Step 9:

[0444] The device then provides the received audio data to the user. The user can then play the audio data on the device and listen to the mnemonics set to the melody. The audio data is also provided in a downloadable format, so it can be listened to repeatedly.

[0445] Step 10:

[0446] By repeatedly listening to audio data, users can efficiently memorize the information they want to remember. This process improves learning effectiveness and is personalized to their emotional state, resulting in more effective learning.

[0447] Example 2

[0448] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0449] Conventional learning support systems lack the ingenuity to help users effectively memorize the information they want to remember, and in particular, they lack individual optimization based on the user's emotional state, which results in a decrease in learning efficiency.In addition, there is a demand for a method that makes learning fun rather than simply memorizing.

[0450] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be memorized from a user and transmitting the information to a natural language processing system; means for the natural language processing system to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic and the user's emotional information to an automatic composition system; means for the automatic composition system to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for analyzing the user's emotions and transmitting the analysis results to the natural language processing system and the automatic composition system; and means for providing the received audio data to the user. This enables individually optimized learning support tailored to the user's emotional state, resulting in enjoyable and efficient memorization.

[0451] "User" refers to anyone who wants to use the system to remember information.

[0452] "Information to be remembered" refers to specific data or facts that the user wants to remember.

[0453] A "natural language processing system" refers to an artificial intelligence-based processing system that analyzes input text and generates appropriate output.

[0454] A mnemonic is a sentence or phrase designed to make it easier to remember information.

[0455] An "automated composition system" refers to an artificial intelligence-based system that can generate melodies and music based on input linguistic data.

[0456] "Audio data" refers to the data of an acoustic signal that combines the generated melody and mnemonic.

[0457] "Emotional information" is data that indicates the user's emotional state during input, and refers to the analysis results obtained from voice tone, facial expression, input content, etc.

[0458] This system supports learning by allowing users to input information they want to remember and converting it into mnemonics and melodies. By combining this with an emotion engine that recognizes the user's emotions, it enables individually optimized learning support. This system is comprised of a terminal, a natural language processing system, an automatic composition system, and an emotion recognition engine.

[0459] First, the user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." In this case, it is desirable for the device to have an interface such as voice input or touch panel input.

[0460] Next, the device sends the input information to a natural language processing system on a server. For example, GPT-3 or GPT-4 is used for the natural language processing system. During this process, the device's built-in emotion recognition engine recognizes and analyzes the user's emotions from voice tone, facial expressions, etc. The analysis results are sent to the natural language processing system or automatic composition system, making it possible to provide output tailored to the user's emotional state.

[0461] The natural language processing system analyzes the received information and generates an appropriate mnemonic. For example, it converts the information "the order of the planets in the solar system" into "tai, water, gold, earth, fire, wood, earth, sky, sea." This mnemonic may then be adjusted according to the user's emotional state based on the analysis results of an emotion recognition engine.

[0462] The generated mnemonic is sent back to the device, which then sends the mnemonic and the user's emotion analysis results to an automatic composition system on the server. Examples of automatic composition systems include AWS DeepComposer and Magenta. The automatic composition system generates a melody that matches the mnemonic and generates audio data that incorporates this melody as lyrics. The tempo and key of the melody are also adjusted based on the analysis results of the emotion recognition engine.

[0463] The audio data generated by the automatic composition system is usually sent to the device in MP3 or WAV format. The device receives this audio data and provides it to the user. The user can play the audio data on their device and listen to the mnemonics that go with the melody, improving memorization efficiency by integrating visual and auditory information. The audio data can also be downloaded and listened to repeatedly.

[0464] Example prompt sentence:

[0465] "Please explain in detail the process by which the system inputs information that the user wants to remember and converts it into a mnemonic and melody."

[0466] In this way, the present invention provides efficient and enjoyable learning support that incorporates customization based on the user's emotions, and can improve upon the problems of conventional learning support systems.

[0467] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0468] Step 1:

[0469] The user inputs the information they want to remember into the terminal.

[0470] Input: Information you want to remember (e.g., "I want to remember the order of the planets in the solar system")

[0471] Output: User input data

[0472] Specific operation: The user inputs text information using the device's interface, and the device stores the input data in its internal memory.

[0473] Step 2:

[0474] The terminal transmits the input information to a natural language processing system.

[0475] Input: User-entered data

[0476] Output: Data (text data) to be sent to the natural language processing system

[0477] Specific operation: The device acquires user input data stored in its internal memory and transmits it to the server's natural language processing system via the Internet.

[0478] Step 3:

[0479] The device analyzes the user's emotional information and sends it to a natural language processing system.

[0480] Input: User's tone of voice, facial expressions, and input

[0481] Output: Sentiment analysis data

[0482] Specific operation: The device uses the emotion engine to analyze emotional information such as voice tone and facial expressions. The analysis results are sent to the server's natural language processing system as emotion analysis data.

[0483] Step 4:

[0484] The server generates mnemonics using a natural language processing system.

[0485] Input: User input data, sentiment analysis data

[0486] Output: Mnemonic data

[0487] Specific operation: The natural language processing system analyzes user input data and generates easy-to-remember mnemonics. The mnemonics are adjusted based on sentiment analysis data. The generated mnemonics are sent back to the device from the server.

[0488] Step 5:

[0489] The terminal again transmits the mnemonic and emotional information to the automatic composition system on the server.

[0490] Input: mnemonic data, sentiment analysis data

[0491] Output: Data to be sent to the automated composition system

[0492] Specific operation: The device combines the mnemonics received from the server with the emotion analysis data and sends it to the automatic composition system.

[0493] Step 6:

[0494] The server generates a melody using an automatic composition system.

[0495] Input: mnemonic data, sentiment analysis data

[0496] Output: Audio data

[0497] Specific operation: The automatic composition system generates a melody that matches the user's emotional state based on the received mnemonic. Audio data is then generated with the mnemonic incorporated into the melody as lyrics.

[0498] Step 7:

[0499] The server transmits the generated voice data to the terminal.

[0500] Input: Audio data

[0501] Output: Audio data sent to the device (e.g. MP3 or WAV format)

[0502] Specific operation: The server sends the generated audio data back to the device, usually in MP3 or WAV format.

[0503] Step 8:

[0504] The terminal provides the audio data to the user.

[0505] Input: Audio data

[0506] Output: Playable audio data

[0507] Specific operation: The device uses playback software to provide the received audio data to the user. The user plays the audio data and listens to the mnemonic set to the melody. The user can also download the audio data and listen to it repeatedly.

[0508] (Application example 2)

[0509] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0510] Conventional learning support systems do not provide a method for users to efficiently memorize information. Furthermore, they are unable to optimize the content of learning support according to the user's emotional state. As a result, learning efficiency is low and it is difficult for users to easily memorize information. The present invention aims to solve these problems by analyzing the user's emotions and providing optimized learning support based on the analysis.

[0511] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be remembered from a user and transmitting the information to a chat-type AI; means for the chat-type AI to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic to an automatic composition-type AI; means for the automatic composition-type AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for providing the received audio data to the user; and means for incorporating an emotion engine that analyzes the user's emotions and optimizing a series of processes based on the emotion information analyzed by the emotion engine. This enables the user to memorize information in a form optimized according to their emotional state.

[0512] A "user" is a person who uses the system to input information and receive learning support.

[0513] "Information to be remembered" is content or data that the user wants to remember.

[0514] "Chat AI" refers to algorithms or programs that generate mnemonics based on the information they receive.

[0515] A mnemonic is a simple word that uses sound and rhythm to make difficult-to-remember information easier to remember.

[0516] "Automatic composition artificial intelligence" refers to algorithms and programs that generate melodies suitable for mnemonics and incorporate them as audio data.

[0517] A "melody" is a musical melody, a series of sounds generated according to a mnemonic.

[0518] "Audio data" refers to digital sound information including the generated melody.

[0519] An "emotion engine" refers to an algorithm or program that analyzes the user's emotional state and optimizes the system's operation based on that information.

[0520] "Playback" refers to the act of letting the user listen to the generated audio data.

[0521] A "downloadable format" is a digital data format that allows a user to save and later play the content.

[0522] This invention relates to a learning support system that inputs information that a user wants to remember and converts it into mnemonics and melodies. The system is composed of the following elements: a terminal, chat AI, automatic composition AI, and an emotion engine. This system provides personalized and optimized learning support according to the user's emotional state.

[0523] First, the user inputs the information they want to remember into a device. This device could be a smartphone, smart glasses, or a head-mounted display. The input information is sent to a chat-based AI on a server. The chat-based AI generates an appropriate mnemonic based on the received information.

[0524] Next, the device is equipped with an emotion engine that analyzes the emotional information as the user inputs. The analysis results are sent to chat AI and automatic composition AI. The emotion engine determines the user's emotions based on voice tone, facial expressions, input content, etc. Emotion engines that can be used include Affectiva and Microsoft Azure Emotional Analysis API.

[0525] The generated mnemonics and emotional information are sent to an automatic composition AI, which then generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonics as lyrics. Possible AIs that could be used include OpenAI's MuseNet.

[0526] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user and plays it back. A smartphone's Pygame library is an effective way to play the audio. Users can also download the audio data and listen to it repeatedly.

[0527] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and the emotion information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0528] An example prompt might have the following format:

[0529] Chat AI prompt:

[0530] User input: I want to remember the order of the planets in the solar system.

[0531] User emotion: Excited

[0532] Prompt for automatic composition AI:

[0533] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0534] User emotion: Excited

[0535] In the above-described manner, the present invention enables learning support that is optimized for the emotional state of the user.

[0536] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0537] Step 1:

[0538] The user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." The device then sends this input information to a chat-based AI on the server.

[0539] Step 2:

[0540] The chat AI on the server analyzes the received information and generates a mnemonic. For example, in response to the input information "I want to remember the order of the planets in the solar system," it generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is then sent back to the device.

[0541] Step 3:

[0542] At the same time, the device uses an emotion engine to analyze the emotional information obtained during the user's input process. For example, it may analyze the user's tone of voice and facial expression as "excited." This emotional information is then sent to the server.

[0543] Step 4:

[0544] The device sends the received mnemonic and emotional information to the automatic composition AI on the server, using the following format as a prompt:

[0545] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0546] User emotion: Excited

[0547] Step 5:

[0548] The automatic composition AI on the server generates an appropriate melody based on mnemonics and emotional information. For example, it creates a bright, fast-paced melody for the mnemonic "tai (sea), water, gold (sea), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), sea (ocean)." This melody is generated as audio data and sent back to the device.

[0549] Step 6:

[0550] The device provides the received audio data to the user. The audio data is usually in MP3 or WAV format and is played on the device. For example, you can play audio using the Pygame library.

[0551] Step 7:

[0552] The user can listen to the audio data played on the device and confirm that the mnemonic is set to the melody. The user can also download the audio data and play it as many times as they like.

[0553] These steps allow the user to remember information in a way that is optimized for their emotional state.

[0554] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0555] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0556] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0557] [Third embodiment]

[0558] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0559] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0560] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0561] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0562] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0563] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0564] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0565] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0566] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0567] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0568] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0569] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0570] This invention describes a system that assists learning by inputting information that a user wants to remember and converting it into mnemonics and melodies. This system is composed of a terminal, a chat system AI, and an automatic composition system AI.

[0571] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. Once input is complete, the device sends this information to a chat-based AI located on a server. The chat-based AI analyzes the received information and generates an appropriate mnemonic. The mnemonic is a concise and easy-to-remember version of the original information (for example, "tai, water, gold, earth, fire, wood, earth, heaven, sea").

[0572] The generated mnemonic is sent back to the device. The device then sends the mnemonic to an automatic composition AI on the server. The automatic composition AI creates an appropriate melody based on the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The audio data is generated in a music file format (e.g., MP3 format) and sent back to the device.

[0573] The terminal provides the received audio data to the user. By playing the audio data, the user can listen to the mnemonic set to the melody, allowing the user to memorize the information they want to remember in an enjoyable and efficient way.

[0574] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the automatic composition AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who will play it back and study.

[0575] This system allows users to easily memorize information, improving learning efficiency. Furthermore, learning becomes more enjoyable because information is more easily retained in memory in the form of music.

[0576] The processing flow will be explained below.

[0577] Step 1:

[0578] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0579] Step 2:

[0580] The device receives the input information and sends it to the chat AI, which can use protocols such as APIs.

[0581] Step 3:

[0582] The chat AI on the server analyzes the information it receives. For example, if the input is about the order of the planets in the solar system, it will extract the necessary data.

[0583] Step 4:

[0584] The chat AI generates simple and easy-to-remember mnemonics based on the extracted data, such as "tai (sea), water, gold (metal), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)."

[0585] Step 5:

[0586] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0587] Step 6:

[0588] The device then sends the received mnemonics to the automatic composition AI on the server. The transmission protocol is the same as in step 2.

[0589] Step 7:

[0590] The automatic composition AI analyzes the received mnemonics and generates a suitable melody that incorporates the mnemonics as lyrics.

[0591] Step 8:

[0592] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0593] Step 9:

[0594] The terminal provides the received audio data to the user, who can then play the audio data on the terminal and listen to the mnemonic set to the melody.

[0595] Step 10:

[0596] By repeatedly listening to the audio data, users can efficiently memorize the information they want to remember, and this process improves learning effectiveness.

[0597] Example 1

[0598] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0599] Conventional learning support systems simply present the information users want to remember as text or images, making it difficult for them to memorize the information effectively and in an enjoyable way. Furthermore, they lacked an appealing method to help users solidify their memories. While memorizing information accompanied by music is expected to help users remember it longer, no such system existed. Therefore, a method was needed to enable users to learn the information they want to memorize efficiently and in an enjoyable way.

[0600] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0601] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a natural language processing AI, means for the natural language processing AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to a music generation AI, means for the music generation AI to generate music suitable for the mnemonic, generate audio data incorporating the music into the mnemonic, and return the audio data, and means for providing the received audio data to the user. This allows the user to learn the information they want to memorize in an enjoyable and effective way.

[0602] "User" refers to the entity that inputs the information they want to remember into the system and learns.

[0603] "Terminal" refers to a device that allows a user to input information or receive voice data. Examples include smartphones, tablets, and PCs.

[0604] "Server" refers to the central part of the computer network where the chat AI and automatic composition AI run.

[0605] "Natural language processing AI" refers to artificial intelligence programs designed to understand and analyze human language. Examples include GPT-3.

[0606] "Generative music AI" refers to artificial intelligence programs that generate music based on given text or information. Examples include Jukedeck and Amper Music.

[0607] "Information to be remembered" refers to specific data or knowledge that the user wants to remember. For example, "the order of the planets in the solar system."

[0608] A mnemonic is a phrase that converts information you want to remember into a concise and easy-to-remember form.

[0609] "Audio Data" refers to data formats containing music generated by automatic music composition AI, such as MP3 format.

[0610] This invention relates to a system that assists learning by converting information that a user wants to remember into mnemonics and melodies. This system is composed of a terminal, a natural language processing AI, and a music generation AI.

[0611] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into a terminal. The terminal can be a smartphone, tablet, PC, or the like. Once the user inputs the information, the terminal has a means for sending this information to a server. This means sends the information to the server using a protocol such as an HTTP request.

[0612] A natural language processing AI (e.g., GPT-3) runs on the server, analyzes the received information, and generates an appropriate mnemonic. For example, if a user inputs, "I want to remember the order of the planets in the solar system," the natural language processing AI generates a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is stored in memory and sent back to the device.

[0613] The device then sends the acquired mnemonic back to the server, this time passing it on to a music generation AI (e.g., Jukedeck or Amper Music). The music generation AI then creates an appropriate melody based on the mnemonic. Specifically, it incorporates the generated mnemonic as lyrics and generates the audio in an audio data format (e.g., MP3 format).

[0614] The generated audio data is sent back from the server to the device. The device has the means to store the audio data in local storage and provide it to the user. The user can play this audio data using the device's music playback application. This allows the user to enjoy learning while listening to mnemonics set to the melody.

[0615] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system" into a device, the natural language processing AI will generate a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the music generation AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who can play it back to study.

[0616] An example of a prompt is as follows:

[0617] User: I want to remember the order of the planets in the solar system.

[0618] This system allows users to memorize information in a fun and efficient way, improving learning efficiency.In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0619] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0620] Step 1:

[0621] The user inputs the information they want to remember into the device. For example, they input "I want to remember the order of the planets in the solar system." This input data is saved in an input field on the device.

[0622] Step 2:

[0623] The device sends the information entered to the server. Specifically, it creates an HTTP request and sends the information entered by the user to the server's natural language processing AI. The input is the information entered by the user, and the output is a notification to the server that transmission has been completed.

[0624] Step 3:

[0625] The server's natural language processing AI analyzes the information sent. A chat-based AI (e.g., GPT-3) analyzes the user's input and generates an appropriate mnemonic. For example, an input such as "I want to remember the order of the planets in the solar system" is converted into a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." The input is the information sent by the user, and the output is the generated mnemonic.

[0626] Step 4:

[0627] The server sends the generated mnemonic to the terminal. Specifically, it returns this mnemonic as an HTTP response. The terminal receives this response. The input is the generated mnemonic, and the output is a notification to the terminal that transmission has been completed.

[0628] Step 5:

[0629] The device sends the received mnemonic to the music generation AI on the server. The generated mnemonic is sent again to the server, this time requesting processing from the music generation AI. Specifically, an HTTP request is created and the mnemonic data is sent to the music generation AI endpoint. The input is the received mnemonic, and the output is a notification to the server that transmission has been completed.

[0630] Step 6:

[0631] The server's music generation AI generates a melody that matches the mnemonic. Specifically, it creates a melody based on this mnemonic and generates audio data that incorporates the mnemonic as lyrics. For example, it generates audio data (MP3 format, etc.) with the lyrics "tai, water, gold, earth, fire, wood, earth, heaven, sea" set to a melody. The input is the mnemonic, and the output is the generated audio data.

[0632] Step 7:

[0633] The server sends the generated audio data to the terminal. Specifically, it returns the generated audio data to the terminal as an HTTP response. The input is the generated audio data, and the output is a transmission completion notification to the terminal.

[0634] Step 8:

[0635] The device provides the received audio data to the user. The audio data stored on the device can be played back upon user request. The user can listen to this audio data using the device's music playback application. The input is the received audio data, and the output is the audio data provided in a playable format.

[0636] This system allows users to memorize information in a fun and efficient way, improving learning efficiency. In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0637] (Application example 1)

[0638] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0639] While there are many methods for efficiently memorizing information, there are only a limited number of ways for users to study in an enjoyable way. Furthermore, there are not enough methods for effectively tracking learning progress. Therefore, there is a need for a method to maintain motivation while solidifying information in memory.

[0640] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0641] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a chat-based AI, means for the chat-based AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to an automatic composition-based AI, means for the automatic composition-based AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data, means for providing the received audio data to the user, means for playing the audio data through a smartphone application, and means for tracking the user's learning progress. This allows the user to memorize information in a fun and efficient way, and enables the user's learning progress to be effectively tracked.

[0642] "User" refers to a person who uses the system of the present invention to learn information.

[0643] "Memorable information" refers to facts or data that a user wishes to remember.

[0644] "Chat AI" is AI that conducts dialogue based on input data, analyzes it, and generates information.

[0645] A mnemonic is a conversion of original information into easy-to-remember words or phrases.

[0646] "Automatic composition artificial intelligence" is an artificial intelligence that has the ability to automatically create music.

[0647] A "melody" is a series of notes combined together to form a section of music that is aurally pleasing.

[0648] "Audio data" is data in the form of a digital file that contains audio information.

[0649] A "smartphone application" is a type of software that runs on a smartphone.

[0650] "Study Progress Tracking" means the act of continuously recording and monitoring a user's learning progress.

[0651] The "Melody Learning Assistant" system of the present invention is designed to help users memorize information they want to remember in a fun and effective way. First, the user inputs the information they want to memorize using a smartphone application. This information is called "information to memorize." Once the user inputs the information, it is sent from the device to a server.

[0652] A "chat AI" is placed on the server, and this AI analyzes the input information and generates easy-to-remember "mnemonics." The chat AI creates mnemonics using a generative AI model. These mnemonics convert the original information into a concise and easy-to-remember format. For example, if a user inputs "I want to remember the order of the planets in the solar system," the chat AI generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi."

[0653] The generated mnemonics are stored on the server and then sent to an "automatic composition AI." This automatic composition AI also uses a generative AI model to create a "melody" that suits the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The generated audio data is generated in a music file format such as MP3 and sent back from the server to the device.

[0654] The smartphone application has the function of providing the received audio data to the user and playing it back. This allows the user to listen to the mnemonics set to the melody, making it fun to memorize the information. The application also has the function of tracking the user's learning progress, recording and monitoring how much progress has been made.

[0655] This system not only enables users to memorize information efficiently, but also enables them to effectively manage their learning progress. Below are some examples of specific prompt sentences.

[0656] (Example of a prompt)

[0657] 1. Chat AI prompt:

[0658] Turn the following information into a mnemonic: The order of the planets in the solar system

[0659] 2. Prompts for automatic composition AI:

[0660] Compose a melody based on the following puns: Tai (tai), Mizu (sui), Kin (kin), Chi (chi), Ka (hi), Moku (ki), Tsuchi (do), Ten (ten), Umi (kai)

[0661] As described above, the system of the present invention allows users to memorize information while having fun and effectively manage their learning progress.

[0662] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0663] Step 1:

[0664] The user enters the information they want to remember.

[0665] The user launches the smartphone application and inputs the information they want to remember. The input information is sent to the server as "information to remember." The input here is in text format, such as "the order of the planets in the solar system."

[0666] Step 2:

[0667] The server sends the information to the chat AI.

[0668] The server sends the received information to a chat-based AI, which analyzes the input information and generates easy-to-remember mnemonics. The AI ​​uses a generative AI model to analyze the data and output mnemonics such as "tai (tai), water (mizu), kin (kin), chi (chi), ka (hi), ki (moku), tsuchi (tsuchi), ten (ten), umi (umi)."

[0669] Step 3:

[0670] The generated mnemonics are sent from the server to an automatic composition AI.

[0671] The server receives the mnemonic returned by the chat AI and sends it to the automatic composition AI. The automatic composition AI uses the mnemonic as input data to generate an appropriate melody. The melody is generated using a prompt sentence.

[0672] Step 4:

[0673] An automatic composition AI generates a melody and returns it as audio data.

[0674] The automatic composition AI creates a melody based on the received mnemonic and generates it as audio data. This audio data is sent back to the server in a playable format such as MP3. The output audio data is provided to the user in a format that allows them to learn in an enjoyable way.

[0675] Step 5:

[0676] The server sends the audio data back to the smartphone application.

[0677] The server sends the audio data received from the automatic composition AI back to the smartphone application, at which point the transmission of the audio data is complete.

[0678] Step 6:

[0679] The smartphone application plays the audio data and provides it to the user.

[0680] The smartphone application plays the received audio data and provides it to the user. By listening to the mnemonics set to the melody, the user can memorize information in a fun and efficient way. Data processing and calculations are performed in real time during playback.

[0681] Step 7:

[0682] A smartphone application tracks learning progress.

[0683] The smartphone application tracks the user's learning progress and records their achievements and progress, allowing them to effectively manage their own learning. Data is periodically sent to a server and stored.

[0684] The above are the specific processing steps of the "Melody Learning Assistant" system.

[0685] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0686] This paper describes a system that supports learning by inputting information that a user wants to remember and converting it into mnemonics and melodies, and further combines this with an emotion engine that recognizes the user's emotions, enabling individually optimized learning support. This system is composed of a terminal, chat AI, automatic composition AI, and the emotion engine.

[0687] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. The device then sends the input information to a chat-based AI on the server. A unique feature of the device is that it is equipped with an emotion engine that analyzes emotional information (for example, voice tone, facial expression, input content, etc.) during the user's input process. The results of this analysis are sent to the chat-based AI and the automatic composition AI, and are used to adjust the generated mnemonics and melodies.

[0688] The chat AI analyzes the received information and generates appropriate mnemonics. The mnemonics are a concise and easy-to-remember version of the original information, such as "tai, water, gold, earth, fire, wood, earth, heaven, sea." Based on the analysis results of the emotion engine, the mnemonics can be adjusted according to the user's emotional state.

[0689] The generated mnemonic is sent back to the device. The device then sends the mnemonic and the analysis results of the emotion engine to an automatic composition AI on the server. The automatic composition AI generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. Furthermore, based on the results of the emotion engine, the tempo and key of the melody are adjusted to match the user's emotional state.

[0690] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user. The user can play the audio data on the device and listen to the mnemonics that go with the melody. The user can also download the data and listen to it repeatedly.

[0691] For example, if a user types "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and emotional information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0692] This system allows users to easily memorize information they want to remember in an individually optimized way, improving learning efficiency. The introduction of an emotion engine makes it possible to further customize the system to match the user's emotional state, resulting in more effective learning support.

[0693] The processing flow will be explained below.

[0694] Step 1:

[0695] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0696] Step 2:

[0697] As soon as the device receives the information entered by the user, it activates an emotion engine to analyze the user's emotional state, based on the user's tone of voice, facial expression, and typing speed.

[0698] Step 3:

[0699] The device sends the input information and the emotional state analyzed by the emotion engine to the chat AI on the server, and the data includes emotional information.

[0700] Step 4:

[0701] The chat AI on the server analyzes the received information and generates appropriate mnemonics. For example, it creates mnemonics such as "tai (sea), water, gold (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)." It also adjusts the expressions of the mnemonics based on emotional information.

[0702] Step 5:

[0703] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0704] Step 6:

[0705] The device then sends the received mnemonic and the analysis results of the emotion engine to the automatic composition AI on the server. The transmission protocol is the same as in step 3.

[0706] Step 7:

[0707] The automatic composition AI on the server analyzes the received mnemonics and emotional information. Based on this, it generates a melody that matches the mnemonics. This melody is generated in a format that incorporates the mnemonics as lyrics. It also adjusts the tempo and key of the melody based on the emotional information.

[0708] Step 8:

[0709] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0710] Step 9:

[0711] The device then provides the received audio data to the user. The user can then play the audio data on the device and listen to the mnemonics set to the melody. The audio data is also provided in a downloadable format, so it can be listened to repeatedly.

[0712] Step 10:

[0713] By repeatedly listening to audio data, users can efficiently memorize the information they want to remember. This process improves learning effectiveness and is personalized to their emotional state, resulting in more effective learning.

[0714] Example 2

[0715] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0716] Conventional learning support systems lack the ingenuity to help users effectively memorize the information they want to remember, and in particular, they lack individual optimization based on the user's emotional state, which results in a decrease in learning efficiency.In addition, there is a demand for a method that makes learning fun rather than simply memorizing.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be memorized from a user and transmitting the information to a natural language processing system; means for the natural language processing system to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic and the user's emotional information to an automatic composition system; means for the automatic composition system to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for analyzing the user's emotions and transmitting the analysis results to the natural language processing system and the automatic composition system; and means for providing the received audio data to the user. This enables individually optimized learning support tailored to the user's emotional state, resulting in enjoyable and efficient memorization.

[0718] "User" refers to anyone who wants to use the system to remember information.

[0719] "Information to be remembered" refers to specific data or facts that the user wants to remember.

[0720] A "natural language processing system" refers to an artificial intelligence-based processing system that analyzes input text and generates appropriate output.

[0721] A mnemonic is a sentence or phrase designed to make it easier to remember information.

[0722] An "automated composition system" refers to an artificial intelligence-based system that can generate melodies and music based on input linguistic data.

[0723] "Audio data" refers to the data of an acoustic signal that combines the generated melody and mnemonic.

[0724] "Emotional information" is data that indicates the user's emotional state during input, and refers to the analysis results obtained from voice tone, facial expression, input content, etc.

[0725] This system supports learning by allowing users to input information they want to remember and converting it into mnemonics and melodies. By combining this with an emotion engine that recognizes the user's emotions, it enables individually optimized learning support. This system is comprised of a terminal, a natural language processing system, an automatic composition system, and an emotion recognition engine.

[0726] First, the user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." In this case, it is desirable for the device to have an interface such as voice input or touch panel input.

[0727] Next, the device sends the input information to a natural language processing system on a server. For example, GPT-3 or GPT-4 is used for the natural language processing system. During this process, the device's built-in emotion recognition engine recognizes and analyzes the user's emotions from voice tone, facial expressions, etc. The analysis results are sent to the natural language processing system or automatic composition system, making it possible to provide output tailored to the user's emotional state.

[0728] The natural language processing system analyzes the received information and generates an appropriate mnemonic. For example, it converts the information "the order of the planets in the solar system" into "tai, water, gold, earth, fire, wood, earth, sky, sea." This mnemonic may then be adjusted according to the user's emotional state based on the analysis results of an emotion recognition engine.

[0729] The generated mnemonic is sent back to the device, which then sends the mnemonic and the user's emotion analysis results to an automatic composition system on the server. Examples of automatic composition systems include AWS DeepComposer and Magenta. The automatic composition system generates a melody that matches the mnemonic and generates audio data that incorporates this melody as lyrics. The tempo and key of the melody are also adjusted based on the analysis results of the emotion recognition engine.

[0730] The audio data generated by the automatic composition system is usually sent to the device in MP3 or WAV format. The device receives this audio data and provides it to the user. The user can play the audio data on their device and listen to the mnemonics that go with the melody, improving memorization efficiency by integrating visual and auditory information. The audio data can also be downloaded and listened to repeatedly.

[0731] Example prompt sentence:

[0732] "Please explain in detail the process by which the system inputs information that the user wants to remember and converts it into a mnemonic and melody."

[0733] In this way, the present invention provides efficient and enjoyable learning support that incorporates customization based on the user's emotions, and can improve upon the problems of conventional learning support systems.

[0734] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0735] Step 1:

[0736] The user inputs the information they want to remember into the terminal.

[0737] Input: Information you want to remember (e.g., "I want to remember the order of the planets in the solar system")

[0738] Output: User input data

[0739] Specific operation: The user inputs text information using the device's interface, and the device stores the input data in its internal memory.

[0740] Step 2:

[0741] The terminal transmits the input information to a natural language processing system.

[0742] Input: User-entered data

[0743] Output: Data (text data) to be sent to the natural language processing system

[0744] Specific operation: The device acquires user input data stored in its internal memory and transmits it to the server's natural language processing system via the Internet.

[0745] Step 3:

[0746] The device analyzes the user's emotional information and sends it to a natural language processing system.

[0747] Input: User's tone of voice, facial expressions, and input

[0748] Output: Sentiment analysis data

[0749] Specific operation: The device uses the emotion engine to analyze emotional information such as voice tone and facial expressions. The analysis results are sent to the server's natural language processing system as emotion analysis data.

[0750] Step 4:

[0751] The server generates mnemonics using a natural language processing system.

[0752] Input: User input data, sentiment analysis data

[0753] Output: Mnemonic data

[0754] Specific operation: The natural language processing system analyzes user input data and generates easy-to-remember mnemonics. The mnemonics are adjusted based on sentiment analysis data. The generated mnemonics are sent back to the device from the server.

[0755] Step 5:

[0756] The terminal again transmits the mnemonic and emotional information to the automatic composition system on the server.

[0757] Input: mnemonic data, sentiment analysis data

[0758] Output: Data to be sent to the automated composition system

[0759] Specific operation: The device combines the mnemonics received from the server with the emotion analysis data and sends it to the automatic composition system.

[0760] Step 6:

[0761] The server generates a melody using an automatic composition system.

[0762] Input: mnemonic data, sentiment analysis data

[0763] Output: Audio data

[0764] Specific operation: The automatic composition system generates a melody that matches the user's emotional state based on the received mnemonic. Audio data is then generated with the mnemonic incorporated into the melody as lyrics.

[0765] Step 7:

[0766] The server transmits the generated voice data to the terminal.

[0767] Input: Audio data

[0768] Output: Audio data sent to the device (e.g. MP3 or WAV format)

[0769] Specific operation: The server sends the generated audio data back to the device, usually in MP3 or WAV format.

[0770] Step 8:

[0771] The terminal provides the audio data to the user.

[0772] Input: Audio data

[0773] Output: Playable audio data

[0774] Specific operation: The device uses playback software to provide the received audio data to the user. The user plays the audio data and listens to the mnemonic set to the melody. The user can also download the audio data and listen to it repeatedly.

[0775] (Application example 2)

[0776] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0777] Conventional learning support systems do not provide a method for users to efficiently memorize information. Furthermore, they are unable to optimize the content of learning support according to the user's emotional state. As a result, learning efficiency is low and it is difficult for users to easily memorize information. The present invention aims to solve these problems by analyzing the user's emotions and providing optimized learning support based on the analysis.

[0778] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be remembered from a user and transmitting the information to a chat-type AI; means for the chat-type AI to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic to an automatic composition-type AI; means for the automatic composition-type AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for providing the received audio data to the user; and means for incorporating an emotion engine that analyzes the user's emotions and optimizing a series of processes based on the emotion information analyzed by the emotion engine. This enables the user to memorize information in a form optimized according to their emotional state.

[0779] A "user" is a person who uses the system to input information and receive learning support.

[0780] "Information to be remembered" is content or data that the user wants to remember.

[0781] "Chat AI" refers to algorithms or programs that generate mnemonics based on the information they receive.

[0782] A mnemonic is a simple word that uses sound and rhythm to make difficult-to-remember information easier to remember.

[0783] "Automatic composition artificial intelligence" refers to algorithms and programs that generate melodies suitable for mnemonics and incorporate them as audio data.

[0784] A "melody" is a musical melody, a series of sounds generated according to a mnemonic.

[0785] "Audio data" refers to digital sound information including the generated melody.

[0786] An "emotion engine" refers to an algorithm or program that analyzes the user's emotional state and optimizes the system's operation based on that information.

[0787] "Playback" refers to the act of letting the user listen to the generated audio data.

[0788] A "downloadable format" is a digital data format that allows a user to save and later play the content.

[0789] This invention relates to a learning support system that inputs information that a user wants to remember and converts it into mnemonics and melodies. The system is composed of the following elements: a terminal, chat AI, automatic composition AI, and an emotion engine. This system provides personalized and optimized learning support according to the user's emotional state.

[0790] First, the user inputs the information they want to remember into a device. This device could be a smartphone, smart glasses, or a head-mounted display. The input information is sent to a chat-based AI on a server. The chat-based AI generates an appropriate mnemonic based on the received information.

[0791] Next, the device is equipped with an emotion engine that analyzes the emotional information as the user inputs. The analysis results are sent to chat AI and automatic composition AI. The emotion engine determines the user's emotions based on voice tone, facial expressions, input content, etc. Emotion engines that can be used include Affectiva and Microsoft Azure Emotional Analysis API.

[0792] The generated mnemonics and emotional information are sent to an automatic composition AI, which then generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonics as lyrics. Possible AIs that could be used include OpenAI's MuseNet.

[0793] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user and plays it back. A smartphone's Pygame library is an effective way to play the audio. Users can also download the audio data and listen to it repeatedly.

[0794] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and the emotion information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0795] An example prompt might have the following format:

[0796] Chat AI prompt:

[0797] User input: I want to remember the order of the planets in the solar system.

[0798] User emotion: Excited

[0799] Prompt for automatic composition AI:

[0800] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0801] User emotion: Excited

[0802] In the above-described manner, the present invention enables learning support that is optimized for the emotional state of the user.

[0803] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0804] Step 1:

[0805] The user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." The device then sends this input information to a chat-based AI on the server.

[0806] Step 2:

[0807] The chat AI on the server analyzes the received information and generates a mnemonic. For example, in response to the input information "I want to remember the order of the planets in the solar system," it generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is then sent back to the device.

[0808] Step 3:

[0809] At the same time, the device uses an emotion engine to analyze the emotional information obtained during the user's input process. For example, it may analyze the user's tone of voice and facial expression as "excited." This emotional information is then sent to the server.

[0810] Step 4:

[0811] The device sends the received mnemonic and emotional information to the automatic composition AI on the server, using the following format as a prompt:

[0812] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[0813] User emotion: Excited

[0814] Step 5:

[0815] The automatic composition AI on the server generates an appropriate melody based on mnemonics and emotional information. For example, it creates a bright, fast-paced melody for the mnemonic "tai (sea), water, gold (sea), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), sea (ocean)." This melody is generated as audio data and sent back to the device.

[0816] Step 6:

[0817] The device provides the received audio data to the user. The audio data is usually in MP3 or WAV format and is played on the device. For example, you can play audio using the Pygame library.

[0818] Step 7:

[0819] The user can listen to the audio data played on the device and confirm that the mnemonic is set to the melody. The user can also download the audio data and play it as many times as they like.

[0820] These steps allow the user to remember information in a way that is optimized for their emotional state.

[0821] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0822] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0823] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0824] [Fourth embodiment]

[0825] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0826] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0827] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0828] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0829] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0830] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0831] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0832] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0833] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0834] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0835] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0836] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0837] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0838] This invention describes a system that assists learning by inputting information that a user wants to remember and converting it into mnemonics and melodies. This system is composed of a terminal, a chat system AI, and an automatic composition system AI.

[0839] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. Once input is complete, the device sends this information to a chat-based AI located on a server. The chat-based AI analyzes the received information and generates an appropriate mnemonic. The mnemonic is a concise and easy-to-remember version of the original information (for example, "tai, water, gold, earth, fire, wood, earth, heaven, sea").

[0840] The generated mnemonic is sent back to the device. The device then sends the mnemonic to an automatic composition AI on the server. The automatic composition AI creates an appropriate melody based on the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The audio data is generated in a music file format (e.g., MP3 format) and sent back to the device.

[0841] The terminal provides the received audio data to the user. By playing the audio data, the user can listen to the mnemonic set to the melody, allowing the user to memorize the information they want to remember in an enjoyable and efficient way.

[0842] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the automatic composition AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who will play it back and study.

[0843] This system allows users to easily memorize information, improving learning efficiency. Furthermore, learning becomes more enjoyable because information is more easily retained in memory in the form of music.

[0844] The processing flow will be explained below.

[0845] Step 1:

[0846] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0847] Step 2:

[0848] The device receives the input information and sends it to the chat AI, which can use protocols such as APIs.

[0849] Step 3:

[0850] The chat AI on the server analyzes the information it receives. For example, if the input is about the order of the planets in the solar system, it will extract the necessary data.

[0851] Step 4:

[0852] The chat AI generates simple and easy-to-remember mnemonics based on the extracted data, such as "tai (sea), water, gold (metal), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)."

[0853] Step 5:

[0854] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0855] Step 6:

[0856] The device then sends the received mnemonics to the automatic composition AI on the server. The transmission protocol is the same as in step 2.

[0857] Step 7:

[0858] The automatic composition AI analyzes the received mnemonics and generates a suitable melody that incorporates the mnemonics as lyrics.

[0859] Step 8:

[0860] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0861] Step 9:

[0862] The terminal provides the received audio data to the user, who can then play the audio data on the terminal and listen to the mnemonic set to the melody.

[0863] Step 10:

[0864] By repeatedly listening to the audio data, users can efficiently memorize the information they want to remember, and this process improves learning effectiveness.

[0865] Example 1

[0866] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0867] Conventional learning support systems simply present the information users want to remember as text or images, making it difficult for them to memorize the information effectively and in an enjoyable way. Furthermore, they lacked an appealing method to help users solidify their memories. While memorizing information accompanied by music is expected to help users remember it longer, no such system existed. Therefore, a method was needed to enable users to learn the information they want to memorize efficiently and in an enjoyable way.

[0868] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0869] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a natural language processing AI, means for the natural language processing AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to a music generation AI, means for the music generation AI to generate music suitable for the mnemonic, generate audio data incorporating the music into the mnemonic, and return the audio data, and means for providing the received audio data to the user. This allows the user to learn the information they want to memorize in an enjoyable and effective way.

[0870] "User" refers to the entity that inputs the information they want to remember into the system and learns.

[0871] "Terminal" refers to a device that allows a user to input information or receive voice data. Examples include smartphones, tablets, and PCs.

[0872] "Server" refers to the central part of the computer network where the chat AI and automatic composition AI run.

[0873] "Natural language processing AI" refers to artificial intelligence programs designed to understand and analyze human language. Examples include GPT-3.

[0874] "Generative music AI" refers to artificial intelligence programs that generate music based on given text or information. Examples include Jukedeck and Amper Music.

[0875] "Information to be remembered" refers to specific data or knowledge that the user wants to remember. For example, "the order of the planets in the solar system."

[0876] A mnemonic is a phrase that converts information you want to remember into a concise and easy-to-remember form.

[0877] "Audio Data" refers to data formats containing music generated by automatic music composition AI, such as MP3 format.

[0878] This invention relates to a system that assists learning by converting information that a user wants to remember into mnemonics and melodies. This system is composed of a terminal, a natural language processing AI, and a music generation AI.

[0879] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into a terminal. The terminal can be a smartphone, tablet, PC, or the like. Once the user inputs the information, the terminal has a means for sending this information to a server. This means sends the information to the server using a protocol such as an HTTP request.

[0880] A natural language processing AI (e.g., GPT-3) runs on the server, analyzes the received information, and generates an appropriate mnemonic. For example, if a user inputs, "I want to remember the order of the planets in the solar system," the natural language processing AI generates a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is stored in memory and sent back to the device.

[0881] The device then sends the acquired mnemonic back to the server, this time passing it on to a music generation AI (e.g., Jukedeck or Amper Music). The music generation AI then creates an appropriate melody based on the mnemonic. Specifically, it incorporates the generated mnemonic as lyrics and generates the audio in an audio data format (e.g., MP3 format).

[0882] The generated audio data is sent back from the server to the device. The device has the means to store the audio data in local storage and provide it to the user. The user can play this audio data using the device's music playback application. This allows the user to enjoy learning while listening to mnemonics set to the melody.

[0883] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system" into a device, the natural language processing AI will generate a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." Next, the music generation AI will create a melody that matches this mnemonic, and generate and return audio data with the lyrics "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi" set to the melody. The device will provide this audio data to the user, who can play it back to study.

[0884] An example of a prompt is as follows:

[0885] User: I want to remember the order of the planets in the solar system.

[0886] This system allows users to memorize information in a fun and efficient way, improving learning efficiency.In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0887] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0888] Step 1:

[0889] The user inputs the information they want to remember into the device. For example, they input "I want to remember the order of the planets in the solar system." This input data is saved in an input field on the device.

[0890] Step 2:

[0891] The device sends the information entered to the server. Specifically, it creates an HTTP request and sends the information entered by the user to the server's natural language processing AI. The input is the information entered by the user, and the output is a notification to the server that transmission has been completed.

[0892] Step 3:

[0893] The server's natural language processing AI analyzes the information sent. A chat-based AI (e.g., GPT-3) analyzes the user's input and generates an appropriate mnemonic. For example, an input such as "I want to remember the order of the planets in the solar system" is converted into a mnemonic such as "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." The input is the information sent by the user, and the output is the generated mnemonic.

[0894] Step 4:

[0895] The server sends the generated mnemonic to the terminal. Specifically, it returns this mnemonic as an HTTP response. The terminal receives this response. The input is the generated mnemonic, and the output is a notification to the terminal that transmission has been completed.

[0896] Step 5:

[0897] The device sends the received mnemonic to the music generation AI on the server. The generated mnemonic is sent again to the server, this time requesting processing from the music generation AI. Specifically, an HTTP request is created and the mnemonic data is sent to the music generation AI endpoint. The input is the received mnemonic, and the output is a notification to the server that transmission has been completed.

[0898] Step 6:

[0899] The server's music generation AI generates a melody that matches the mnemonic. Specifically, it creates a melody based on this mnemonic and generates audio data that incorporates the mnemonic as lyrics. For example, it generates audio data (MP3 format, etc.) with the lyrics "tai, water, gold, earth, fire, wood, earth, heaven, sea" set to a melody. The input is the mnemonic, and the output is the generated audio data.

[0900] Step 7:

[0901] The server sends the generated audio data to the terminal. Specifically, it returns the generated audio data to the terminal as an HTTP response. The input is the generated audio data, and the output is a transmission completion notification to the terminal.

[0902] Step 8:

[0903] The device provides the received audio data to the user. The audio data stored on the device can be played back upon user request. The user can listen to this audio data using the device's music playback application. The input is the received audio data, and the output is the audio data provided in a playable format.

[0904] This system allows users to memorize information in a fun and efficient way, improving learning efficiency. In addition, since information is more easily retained in memory in the form of music, it is expected to make learning more enjoyable.

[0905] (Application example 1)

[0906] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0907] While there are many methods for efficiently memorizing information, there are only a limited number of ways for users to study in an enjoyable way. Furthermore, there are not enough methods for effectively tracking learning progress. Therefore, there is a need for a method to maintain motivation while solidifying information in memory.

[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0909] In this invention, the server includes means for receiving information to be memorized from a user and transmitting the information to a chat-based AI, means for the chat-based AI to generate a mnemonic based on the information and return the mnemonic, means for transmitting the received mnemonic to an automatic composition-based AI, means for the automatic composition-based AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data, means for providing the received audio data to the user, means for playing the audio data through a smartphone application, and means for tracking the user's learning progress. This allows the user to memorize information in a fun and efficient way, and enables the user's learning progress to be effectively tracked.

[0910] "User" refers to a person who uses the system of the present invention to learn information.

[0911] "Memorable information" refers to facts or data that a user wishes to remember.

[0912] "Chat AI" is AI that conducts dialogue based on input data, analyzes it, and generates information.

[0913] A mnemonic is a conversion of original information into easy-to-remember words or phrases.

[0914] "Automatic composition artificial intelligence" is an artificial intelligence that has the ability to automatically create music.

[0915] A "melody" is a series of notes combined together to form a section of music that is aurally pleasing.

[0916] "Audio data" is data in the form of a digital file that contains audio information.

[0917] A "smartphone application" is a type of software that runs on a smartphone.

[0918] "Study Progress Tracking" means the act of continuously recording and monitoring a user's learning progress.

[0919] The "Melody Learning Assistant" system of the present invention is designed to help users memorize information they want to remember in a fun and effective way. First, the user inputs the information they want to memorize using a smartphone application. This information is called "information to memorize." Once the user inputs the information, it is sent from the device to a server.

[0920] A "chat AI" is placed on the server, and this AI analyzes the input information and generates easy-to-remember "mnemonics." The chat AI creates mnemonics using a generative AI model. These mnemonics convert the original information into a concise and easy-to-remember format. For example, if a user inputs "I want to remember the order of the planets in the solar system," the chat AI generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi."

[0921] The generated mnemonics are stored on the server and then sent to an "automatic composition AI." This automatic composition AI also uses a generative AI model to create a "melody" that suits the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. The generated audio data is generated in a music file format such as MP3 and sent back from the server to the device.

[0922] The smartphone application has the function of providing the received audio data to the user and playing it back. This allows the user to listen to the mnemonics set to the melody, making it fun to memorize the information. The application also has the function of tracking the user's learning progress, recording and monitoring how much progress has been made.

[0923] This system not only enables users to memorize information efficiently, but also enables them to effectively manage their learning progress. Below are some examples of specific prompt sentences.

[0924] (Example of a prompt)

[0925] 1. Chat AI prompt:

[0926] Turn the following information into a mnemonic: The order of the planets in the solar system

[0927] 2. Prompts for automatic composition AI:

[0928] Compose a melody based on the following puns: Tai (tai), Mizu (sui), Kin (kin), Chi (chi), Ka (hi), Moku (ki), Tsuchi (do), Ten (ten), Umi (kai)

[0929] As described above, the system of the present invention allows users to memorize information while having fun and effectively manage their learning progress.

[0930] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0931] Step 1:

[0932] The user enters the information they want to remember.

[0933] The user launches the smartphone application and inputs the information they want to remember. The input information is sent to the server as "information to remember." The input here is in text format, such as "the order of the planets in the solar system."

[0934] Step 2:

[0935] The server sends the information to the chat AI.

[0936] The server sends the received information to a chat-based AI, which analyzes the input information and generates easy-to-remember mnemonics. The AI ​​uses a generative AI model to analyze the data and output mnemonics such as "tai (tai), water (mizu), kin (kin), chi (chi), ka (hi), ki (moku), tsuchi (tsuchi), ten (ten), umi (umi)."

[0937] Step 3:

[0938] The generated mnemonics are sent from the server to an automatic composition AI.

[0939] The server receives the mnemonic returned by the chat AI and sends it to the automatic composition AI. The automatic composition AI uses the mnemonic as input data to generate an appropriate melody. The melody is generated using a prompt sentence.

[0940] Step 4:

[0941] An automatic composition AI generates a melody and returns it as audio data.

[0942] The automatic composition AI creates a melody based on the received mnemonic and generates it as audio data. This audio data is sent back to the server in a playable format such as MP3. The output audio data is provided to the user in a format that allows them to learn in an enjoyable way.

[0943] Step 5:

[0944] The server sends the audio data back to the smartphone application.

[0945] The server sends the audio data received from the automatic composition AI back to the smartphone application, at which point the transmission of the audio data is complete.

[0946] Step 6:

[0947] The smartphone application plays the audio data and provides it to the user.

[0948] The smartphone application plays the received audio data and provides it to the user. By listening to the mnemonics set to the melody, the user can memorize information in a fun and efficient way. Data processing and calculations are performed in real time during playback.

[0949] Step 7:

[0950] A smartphone application tracks learning progress.

[0951] The smartphone application tracks the user's learning progress and records their achievements and progress, allowing them to effectively manage their own learning. Data is periodically sent to a server and stored.

[0952] The above are the specific processing steps of the "Melody Learning Assistant" system.

[0953] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0954] This paper describes a system that supports learning by inputting information that a user wants to remember and converting it into mnemonics and melodies, and further combines this with an emotion engine that recognizes the user's emotions, enabling individually optimized learning support. This system is composed of a terminal, chat AI, automatic composition AI, and the emotion engine.

[0955] First, the user inputs the information they want to remember (for example, "the order of the planets in the solar system") into the device. The device then sends the input information to a chat-based AI on the server. A unique feature of the device is that it is equipped with an emotion engine that analyzes emotional information (for example, voice tone, facial expression, input content, etc.) during the user's input process. The results of this analysis are sent to the chat-based AI and the automatic composition AI, and are used to adjust the generated mnemonics and melodies.

[0956] The chat AI analyzes the received information and generates appropriate mnemonics. The mnemonics are a concise and easy-to-remember version of the original information, such as "tai, water, gold, earth, fire, wood, earth, heaven, sea." Based on the analysis results of the emotion engine, the mnemonics can be adjusted according to the user's emotional state.

[0957] The generated mnemonic is sent back to the device. The device then sends the mnemonic and the analysis results of the emotion engine to an automatic composition AI on the server. The automatic composition AI generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonic as lyrics. Furthermore, based on the results of the emotion engine, the tempo and key of the melody are adjusted to match the user's emotional state.

[0958] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user. The user can play the audio data on the device and listen to the mnemonics that go with the melody. The user can also download the data and listen to it repeatedly.

[0959] For example, if a user types "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and emotional information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[0960] This system allows users to easily memorize information they want to remember in an individually optimized way, improving learning efficiency. The introduction of an emotion engine makes it possible to further customize the system to match the user's emotional state, resulting in more effective learning support.

[0961] The processing flow will be explained below.

[0962] Step 1:

[0963] The user inputs the information they want to remember into the terminal. For example, the user inputs, "I want to remember the order of the planets in the solar system."

[0964] Step 2:

[0965] As soon as the device receives the information entered by the user, it activates an emotion engine to analyze the user's emotional state, based on the user's tone of voice, facial expression, and typing speed.

[0966] Step 3:

[0967] The device sends the input information and the emotional state analyzed by the emotion engine to the chat AI on the server, and the data includes emotional information.

[0968] Step 4:

[0969] The chat AI on the server analyzes the received information and generates appropriate mnemonics. For example, it creates mnemonics such as "tai (sea), water, gold (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), and sea (ocean)." It also adjusts the expressions of the mnemonics based on emotional information.

[0970] Step 5:

[0971] The chat AI sends the generated mnemonics back to the device, which then receives them.

[0972] Step 6:

[0973] The device then sends the received mnemonic and the analysis results of the emotion engine to the automatic composition AI on the server. The transmission protocol is the same as in step 3.

[0974] Step 7:

[0975] The automatic composition AI on the server analyzes the received mnemonics and emotional information. Based on this, it generates a melody that matches the mnemonics. This melody is generated in a format that incorporates the mnemonics as lyrics. It also adjusts the tempo and key of the melody based on the emotional information.

[0976] Step 8:

[0977] The automatic composition AI combines the generated melody and the mnemonic to create audio data, which is then sent back to the device. This audio data is usually in MP3 or WAV format.

[0978] Step 9:

[0979] The device then provides the received audio data to the user. The user can then play the audio data on the device and listen to the mnemonics set to the melody. The audio data is also provided in a downloadable format, so it can be listened to repeatedly.

[0980] Step 10:

[0981] By repeatedly listening to audio data, users can efficiently memorize the information they want to remember. This process improves learning effectiveness and is personalized to their emotional state, resulting in more effective learning.

[0982] Example 2

[0983] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0984] Conventional learning support systems lack the ingenuity to help users effectively memorize the information they want to remember, and in particular, they lack individual optimization based on the user's emotional state, which results in a decrease in learning efficiency.In addition, there is a demand for a method that makes learning fun rather than simply memorizing.

[0985] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be memorized from a user and transmitting the information to a natural language processing system; means for the natural language processing system to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic and the user's emotional information to an automatic composition system; means for the automatic composition system to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for analyzing the user's emotions and transmitting the analysis results to the natural language processing system and the automatic composition system; and means for providing the received audio data to the user. This enables individually optimized learning support tailored to the user's emotional state, resulting in enjoyable and efficient memorization.

[0986] "User" refers to anyone who wants to use the system to remember information.

[0987] "Information to be remembered" refers to specific data or facts that the user wants to remember.

[0988] A "natural language processing system" refers to an artificial intelligence-based processing system that analyzes input text and generates appropriate output.

[0989] A mnemonic is a sentence or phrase designed to make it easier to remember information.

[0990] An "automated composition system" refers to an artificial intelligence-based system that can generate melodies and music based on input linguistic data.

[0991] "Audio data" refers to the data of an acoustic signal that combines the generated melody and mnemonic.

[0992] "Emotional information" is data that indicates the user's emotional state during input, and refers to the analysis results obtained from voice tone, facial expression, input content, etc.

[0993] This system supports learning by allowing users to input information they want to remember and converting it into mnemonics and melodies. By combining this with an emotion engine that recognizes the user's emotions, it enables individually optimized learning support. This system is comprised of a terminal, a natural language processing system, an automatic composition system, and an emotion recognition engine.

[0994] First, the user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." In this case, it is desirable for the device to have an interface such as voice input or touch panel input.

[0995] Next, the device sends the input information to a natural language processing system on a server. For example, GPT-3 or GPT-4 is used for the natural language processing system. During this process, the device's built-in emotion recognition engine recognizes and analyzes the user's emotions from voice tone, facial expressions, etc. The analysis results are sent to the natural language processing system or automatic composition system, making it possible to provide output tailored to the user's emotional state.

[0996] The natural language processing system analyzes the received information and generates an appropriate mnemonic. For example, it converts the information "the order of the planets in the solar system" into "tai, water, gold, earth, fire, wood, earth, sky, sea." This mnemonic may then be adjusted according to the user's emotional state based on the analysis results of an emotion recognition engine.

[0997] The generated mnemonic is sent back to the device, which then sends the mnemonic and the user's emotion analysis results to an automatic composition system on the server. Examples of automatic composition systems include AWS DeepComposer and Magenta. The automatic composition system generates a melody that matches the mnemonic and generates audio data that incorporates this melody as lyrics. The tempo and key of the melody are also adjusted based on the analysis results of the emotion recognition engine.

[0998] The audio data generated by the automatic composition system is usually sent to the device in MP3 or WAV format. The device receives this audio data and provides it to the user. The user can play the audio data on their device and listen to the mnemonics that go with the melody, improving memorization efficiency by integrating visual and auditory information. The audio data can also be downloaded and listened to repeatedly.

[0999] Example prompt sentence:

[1000] "Please explain in detail the process by which the system inputs information that the user wants to remember and converts it into a mnemonic and melody."

[1001] In this way, the present invention provides efficient and enjoyable learning support that incorporates customization based on the user's emotions, and can improve upon the problems of conventional learning support systems.

[1002] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1003] Step 1:

[1004] The user inputs the information they want to remember into the terminal.

[1005] Input: Information you want to remember (e.g., "I want to remember the order of the planets in the solar system")

[1006] Output: User input data

[1007] Specific operation: The user inputs text information using the device's interface, and the device stores the input data in its internal memory.

[1008] Step 2:

[1009] The terminal transmits the input information to a natural language processing system.

[1010] Input: User-entered data

[1011] Output: Data (text data) to be sent to the natural language processing system

[1012] Specific operation: The device acquires user input data stored in its internal memory and transmits it to the server's natural language processing system via the Internet.

[1013] Step 3:

[1014] The device analyzes the user's emotional information and sends it to a natural language processing system.

[1015] Input: User's tone of voice, facial expressions, and input

[1016] Output: Sentiment analysis data

[1017] Specific operation: The device uses the emotion engine to analyze emotional information such as voice tone and facial expressions. The analysis results are sent to the server's natural language processing system as emotion analysis data.

[1018] Step 4:

[1019] The server generates mnemonics using a natural language processing system.

[1020] Input: User input data, sentiment analysis data

[1021] Output: Mnemonic data

[1022] Specific operation: The natural language processing system analyzes user input data and generates easy-to-remember mnemonics. The mnemonics are adjusted based on sentiment analysis data. The generated mnemonics are sent back to the device from the server.

[1023] Step 5:

[1024] The terminal again transmits the mnemonic and emotional information to the automatic composition system on the server.

[1025] Input: mnemonic data, sentiment analysis data

[1026] Output: Data to be sent to the automated composition system

[1027] Specific operation: The device combines the mnemonics received from the server with the emotion analysis data and sends it to the automatic composition system.

[1028] Step 6:

[1029] The server generates a melody using an automatic composition system.

[1030] Input: mnemonic data, sentiment analysis data

[1031] Output: Audio data

[1032] Specific operation: The automatic composition system generates a melody that matches the user's emotional state based on the received mnemonic. Audio data is then generated with the mnemonic incorporated into the melody as lyrics.

[1033] Step 7:

[1034] The server transmits the generated voice data to the terminal.

[1035] Input: Audio data

[1036] Output: Audio data sent to the device (e.g. MP3 or WAV format)

[1037] Specific operation: The server sends the generated audio data back to the device, usually in MP3 or WAV format.

[1038] Step 8:

[1039] The terminal provides the audio data to the user.

[1040] Input: Audio data

[1041] Output: Playable audio data

[1042] Specific operation: The device uses playback software to provide the received audio data to the user. The user plays the audio data and listens to the mnemonic set to the melody. The user can also download the audio data and listen to it repeatedly.

[1043] (Application example 2)

[1044] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1045] Conventional learning support systems do not provide a method for users to efficiently memorize information. Furthermore, they are unable to optimize the content of learning support according to the user's emotional state. As a result, learning efficiency is low and it is difficult for users to easily memorize information. The present invention aims to solve these problems by analyzing the user's emotions and providing optimized learning support based on the analysis.

[1046] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving information to be remembered from a user and transmitting the information to a chat-type AI; means for the chat-type AI to generate a mnemonic based on the information and return the mnemonic; means for transmitting the received mnemonic to an automatic composition-type AI; means for the automatic composition-type AI to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for providing the received audio data to the user; and means for incorporating an emotion engine that analyzes the user's emotions and optimizing a series of processes based on the emotion information analyzed by the emotion engine. This enables the user to memorize information in a form optimized according to their emotional state.

[1047] A "user" is a person who uses the system to input information and receive learning support.

[1048] "Information to be remembered" is content or data that the user wants to remember.

[1049] "Chat AI" refers to algorithms or programs that generate mnemonics based on the information they receive.

[1050] A mnemonic is a simple word that uses sound and rhythm to make difficult-to-remember information easier to remember.

[1051] "Automatic composition artificial intelligence" refers to algorithms and programs that generate melodies suitable for mnemonics and incorporate them as audio data.

[1052] A "melody" is a musical melody, a series of sounds generated according to a mnemonic.

[1053] "Audio data" refers to digital sound information including the generated melody.

[1054] An "emotion engine" refers to an algorithm or program that analyzes the user's emotional state and optimizes the system's operation based on that information.

[1055] "Playback" refers to the act of letting the user listen to the generated audio data.

[1056] A "downloadable format" is a digital data format that allows a user to save and later play the content.

[1057] This invention relates to a learning support system that inputs information that a user wants to remember and converts it into mnemonics and melodies. The system is composed of the following elements: a terminal, chat AI, automatic composition AI, and an emotion engine. This system provides personalized and optimized learning support according to the user's emotional state.

[1058] First, the user inputs the information they want to remember into a device. This device could be a smartphone, smart glasses, or a head-mounted display. The input information is sent to a chat-based AI on a server. The chat-based AI generates an appropriate mnemonic based on the received information.

[1059] Next, the device is equipped with an emotion engine that analyzes the emotional information as the user inputs. The analysis results are sent to chat AI and automatic composition AI. The emotion engine determines the user's emotions based on voice tone, facial expressions, input content, etc. Emotion engines that can be used include Affectiva and Microsoft Azure Emotional Analysis API.

[1060] The generated mnemonics and emotional information are sent to an automatic composition AI, which then generates a melody that matches the mnemonic. This melody is generated as audio data incorporating the mnemonics as lyrics. Possible AIs that could be used include OpenAI's MuseNet.

[1061] The automatic composition AI sends the generated audio data back to the device. This audio data is usually generated in MP3 or WAV format. The device then provides the received audio data to the user and plays it back. A smartphone's Pygame library is an effective way to play the audio. Users can also download the audio data and listen to it repeatedly.

[1062] As a concrete example, if a user inputs "I want to memorize the order of the planets in the solar system," and the emotion they are feeling while typing is analyzed as "excited," the chat AI will generate a mnemonic: "Tai, Sui, Kin, Chi, Ka, Ki, Tsuchi, Ten, Umi." This mnemonic and the emotion information are sent to the automatic composition AI, which generates a bright, fast-paced melody. This allows the user to memorize the information they want to remember in a fun and efficient way.

[1063] An example prompt might have the following format:

[1064] Chat AI prompt:

[1065] User input: I want to remember the order of the planets in the solar system.

[1066] User emotion: Excited

[1067] Prompt for automatic composition AI:

[1068] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[1069] User emotion: Excited

[1070] In the above-described manner, the present invention enables learning support that is optimized for the emotional state of the user.

[1071] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1072] Step 1:

[1073] The user inputs the information they want to remember into the device. For example, they might input, "I want to remember the order of the planets in the solar system." The device then sends this input information to a chat-based AI on the server.

[1074] Step 2:

[1075] The chat AI on the server analyzes the received information and generates a mnemonic. For example, in response to the input information "I want to remember the order of the planets in the solar system," it generates the mnemonic "Tai, Sui, Kin, Chi, Ka, Moku, Tsuchi, Ten, Umi." This mnemonic is then sent back to the device.

[1076] Step 3:

[1077] At the same time, the device uses an emotion engine to analyze the emotional information obtained during the user's input process. For example, it may analyze the user's tone of voice and facial expression as "excited." This emotional information is then sent to the server.

[1078] Step 4:

[1079] The device sends the received mnemonic and emotional information to the automatic composition AI on the server, using the following format as a prompt:

[1080] Goro: Tai, Water, Metal, Earth, Fire, Wood, Earth, Heaven, Sea

[1081] User emotion: Excited

[1082] Step 5:

[1083] The automatic composition AI on the server generates an appropriate melody based on mnemonics and emotional information. For example, it creates a bright, fast-paced melody for the mnemonic "tai (sea), water, gold (sea), earth (earth), fire (fire), wood (wood), earth (earth), heaven (heaven), sea (ocean)." This melody is generated as audio data and sent back to the device.

[1084] Step 6:

[1085] The device provides the received audio data to the user. The audio data is usually in MP3 or WAV format and is played on the device. For example, you can play audio using the Pygame library.

[1086] Step 7:

[1087] The user can listen to the audio data played on the device and confirm that the mnemonic is set to the melody. The user can also download the audio data and play it as many times as they like.

[1088] These steps allow the user to remember information in a way that is optimized for their emotional state.

[1089] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1090] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1091] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1092] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1093] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1094] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1095] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1096] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1097] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1098] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1099] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1100] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1101] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1102] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1103] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1104] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1105] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1106] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1107] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1108] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1109] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1110] The following is further disclosed regarding the above embodiment.

[1111] (Claim 1)

[1112] A means for receiving information to be remembered from a user and transmitting the information to a chat-based AI;

[1113] a means for the chat AI to generate a mnemonic based on the information and return the mnemonic;

[1114] A means for sending the received mnemonics to an automatic composition AI;

[1115] means for the automatic composition system artificial intelligence to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data;

[1116] means for providing the received audio data to a user;

[1117] A system including:

[1118] (Claim 2)

[1119] 10. The system of claim 1, further comprising: means for playing the received audio data to a user.

[1120] (Claim 3)

[1121] 10. The system of claim 1, further comprising means for providing the received audio data in a user downloadable format.

[1122] "Example 1"

[1123] (Claim 1)

[1124] A means for receiving information to be memorized from a user and transmitting the information to a natural language processing AI;

[1125] A means for the natural language processing AI to generate a mnemonic based on the information and return the mnemonic;

[1126] A means to send the received mnemonics to a music generation AI,

[1127] a means for the music generation AI to generate music suitable for the mnemonic, generate audio data incorporating the music into the mnemonic, and return the audio data;

[1128] means for providing the received audio data to a user;

[1129] A system including:

[1130] (Claim 2)

[1131] 10. The system of claim 1, further comprising: means for playing the received audio data.

[1132] (Claim 3)

[1133] 10. The system of claim 1, further comprising: means for providing the received audio data in a downloadable format.

[1134] "Application Example 1"

[1135] (Claim 1)

[1136] A means for receiving information to be remembered from a user and transmitting the information to a chat-based AI;

[1137] a means for the chat AI to generate a mnemonic based on the information and return the mnemonic;

[1138] A means for sending the received mnemonics to an automatic composition AI;

[1139] means for the automatic composition system artificial intelligence to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data;

[1140] means for providing the received audio data to a user;

[1141] means for playing the audio data through a smartphone application;

[1142] means for tracking a user's learning progress;

[1143] A system including:

[1144] (Claim 2)

[1145] 10. The system of claim 1, further comprising: means for playing the received audio data to a user.

[1146] (Claim 3)

[1147] 10. The system of claim 1, further comprising means for providing the received audio data in a user downloadable format.

[1148] "Example 2: Combining Emotion Engines"

[1149] (Claim 1)

[1150] means for receiving information to be remembered from a user and transmitting the information to a natural language processing system;

[1151] means for the natural language processing system to generate a mnemonic based on the information and return the mnemonic;

[1152] means for transmitting the received mnemonics and the user's emotional information to an automatic composition system;

[1153] means for the automatic composition system to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data;

[1154] means for analyzing the user's emotions and transmitting the analysis results to the natural language processing system and the automatic composition system;

[1155] means for providing the received audio data to a user;

[1156] A system including:

[1157] (Claim 2)

[1158] 10. The system of claim 1, further comprising: means for playing the received audio data to a user.

[1159] (Claim 3)

[1160] 10. The system of claim 1, further comprising means for providing the received audio data in a user downloadable format.

[1161] "Application example 2 when combining emotion engines"

[1162] (Claim 1)

[1163] A means for receiving information to be remembered from a user and transmitting the information to a chat-based AI;

[1164] a means for the chat AI to generate a mnemonic based on the information and return the mnemonic;

[1165] A means for sending the received mnemonics to an automatic composition AI;

[1166] means for the automatic composition system artificial intelligence to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data;

[1167] means for providing the received audio data to a user;

[1168] a means for optimizing a series of processes based on emotion information analyzed by an emotion engine that is equipped with the device and analyzes the emotion of a user;

[1169] A system including:

[1170] (Claim 2)

[1171] 10. The system of claim 1, further comprising: means for playing the received audio data to a user.

[1172] (Claim 3)

[1173] 10. The system of claim 1, further comprising means for providing the received audio data in a user downloadable format. [Explanation of symbols]

[1174] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving information to be remembered from a user and transmitting the information to a chat-based AI; a means for the chat AI to generate a mnemonic based on the information and return the mnemonic; A means for sending the received mnemonics to an automatic composition AI; a means for the automatic composition system artificial intelligence to generate a melody suitable for the mnemonic, generate audio data incorporating the melody into the mnemonic, and return the audio data; means for providing the received audio data to a user; A system including:

2. 2. The system of claim 1, further comprising: means for playing said received audio data to a user.

3. 10. The system of claim 1, further comprising means for providing said received audio data in a user downloadable format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A