System

The system addresses the inadequacy of existing language learning materials by recording and analyzing user conversations to generate personalized phrasebooks, improving learning efficiency and relevance.

JP2026030578APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024133562
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Existing foreign language learning materials fail to meet individual needs, particularly in prioritizing frequently used phrases and vocabulary in everyday conversation, and lack efficient methods for extracting and learning these expressions.

Method used

A system that records user conversations, analyzes speech data using natural language processing to extract frequently occurring expressions, translates them into other languages, and sorts them by frequency for generating personalized phrasebooks.

Benefits of technology

Enables efficient foreign language learning tailored to individual needs by providing personalized phrasebooks based on daily conversations, enhancing learning efficiency and practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026030578000001_ABST
    Figure 2026030578000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for recording a user's conversation; means for transmitting the recorded speech data to a server; means for analyzing the speech data on the server using natural language processing and extracting frequent expressions; means for translating the extracted expressions into another language; and means for arranging the translated expressions in order of frequency of use and generating a phrase collection.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In foreign language learning, there is a problem in that general learning materials cannot fully meet the needs of each individual. In particular, there is a demand for prioritizing the learning of phrases and vocabulary frequently used in everyday conversation, but there are few learning materials that meet this demand. In addition, there is no established method for extracting frequently used expressions individually and efficiently learning them. Therefore, it is necessary to provide a system that allows users to identify expressions frequently used in everyday conversation and efficiently study foreign languages. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for recording a user's conversation, a means for transmitting the recorded speech data to a server, a means for analyzing the speech data on the server using natural language processing to extract frequently occurring expressions, a means for translating the extracted expressions into other languages, and a means for sorting the translated expressions in order of frequency of use to generate a phrasebook. The system also includes a means for notifying the user of the generated phrasebook and a means for recording and managing the frequency of use of the extracted expressions. This system, configured in this way, enables efficient foreign language learning tailored to the individual needs of each user.

[0006] "User" refers to a person who uses this system.

[0007] "Means for recording conversations" refers to a device or software for recording the user's voice.

[0008] "Voice data" refers to data that is a digital recording of a user's conversation.

[0009] "Server" refers to a computer system for analyzing and storing audio data.

[0010] "Natural language processing" refers to the technology for converting voice data into text data and analyzing that text data.

[0011] "Frequent expressions" refer to words, phrases, short sentences, etc. that appear repeatedly in user conversations.

[0012] "Other languages" refers to foreign languages ​​that the user wants to learn.

[0013] "Means for translation" refers to software or services for translating extracted frequent expressions into other languages.

[0014] A "phrase book" refers to data that collects translated frequently occurring expressions and provides them to users as a single collection.

[0015] "Means for notifying" refers to a device or software for communicating the generated phrasebook to the user.

[0016] "Means for recording and managing usage frequency" refers to software or systems for tracking and managing the number of times an extracted expression is used. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] An embodiment of this invention is a system for recording user conversations and using them to help with foreign language learning. This system includes a means for constantly recording user conversations, a means for analyzing the recorded voice data using natural language processing, a means for extracting frequently occurring expressions and translating them into other languages, and a means for sorting the translated expressions in order of frequency of use to generate a phrasebook.

[0039] Specific operation of the system

[0040] Conversation recording and data storage

[0041] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[0042] The recorded audio data is sent to the server as a file at regular intervals, where it is saved in an appropriate format (e.g., WAV, MP3).

[0043] Analysis of audio data

[0044] The server analyzes the received voice data using natural language processing (NLP) technology. Specifically, it converts the voice data into text data using speech recognition technology.

[0045] The converted text data undergoes morphological and grammatical analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0046] Extraction and translation of frequently occurring expressions

[0047] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[0048] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[0049] Phrasebook generation and provision

[0050] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[0051] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[0052] Specific examples

[0053] 1. Record and send conversations

[0054] User: "What are you having for dinner tonight?"

[0055] The terminal records this conversation and transmits it to the server as audio data.

[0056] 2. Analysis of audio data

[0057] The server analyzes the recorded voice data and obtains the text data "What are you having for dinner tonight?"

[0058] Extract individual words from text data: "Tonight," "What," and "Would you like to eat?"

[0059] 3. Extraction and translation of frequently occurring expressions

[0060] The server calculates the frequency of occurrence of all words extracted from this conversation, and identifies "Tonight" and "Shall we eat?" as frequently occurring expressions.

[0061] Translate these expressions into phrases such as "What will we eat tonight?" or "What's for dinner tonight?"

[0062] 4. Phrasebook generation and provision

[0063] The translated expressions are arranged in a phrasebook, such as "What will we eat tonight?" or "What's for dinner tonight?"

[0064] These phrasebooks are then notified to users, who can easily access them via an app or web portal to begin learning.

[0065] This system enables efficient foreign language learning tailored to individual user needs.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone to capture the user's speech and saves it as a digital audio file (e.g., WAV format).

[0069] Step 2:

[0070] The device periodically (e.g., immediately after the conversation ends) sends the recorded voice data to the server, which is then securely transmitted over the Internet.

[0071] Step 3:

[0072] The server stores the received voice data in a storage device, such as a cloud storage device.

[0073] Step 4:

[0074] The server converts the stored voice data into text data using voice recognition technology, specifically by calling a voice recognition API.

[0075] Step 5:

[0076] The converted text data is then analyzed by the server using natural language processing (NLP) technology, which involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0077] Step 6:

[0078] The server calculates the frequency of use of each extracted expression by comparing it with past analysis data stored in a database.

[0079] Step 7:

[0080] The server translates frequently used expressions into other languages ​​using a translation API (e.g., Google Translate API).

[0081] Step 8:

[0082] The translated expressions are sorted by frequency of use by the server and generated into a phrasebook.

[0083] Step 9:

[0084] The generated phrasebook is then sent to the user by the server via push notifications within the app or email.

[0085] Step 10:

[0086] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[0087] In this way, a process for efficiently learning a foreign language based on the user's conversation data is realized.

[0088] Example 1

[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0090] In foreign language learning, there is a problem that it is difficult for individual users to efficiently learn expressions that they use on a daily basis. Conventional learning methods only allow learning of general phrases and words, making it difficult to learn in accordance with the needs of individual users. In addition, there is a problem that learning efficiency is reduced because there is no adequate system in place to automatically extract frequently used expressions and provide them as learning materials.

[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0092] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data, means for converting the voice data into text data using natural language processing on the server, means for analyzing the converted text data and extracting words, phrases, short sentences, and sentence structure patterns, means for calculating the frequency of appearance of the extracted expressions and identifying frequently occurring expressions, means for translating the identified frequently occurring expressions into other languages, and means for arranging the translated expressions in order of frequency of use and generating a phrasebook, thereby enabling users to efficiently learn specific expressions that they use on a daily basis.

[0093] "User" refers to the person who uses the system to record conversations and learn translation phrases.

[0094] "Means for recording conversations" refers to a function that uses a microphone built into the device to record the user's speech in real time.

[0095] "Audio data" refers to the electronic data format in which a user's conversation is recorded, typically in file formats such as WAV or MP3.

[0096] "Server" refers to a remote computer system that processes received voice data and performs advanced calculations such as analysis and translation.

[0097] "Natural language processing" refers to the technology of analyzing voice data and text data to understand meaning and analyze grammar.

[0098] "Means for converting into text data" refers to speech recognition technology for converting voice data into text data.

[0099] "Analysis" refers to the process of breaking down text data and extracting words, phrases, short sentences, sentence structure patterns, etc.

[0100] "Means of extraction" refers to techniques for finding specific words or phrases from analyzed text data.

[0101] "Means for calculating frequency of occurrence and identifying frequently occurring expressions" refers to a technology that statistically calculates words and phrases that are particularly frequently used within text data and identifies important expressions.

[0102] "Translation means" refers to the function for converting frequently occurring expressions into the foreign language that the user wishes to learn.

[0103] "Means for sorting by frequency of use and generating a phrasebook" refers to a technology that sorts translated expressions based on their frequency of use to create a phrasebook for study.

[0104] "Means of notification" refers to the system used to provide the generated phrasebook to the user, including in-app notifications and email notifications.

[0105] The basic flow of an embodiment of this invention is as follows: First, a user's conversation is recorded in real time, and the recorded data is sent to a server. The server analyzes the received voice data using natural language processing technology, extracts and translates frequently occurring expressions, and then sorts the translated expressions in order of frequency of use to generate a phrasebook. Finally, this phrasebook is provided to the user.

[0106] Hardware and software used

[0107] The following hardware and software is recommended for implementing this system:

[0108] Hardware: Smartphones, smartwatches, smart glasses

[0109] Software: speech recognition technology (e.g., Google Speech-to-Text API), natural language processing technology, translation API (e.g., Google Translate API)

[0110] Specific operation of the system

[0111] Recording conversations and storing data

[0112] The device records the user's conversation in real time using the microphone built into the smartphone, smartwatch, or smart glasses. The recorded data is saved in WAV or MP3 format and sent to the server at regular intervals.

[0113] Analysis of audio data

[0114] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses technologies such as the Google Speech-to-Text API. The converted text data is then analyzed using natural language processing technology. The analysis process includes morphological and grammatical analysis, which allows for the extraction of words, phrases, short sentences, and sentence structure patterns.

[0115] Extraction and translation of frequently occurring expressions

[0116] The server identifies frequently used words and phrases from the analyzed text data by calculating their frequency of occurrence.These frequently used expressions are then translated into the foreign language the user wishes to learn using a translation API (e.g., Google Translate API).

[0117] Phrasebook generation and provision

[0118] The server then sorts the translated phrases into a phrasebook by frequency of use, and provides the phrasebook to users via a notification system, possibly via in-app notifications or email notifications.

[0119] Specific examples

[0120] The specific operation sequence is shown below.

[0121] Prompt Sentence Examples

[0122] User: "I want to eat curry rice today."

[0123] On the device: Records audio in real time and sends it to the server in the appropriate format (WAV or MP3).

[0124] Server: Converts the voice data into text and obtains the text data "I want to eat curry rice today."

[0125] Server: Analyzes using natural language processing technology and extracts "today," "curry rice," and "want to eat."

[0126] Server: Identify the frequently occurring expressions "today," "curry rice," and "want to eat" and translate them into English: "Today," "Curry rice," and "want to eat."

[0127] Server: Compiles the translated phrases into a phrasebook and notifies the user.

[0128] This system allows users to efficiently learn specific expressions used in daily life. In addition, because the generated phrasebooks are based on frequently occurring expressions, it provides practical and effective foreign language learning for users.

[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0130] Step 1: Record the conversation

[0131] The device records the user's conversation in real time. The input is the user's voice, and the output is audio data. Specifically, when the user says, "What will you eat today?", the device's microphone picks up the audio and saves it as an audio data file (e.g., WAV format).

[0132] Step 2: Sending audio data

[0133] The device sends the recorded audio data to the server at regular intervals. The input is the audio data stored on the device, and the output is the audio data sent to the server. Specifically, the recorded data is uploaded to the server every five minutes.

[0134] Step 3: Convert audio data to text

[0135] The server converts the received voice data into text data using voice recognition technology (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data. Specifically, the server receives the voice saying "What will you eat today?" and converts it into the text "What will you eat today?"

[0136] Step 4: Analyzing the text data

[0137] The server analyzes the text data using natural language processing (NLP) technology to extract words, phrases, short sentences, and sentence structure patterns. The input is text data, and the output is the analyzed elements (words, phrases, short sentences, sentence structure patterns). Specifically, the server performs morphological analysis on the text "What will you eat today?" to extract the words "today," "what," and "will you eat."

[0138] Step 5: Extracting frequent expressions

[0139] The server identifies frequently used words and phrases from the analyzed text data. The input is the analyzed text data, and the output is frequently used expressions. Specifically, the server synthesizes past conversation data and creates a list of frequently used expressions such as "Today" and "Shall we eat?"

[0140] Step 6: Translating expressions

[0141] The server uses a translation API (e.g., Google Translate API) to translate frequently occurring expressions into the foreign language the user wishes to learn. The input is the frequently occurring expression, and the output is the translated expression. Specifically, the server translates "Kyou" into "Today" and "Taberu masuka" into "What will we eat?"

[0142] Step 7: Generate phrasebooks

[0143] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. The input is the translated expression, and the output is a phrasebook sorted in order of frequency of use. Specifically, the server lists phrases such as "What will we eat today?" and "What's for dinner tonight?" in order of frequency.

[0144] Step 8: Provide a phrasebook

[0145] The server provides the generated phrasebook to the user through a notification system. The input is the phrasebook, and the output is the notified phrasebook. Specifically, the server sends the generated phrasebook to the user via app notification or email notification, and the user confirms it.

[0146] (Application example 1)

[0147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0148] In today's world, improving the efficiency of foreign language learning is an important challenge for many individuals. However, there are currently no efficient and personalized learning systems based on individual users' conversational content. Furthermore, there is a lack of systems that combine real-time conversation recording, analysis, translation, and notification functions. Therefore, there is a need for a system that encourages users to learn foreign languages ​​naturally in their daily lives.

[0149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0150] In this invention, the server includes means for recording user conversations in real time and saving the voice data in an appropriate format, means for analyzing text data obtained from the recorded voice data and extracting words and short sentences, means for translating the extracted frequently occurring expressions into a foreign language in cooperation with a translation API via the server, and means for providing the generated phrasebook to the user via in-app notifications or email notifications, thereby enabling efficient and personalized foreign language learning based on the individual user's conversations.

[0151] "User" refers to a person who uses the system to record conversations and learn foreign languages.

[0152] "Means for recording conversations" refers to a function that uses a microphone to record a user's conversations in real time.

[0153] "Audio data" refers to data that is a recorded user's conversation stored in digital format.

[0154] "Means for transmitting to a server" refers to the technology for transferring recorded audio data to a remote server via the Internet.

[0155] "Natural language processing" refers to the general technology of analyzing voice data and converting it into text data.

[0156] "Frequent expressions" refer to words and phrases that are used particularly frequently in user conversations.

[0157] "Translation methods" refers to techniques and procedures for converting frequently occurring expressions into other languages.

[0158] "Means for generating phrasebooks" refers to a technology that organizes translated expressions in order of frequency of use and creates a phrasebook for study.

[0159] "Means for storage" refers to the ability to digitally store recorded audio data in an appropriate format (e.g., WAV, MP3).

[0160] "Text data" refers to the textual information obtained from the analyzed voice data.

[0161] "Translation API" refers to an interface that allows a program to provide translation services to other programs.

[0162] "In-app notifications" refers to a function that notifies users of information through smartphone applications.

[0163] "Email notification" refers to the means of providing information to users via email.

[0164] The present invention is a system for recording a user's conversation and using the recorded conversation to help the user learn a foreign language. An embodiment of the system will be described below.

[0165] System Overview

[0166] Users use smartphones or other devices (such as smart glasses or smart watches) that have built-in microphones and can record conversations continuously.

[0167] Server Roles

[0168] The server first receives the recorded audio data, which is periodically sent from the device via the Internet. The audio data is in WAV or MP3 format.

[0169] The server then uses speech recognition technology to convert the audio data into text data. Specifically, it uses a library called speech_recognition. After the audio data is converted into text, it analyzes the text data using natural language processing (NLP) technology to extract frequently occurring expressions. Words and short sentences are identified through morphological and grammatical analysis.

[0170] The server then translates the collected frequently used expressions into other languages, using the DeepL API, for example. The translated phrases are sorted by frequency of use and generated as a learning phrasebook.

[0171] Device Role

[0172] The device records the user's voice and sends the voice data to the server. The recording is done in real time and saved in an appropriate format. The server provides frequently used translation phrases to the user via in-app notifications and email notifications.

[0173] User Roles

[0174] By utilizing this system through everyday conversation, users receive a foreign language phrasebook tailored to their individual needs and progress with their learning.

[0175] Specific examples

[0176] For example, if a user says "I want to order from the virtual store" in a virtual store, this conversation is recorded and saved as audio data. The server then analyzes the text data "I want to order from the virtual store" and extracts the frequently occurring phrases "virtual store" and "I want to order." These phrases are translated into English via a translation API, resulting in "Virtual store" and "I want to order." These translation results are organized into a phrase collection and provided to the user via the smartphone app's notification function.

[0177] Prompt Sentence Examples

[0178] "Please translate the user-mentioned phrase "I want to order from a virtual store" into English."

[0179] This invention enables users to effectively learn a foreign language through everyday conversations. The overall system flow consists of a series of processes: recording the user's conversation, converting it into text, analyzing, translating, and notifying. The system provided by this invention aims to dramatically improve the efficiency of foreign language learning.

[0180] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0181] Step 1:

[0182] recording

[0183] The device records the user's conversation in real time using the device's built-in microphone to capture the user's speech, generating recorded voice data.

[0184] Input: User conversation

[0185] Output: Recorded audio data (e.g. WAV, MP3 format)

[0186] Step 2:

[0187] Sending audio data

[0188] The device transmits the recorded audio data to a server via the Internet at regular intervals. At this time, the data is encrypted for security purposes.

[0189] Input: Recorded audio data

[0190] Output: Audio data sent to the server

[0191] Step 3:

[0192] Voice Recognition

[0193] The server converts the received voice data into text data using natural language processing (NLP) techniques. Specifically, it uses the speech_recognition library to analyze the voice and extract the corresponding text information.

[0194] Input: Audio data sent to the server

[0195] Output: Text data (e.g., "I would like to order from a virtual store")

[0196] Step 4:

[0197] Text data analysis

[0198] The server analyzes the converted text data through morphological and grammatical analysis. It extracts words, phrases, and short sentences from the text data and calculates their frequency of occurrence. This process uses an NLP model.

[0199] Input: Text data

[0200] Output: Analyzed words and phrases and their frequency

[0201] Step 5:

[0202] Extraction of frequent expressions

[0203] The server identifies frequently used words and phrases from the analysis results. By extracting frequently occurring expressions, it identifies phrases and words that users frequently use.

[0204] Input: The words or phrases to be analyzed, and their frequency of occurrence

[0205] Output:Frequent expressions

[0206] Step 6:

[0207] translation

[0208] The server translates the extracted frequently occurring expressions into other languages ​​using a translation API (e.g., DeepL's API). By providing the frequently occurring expressions as prompts to the translation API, the server obtains the corresponding translation results.

[0209] Input:Frequent expressions

[0210] Output: The translated expression

[0211] Step 7:

[0212] Phrasebook generation

[0213] The server sorts the translated frequently occurring expressions in order of frequency of use to generate a phrasebook, which prioritizes important phrases to help users study efficiently.

[0214] Input: translated expression

[0215] Output: A collection of phrases sorted by frequency of use

[0216] Step 8:

[0217] notification

[0218] The server implements a means for notifying the user of the generated phrasebook, for example, by providing an in-app notification using a smartphone app or an email notification.

[0219] Input: A collection of phrases sorted by frequency of use

[0220] Output: User notification (e.g. in-app notification, email notification)

[0221] Through these processing steps, users can learn a foreign language in a personalized way based on everyday conversations. This series of processes aims to increase user convenience and maximize learning effectiveness.

[0222] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0223] An embodiment of the present invention is a system for recording a user's conversation and utilizing the recorded conversation for foreign language learning. The system also includes an emotion engine that recognizes the user's emotions, thereby further enhancing the learning experience. The system includes means for continuously recording the user's conversation, means for analyzing the recorded speech data using natural language processing, and means for extracting frequently occurring expressions and translating them into other languages. The system further includes means for sorting the translated expressions in order of frequency of use to generate a phrasebook, the emotion engine, and means for notifying the user of the generated phrasebook.

[0224] Specific operation of the system

[0225] Conversation recording and data storage

[0226] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[0227] The recorded audio data is sent to the server as a file at regular intervals, and is saved in an appropriate format (e.g., WAV format).

[0228] Analysis of audio data

[0229] The server converts the received voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API to perform the data conversion.

[0230] The converted text data is then analyzed using natural language processing (NLP) techniques, which involve morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0231] Emotion recognition

[0232] The server runs an emotion engine to recognize the user's emotions from the converted text and voice data. This engine detects emotions from the tone, pitch, and word choice of the voice.

[0233] For example, if a user is talking excitedly, this is recognized as "joy."

[0234] Extraction and translation of frequently occurring expressions

[0235] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[0236] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[0237] Phrasebook generation and refinement

[0238] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[0239] The server adjusts the content of the phrase book based on the perceived user emotion, for example, adding more challenging phrases if the user is perceived as relaxed.

[0240] Providing generated phrasebooks

[0241] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[0242] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[0243] Specific examples

[0244] Conversation recording and analysis

[0245] User: "It's such a beautiful day today!"

[0246] The terminal records this conversation and transmits it to the server as audio data.

[0247] The server converts the voice data into text and obtains the text data "What a beautiful day today!"

[0248] Emotion recognition and frequent expression extraction

[0249] The server's emotion engine detects the emotion "joy" from this text and voice tone.

[0250] The server extracts the frequently occurring expression "ii tenki" (good weather) from the text data and translates it to "Good weather."

[0251] Phrasebook generation and provision

[0252] The server adjusts the phrase collection based on the sentiment and generates a phrase collection that includes "Good weather."

[0253] The generated phrasebook is provided to the user as an in-app notification, and the user learns it through the app.

[0254] In this way, efficient foreign language learning according to the user's emotional state is realized.

[0255] The processing flow will be explained below.

[0256] Step 1:

[0257] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone and saves the user's speech as a digital audio file (e.g., WAV format).

[0258] Step 2:

[0259] The device periodically (for example, when the conversation ends) sends the recorded voice data to the server, which is then securely transferred over the Internet.

[0260] Step 3:

[0261] The server stores the received audio data in storage, which is expected to be cloud storage (e.g., AWS S3, Google Cloud Storage).

[0262] Step 4:

[0263] The server converts the stored voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[0264] Step 5:

[0265] The server analyzes the converted text data using natural language processing (NLP) techniques, including morphological and grammatical analysis, to extract words, phrases, short sentences, and sentence structure patterns.

[0266] Step 6:

[0267] The server uses the extracted text data to run an emotion engine to recognize the user's emotions, which analyzes the text's wording and the tone and pitch of the voice.

[0268] Step 7:

[0269] The server extracts frequently used expressions from the speech data and sentiment analysis data, specifically by calculating the frequency of occurrence by comparing them with past data stored in a database, and identifying frequently occurring expressions.

[0270] Step 8:

[0271] The server translates the extracted frequently occurring expressions into foreign languages ​​(e.g., English, Korean) using a translation API, such as Google Translate API.

[0272] Step 9:

[0273] The server sorts the translated expressions by frequency of use to generate a phrasebook, which is also adjusted based on the user's emotions—for example, if the user is tired, simpler phrases are prioritized.

[0274] Step 10:

[0275] The server notifies the user of the generated phrasebook via in-app notifications, push notifications, email, etc.

[0276] Step 11:

[0277] Users can access the phrasebooks provided through their devices and efficiently study foreign languages. The phrasebooks can be viewed through an app or web portal.

[0278] Specific examples

[0279] 1. Conversation recording: A user says, "What a beautiful day today!"

[0280] 2. Sending audio: The device records the conversation and sends the audio data to the server.

[0281] 3. Analysis of voice data: The server converts the voice data into text "What a beautiful day today!"

[0282] 4. Emotion recognition: The server uses the emotion engine to recognize the emotion of "joy."

[0283] 5. Extraction of frequently occurring expressions: The server extracts the frequently occurring expression "ii tenki" and translates it as "Good weather."

[0284] 6. Phrasebook generation and adjustment: The server generates a phrasebook adjusted based on emotions to help the user relax.

[0285] 7. Phrasebook notification and learning: The server notifies the user of the generated phrasebook via an in-app notification, and the user learns it.

[0286] This system enables efficient foreign language learning based on the user's emotions.

[0287] Example 2

[0288] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0289] Conventional foreign language learning systems have the problem of reducing learning efficiency because they provide uniform learning content without considering the user's emotional state. Furthermore, there is no established method for effectively extracting and translating expressions frequently used by users in real conversations and generating a learning phrasebook based on them. This often leads to a decrease in users' motivation to learn and results in poor learning outcomes.

[0290] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and text data to recognize the user's emotions, means for adjusting the content of the phrase book based on the emotions, and means for notifying the user of the generated phrase book. This enables efficient foreign language learning according to the user's emotional state.

[0291] "User" refers to an individual who uses the system to learn.

[0292] "Means for recording conversations" refers to a device or system that records the user's voice in real time.

[0293] "Means for transmitting audio data to a server" refers to the technology or protocol for transferring recorded audio data to a server over a network.

[0294] "Natural language processing" refers to computer technology that analyzes voice and text data to understand and process their content.

[0295] "Means for extracting frequent expressions" refers to a technique for selecting specific words or phrases from the analyzed data based on their frequency.

[0296] "Means for translating into other languages" refers to a technology or system that translates the extracted expressions into a different language.

[0297] "Means for generating phrasebooks" refers to techniques for organizing translated expressions into a single collection for study.

[0298] "Means for recognizing user emotions" refers to technologies and systems that analyze voice data and text data to detect the user's emotional state.

[0299] "Means for adjusting phrase book content based on emotion" refers to a technique for changing or adjusting the content or difficulty of a phrase book in accordance with a detected emotional state.

[0300] "Means of notification" refers to the technology or system used to notify users of the generated phrasebook.

[0301] The present invention relates to a system that records user conversations in real time and analyzes and translates the data for foreign language learning. The system is designed to recognize the user's emotions and further enhance the learning experience.

[0302] Conversation recording and data storage

[0303] The device records the user's conversation in real time using a microphone built into a smartphone, smartwatch, smart glasses, etc. The recorded voice data is sent to a server as a file in an appropriate format such as WAV at regular intervals.

[0304] Examples:

[0305] A user says, "What a beautiful day today!"

[0306] The device records this conversation, saves the audio data with the file name "2023-10-01-1230.wav", and sends it to the server.

[0307] Analysis of audio data

[0308] The server converts the received voice data into text data using voice recognition technology. Specifically, it converts the voice data into text using a common voice recognition API (e.g., Google Cloud Speech-to-Text API). The converted text data is then analyzed using natural language processing (NLP) technology. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0309] Examples:

[0310] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API and receives the text data "What a beautiful day today!" as a result.

[0311] The server performs morphological analysis on the text data "The weather is so nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[0312] Emotion recognition

[0313] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. For example, it uses a common emotion analysis tool (e.g., IBM Watson Tone Analyzer) to detect emotions from the tone, pitch, and choice of words of the voice.

[0314] Examples:

[0315] The server detects the emotion of "joy" from the text data and tone of voice.

[0316] Extraction and translation of frequently occurring expressions

[0317] The server identifies frequently used words and phrases from the analysis results, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using common translation tools (e.g., Google Cloud Translation API).

[0318] Examples:

[0319] The server extracts the frequently occurring expression "nice weather" based on past analysis results.

[0320] The server translates "ii tenki" into "Good weather."

[0321] Phrasebook generation and provision

[0322] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook. The server adjusts the contents of the phrasebook based on the user's recognized emotions. For example, if the server recognizes that the user is relaxed, more difficult phrases will be added. The generated phrasebook is provided to the user through a notification system. For example, this could be an in-app notification or an email notification.

[0323] Examples:

[0324] The server generates a collection of phrases containing "Good weather" based on the sentiment analysis results.

[0325] The server sends the generated phrasebook as an in-app notification.

[0326] Users can browse and learn from the phrasebook "Good weather" provided through the app.

[0327] Prompt Sentence Examples

[0328] If a user says "What a beautiful day today!" during a conversation, use the following prompt example:

[0329] "The user says, 'The weather is so nice today!' Based on this conversation, extract frequently occurring expressions, perform emotion recognition, and generate a phrasebook for foreign language learning based on the results."

[0330] This enables efficient foreign language learning according to the user's emotional state.

[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0332] Step 1:

[0333] The device records the user's conversation in real time. Specific devices include smartphones, smart watches, and smart glasses. The input is the user's conversational voice, and the output is the recorded voice data.

[0334] Specific behavior:

[0335] A user says, "What a beautiful day today!"

[0336] The device records this conversation and saves it as audio data.

[0337] Step 2:

[0338] The device sends the recorded audio data to the server as a file at regular intervals. The audio data format is WAV, etc. The input is the recorded audio data, and the output is the file sent to the server.

[0339] Specific behavior:

[0340] The device saves the audio data with the file name "2023-10-01-1230.wav" and sends it to the server.

[0341] Step 3:

[0342] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API, etc. The input is voice data, and the output is the converted text data.

[0343] Specific behavior:

[0344] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API, and receives the text data "What a beautiful day today!" as a result.

[0345] Step 4:

[0346] The server analyzes the converted text data using natural language processing (NLP) technology. As a result of the analysis, morphological analysis is performed to extract words, phrases, and sentence structure patterns. The input is text data, and the output is analyzed information (extracted words and phrases).

[0347] Specific behavior:

[0348] The server performs morphological analysis on the text data "The weather is very nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[0349] Step 5:

[0350] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. Specifically, emotions are detected from the tone, pitch, and word choice of the voice. The input is text and voice data, and the output is the recognized emotion.

[0351] Specific behavior:

[0352] The server detects the emotion of "joy" from the text data and tone of voice.

[0353] Step 6:

[0354] The server identifies frequently used words and phrases from the analysis results. Specifically, it calculates their frequency of occurrence. The input is the analyzed information (words and phrases), and the output is frequently used expressions.

[0355] Specific behavior:

[0356] The server extracts the frequently occurring expression "nice weather" based on the analysis results.

[0357] Step 7:

[0358] The server translates the identified frequent expressions into other languages ​​using a translation API. Specifically, it uses the Google Cloud Translation API. The input is the frequent expression, and the output is the translated expression.

[0359] Specific behavior:

[0360] The server translates "ii tenki" into "Good weather."

[0361] Step 8:

[0362] The server sorts the translated expressions by frequency of use to generate a phrasebook. The content of the phrasebook is adjusted based on the recognized emotions. The input is the translated expressions and emotional information, and the output is an adjusted phrasebook.

[0363] Specific behavior:

[0364] The server generates a list of phrases containing "Good weather" based on the results of sentiment analysis.

[0365] Step 9:

[0366] The server provides the generated phrasebook to the user through a notification system: the input is the generated phrasebook, and the output is the notified phrasebook.

[0367] Specific behavior:

[0368] The server sends the phrasebook as an in-app notification.

[0369] Users can browse and learn from a collection of phrases called "Good weather" provided through the app.

[0370] (Application example 2)

[0371] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0372] Conventional foreign language learning systems have the problem of being difficult to use efficiently because they do not take into account the user's emotional state or the context of real-life conversations. In particular, there is a need for a system that can analyze frequently used expressions in real time and provide appropriate learning content according to the user's emotions.

[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0374] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data to the server, means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions, means for translating the extracted expressions into other languages, means for arranging the translated expressions in order of frequency of use and generating a phrasebook, means for displaying the phrasebook to the user via smart glasses, and means for recognizing the user's emotions from the analyzed data. This makes it possible to provide emotion-sensitive foreign language learning in real time based on the user's real-life conversation data.

[0375] "Means for recording user conversations" refers to devices or technologies for collecting voice data and storing user utterances.

[0376] The "means for transmitting recorded voice data to a server" refers to a device or method for transferring collected voice data to a server via a network.

[0377] "Means of analyzing voice data on a server using natural language processing and extracting frequently occurring expressions" refers to a technology that converts voice data into text data and identifies frequently used words and phrases.

[0378] The "means for translating extracted expressions into other languages" is a method for converting identified words and phrases into the foreign language that the user wishes to learn.

[0379] "Means for sorting translated expressions in order of frequency of use and generating a phrasebook" refers to a technology that sorts translated words and phrases in order of frequency and provides them as a single, comprehensive learning content.

[0380] The "means for displaying a phrasebook to a user via smart glasses" refers to a technique for displaying the generated phrasebook on the display of smart glasses worn by the user.

[0381] "Means for recognizing user emotions from analyzed data" refers to engines or technologies for identifying a user's emotional state based on voice data and text data.

[0382] The present invention provides a system for recording and analyzing a user's conversations and supporting foreign language learning based on the user's emotional state. The system includes: a means for recording the user's conversations; a means for transmitting the recorded voice data to a server; a means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions; a means for translating the extracted expressions into other languages; a means for sorting the translated expressions in order of frequency of use and generating a phrasebook; a means for displaying the phrasebook to the user via smart glasses; and a means for recognizing the user's emotions from the analyzed data.

[0383] First, when a user wears smart glasses and engages in everyday conversation, the microphone built into the smart glasses collects voice data. This voice data is recorded in real time and sent to a server via a network. The voice data is then saved in an appropriate format (e.g., WAV format).

[0384] Next, on the server, the received voice data is converted into text data using speech recognition technology. Specifically, data conversion is performed using a speech recognition API such as the Google Speech-to-Text API. This converted text data is then analyzed using a natural language processing (NLP) engine, such as spaCy. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0385] Next, based on the analyzed text and voice data, an emotion recognition engine, such as IBM Watson Tone Analyzer, is used to recognize the user's emotions. This engine detects emotions from the tone, pitch, and word choice of the voice. For example, if the user is speaking with an emotion of joy, it will be recognized as "joy."

[0386] The server then uses the analysis results to identify frequently used words and phrases, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using a translation API, such as the Google Cloud Translation API.

[0387] The translated frequently occurring expressions are organized in order of frequency and generated into a phrasebook. This phrasebook is sent in real time from the server to the smart glasses display and displayed to the user. This allows users to learn a foreign language naturally through everyday conversation.

[0388] As a specific example of use, if a user says, "What a beautiful day today!", this voice is recorded by the smart glasses and sent to the server. The speech recognition system on the server converts this voice data into text, obtaining the text data "What a beautiful day today!". Furthermore, the emotion recognition engine detects the emotion of "joy," extracts the frequently occurring expression "nice weather," and translates it into "Good weather." The translated phrases are organized in order of frequency and displayed on the smart glasses.

[0389] Through this process, efficient foreign language learning that responds to emotions can be realized based on the user's daily conversation data.

[0390] An example prompt might look like this:

[0391] "Your smart glasses are currently recording everyday conversations. Please explain the process of extracting common expressions from these conversations, translating them into English, and creating a phrasebook based on emotions."

[0392] As described above, the present invention is a system that provides emotion-responsive foreign language learning in real time based on conversation data from the user's real life.

[0393] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0394] Step 1:

[0395] The device records the user's conversation in real time. The device's built-in microphone collects audio data and saves it in a file format (e.g., WAV format). This outputs an audio file, which is raw data.

[0396] Step 2:

[0397] The device sends the recorded audio data to the server at regular intervals. The audio file is uploaded to the server via the network. The audio file received as input is transferred to the server, where it is saved.

[0398] Step 3:

[0399] The server converts the received voice data into text data using a speech recognition API (for example, Google Speech-to-Text API). It receives the voice file as input and performs data conversion processing to generate text data as output.

[0400] Step 4:

[0401] The server analyzes the generated text data using a natural language processing (NLP) engine (such as spaCy). During the analysis, morphological analysis is performed to extract words, phrases, short sentences, and sentence structure patterns. By receiving the text data as input and performing data analysis, the analysis results are obtained as output.

[0402] Step 5:

[0403] The server uses an emotion recognition engine (such as IBM Watson Tone Analyzer) to recognize the user's emotions based on the analyzed text data and voice data. It receives text data and voice parameters as input, performs emotion recognition processing, and generates emotion data as output.

[0404] Step 6:

[0405] The server identifies frequently used words and phrases from the analysis results, which includes calculating their frequency of occurrence. It receives the analysis results as input, extracts frequently occurring expressions, and outputs a list of frequently occurring expressions.

[0406] Step 7:

[0407] The server translates the extracted frequent expressions into other languages ​​using a translation API (e.g., Google Cloud Translation API). It receives a list of frequent expressions as input and performs translation processing, generating a list of translated expressions as output.

[0408] Step 8:

[0409] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. It receives the list of translated expressions as input, performs sorting, and obtains a phrasebook as output.

[0410] Step 9:

[0411] The server sends the generated phrase book to the smart glasses display in real time and displays it to the user. The server receives the phrase book as input and performs communication processing, and the phrase book is displayed on the smart glasses as output.

[0412] This series of processes enables users to efficiently learn a foreign language through real-life conversations.

[0413] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0414] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0415] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0416] [Second embodiment]

[0417] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0418] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0419] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0420] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0421] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0422] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0423] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0424] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0425] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0426] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0427] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0428] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0429] An embodiment of this invention is a system for recording user conversations and using them to help with foreign language learning. This system includes a means for constantly recording user conversations, a means for analyzing the recorded voice data using natural language processing, a means for extracting frequently occurring expressions and translating them into other languages, and a means for sorting the translated expressions in order of frequency of use to generate a phrasebook.

[0430] Specific operation of the system

[0431] Conversation recording and data storage

[0432] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[0433] The recorded audio data is sent to the server as a file at regular intervals, where it is saved in an appropriate format (e.g., WAV, MP3).

[0434] Analysis of audio data

[0435] The server analyzes the received voice data using natural language processing (NLP) technology. Specifically, it converts the voice data into text data using speech recognition technology.

[0436] The converted text data undergoes morphological and grammatical analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0437] Extraction and translation of frequently occurring expressions

[0438] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[0439] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[0440] Phrasebook generation and provision

[0441] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[0442] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[0443] Specific examples

[0444] 1. Record and send conversations

[0445] User: "What are you having for dinner tonight?"

[0446] The terminal records this conversation and transmits it to the server as audio data.

[0447] 2. Analysis of audio data

[0448] The server analyzes the recorded voice data and obtains the text data "What are you having for dinner tonight?"

[0449] Extract individual words from text data: "Tonight," "What," and "Would you like to eat?"

[0450] 3. Extraction and translation of frequently occurring expressions

[0451] The server calculates the frequency of occurrence of all words extracted from this conversation, and identifies "Tonight" and "Shall we eat?" as frequently occurring expressions.

[0452] Translate these expressions into phrases such as "What will we eat tonight?" or "What's for dinner tonight?"

[0453] 4. Phrasebook generation and provision

[0454] The translated expressions are arranged in a phrasebook, such as "What will we eat tonight?" or "What's for dinner tonight?"

[0455] These phrasebooks are then notified to users, who can easily access them via an app or web portal to begin learning.

[0456] This system enables efficient foreign language learning tailored to individual user needs.

[0457] The processing flow will be explained below.

[0458] Step 1:

[0459] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone to capture the user's speech and saves it as a digital audio file (e.g., WAV format).

[0460] Step 2:

[0461] The device periodically (e.g., immediately after the conversation ends) sends the recorded voice data to the server, which is then securely transmitted over the Internet.

[0462] Step 3:

[0463] The server stores the received voice data in a storage device, such as a cloud storage device.

[0464] Step 4:

[0465] The server converts the stored voice data into text data using voice recognition technology, specifically by calling a voice recognition API.

[0466] Step 5:

[0467] The converted text data is then analyzed by the server using natural language processing (NLP) technology, which involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0468] Step 6:

[0469] The server calculates the frequency of use of each extracted expression by comparing it with past analysis data stored in a database.

[0470] Step 7:

[0471] The server translates frequently used expressions into other languages ​​using a translation API (e.g., Google Translate API).

[0472] Step 8:

[0473] The translated expressions are sorted by frequency of use by the server and generated into a phrasebook.

[0474] Step 9:

[0475] The generated phrasebook is then sent to the user by the server via push notifications within the app or email.

[0476] Step 10:

[0477] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[0478] In this way, a process for efficiently learning a foreign language based on the user's conversation data is realized.

[0479] Example 1

[0480] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0481] In foreign language learning, there is a problem that it is difficult for individual users to efficiently learn expressions that they use on a daily basis. Conventional learning methods only allow learning of general phrases and words, making it difficult to learn in accordance with the needs of individual users. In addition, there is a problem that learning efficiency is reduced because there is no adequate system in place to automatically extract frequently used expressions and provide them as learning materials.

[0482] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0483] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data, means for converting the voice data into text data using natural language processing on the server, means for analyzing the converted text data and extracting words, phrases, short sentences, and sentence structure patterns, means for calculating the frequency of appearance of the extracted expressions and identifying frequently occurring expressions, means for translating the identified frequently occurring expressions into other languages, and means for arranging the translated expressions in order of frequency of use and generating a phrasebook, thereby enabling users to efficiently learn specific expressions that they use on a daily basis.

[0484] "User" refers to the person who uses the system to record conversations and learn translation phrases.

[0485] "Means for recording conversations" refers to a function that uses a microphone built into the device to record the user's speech in real time.

[0486] "Audio data" refers to the electronic data format in which a user's conversation is recorded, typically in file formats such as WAV or MP3.

[0487] "Server" refers to a remote computer system that processes received voice data and performs advanced calculations such as analysis and translation.

[0488] "Natural language processing" refers to the technology of analyzing voice data and text data to understand meaning and analyze grammar.

[0489] "Means for converting into text data" refers to speech recognition technology for converting voice data into text data.

[0490] "Analysis" refers to the process of breaking down text data and extracting words, phrases, short sentences, sentence structure patterns, etc.

[0491] "Means of extraction" refers to techniques for finding specific words or phrases from analyzed text data.

[0492] "Means for calculating frequency of occurrence and identifying frequently occurring expressions" refers to a technology that statistically calculates words and phrases that are particularly frequently used within text data and identifies important expressions.

[0493] "Translation means" refers to the function for converting frequently occurring expressions into the foreign language that the user wishes to learn.

[0494] "Means for sorting by frequency of use and generating a phrasebook" refers to a technology that sorts translated expressions based on their frequency of use to create a phrasebook for study.

[0495] "Means of notification" refers to the system used to provide the generated phrasebook to the user, including in-app notifications and email notifications.

[0496] The basic flow of an embodiment of this invention is as follows: First, a user's conversation is recorded in real time, and the recorded data is sent to a server. The server analyzes the received voice data using natural language processing technology, extracts and translates frequently occurring expressions, and then sorts the translated expressions in order of frequency of use to generate a phrasebook. Finally, this phrasebook is provided to the user.

[0497] Hardware and software used

[0498] The following hardware and software is recommended for implementing this system:

[0499] Hardware: Smartphones, smartwatches, smart glasses

[0500] Software: speech recognition technology (e.g., Google Speech-to-Text API), natural language processing technology, translation API (e.g., Google Translate API)

[0501] Specific operation of the system

[0502] Recording conversations and storing data

[0503] The device records the user's conversation in real time using the microphone built into the smartphone, smartwatch, or smart glasses. The recorded data is saved in WAV or MP3 format and sent to the server at regular intervals.

[0504] Analysis of audio data

[0505] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses technologies such as the Google Speech-to-Text API. The converted text data is then analyzed using natural language processing technology. The analysis process includes morphological and grammatical analysis, which allows for the extraction of words, phrases, short sentences, and sentence structure patterns.

[0506] Extraction and translation of frequently occurring expressions

[0507] The server identifies frequently used words and phrases from the analyzed text data by calculating their frequency of occurrence.These frequently used expressions are then translated into the foreign language the user wishes to learn using a translation API (e.g., Google Translate API).

[0508] Phrasebook generation and provision

[0509] The server then sorts the translated phrases into a phrasebook by frequency of use, and provides the phrasebook to users via a notification system, possibly via in-app notifications or email notifications.

[0510] Specific examples

[0511] The specific operation sequence is shown below.

[0512] Prompt Sentence Examples

[0513] User: "I want to eat curry rice today."

[0514] On the device: Records audio in real time and sends it to the server in the appropriate format (WAV or MP3).

[0515] Server: Converts the voice data into text and obtains the text data "I want to eat curry rice today."

[0516] Server: Analyzes using natural language processing technology and extracts "today," "curry rice," and "want to eat."

[0517] Server: Identify the frequently occurring expressions "today," "curry rice," and "want to eat" and translate them into English: "Today," "Curry rice," and "want to eat."

[0518] Server: Compiles the translated phrases into a phrasebook and notifies the user.

[0519] This system allows users to efficiently learn specific expressions used in daily life. In addition, because the generated phrasebooks are based on frequently occurring expressions, it provides practical and effective foreign language learning for users.

[0520] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0521] Step 1: Record the conversation

[0522] The device records the user's conversation in real time. The input is the user's voice, and the output is audio data. Specifically, when the user says, "What will you eat today?", the device's microphone picks up the audio and saves it as an audio data file (e.g., WAV format).

[0523] Step 2: Sending audio data

[0524] The device sends the recorded audio data to the server at regular intervals. The input is the audio data stored on the device, and the output is the audio data sent to the server. Specifically, the recorded data is uploaded to the server every five minutes.

[0525] Step 3: Convert audio data to text

[0526] The server converts the received voice data into text data using voice recognition technology (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data. Specifically, the server receives the voice saying "What will you eat today?" and converts it into the text "What will you eat today?"

[0527] Step 4: Analyzing the text data

[0528] The server analyzes the text data using natural language processing (NLP) technology to extract words, phrases, short sentences, and sentence structure patterns. The input is text data, and the output is the analyzed elements (words, phrases, short sentences, sentence structure patterns). Specifically, the server performs morphological analysis on the text "What will you eat today?" to extract the words "today," "what," and "will you eat."

[0529] Step 5: Extracting frequent expressions

[0530] The server identifies frequently used words and phrases from the analyzed text data. The input is the analyzed text data, and the output is frequently used expressions. Specifically, the server synthesizes past conversation data and creates a list of frequently used expressions such as "Today" and "Shall we eat?"

[0531] Step 6: Translating expressions

[0532] The server uses a translation API (e.g., Google Translate API) to translate frequently occurring expressions into the foreign language the user wishes to learn. The input is the frequently occurring expression, and the output is the translated expression. Specifically, the server translates "Kyou" into "Today" and "Taberu masuka" into "What will we eat?"

[0533] Step 7: Generate phrasebooks

[0534] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. The input is the translated expression, and the output is a phrasebook sorted in order of frequency of use. Specifically, the server lists phrases such as "What will we eat today?" and "What's for dinner tonight?" in order of frequency.

[0535] Step 8: Provide a phrasebook

[0536] The server provides the generated phrasebook to the user through a notification system. The input is the phrasebook, and the output is the notified phrasebook. Specifically, the server sends the generated phrasebook to the user via app notification or email notification, and the user confirms it.

[0537] (Application example 1)

[0538] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0539] In today's world, improving the efficiency of foreign language learning is an important challenge for many individuals. However, there are currently no efficient and personalized learning systems based on individual users' conversational content. Furthermore, there is a lack of systems that combine real-time conversation recording, analysis, translation, and notification functions. Therefore, there is a need for a system that encourages users to learn foreign languages ​​naturally in their daily lives.

[0540] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0541] In this invention, the server includes means for recording user conversations in real time and saving the voice data in an appropriate format, means for analyzing text data obtained from the recorded voice data and extracting words and short sentences, means for translating the extracted frequently occurring expressions into a foreign language in cooperation with a translation API via the server, and means for providing the generated phrasebook to the user via in-app notifications or email notifications, thereby enabling efficient and personalized foreign language learning based on the individual user's conversations.

[0542] "User" refers to a person who uses the system to record conversations and learn foreign languages.

[0543] "Means for recording conversations" refers to a function that uses a microphone to record a user's conversations in real time.

[0544] "Audio data" refers to data that is a recorded user's conversation stored in digital format.

[0545] "Means for transmitting to a server" refers to the technology for transferring recorded audio data to a remote server via the Internet.

[0546] "Natural language processing" refers to the general technology of analyzing voice data and converting it into text data.

[0547] "Frequent expressions" refer to words and phrases that are used particularly frequently in user conversations.

[0548] "Translation methods" refers to techniques and procedures for converting frequently occurring expressions into other languages.

[0549] "Means for generating phrasebooks" refers to a technology that organizes translated expressions in order of frequency of use and creates a phrasebook for study.

[0550] "Means for storage" refers to the ability to digitally store recorded audio data in an appropriate format (e.g., WAV, MP3).

[0551] "Text data" refers to the textual information obtained from the analyzed voice data.

[0552] "Translation API" refers to an interface that allows a program to provide translation services to other programs.

[0553] "In-app notifications" refers to a function that notifies users of information through smartphone applications.

[0554] "Email notification" refers to the means of providing information to users via email.

[0555] The present invention is a system for recording a user's conversation and using the recorded conversation to help the user learn a foreign language. An embodiment of the system will be described below.

[0556] System Overview

[0557] Users use smartphones or other devices (such as smart glasses or smart watches) that have built-in microphones and can record conversations continuously.

[0558] Server Roles

[0559] The server first receives the recorded audio data, which is periodically sent from the device via the Internet. The audio data is in WAV or MP3 format.

[0560] The server then uses speech recognition technology to convert the audio data into text data. Specifically, it uses a library called speech_recognition. After the audio data is converted into text, it analyzes the text data using natural language processing (NLP) technology to extract frequently occurring expressions. Words and short sentences are identified through morphological and grammatical analysis.

[0561] The server then translates the collected frequently used expressions into other languages, using the DeepL API, for example. The translated phrases are sorted by frequency of use and generated as a learning phrasebook.

[0562] Device Role

[0563] The device records the user's voice and sends the voice data to the server. The recording is done in real time and saved in an appropriate format. The server provides frequently used translation phrases to the user via in-app notifications and email notifications.

[0564] User Roles

[0565] By utilizing this system through everyday conversation, users receive a foreign language phrasebook tailored to their individual needs and progress with their learning.

[0566] Specific examples

[0567] For example, if a user says "I want to order from the virtual store" in a virtual store, this conversation is recorded and saved as audio data. The server then analyzes the text data "I want to order from the virtual store" and extracts the frequently occurring phrases "virtual store" and "I want to order." These phrases are translated into English via a translation API, resulting in "Virtual store" and "I want to order." These translation results are organized into a phrase collection and provided to the user via the smartphone app's notification function.

[0568] Prompt Sentence Examples

[0569] "Please translate the user-mentioned phrase "I want to order from a virtual store" into English."

[0570] This invention enables users to effectively learn a foreign language through everyday conversations. The overall system flow consists of a series of processes: recording the user's conversation, converting it into text, analyzing, translating, and notifying. The system provided by this invention aims to dramatically improve the efficiency of foreign language learning.

[0571] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0572] Step 1:

[0573] recording

[0574] The device records the user's conversation in real time using the device's built-in microphone to capture the user's speech, generating recorded voice data.

[0575] Input: User conversation

[0576] Output: Recorded audio data (e.g. WAV, MP3 format)

[0577] Step 2:

[0578] Sending audio data

[0579] The device transmits the recorded audio data to a server via the Internet at regular intervals. At this time, the data is encrypted for security purposes.

[0580] Input: Recorded audio data

[0581] Output: Audio data sent to the server

[0582] Step 3:

[0583] Voice Recognition

[0584] The server converts the received voice data into text data using natural language processing (NLP) techniques. Specifically, it uses the speech_recognition library to analyze the voice and extract the corresponding text information.

[0585] Input: Audio data sent to the server

[0586] Output: Text data (e.g., "I would like to order from a virtual store")

[0587] Step 4:

[0588] Text data analysis

[0589] The server analyzes the converted text data through morphological and grammatical analysis. It extracts words, phrases, and short sentences from the text data and calculates their frequency of occurrence. This process uses an NLP model.

[0590] Input: Text data

[0591] Output: Analyzed words and phrases and their frequency

[0592] Step 5:

[0593] Extraction of frequent expressions

[0594] The server identifies frequently used words and phrases from the analysis results. By extracting frequently occurring expressions, it identifies phrases and words that users frequently use.

[0595] Input: The words or phrases to be analyzed, and their frequency of occurrence

[0596] Output:Frequent expressions

[0597] Step 6:

[0598] translation

[0599] The server translates the extracted frequently occurring expressions into other languages ​​using a translation API (e.g., DeepL's API). By providing the frequently occurring expressions as prompts to the translation API, the server obtains the corresponding translation results.

[0600] Input:Frequent expressions

[0601] Output: The translated expression

[0602] Step 7:

[0603] Phrasebook generation

[0604] The server sorts the translated frequently occurring expressions in order of frequency of use to generate a phrasebook, which prioritizes important phrases to help users study efficiently.

[0605] Input: translated expression

[0606] Output: A collection of phrases sorted by frequency of use

[0607] Step 8:

[0608] notification

[0609] The server implements a means for notifying the user of the generated phrasebook, for example, by providing an in-app notification using a smartphone app or an email notification.

[0610] Input: A collection of phrases sorted by frequency of use

[0611] Output: User notification (e.g. in-app notification, email notification)

[0612] Through these processing steps, users can learn a foreign language in a personalized way based on everyday conversations. This series of processes aims to increase user convenience and maximize learning effectiveness.

[0613] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0614] An embodiment of the present invention is a system for recording a user's conversation and utilizing the recorded conversation for foreign language learning. The system also includes an emotion engine that recognizes the user's emotions, thereby further enhancing the learning experience. The system includes means for continuously recording the user's conversation, means for analyzing the recorded speech data using natural language processing, and means for extracting frequently occurring expressions and translating them into other languages. The system further includes means for sorting the translated expressions in order of frequency of use to generate a phrasebook, the emotion engine, and means for notifying the user of the generated phrasebook.

[0615] Specific operation of the system

[0616] Conversation recording and data storage

[0617] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[0618] The recorded audio data is sent to the server as a file at regular intervals, and is saved in an appropriate format (e.g., WAV format).

[0619] Analysis of audio data

[0620] The server converts the received voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API to perform the data conversion.

[0621] The converted text data is then analyzed using natural language processing (NLP) techniques, which involve morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0622] Emotion recognition

[0623] The server runs an emotion engine to recognize the user's emotions from the converted text and voice data. This engine detects emotions from the tone, pitch, and word choice of the voice.

[0624] For example, if a user is talking excitedly, this is recognized as "joy."

[0625] Extraction and translation of frequently occurring expressions

[0626] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[0627] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[0628] Phrasebook generation and refinement

[0629] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[0630] The server adjusts the content of the phrase book based on the perceived user emotion, for example, adding more challenging phrases if the user is perceived as relaxed.

[0631] Providing generated phrasebooks

[0632] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[0633] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[0634] Specific examples

[0635] Conversation recording and analysis

[0636] User: "It's such a beautiful day today!"

[0637] The terminal records this conversation and transmits it to the server as audio data.

[0638] The server converts the voice data into text and obtains the text data "What a beautiful day today!"

[0639] Emotion recognition and frequent expression extraction

[0640] The server's emotion engine detects the emotion "joy" from this text and voice tone.

[0641] The server extracts the frequently occurring expression "ii tenki" (good weather) from the text data and translates it to "Good weather."

[0642] Phrasebook generation and provision

[0643] The server adjusts the phrase collection based on the sentiment and generates a phrase collection that includes "Good weather."

[0644] The generated phrasebook is provided to the user as an in-app notification, and the user learns it through the app.

[0645] In this way, efficient foreign language learning according to the user's emotional state is realized.

[0646] The processing flow will be explained below.

[0647] Step 1:

[0648] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone and saves the user's speech as a digital audio file (e.g., WAV format).

[0649] Step 2:

[0650] The device periodically (for example, when the conversation ends) sends the recorded voice data to the server, which is then securely transferred over the Internet.

[0651] Step 3:

[0652] The server stores the received audio data in storage, which is expected to be cloud storage (e.g., AWS S3, Google Cloud Storage).

[0653] Step 4:

[0654] The server converts the stored voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[0655] Step 5:

[0656] The server analyzes the converted text data using natural language processing (NLP) techniques, including morphological and grammatical analysis, to extract words, phrases, short sentences, and sentence structure patterns.

[0657] Step 6:

[0658] The server uses the extracted text data to run an emotion engine to recognize the user's emotions, which analyzes the text's wording and the tone and pitch of the voice.

[0659] Step 7:

[0660] The server extracts frequently used expressions from the speech data and sentiment analysis data, specifically by calculating the frequency of occurrence by comparing them with past data stored in a database, and identifying frequently occurring expressions.

[0661] Step 8:

[0662] The server translates the extracted frequently occurring expressions into foreign languages ​​(e.g., English, Korean) using a translation API, such as Google Translate API.

[0663] Step 9:

[0664] The server sorts the translated expressions by frequency of use to generate a phrasebook, which is also adjusted based on the user's emotions—for example, if the user is tired, simpler phrases are prioritized.

[0665] Step 10:

[0666] The server notifies the user of the generated phrasebook via in-app notifications, push notifications, email, etc.

[0667] Step 11:

[0668] Users can access the phrasebooks provided through their devices and efficiently study foreign languages. The phrasebooks can be viewed through an app or web portal.

[0669] Specific examples

[0670] 1. Conversation recording: A user says, "What a beautiful day today!"

[0671] 2. Sending audio: The device records the conversation and sends the audio data to the server.

[0672] 3. Analysis of voice data: The server converts the voice data into text "What a beautiful day today!"

[0673] 4. Emotion recognition: The server uses the emotion engine to recognize the emotion of "joy."

[0674] 5. Extraction of frequently occurring expressions: The server extracts the frequently occurring expression "ii tenki" and translates it as "Good weather."

[0675] 6. Phrasebook generation and adjustment: The server generates a phrasebook adjusted based on emotions to help the user relax.

[0676] 7. Phrasebook notification and learning: The server notifies the user of the generated phrasebook via an in-app notification, and the user learns it.

[0677] This system enables efficient foreign language learning based on the user's emotions.

[0678] Example 2

[0679] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0680] Conventional foreign language learning systems have the problem of reducing learning efficiency because they provide uniform learning content without considering the user's emotional state. Furthermore, there is no established method for effectively extracting and translating expressions frequently used by users in real conversations and generating a learning phrasebook based on them. This often leads to a decrease in users' motivation to learn and results in poor learning outcomes.

[0681] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and text data to recognize the user's emotions, means for adjusting the content of the phrase book based on the emotions, and means for notifying the user of the generated phrase book. This enables efficient foreign language learning according to the user's emotional state.

[0682] "User" refers to an individual who uses the system to learn.

[0683] "Means for recording conversations" refers to a device or system that records the user's voice in real time.

[0684] "Means for transmitting audio data to a server" refers to the technology or protocol for transferring recorded audio data to a server over a network.

[0685] "Natural language processing" refers to computer technology that analyzes voice and text data to understand and process their content.

[0686] "Means for extracting frequent expressions" refers to a technique for selecting specific words or phrases from the analyzed data based on their frequency.

[0687] "Means for translating into other languages" refers to a technology or system that translates the extracted expressions into a different language.

[0688] "Means for generating phrasebooks" refers to techniques for organizing translated expressions into a single collection for study.

[0689] "Means for recognizing user emotions" refers to technologies and systems that analyze voice data and text data to detect the user's emotional state.

[0690] "Means for adjusting phrase book content based on emotion" refers to a technique for changing or adjusting the content or difficulty of a phrase book in accordance with a detected emotional state.

[0691] "Means of notification" refers to the technology or system used to notify users of the generated phrasebook.

[0692] The present invention relates to a system that records user conversations in real time and analyzes and translates the data for foreign language learning. The system is designed to recognize the user's emotions and further enhance the learning experience.

[0693] Conversation recording and data storage

[0694] The device records the user's conversation in real time using a microphone built into a smartphone, smartwatch, smart glasses, etc. The recorded voice data is sent to a server as a file in an appropriate format such as WAV at regular intervals.

[0695] Examples:

[0696] A user says, "What a beautiful day today!"

[0697] The device records this conversation, saves the audio data with the file name "2023-10-01-1230.wav", and sends it to the server.

[0698] Analysis of audio data

[0699] The server converts the received voice data into text data using voice recognition technology. Specifically, it converts the voice data into text using a common voice recognition API (e.g., Google Cloud Speech-to-Text API). The converted text data is then analyzed using natural language processing (NLP) technology. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0700] Examples:

[0701] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API and receives the text data "What a beautiful day today!" as a result.

[0702] The server performs morphological analysis on the text data "The weather is so nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[0703] Emotion recognition

[0704] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. For example, it uses a common emotion analysis tool (e.g., IBM Watson Tone Analyzer) to detect emotions from the tone, pitch, and choice of words of the voice.

[0705] Examples:

[0706] The server detects the emotion of "joy" from the text data and tone of voice.

[0707] Extraction and translation of frequently occurring expressions

[0708] The server identifies frequently used words and phrases from the analysis results, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using common translation tools (e.g., Google Cloud Translation API).

[0709] Examples:

[0710] The server extracts the frequently occurring expression "nice weather" based on past analysis results.

[0711] The server translates "ii tenki" into "Good weather."

[0712] Phrasebook generation and provision

[0713] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook. The server adjusts the contents of the phrasebook based on the user's recognized emotions. For example, if the server recognizes that the user is relaxed, more difficult phrases will be added. The generated phrasebook is provided to the user through a notification system. For example, this could be an in-app notification or an email notification.

[0714] Examples:

[0715] The server generates a collection of phrases containing "Good weather" based on the sentiment analysis results.

[0716] The server sends the generated phrasebook as an in-app notification.

[0717] Users can browse and learn from the phrasebook "Good weather" provided through the app.

[0718] Prompt Sentence Examples

[0719] If a user says "What a beautiful day today!" during a conversation, use the following prompt example:

[0720] "The user says, 'The weather is so nice today!' Based on this conversation, extract frequently occurring expressions, perform emotion recognition, and generate a phrasebook for foreign language learning based on the results."

[0721] This enables efficient foreign language learning according to the user's emotional state.

[0722] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0723] Step 1:

[0724] The device records the user's conversation in real time. Specific devices include smartphones, smart watches, and smart glasses. The input is the user's conversational voice, and the output is the recorded voice data.

[0725] Specific behavior:

[0726] A user says, "What a beautiful day today!"

[0727] The device records this conversation and saves it as audio data.

[0728] Step 2:

[0729] The device sends the recorded audio data to the server as a file at regular intervals. The audio data format is WAV, etc. The input is the recorded audio data, and the output is the file sent to the server.

[0730] Specific behavior:

[0731] The device saves the audio data with the file name "2023-10-01-1230.wav" and sends it to the server.

[0732] Step 3:

[0733] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API, etc. The input is voice data, and the output is the converted text data.

[0734] Specific behavior:

[0735] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API, and receives the text data "What a beautiful day today!" as a result.

[0736] Step 4:

[0737] The server analyzes the converted text data using natural language processing (NLP) technology. As a result of the analysis, morphological analysis is performed to extract words, phrases, and sentence structure patterns. The input is text data, and the output is analyzed information (extracted words and phrases).

[0738] Specific behavior:

[0739] The server performs morphological analysis on the text data "The weather is very nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[0740] Step 5:

[0741] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. Specifically, emotions are detected from the tone, pitch, and word choice of the voice. The input is text and voice data, and the output is the recognized emotion.

[0742] Specific behavior:

[0743] The server detects the emotion of "joy" from the text data and tone of voice.

[0744] Step 6:

[0745] The server identifies frequently used words and phrases from the analysis results. Specifically, it calculates their frequency of occurrence. The input is the analyzed information (words and phrases), and the output is frequently used expressions.

[0746] Specific behavior:

[0747] The server extracts the frequently occurring expression "nice weather" based on the analysis results.

[0748] Step 7:

[0749] The server translates the identified frequent expressions into other languages ​​using a translation API. Specifically, it uses the Google Cloud Translation API. The input is the frequent expression, and the output is the translated expression.

[0750] Specific behavior:

[0751] The server translates "ii tenki" into "Good weather."

[0752] Step 8:

[0753] The server sorts the translated expressions by frequency of use to generate a phrasebook. The content of the phrasebook is adjusted based on the recognized emotions. The input is the translated expressions and emotional information, and the output is an adjusted phrasebook.

[0754] Specific behavior:

[0755] The server generates a list of phrases containing "Good weather" based on the results of sentiment analysis.

[0756] Step 9:

[0757] The server provides the generated phrasebook to the user through a notification system: the input is the generated phrasebook, and the output is the notified phrasebook.

[0758] Specific behavior:

[0759] The server sends the phrasebook as an in-app notification.

[0760] Users can browse and learn from a collection of phrases called "Good weather" provided through the app.

[0761] (Application example 2)

[0762] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0763] Conventional foreign language learning systems have the problem of being difficult to use efficiently because they do not take into account the user's emotional state or the context of real-life conversations. In particular, there is a need for a system that can analyze frequently used expressions in real time and provide appropriate learning content according to the user's emotions.

[0764] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0765] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data to the server, means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions, means for translating the extracted expressions into other languages, means for arranging the translated expressions in order of frequency of use and generating a phrasebook, means for displaying the phrasebook to the user via smart glasses, and means for recognizing the user's emotions from the analyzed data. This makes it possible to provide emotion-sensitive foreign language learning in real time based on the user's real-life conversation data.

[0766] "Means for recording user conversations" refers to devices or technologies for collecting voice data and storing user utterances.

[0767] The "means for transmitting recorded voice data to a server" refers to a device or method for transferring collected voice data to a server via a network.

[0768] "Means of analyzing voice data on a server using natural language processing and extracting frequently occurring expressions" refers to a technology that converts voice data into text data and identifies frequently used words and phrases.

[0769] The "means for translating extracted expressions into other languages" is a method for converting identified words and phrases into the foreign language that the user wishes to learn.

[0770] "Means for sorting translated expressions in order of frequency of use and generating a phrasebook" refers to a technology that sorts translated words and phrases in order of frequency and provides them as a single, comprehensive learning content.

[0771] The "means for displaying a phrasebook to a user via smart glasses" refers to a technique for displaying the generated phrasebook on the display of smart glasses worn by the user.

[0772] "Means for recognizing user emotions from analyzed data" refers to engines or technologies for identifying a user's emotional state based on voice data and text data.

[0773] The present invention provides a system for recording and analyzing a user's conversations and supporting foreign language learning based on the user's emotional state. The system includes: a means for recording the user's conversations; a means for transmitting the recorded voice data to a server; a means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions; a means for translating the extracted expressions into other languages; a means for sorting the translated expressions in order of frequency of use and generating a phrasebook; a means for displaying the phrasebook to the user via smart glasses; and a means for recognizing the user's emotions from the analyzed data.

[0774] First, when a user wears smart glasses and engages in everyday conversation, the microphone built into the smart glasses collects voice data. This voice data is recorded in real time and sent to a server via a network. The voice data is then saved in an appropriate format (e.g., WAV format).

[0775] Next, on the server, the received voice data is converted into text data using speech recognition technology. Specifically, data conversion is performed using a speech recognition API such as the Google Speech-to-Text API. This converted text data is then analyzed using a natural language processing (NLP) engine, such as spaCy. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0776] Next, based on the analyzed text and voice data, an emotion recognition engine, such as IBM Watson Tone Analyzer, is used to recognize the user's emotions. This engine detects emotions from the tone, pitch, and word choice of the voice. For example, if the user is speaking with an emotion of joy, it will be recognized as "joy."

[0777] The server then uses the analysis results to identify frequently used words and phrases, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using a translation API, such as the Google Cloud Translation API.

[0778] The translated frequently occurring expressions are organized in order of frequency and generated into a phrasebook. This phrasebook is sent in real time from the server to the smart glasses display and displayed to the user. This allows users to learn a foreign language naturally through everyday conversation.

[0779] As a specific example of use, if a user says, "What a beautiful day today!", this voice is recorded by the smart glasses and sent to the server. The speech recognition system on the server converts this voice data into text, obtaining the text data "What a beautiful day today!". Furthermore, the emotion recognition engine detects the emotion of "joy," extracts the frequently occurring expression "nice weather," and translates it into "Good weather." The translated phrases are organized in order of frequency and displayed on the smart glasses.

[0780] Through this process, efficient foreign language learning that responds to emotions can be realized based on the user's daily conversation data.

[0781] An example prompt might look like this:

[0782] "Your smart glasses are currently recording everyday conversations. Please explain the process of extracting common expressions from these conversations, translating them into English, and creating a phrasebook based on emotions."

[0783] As described above, the present invention is a system that provides emotion-responsive foreign language learning in real time based on conversation data from the user's real life.

[0784] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0785] Step 1:

[0786] The device records the user's conversation in real time. The device's built-in microphone collects audio data and saves it in a file format (e.g., WAV format). This outputs an audio file, which is raw data.

[0787] Step 2:

[0788] The device sends the recorded audio data to the server at regular intervals. The audio file is uploaded to the server via the network. The audio file received as input is transferred to the server, where it is saved.

[0789] Step 3:

[0790] The server converts the received voice data into text data using a speech recognition API (for example, Google Speech-to-Text API). It receives the voice file as input and performs data conversion processing to generate text data as output.

[0791] Step 4:

[0792] The server analyzes the generated text data using a natural language processing (NLP) engine (such as spaCy). During the analysis, morphological analysis is performed to extract words, phrases, short sentences, and sentence structure patterns. By receiving the text data as input and performing data analysis, the analysis results are obtained as output.

[0793] Step 5:

[0794] The server uses an emotion recognition engine (such as IBM Watson Tone Analyzer) to recognize the user's emotions based on the analyzed text data and voice data. It receives text data and voice parameters as input, performs emotion recognition processing, and generates emotion data as output.

[0795] Step 6:

[0796] The server identifies frequently used words and phrases from the analysis results, which includes calculating their frequency of occurrence. It receives the analysis results as input, extracts frequently occurring expressions, and outputs a list of frequently occurring expressions.

[0797] Step 7:

[0798] The server translates the extracted frequent expressions into other languages ​​using a translation API (e.g., Google Cloud Translation API). It receives a list of frequent expressions as input and performs translation processing, generating a list of translated expressions as output.

[0799] Step 8:

[0800] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. It receives the list of translated expressions as input, performs sorting, and obtains a phrasebook as output.

[0801] Step 9:

[0802] The server sends the generated phrase book to the smart glasses display in real time and displays it to the user. The server receives the phrase book as input and performs communication processing, and the phrase book is displayed on the smart glasses as output.

[0803] This series of processes enables users to efficiently learn a foreign language through real-life conversations.

[0804] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0805] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0806] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0807] [Third embodiment]

[0808] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0809] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0810] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0811] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0812] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0813] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0814] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0815] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0816] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0817] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0818] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0819] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0820] An embodiment of this invention is a system for recording user conversations and using them to help with foreign language learning. This system includes a means for constantly recording user conversations, a means for analyzing the recorded voice data using natural language processing, a means for extracting frequently occurring expressions and translating them into other languages, and a means for sorting the translated expressions in order of frequency of use to generate a phrasebook.

[0821] Specific operation of the system

[0822] Conversation recording and data storage

[0823] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[0824] The recorded audio data is sent to the server as a file at regular intervals, where it is saved in an appropriate format (e.g., WAV, MP3).

[0825] Analysis of audio data

[0826] The server analyzes the received voice data using natural language processing (NLP) technology. Specifically, it converts the voice data into text data using speech recognition technology.

[0827] The converted text data undergoes morphological and grammatical analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0828] Extraction and translation of frequently occurring expressions

[0829] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[0830] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[0831] Phrasebook generation and provision

[0832] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[0833] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[0834] Specific examples

[0835] 1. Record and send conversations

[0836] User: "What are you having for dinner tonight?"

[0837] The terminal records this conversation and transmits it to the server as audio data.

[0838] 2. Analysis of audio data

[0839] The server analyzes the recorded voice data and obtains the text data "What are you having for dinner tonight?"

[0840] Extract individual words from text data: "Tonight," "What," and "Would you like to eat?"

[0841] 3. Extraction and translation of frequently occurring expressions

[0842] The server calculates the frequency of occurrence of all words extracted from this conversation, and identifies "Tonight" and "Shall we eat?" as frequently occurring expressions.

[0843] Translate these expressions into phrases such as "What will we eat tonight?" or "What's for dinner tonight?"

[0844] 4. Phrasebook generation and provision

[0845] The translated expressions are arranged in a phrasebook, such as "What will we eat tonight?" or "What's for dinner tonight?"

[0846] These phrasebooks are then notified to users, who can easily access them via an app or web portal to begin learning.

[0847] This system enables efficient foreign language learning tailored to individual user needs.

[0848] The processing flow will be explained below.

[0849] Step 1:

[0850] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone to capture the user's speech and saves it as a digital audio file (e.g., WAV format).

[0851] Step 2:

[0852] The device periodically (e.g., immediately after the conversation ends) sends the recorded voice data to the server, which is then securely transmitted over the Internet.

[0853] Step 3:

[0854] The server stores the received voice data in a storage device, such as a cloud storage device.

[0855] Step 4:

[0856] The server converts the stored voice data into text data using voice recognition technology, specifically by calling a voice recognition API.

[0857] Step 5:

[0858] The converted text data is then analyzed by the server using natural language processing (NLP) technology, which involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[0859] Step 6:

[0860] The server calculates the frequency of use of each extracted expression by comparing it with past analysis data stored in a database.

[0861] Step 7:

[0862] The server translates frequently used expressions into other languages ​​using a translation API (e.g., Google Translate API).

[0863] Step 8:

[0864] The translated expressions are sorted by frequency of use by the server and generated into a phrasebook.

[0865] Step 9:

[0866] The generated phrasebook is then sent to the user by the server via push notifications within the app or email.

[0867] Step 10:

[0868] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[0869] In this way, a process for efficiently learning a foreign language based on the user's conversation data is realized.

[0870] Example 1

[0871] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0872] In foreign language learning, there is a problem that it is difficult for individual users to efficiently learn expressions that they use on a daily basis. Conventional learning methods only allow learning of general phrases and words, making it difficult to learn in accordance with the needs of individual users. In addition, there is a problem that learning efficiency is reduced because there is no adequate system in place to automatically extract frequently used expressions and provide them as learning materials.

[0873] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0874] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data, means for converting the voice data into text data using natural language processing on the server, means for analyzing the converted text data and extracting words, phrases, short sentences, and sentence structure patterns, means for calculating the frequency of appearance of the extracted expressions and identifying frequently occurring expressions, means for translating the identified frequently occurring expressions into other languages, and means for arranging the translated expressions in order of frequency of use and generating a phrasebook, thereby enabling users to efficiently learn specific expressions that they use on a daily basis.

[0875] "User" refers to the person who uses the system to record conversations and learn translation phrases.

[0876] "Means for recording conversations" refers to a function that uses a microphone built into the device to record the user's speech in real time.

[0877] "Audio data" refers to the electronic data format in which a user's conversation is recorded, typically in file formats such as WAV or MP3.

[0878] "Server" refers to a remote computer system that processes received voice data and performs advanced calculations such as analysis and translation.

[0879] "Natural language processing" refers to the technology of analyzing voice data and text data to understand meaning and analyze grammar.

[0880] "Means for converting into text data" refers to speech recognition technology for converting voice data into text data.

[0881] "Analysis" refers to the process of breaking down text data and extracting words, phrases, short sentences, sentence structure patterns, etc.

[0882] "Means of extraction" refers to techniques for finding specific words or phrases from analyzed text data.

[0883] "Means for calculating frequency of occurrence and identifying frequently occurring expressions" refers to a technology that statistically calculates words and phrases that are particularly frequently used within text data and identifies important expressions.

[0884] "Translation means" refers to the function for converting frequently occurring expressions into the foreign language that the user wishes to learn.

[0885] "Means for sorting by frequency of use and generating a phrasebook" refers to a technology that sorts translated expressions based on their frequency of use to create a phrasebook for study.

[0886] "Means of notification" refers to the system used to provide the generated phrasebook to the user, including in-app notifications and email notifications.

[0887] The basic flow of an embodiment of this invention is as follows: First, a user's conversation is recorded in real time, and the recorded data is sent to a server. The server analyzes the received voice data using natural language processing technology, extracts and translates frequently occurring expressions, and then sorts the translated expressions in order of frequency of use to generate a phrasebook. Finally, this phrasebook is provided to the user.

[0888] Hardware and software used

[0889] The following hardware and software is recommended for implementing this system:

[0890] Hardware: Smartphones, smartwatches, smart glasses

[0891] Software: speech recognition technology (e.g., Google Speech-to-Text API), natural language processing technology, translation API (e.g., Google Translate API)

[0892] Specific operation of the system

[0893] Recording conversations and storing data

[0894] The device records the user's conversation in real time using the microphone built into the smartphone, smartwatch, or smart glasses. The recorded data is saved in WAV or MP3 format and sent to the server at regular intervals.

[0895] Analysis of audio data

[0896] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses technologies such as the Google Speech-to-Text API. The converted text data is then analyzed using natural language processing technology. The analysis process includes morphological and grammatical analysis, which allows for the extraction of words, phrases, short sentences, and sentence structure patterns.

[0897] Extraction and translation of frequently occurring expressions

[0898] The server identifies frequently used words and phrases from the analyzed text data by calculating their frequency of occurrence.These frequently used expressions are then translated into the foreign language the user wishes to learn using a translation API (e.g., Google Translate API).

[0899] Phrasebook generation and provision

[0900] The server then sorts the translated phrases into a phrasebook by frequency of use, and provides the phrasebook to users via a notification system, possibly via in-app notifications or email notifications.

[0901] Specific examples

[0902] The specific operation sequence is shown below.

[0903] Prompt Sentence Examples

[0904] User: "I want to eat curry rice today."

[0905] On the device: Records audio in real time and sends it to the server in the appropriate format (WAV or MP3).

[0906] Server: Converts the voice data into text and obtains the text data "I want to eat curry rice today."

[0907] Server: Analyzes using natural language processing technology and extracts "today," "curry rice," and "want to eat."

[0908] Server: Identify the frequently occurring expressions "today," "curry rice," and "want to eat" and translate them into English: "Today," "Curry rice," and "want to eat."

[0909] Server: Compiles the translated phrases into a phrasebook and notifies the user.

[0910] This system allows users to efficiently learn specific expressions used in daily life. In addition, because the generated phrasebooks are based on frequently occurring expressions, it provides practical and effective foreign language learning for users.

[0911] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0912] Step 1: Record the conversation

[0913] The device records the user's conversation in real time. The input is the user's voice, and the output is audio data. Specifically, when the user says, "What will you eat today?", the device's microphone picks up the audio and saves it as an audio data file (e.g., WAV format).

[0914] Step 2: Sending audio data

[0915] The device sends the recorded audio data to the server at regular intervals. The input is the audio data stored on the device, and the output is the audio data sent to the server. Specifically, the recorded data is uploaded to the server every five minutes.

[0916] Step 3: Convert audio data to text

[0917] The server converts the received voice data into text data using voice recognition technology (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data. Specifically, the server receives the voice saying "What will you eat today?" and converts it into the text "What will you eat today?"

[0918] Step 4: Analyzing the text data

[0919] The server analyzes the text data using natural language processing (NLP) technology to extract words, phrases, short sentences, and sentence structure patterns. The input is text data, and the output is the analyzed elements (words, phrases, short sentences, sentence structure patterns). Specifically, the server performs morphological analysis on the text "What will you eat today?" to extract the words "today," "what," and "will you eat."

[0920] Step 5: Extracting frequent expressions

[0921] The server identifies frequently used words and phrases from the analyzed text data. The input is the analyzed text data, and the output is frequently used expressions. Specifically, the server synthesizes past conversation data and creates a list of frequently used expressions such as "Today" and "Shall we eat?"

[0922] Step 6: Translating expressions

[0923] The server uses a translation API (e.g., Google Translate API) to translate frequently occurring expressions into the foreign language the user wishes to learn. The input is the frequently occurring expression, and the output is the translated expression. Specifically, the server translates "Kyou" into "Today" and "Taberu masuka" into "What will we eat?"

[0924] Step 7: Generate phrasebooks

[0925] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. The input is the translated expression, and the output is a phrasebook sorted in order of frequency of use. Specifically, the server lists phrases such as "What will we eat today?" and "What's for dinner tonight?" in order of frequency.

[0926] Step 8: Provide a phrasebook

[0927] The server provides the generated phrasebook to the user through a notification system. The input is the phrasebook, and the output is the notified phrasebook. Specifically, the server sends the generated phrasebook to the user via app notification or email notification, and the user confirms it.

[0928] (Application example 1)

[0929] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0930] In today's world, improving the efficiency of foreign language learning is an important challenge for many individuals. However, there are currently no efficient and personalized learning systems based on individual users' conversational content. Furthermore, there is a lack of systems that combine real-time conversation recording, analysis, translation, and notification functions. Therefore, there is a need for a system that encourages users to learn foreign languages ​​naturally in their daily lives.

[0931] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0932] In this invention, the server includes means for recording user conversations in real time and saving the voice data in an appropriate format, means for analyzing text data obtained from the recorded voice data and extracting words and short sentences, means for translating the extracted frequently occurring expressions into a foreign language in cooperation with a translation API via the server, and means for providing the generated phrasebook to the user via in-app notifications or email notifications, thereby enabling efficient and personalized foreign language learning based on the individual user's conversations.

[0933] "User" refers to a person who uses the system to record conversations and learn foreign languages.

[0934] "Means for recording conversations" refers to a function that uses a microphone to record a user's conversations in real time.

[0935] "Audio data" refers to data that is a recorded user's conversation stored in digital format.

[0936] "Means for transmitting to a server" refers to the technology for transferring recorded audio data to a remote server via the Internet.

[0937] "Natural language processing" refers to the general technology of analyzing voice data and converting it into text data.

[0938] "Frequent expressions" refer to words and phrases that are used particularly frequently in user conversations.

[0939] "Translation methods" refers to techniques and procedures for converting frequently occurring expressions into other languages.

[0940] "Means for generating phrasebooks" refers to a technology that organizes translated expressions in order of frequency of use and creates a phrasebook for study.

[0941] "Means for storage" refers to the ability to digitally store recorded audio data in an appropriate format (e.g., WAV, MP3).

[0942] "Text data" refers to the textual information obtained from the analyzed voice data.

[0943] "Translation API" refers to an interface that allows a program to provide translation services to other programs.

[0944] "In-app notifications" refers to a function that notifies users of information through smartphone applications.

[0945] "Email notification" refers to the means of providing information to users via email.

[0946] The present invention is a system for recording a user's conversation and using the recorded conversation to help the user learn a foreign language. An embodiment of the system will be described below.

[0947] System Overview

[0948] Users use smartphones or other devices (such as smart glasses or smart watches) that have built-in microphones and can record conversations continuously.

[0949] Server Roles

[0950] The server first receives the recorded audio data, which is periodically sent from the device via the Internet. The audio data is in WAV or MP3 format.

[0951] The server then uses speech recognition technology to convert the audio data into text data. Specifically, it uses a library called speech_recognition. After the audio data is converted into text, it analyzes the text data using natural language processing (NLP) technology to extract frequently occurring expressions. Words and short sentences are identified through morphological and grammatical analysis.

[0952] The server then translates the collected frequently used expressions into other languages, using the DeepL API, for example. The translated phrases are sorted by frequency of use and generated as a learning phrasebook.

[0953] Device Role

[0954] The device records the user's voice and sends the voice data to the server. The recording is done in real time and saved in an appropriate format. The server provides frequently used translation phrases to the user via in-app notifications and email notifications.

[0955] User Roles

[0956] By utilizing this system through everyday conversation, users receive a foreign language phrasebook tailored to their individual needs and progress with their learning.

[0957] Specific examples

[0958] For example, if a user says "I want to order from the virtual store" in a virtual store, this conversation is recorded and saved as audio data. The server then analyzes the text data "I want to order from the virtual store" and extracts the frequently occurring phrases "virtual store" and "I want to order." These phrases are translated into English via a translation API, resulting in "Virtual store" and "I want to order." These translation results are organized into a phrase collection and provided to the user via the smartphone app's notification function.

[0959] Prompt Sentence Examples

[0960] "Please translate the user-mentioned phrase "I want to order from a virtual store" into English."

[0961] This invention enables users to effectively learn a foreign language through everyday conversations. The overall system flow consists of a series of processes: recording the user's conversation, converting it into text, analyzing, translating, and notifying. The system provided by this invention aims to dramatically improve the efficiency of foreign language learning.

[0962] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0963] Step 1:

[0964] recording

[0965] The device records the user's conversation in real time using the device's built-in microphone to capture the user's speech, generating recorded voice data.

[0966] Input: User conversation

[0967] Output: Recorded audio data (e.g. WAV, MP3 format)

[0968] Step 2:

[0969] Sending audio data

[0970] The device transmits the recorded audio data to a server via the Internet at regular intervals. At this time, the data is encrypted for security purposes.

[0971] Input: Recorded audio data

[0972] Output: Audio data sent to the server

[0973] Step 3:

[0974] Voice Recognition

[0975] The server converts the received voice data into text data using natural language processing (NLP) techniques. Specifically, it uses the speech_recognition library to analyze the voice and extract the corresponding text information.

[0976] Input: Audio data sent to the server

[0977] Output: Text data (e.g., "I would like to order from a virtual store")

[0978] Step 4:

[0979] Text data analysis

[0980] The server analyzes the converted text data through morphological and grammatical analysis. It extracts words, phrases, and short sentences from the text data and calculates their frequency of occurrence. This process uses an NLP model.

[0981] Input: Text data

[0982] Output: Analyzed words and phrases and their frequency

[0983] Step 5:

[0984] Extraction of frequent expressions

[0985] The server identifies frequently used words and phrases from the analysis results. By extracting frequently occurring expressions, it identifies phrases and words that users frequently use.

[0986] Input: The words or phrases to be analyzed, and their frequency of occurrence

[0987] Output:Frequent expressions

[0988] Step 6:

[0989] translation

[0990] The server translates the extracted frequently occurring expressions into other languages ​​using a translation API (e.g., DeepL's API). By providing the frequently occurring expressions as prompts to the translation API, the server obtains the corresponding translation results.

[0991] Input:Frequent expressions

[0992] Output: The translated expression

[0993] Step 7:

[0994] Phrasebook generation

[0995] The server sorts the translated frequently occurring expressions in order of frequency of use to generate a phrasebook, which prioritizes important phrases to help users study efficiently.

[0996] Input: translated expression

[0997] Output: A collection of phrases sorted by frequency of use

[0998] Step 8:

[0999] notification

[1000] The server implements a means for notifying the user of the generated phrasebook, for example, by providing an in-app notification using a smartphone app or an email notification.

[1001] Input: A collection of phrases sorted by frequency of use

[1002] Output: User notification (e.g. in-app notification, email notification)

[1003] Through these processing steps, users can learn a foreign language in a personalized way based on everyday conversations. This series of processes aims to increase user convenience and maximize learning effectiveness.

[1004] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1005] An embodiment of the present invention is a system for recording a user's conversation and utilizing the recorded conversation for foreign language learning. The system also includes an emotion engine that recognizes the user's emotions, thereby further enhancing the learning experience. The system includes means for continuously recording the user's conversation, means for analyzing the recorded speech data using natural language processing, and means for extracting frequently occurring expressions and translating them into other languages. The system further includes means for sorting the translated expressions in order of frequency of use to generate a phrasebook, the emotion engine, and means for notifying the user of the generated phrasebook.

[1006] Specific operation of the system

[1007] Conversation recording and data storage

[1008] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[1009] The recorded audio data is sent to the server as a file at regular intervals, and is saved in an appropriate format (e.g., WAV format).

[1010] Analysis of audio data

[1011] The server converts the received voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API to perform the data conversion.

[1012] The converted text data is then analyzed using natural language processing (NLP) techniques, which involve morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1013] Emotion recognition

[1014] The server runs an emotion engine to recognize the user's emotions from the converted text and voice data. This engine detects emotions from the tone, pitch, and word choice of the voice.

[1015] For example, if a user is talking excitedly, this is recognized as "joy."

[1016] Extraction and translation of frequently occurring expressions

[1017] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[1018] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[1019] Phrasebook generation and refinement

[1020] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[1021] The server adjusts the content of the phrase book based on the perceived user emotion, for example, adding more challenging phrases if the user is perceived as relaxed.

[1022] Providing generated phrasebooks

[1023] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[1024] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[1025] Specific examples

[1026] Conversation recording and analysis

[1027] User: "It's such a beautiful day today!"

[1028] The terminal records this conversation and transmits it to the server as audio data.

[1029] The server converts the voice data into text and obtains the text data "What a beautiful day today!"

[1030] Emotion recognition and frequent expression extraction

[1031] The server's emotion engine detects the emotion "joy" from this text and voice tone.

[1032] The server extracts the frequently occurring expression "ii tenki" (good weather) from the text data and translates it to "Good weather."

[1033] Phrasebook generation and provision

[1034] The server adjusts the phrase collection based on the sentiment and generates a phrase collection that includes "Good weather."

[1035] The generated phrasebook is provided to the user as an in-app notification, and the user learns it through the app.

[1036] In this way, efficient foreign language learning according to the user's emotional state is realized.

[1037] The processing flow will be explained below.

[1038] Step 1:

[1039] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone and saves the user's speech as a digital audio file (e.g., WAV format).

[1040] Step 2:

[1041] The device periodically (for example, when the conversation ends) sends the recorded voice data to the server, which is then securely transferred over the Internet.

[1042] Step 3:

[1043] The server stores the received audio data in storage, which is expected to be cloud storage (e.g., AWS S3, Google Cloud Storage).

[1044] Step 4:

[1045] The server converts the stored voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[1046] Step 5:

[1047] The server analyzes the converted text data using natural language processing (NLP) techniques, including morphological and grammatical analysis, to extract words, phrases, short sentences, and sentence structure patterns.

[1048] Step 6:

[1049] The server uses the extracted text data to run an emotion engine to recognize the user's emotions, which analyzes the text's wording and the tone and pitch of the voice.

[1050] Step 7:

[1051] The server extracts frequently used expressions from the speech data and sentiment analysis data, specifically by calculating the frequency of occurrence by comparing them with past data stored in a database, and identifying frequently occurring expressions.

[1052] Step 8:

[1053] The server translates the extracted frequently occurring expressions into foreign languages ​​(e.g., English, Korean) using a translation API, such as Google Translate API.

[1054] Step 9:

[1055] The server sorts the translated expressions by frequency of use to generate a phrasebook, which is also adjusted based on the user's emotions—for example, if the user is tired, simpler phrases are prioritized.

[1056] Step 10:

[1057] The server notifies the user of the generated phrasebook via in-app notifications, push notifications, email, etc.

[1058] Step 11:

[1059] Users can access the phrasebooks provided through their devices and efficiently study foreign languages. The phrasebooks can be viewed through an app or web portal.

[1060] Specific examples

[1061] 1. Conversation recording: A user says, "What a beautiful day today!"

[1062] 2. Sending audio: The device records the conversation and sends the audio data to the server.

[1063] 3. Analysis of voice data: The server converts the voice data into text "What a beautiful day today!"

[1064] 4. Emotion recognition: The server uses the emotion engine to recognize the emotion of "joy."

[1065] 5. Extraction of frequently occurring expressions: The server extracts the frequently occurring expression "ii tenki" and translates it as "Good weather."

[1066] 6. Phrasebook generation and adjustment: The server generates a phrasebook adjusted based on emotions to help the user relax.

[1067] 7. Phrasebook notification and learning: The server notifies the user of the generated phrasebook via an in-app notification, and the user learns it.

[1068] This system enables efficient foreign language learning based on the user's emotions.

[1069] Example 2

[1070] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1071] Conventional foreign language learning systems have the problem of reducing learning efficiency because they provide uniform learning content without considering the user's emotional state. Furthermore, there is no established method for effectively extracting and translating expressions frequently used by users in real conversations and generating a learning phrasebook based on them. This often leads to a decrease in users' motivation to learn and results in poor learning outcomes.

[1072] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and text data to recognize the user's emotions, means for adjusting the content of the phrase book based on the emotions, and means for notifying the user of the generated phrase book. This enables efficient foreign language learning according to the user's emotional state.

[1073] "User" refers to an individual who uses the system to learn.

[1074] "Means for recording conversations" refers to a device or system that records the user's voice in real time.

[1075] "Means for transmitting audio data to a server" refers to the technology or protocol for transferring recorded audio data to a server over a network.

[1076] "Natural language processing" refers to computer technology that analyzes voice and text data to understand and process their content.

[1077] "Means for extracting frequent expressions" refers to a technique for selecting specific words or phrases from the analyzed data based on their frequency.

[1078] "Means for translating into other languages" refers to a technology or system that translates the extracted expressions into a different language.

[1079] "Means for generating phrasebooks" refers to techniques for organizing translated expressions into a single collection for study.

[1080] "Means for recognizing user emotions" refers to technologies and systems that analyze voice data and text data to detect the user's emotional state.

[1081] "Means for adjusting phrase book content based on emotion" refers to a technique for changing or adjusting the content or difficulty of a phrase book in accordance with a detected emotional state.

[1082] "Means of notification" refers to the technology or system used to notify users of the generated phrasebook.

[1083] The present invention relates to a system that records user conversations in real time and analyzes and translates the data for foreign language learning. The system is designed to recognize the user's emotions and further enhance the learning experience.

[1084] Conversation recording and data storage

[1085] The device records the user's conversation in real time using a microphone built into a smartphone, smartwatch, smart glasses, etc. The recorded voice data is sent to a server as a file in an appropriate format such as WAV at regular intervals.

[1086] Examples:

[1087] A user says, "What a beautiful day today!"

[1088] The device records this conversation, saves the audio data with the file name "2023-10-01-1230.wav", and sends it to the server.

[1089] Analysis of audio data

[1090] The server converts the received voice data into text data using voice recognition technology. Specifically, it converts the voice data into text using a common voice recognition API (e.g., Google Cloud Speech-to-Text API). The converted text data is then analyzed using natural language processing (NLP) technology. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1091] Examples:

[1092] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API and receives the text data "What a beautiful day today!" as a result.

[1093] The server performs morphological analysis on the text data "The weather is so nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[1094] Emotion recognition

[1095] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. For example, it uses a common emotion analysis tool (e.g., IBM Watson Tone Analyzer) to detect emotions from the tone, pitch, and choice of words of the voice.

[1096] Examples:

[1097] The server detects the emotion of "joy" from the text data and tone of voice.

[1098] Extraction and translation of frequently occurring expressions

[1099] The server identifies frequently used words and phrases from the analysis results, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using common translation tools (e.g., Google Cloud Translation API).

[1100] Examples:

[1101] The server extracts the frequently occurring expression "nice weather" based on past analysis results.

[1102] The server translates "ii tenki" into "Good weather."

[1103] Phrasebook generation and provision

[1104] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook. The server adjusts the contents of the phrasebook based on the user's recognized emotions. For example, if the server recognizes that the user is relaxed, more difficult phrases will be added. The generated phrasebook is provided to the user through a notification system. For example, this could be an in-app notification or an email notification.

[1105] Examples:

[1106] The server generates a collection of phrases containing "Good weather" based on the sentiment analysis results.

[1107] The server sends the generated phrasebook as an in-app notification.

[1108] Users can browse and learn from the phrasebook "Good weather" provided through the app.

[1109] Prompt Sentence Examples

[1110] If a user says "What a beautiful day today!" during a conversation, use the following prompt example:

[1111] "The user says, 'The weather is so nice today!' Based on this conversation, extract frequently occurring expressions, perform emotion recognition, and generate a phrasebook for foreign language learning based on the results."

[1112] This enables efficient foreign language learning according to the user's emotional state.

[1113] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1114] Step 1:

[1115] The device records the user's conversation in real time. Specific devices include smartphones, smart watches, and smart glasses. The input is the user's conversational voice, and the output is the recorded voice data.

[1116] Specific behavior:

[1117] A user says, "What a beautiful day today!"

[1118] The device records this conversation and saves it as audio data.

[1119] Step 2:

[1120] The device sends the recorded audio data to the server as a file at regular intervals. The audio data format is WAV, etc. The input is the recorded audio data, and the output is the file sent to the server.

[1121] Specific behavior:

[1122] The device saves the audio data with the file name "2023-10-01-1230.wav" and sends it to the server.

[1123] Step 3:

[1124] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API, etc. The input is voice data, and the output is the converted text data.

[1125] Specific behavior:

[1126] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API, and receives the text data "What a beautiful day today!" as a result.

[1127] Step 4:

[1128] The server analyzes the converted text data using natural language processing (NLP) technology. As a result of the analysis, morphological analysis is performed to extract words, phrases, and sentence structure patterns. The input is text data, and the output is analyzed information (extracted words and phrases).

[1129] Specific behavior:

[1130] The server performs morphological analysis on the text data "The weather is very nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[1131] Step 5:

[1132] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. Specifically, emotions are detected from the tone, pitch, and word choice of the voice. The input is text and voice data, and the output is the recognized emotion.

[1133] Specific behavior:

[1134] The server detects the emotion of "joy" from the text data and tone of voice.

[1135] Step 6:

[1136] The server identifies frequently used words and phrases from the analysis results. Specifically, it calculates their frequency of occurrence. The input is the analyzed information (words and phrases), and the output is frequently used expressions.

[1137] Specific behavior:

[1138] The server extracts the frequently occurring expression "nice weather" based on the analysis results.

[1139] Step 7:

[1140] The server translates the identified frequent expressions into other languages ​​using a translation API. Specifically, it uses the Google Cloud Translation API. The input is the frequent expression, and the output is the translated expression.

[1141] Specific behavior:

[1142] The server translates "ii tenki" into "Good weather."

[1143] Step 8:

[1144] The server sorts the translated expressions by frequency of use to generate a phrasebook. The content of the phrasebook is adjusted based on the recognized emotions. The input is the translated expressions and emotional information, and the output is an adjusted phrasebook.

[1145] Specific behavior:

[1146] The server generates a list of phrases containing "Good weather" based on the results of sentiment analysis.

[1147] Step 9:

[1148] The server provides the generated phrasebook to the user through a notification system: the input is the generated phrasebook, and the output is the notified phrasebook.

[1149] Specific behavior:

[1150] The server sends the phrasebook as an in-app notification.

[1151] Users can browse and learn from a collection of phrases called "Good weather" provided through the app.

[1152] (Application example 2)

[1153] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1154] Conventional foreign language learning systems have the problem of being difficult to use efficiently because they do not take into account the user's emotional state or the context of real-life conversations. In particular, there is a need for a system that can analyze frequently used expressions in real time and provide appropriate learning content according to the user's emotions.

[1155] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1156] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data to the server, means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions, means for translating the extracted expressions into other languages, means for arranging the translated expressions in order of frequency of use and generating a phrasebook, means for displaying the phrasebook to the user via smart glasses, and means for recognizing the user's emotions from the analyzed data. This makes it possible to provide emotion-sensitive foreign language learning in real time based on the user's real-life conversation data.

[1157] "Means for recording user conversations" refers to devices or technologies for collecting voice data and storing user utterances.

[1158] The "means for transmitting recorded voice data to a server" refers to a device or method for transferring collected voice data to a server via a network.

[1159] "Means of analyzing voice data on a server using natural language processing and extracting frequently occurring expressions" refers to a technology that converts voice data into text data and identifies frequently used words and phrases.

[1160] The "means for translating extracted expressions into other languages" is a method for converting identified words and phrases into the foreign language that the user wishes to learn.

[1161] "Means for sorting translated expressions in order of frequency of use and generating a phrasebook" refers to a technology that sorts translated words and phrases in order of frequency and provides them as a single, comprehensive learning content.

[1162] The "means for displaying a phrasebook to a user via smart glasses" refers to a technique for displaying the generated phrasebook on the display of smart glasses worn by the user.

[1163] "Means for recognizing user emotions from analyzed data" refers to engines or technologies for identifying a user's emotional state based on voice data and text data.

[1164] The present invention provides a system for recording and analyzing a user's conversations and supporting foreign language learning based on the user's emotional state. The system includes: a means for recording the user's conversations; a means for transmitting the recorded voice data to a server; a means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions; a means for translating the extracted expressions into other languages; a means for sorting the translated expressions in order of frequency of use and generating a phrasebook; a means for displaying the phrasebook to the user via smart glasses; and a means for recognizing the user's emotions from the analyzed data.

[1165] First, when a user wears smart glasses and engages in everyday conversation, the microphone built into the smart glasses collects voice data. This voice data is recorded in real time and sent to a server via a network. The voice data is then saved in an appropriate format (e.g., WAV format).

[1166] Next, on the server, the received voice data is converted into text data using speech recognition technology. Specifically, data conversion is performed using a speech recognition API such as the Google Speech-to-Text API. This converted text data is then analyzed using a natural language processing (NLP) engine, such as spaCy. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1167] Next, based on the analyzed text and voice data, an emotion recognition engine, such as IBM Watson Tone Analyzer, is used to recognize the user's emotions. This engine detects emotions from the tone, pitch, and word choice of the voice. For example, if the user is speaking with an emotion of joy, it will be recognized as "joy."

[1168] The server then uses the analysis results to identify frequently used words and phrases, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using a translation API, such as the Google Cloud Translation API.

[1169] The translated frequently occurring expressions are organized in order of frequency and generated into a phrasebook. This phrasebook is sent in real time from the server to the smart glasses display and displayed to the user. This allows users to learn a foreign language naturally through everyday conversation.

[1170] As a specific example of use, if a user says, "What a beautiful day today!", this voice is recorded by the smart glasses and sent to the server. The speech recognition system on the server converts this voice data into text, obtaining the text data "What a beautiful day today!". Furthermore, the emotion recognition engine detects the emotion of "joy," extracts the frequently occurring expression "nice weather," and translates it into "Good weather." The translated phrases are organized in order of frequency and displayed on the smart glasses.

[1171] Through this process, efficient foreign language learning that responds to emotions can be realized based on the user's daily conversation data.

[1172] An example prompt might look like this:

[1173] "Your smart glasses are currently recording everyday conversations. Please explain the process of extracting common expressions from these conversations, translating them into English, and creating a phrasebook based on emotions."

[1174] As described above, the present invention is a system that provides emotion-responsive foreign language learning in real time based on conversation data from the user's real life.

[1175] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1176] Step 1:

[1177] The device records the user's conversation in real time. The device's built-in microphone collects audio data and saves it in a file format (e.g., WAV format). This outputs an audio file, which is raw data.

[1178] Step 2:

[1179] The device sends the recorded audio data to the server at regular intervals. The audio file is uploaded to the server via the network. The audio file received as input is transferred to the server, where it is saved.

[1180] Step 3:

[1181] The server converts the received voice data into text data using a speech recognition API (for example, Google Speech-to-Text API). It receives the voice file as input and performs data conversion processing to generate text data as output.

[1182] Step 4:

[1183] The server analyzes the generated text data using a natural language processing (NLP) engine (such as spaCy). During the analysis, morphological analysis is performed to extract words, phrases, short sentences, and sentence structure patterns. By receiving the text data as input and performing data analysis, the analysis results are obtained as output.

[1184] Step 5:

[1185] The server uses an emotion recognition engine (such as IBM Watson Tone Analyzer) to recognize the user's emotions based on the analyzed text data and voice data. It receives text data and voice parameters as input, performs emotion recognition processing, and generates emotion data as output.

[1186] Step 6:

[1187] The server identifies frequently used words and phrases from the analysis results, which includes calculating their frequency of occurrence. It receives the analysis results as input, extracts frequently occurring expressions, and outputs a list of frequently occurring expressions.

[1188] Step 7:

[1189] The server translates the extracted frequent expressions into other languages ​​using a translation API (e.g., Google Cloud Translation API). It receives a list of frequent expressions as input and performs translation processing, generating a list of translated expressions as output.

[1190] Step 8:

[1191] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. It receives the list of translated expressions as input, performs sorting, and obtains a phrasebook as output.

[1192] Step 9:

[1193] The server sends the generated phrase book to the smart glasses display in real time and displays it to the user. The server receives the phrase book as input and performs communication processing, and the phrase book is displayed on the smart glasses as output.

[1194] This series of processes enables users to efficiently learn a foreign language through real-life conversations.

[1195] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1196] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1197] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1198] [Fourth embodiment]

[1199] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1200] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1201] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1202] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1203] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1205] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1206] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1207] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1208] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1209] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1210] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1211] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1212] An embodiment of this invention is a system for recording user conversations and using them to help with foreign language learning. This system includes a means for constantly recording user conversations, a means for analyzing the recorded voice data using natural language processing, a means for extracting frequently occurring expressions and translating them into other languages, and a means for sorting the translated expressions in order of frequency of use to generate a phrasebook.

[1213] Specific operation of the system

[1214] Conversation recording and data storage

[1215] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[1216] The recorded audio data is sent to the server as a file at regular intervals, where it is saved in an appropriate format (e.g., WAV, MP3).

[1217] Analysis of audio data

[1218] The server analyzes the received voice data using natural language processing (NLP) technology. Specifically, it converts the voice data into text data using speech recognition technology.

[1219] The converted text data undergoes morphological and grammatical analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1220] Extraction and translation of frequently occurring expressions

[1221] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[1222] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[1223] Phrasebook generation and provision

[1224] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[1225] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[1226] Specific examples

[1227] 1. Record and send conversations

[1228] User: "What are you having for dinner tonight?"

[1229] The terminal records this conversation and transmits it to the server as audio data.

[1230] 2. Analysis of audio data

[1231] The server analyzes the recorded voice data and obtains the text data "What are you having for dinner tonight?"

[1232] Extract individual words from text data: "Tonight," "What," and "Would you like to eat?"

[1233] 3. Extraction and translation of frequently occurring expressions

[1234] The server calculates the frequency of occurrence of all words extracted from this conversation, and identifies "Tonight" and "Shall we eat?" as frequently occurring expressions.

[1235] Translate these expressions into phrases such as "What will we eat tonight?" or "What's for dinner tonight?"

[1236] 4. Phrasebook generation and provision

[1237] The translated expressions are arranged in a phrasebook, such as "What will we eat tonight?" or "What's for dinner tonight?"

[1238] These phrasebooks are then notified to users, who can easily access them via an app or web portal to begin learning.

[1239] This system enables efficient foreign language learning tailored to individual user needs.

[1240] The processing flow will be explained below.

[1241] Step 1:

[1242] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone to capture the user's speech and saves it as a digital audio file (e.g., WAV format).

[1243] Step 2:

[1244] The device periodically (e.g., immediately after the conversation ends) sends the recorded voice data to the server, which is then securely transmitted over the Internet.

[1245] Step 3:

[1246] The server stores the received voice data in a storage device, such as a cloud storage device.

[1247] Step 4:

[1248] The server converts the stored voice data into text data using voice recognition technology, specifically by calling a voice recognition API.

[1249] Step 5:

[1250] The converted text data is then analyzed by the server using natural language processing (NLP) technology, which involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1251] Step 6:

[1252] The server calculates the frequency of use of each extracted expression by comparing it with past analysis data stored in a database.

[1253] Step 7:

[1254] The server translates frequently used expressions into other languages ​​using a translation API (e.g., Google Translate API).

[1255] Step 8:

[1256] The translated expressions are sorted by frequency of use by the server and generated into a phrasebook.

[1257] Step 9:

[1258] The generated phrasebook is then sent to the user by the server via push notifications within the app or email.

[1259] Step 10:

[1260] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[1261] In this way, a process for efficiently learning a foreign language based on the user's conversation data is realized.

[1262] Example 1

[1263] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1264] In foreign language learning, there is a problem that it is difficult for individual users to efficiently learn expressions that they use on a daily basis. Conventional learning methods only allow learning of general phrases and words, making it difficult to learn in accordance with the needs of individual users. In addition, there is a problem that learning efficiency is reduced because there is no adequate system in place to automatically extract frequently used expressions and provide them as learning materials.

[1265] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1266] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data, means for converting the voice data into text data using natural language processing on the server, means for analyzing the converted text data and extracting words, phrases, short sentences, and sentence structure patterns, means for calculating the frequency of appearance of the extracted expressions and identifying frequently occurring expressions, means for translating the identified frequently occurring expressions into other languages, and means for arranging the translated expressions in order of frequency of use and generating a phrasebook, thereby enabling users to efficiently learn specific expressions that they use on a daily basis.

[1267] "User" refers to the person who uses the system to record conversations and learn translation phrases.

[1268] "Means for recording conversations" refers to a function that uses a microphone built into the device to record the user's speech in real time.

[1269] "Audio data" refers to the electronic data format in which a user's conversation is recorded, typically in file formats such as WAV or MP3.

[1270] "Server" refers to a remote computer system that processes received voice data and performs advanced calculations such as analysis and translation.

[1271] "Natural language processing" refers to the technology of analyzing voice data and text data to understand meaning and analyze grammar.

[1272] "Means for converting into text data" refers to speech recognition technology for converting voice data into text data.

[1273] "Analysis" refers to the process of breaking down text data and extracting words, phrases, short sentences, sentence structure patterns, etc.

[1274] "Means of extraction" refers to techniques for finding specific words or phrases from analyzed text data.

[1275] "Means for calculating frequency of occurrence and identifying frequently occurring expressions" refers to a technology that statistically calculates words and phrases that are particularly frequently used within text data and identifies important expressions.

[1276] "Translation means" refers to the function for converting frequently occurring expressions into the foreign language that the user wishes to learn.

[1277] "Means for sorting by frequency of use and generating a phrasebook" refers to a technology that sorts translated expressions based on their frequency of use to create a phrasebook for study.

[1278] "Means of notification" refers to the system used to provide the generated phrasebook to the user, including in-app notifications and email notifications.

[1279] The basic flow of an embodiment of this invention is as follows: First, a user's conversation is recorded in real time, and the recorded data is sent to a server. The server analyzes the received voice data using natural language processing technology, extracts and translates frequently occurring expressions, and then sorts the translated expressions in order of frequency of use to generate a phrasebook. Finally, this phrasebook is provided to the user.

[1280] Hardware and software used

[1281] The following hardware and software is recommended for implementing this system:

[1282] Hardware: Smartphones, smartwatches, smart glasses

[1283] Software: speech recognition technology (e.g., Google Speech-to-Text API), natural language processing technology, translation API (e.g., Google Translate API)

[1284] Specific operation of the system

[1285] Recording conversations and storing data

[1286] The device records the user's conversation in real time using the microphone built into the smartphone, smartwatch, or smart glasses. The recorded data is saved in WAV or MP3 format and sent to the server at regular intervals.

[1287] Analysis of audio data

[1288] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses technologies such as the Google Speech-to-Text API. The converted text data is then analyzed using natural language processing technology. The analysis process includes morphological and grammatical analysis, which allows for the extraction of words, phrases, short sentences, and sentence structure patterns.

[1289] Extraction and translation of frequently occurring expressions

[1290] The server identifies frequently used words and phrases from the analyzed text data by calculating their frequency of occurrence.These frequently used expressions are then translated into the foreign language the user wishes to learn using a translation API (e.g., Google Translate API).

[1291] Phrasebook generation and provision

[1292] The server then sorts the translated phrases into a phrasebook by frequency of use, and provides the phrasebook to users via a notification system, possibly via in-app notifications or email notifications.

[1293] Specific examples

[1294] The specific operation sequence is shown below.

[1295] Prompt Sentence Examples

[1296] User: "I want to eat curry rice today."

[1297] On the device: Records audio in real time and sends it to the server in the appropriate format (WAV or MP3).

[1298] Server: Converts the voice data into text and obtains the text data "I want to eat curry rice today."

[1299] Server: Analyzes using natural language processing technology and extracts "today," "curry rice," and "want to eat."

[1300] Server: Identify the frequently occurring expressions "today," "curry rice," and "want to eat" and translate them into English: "Today," "Curry rice," and "want to eat."

[1301] Server: Compiles the translated phrases into a phrasebook and notifies the user.

[1302] This system allows users to efficiently learn specific expressions used in daily life. In addition, because the generated phrasebooks are based on frequently occurring expressions, it provides practical and effective foreign language learning for users.

[1303] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1304] Step 1: Record the conversation

[1305] The device records the user's conversation in real time. The input is the user's voice, and the output is audio data. Specifically, when the user says, "What will you eat today?", the device's microphone picks up the audio and saves it as an audio data file (e.g., WAV format).

[1306] Step 2: Sending audio data

[1307] The device sends the recorded audio data to the server at regular intervals. The input is the audio data stored on the device, and the output is the audio data sent to the server. Specifically, the recorded data is uploaded to the server every five minutes.

[1308] Step 3: Convert audio data to text

[1309] The server converts the received voice data into text data using voice recognition technology (e.g., Google Speech-to-Text API). The input is voice data, and the output is text data. Specifically, the server receives the voice saying "What will you eat today?" and converts it into the text "What will you eat today?"

[1310] Step 4: Analyzing the text data

[1311] The server analyzes the text data using natural language processing (NLP) technology to extract words, phrases, short sentences, and sentence structure patterns. The input is text data, and the output is the analyzed elements (words, phrases, short sentences, sentence structure patterns). Specifically, the server performs morphological analysis on the text "What will you eat today?" to extract the words "today," "what," and "will you eat."

[1312] Step 5: Extracting frequent expressions

[1313] The server identifies frequently used words and phrases from the analyzed text data. The input is the analyzed text data, and the output is frequently used expressions. Specifically, the server synthesizes past conversation data and creates a list of frequently used expressions such as "Today" and "Shall we eat?"

[1314] Step 6: Translating expressions

[1315] The server uses a translation API (e.g., Google Translate API) to translate frequently occurring expressions into the foreign language the user wishes to learn. The input is the frequently occurring expression, and the output is the translated expression. Specifically, the server translates "Kyou" into "Today" and "Taberu masuka" into "What will we eat?"

[1316] Step 7: Generate phrasebooks

[1317] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. The input is the translated expression, and the output is a phrasebook sorted in order of frequency of use. Specifically, the server lists phrases such as "What will we eat today?" and "What's for dinner tonight?" in order of frequency.

[1318] Step 8: Provide a phrasebook

[1319] The server provides the generated phrasebook to the user through a notification system. The input is the phrasebook, and the output is the notified phrasebook. Specifically, the server sends the generated phrasebook to the user via app notification or email notification, and the user confirms it.

[1320] (Application example 1)

[1321] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1322] In today's world, improving the efficiency of foreign language learning is an important challenge for many individuals. However, there are currently no efficient and personalized learning systems based on individual users' conversational content. Furthermore, there is a lack of systems that combine real-time conversation recording, analysis, translation, and notification functions. Therefore, there is a need for a system that encourages users to learn foreign languages ​​naturally in their daily lives.

[1323] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1324] In this invention, the server includes means for recording user conversations in real time and saving the voice data in an appropriate format, means for analyzing text data obtained from the recorded voice data and extracting words and short sentences, means for translating the extracted frequently occurring expressions into a foreign language in cooperation with a translation API via the server, and means for providing the generated phrasebook to the user via in-app notifications or email notifications, thereby enabling efficient and personalized foreign language learning based on the individual user's conversations.

[1325] "User" refers to a person who uses the system to record conversations and learn foreign languages.

[1326] "Means for recording conversations" refers to a function that uses a microphone to record a user's conversations in real time.

[1327] "Audio data" refers to data that is a recorded user's conversation stored in digital format.

[1328] "Means for transmitting to a server" refers to the technology for transferring recorded audio data to a remote server via the Internet.

[1329] "Natural language processing" refers to the general technology of analyzing voice data and converting it into text data.

[1330] "Frequent expressions" refer to words and phrases that are used particularly frequently in user conversations.

[1331] "Translation methods" refers to techniques and procedures for converting frequently occurring expressions into other languages.

[1332] "Means for generating phrasebooks" refers to a technology that organizes translated expressions in order of frequency of use and creates a phrasebook for study.

[1333] "Means for storage" refers to the ability to digitally store recorded audio data in an appropriate format (e.g., WAV, MP3).

[1334] "Text data" refers to the textual information obtained from the analyzed voice data.

[1335] "Translation API" refers to an interface that allows a program to provide translation services to other programs.

[1336] "In-app notifications" refers to a function that notifies users of information through smartphone applications.

[1337] "Email notification" refers to the means of providing information to users via email.

[1338] The present invention is a system for recording a user's conversation and using the recorded conversation to help the user learn a foreign language. An embodiment of the system will be described below.

[1339] System Overview

[1340] Users use smartphones or other devices (such as smart glasses or smart watches) that have built-in microphones and can record conversations continuously.

[1341] Server Roles

[1342] The server first receives the recorded audio data, which is periodically sent from the device via the Internet. The audio data is in WAV or MP3 format.

[1343] The server then uses speech recognition technology to convert the audio data into text data. Specifically, it uses a library called speech_recognition. After the audio data is converted into text, it analyzes the text data using natural language processing (NLP) technology to extract frequently occurring expressions. Words and short sentences are identified through morphological and grammatical analysis.

[1344] The server then translates the collected frequently used expressions into other languages, using the DeepL API, for example. The translated phrases are sorted by frequency of use and generated as a learning phrasebook.

[1345] Device Role

[1346] The device records the user's voice and sends the voice data to the server. The recording is done in real time and saved in an appropriate format. The server provides frequently used translation phrases to the user via in-app notifications and email notifications.

[1347] User Roles

[1348] By utilizing this system through everyday conversation, users receive a foreign language phrasebook tailored to their individual needs and progress with their learning.

[1349] Specific examples

[1350] For example, if a user says "I want to order from the virtual store" in a virtual store, this conversation is recorded and saved as audio data. The server then analyzes the text data "I want to order from the virtual store" and extracts the frequently occurring phrases "virtual store" and "I want to order." These phrases are translated into English via a translation API, resulting in "Virtual store" and "I want to order." These translation results are organized into a phrase collection and provided to the user via the smartphone app's notification function.

[1351] Prompt Sentence Examples

[1352] "Please translate the user-mentioned phrase "I want to order from a virtual store" into English."

[1353] This invention enables users to effectively learn a foreign language through everyday conversations. The overall system flow consists of a series of processes: recording the user's conversation, converting it into text, analyzing, translating, and notifying. The system provided by this invention aims to dramatically improve the efficiency of foreign language learning.

[1354] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1355] Step 1:

[1356] recording

[1357] The device records the user's conversation in real time using the device's built-in microphone to capture the user's speech, generating recorded voice data.

[1358] Input: User conversation

[1359] Output: Recorded audio data (e.g. WAV, MP3 format)

[1360] Step 2:

[1361] Sending audio data

[1362] The device transmits the recorded audio data to a server via the Internet at regular intervals. At this time, the data is encrypted for security purposes.

[1363] Input: Recorded audio data

[1364] Output: Audio data sent to the server

[1365] Step 3:

[1366] Voice Recognition

[1367] The server converts the received voice data into text data using natural language processing (NLP) techniques. Specifically, it uses the speech_recognition library to analyze the voice and extract the corresponding text information.

[1368] Input: Audio data sent to the server

[1369] Output: Text data (e.g., "I would like to order from a virtual store")

[1370] Step 4:

[1371] Text data analysis

[1372] The server analyzes the converted text data through morphological and grammatical analysis. It extracts words, phrases, and short sentences from the text data and calculates their frequency of occurrence. This process uses an NLP model.

[1373] Input: Text data

[1374] Output: Analyzed words and phrases and their frequency

[1375] Step 5:

[1376] Extraction of frequent expressions

[1377] The server identifies frequently used words and phrases from the analysis results. By extracting frequently occurring expressions, it identifies phrases and words that users frequently use.

[1378] Input: The words or phrases to be analyzed, and their frequency of occurrence

[1379] Output:Frequent expressions

[1380] Step 6:

[1381] translation

[1382] The server translates the extracted frequently occurring expressions into other languages ​​using a translation API (e.g., DeepL's API). By providing the frequently occurring expressions as prompts to the translation API, the server obtains the corresponding translation results.

[1383] Input:Frequent expressions

[1384] Output: The translated expression

[1385] Step 7:

[1386] Phrasebook generation

[1387] The server sorts the translated frequently occurring expressions in order of frequency of use to generate a phrasebook, which prioritizes important phrases to help users study efficiently.

[1388] Input: translated expression

[1389] Output: A collection of phrases sorted by frequency of use

[1390] Step 8:

[1391] notification

[1392] The server implements a means for notifying the user of the generated phrasebook, for example, by providing an in-app notification using a smartphone app or an email notification.

[1393] Input: A collection of phrases sorted by frequency of use

[1394] Output: User notification (e.g. in-app notification, email notification)

[1395] Through these processing steps, users can learn a foreign language in a personalized way based on everyday conversations. This series of processes aims to increase user convenience and maximize learning effectiveness.

[1396] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1397] An embodiment of the present invention is a system for recording a user's conversation and utilizing the recorded conversation for foreign language learning. The system also includes an emotion engine that recognizes the user's emotions, thereby further enhancing the learning experience. The system includes means for continuously recording the user's conversation, means for analyzing the recorded speech data using natural language processing, and means for extracting frequently occurring expressions and translating them into other languages. The system further includes means for sorting the translated expressions in order of frequency of use to generate a phrasebook, the emotion engine, and means for notifying the user of the generated phrasebook.

[1398] Specific operation of the system

[1399] Conversation recording and data storage

[1400] The device records users' conversations in real time using microphones built into smartphones, smartwatches, and smart glasses.

[1401] The recorded audio data is sent to the server as a file at regular intervals, and is saved in an appropriate format (e.g., WAV format).

[1402] Analysis of audio data

[1403] The server converts the received voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API to perform the data conversion.

[1404] The converted text data is then analyzed using natural language processing (NLP) techniques, which involve morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1405] Emotion recognition

[1406] The server runs an emotion engine to recognize the user's emotions from the converted text and voice data. This engine detects emotions from the tone, pitch, and word choice of the voice.

[1407] For example, if a user is talking excitedly, this is recognized as "joy."

[1408] Extraction and translation of frequently occurring expressions

[1409] The server uses the analysis to identify frequently used words and phrases, which includes calculating frequency of occurrence.

[1410] The identified frequent expressions are translated into the foreign language (e.g., English, Korean) that the user wishes to learn using a translation API.

[1411] Phrasebook generation and refinement

[1412] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook.

[1413] The server adjusts the content of the phrase book based on the perceived user emotion, for example, adding more challenging phrases if the user is perceived as relaxed.

[1414] Providing generated phrasebooks

[1415] The generated phrasebook is then provided to the user through a notification system, for example, an in-app notification or an email notification.

[1416] Users can access the phrasebooks provided through their devices to efficiently study foreign languages. The phrasebooks are made available for viewing through apps and web portals.

[1417] Specific examples

[1418] Conversation recording and analysis

[1419] User: "It's such a beautiful day today!"

[1420] The terminal records this conversation and transmits it to the server as audio data.

[1421] The server converts the voice data into text and obtains the text data "What a beautiful day today!"

[1422] Emotion recognition and frequent expression extraction

[1423] The server's emotion engine detects the emotion "joy" from this text and voice tone.

[1424] The server extracts the frequently occurring expression "ii tenki" (good weather) from the text data and translates it to "Good weather."

[1425] Phrasebook generation and provision

[1426] The server adjusts the phrase collection based on the sentiment and generates a phrase collection that includes "Good weather."

[1427] The generated phrasebook is provided to the user as an in-app notification, and the user learns it through the app.

[1428] In this way, efficient foreign language learning according to the user's emotional state is realized.

[1429] The processing flow will be explained below.

[1430] Step 1:

[1431] The device records the user's conversation as audio. Specifically, it uses the device's built-in microphone and saves the user's speech as a digital audio file (e.g., WAV format).

[1432] Step 2:

[1433] The device periodically (for example, when the conversation ends) sends the recorded voice data to the server, which is then securely transferred over the Internet.

[1434] Step 3:

[1435] The server stores the received audio data in storage, which is expected to be cloud storage (e.g., AWS S3, Google Cloud Storage).

[1436] Step 4:

[1437] The server converts the stored voice data into text data using voice recognition technology. Specifically, it calls a voice recognition API (e.g., Google Speech-to-Text API) to convert the voice data into text data.

[1438] Step 5:

[1439] The server analyzes the converted text data using natural language processing (NLP) techniques, including morphological and grammatical analysis, to extract words, phrases, short sentences, and sentence structure patterns.

[1440] Step 6:

[1441] The server uses the extracted text data to run an emotion engine to recognize the user's emotions, which analyzes the text's wording and the tone and pitch of the voice.

[1442] Step 7:

[1443] The server extracts frequently used expressions from the speech data and sentiment analysis data, specifically by calculating the frequency of occurrence by comparing them with past data stored in a database, and identifying frequently occurring expressions.

[1444] Step 8:

[1445] The server translates the extracted frequently occurring expressions into foreign languages ​​(e.g., English, Korean) using a translation API, such as Google Translate API.

[1446] Step 9:

[1447] The server sorts the translated expressions by frequency of use to generate a phrasebook, which is also adjusted based on the user's emotions—for example, if the user is tired, simpler phrases are prioritized.

[1448] Step 10:

[1449] The server notifies the user of the generated phrasebook via in-app notifications, push notifications, email, etc.

[1450] Step 11:

[1451] Users can access the phrasebooks provided through their devices and efficiently study foreign languages. The phrasebooks can be viewed through an app or web portal.

[1452] Specific examples

[1453] 1. Conversation recording: A user says, "What a beautiful day today!"

[1454] 2. Sending audio: The device records the conversation and sends the audio data to the server.

[1455] 3. Analysis of voice data: The server converts the voice data into text "What a beautiful day today!"

[1456] 4. Emotion recognition: The server uses the emotion engine to recognize the emotion of "joy."

[1457] 5. Extraction of frequently occurring expressions: The server extracts the frequently occurring expression "ii tenki" and translates it as "Good weather."

[1458] 6. Phrasebook generation and adjustment: The server generates a phrasebook adjusted based on emotions to help the user relax.

[1459] 7. Phrasebook notification and learning: The server notifies the user of the generated phrasebook via an in-app notification, and the user learns it.

[1460] This system enables efficient foreign language learning based on the user's emotions.

[1461] Example 2

[1462] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1463] Conventional foreign language learning systems have the problem of reducing learning efficiency because they provide uniform learning content without considering the user's emotional state. Furthermore, there is no established method for effectively extracting and translating expressions frequently used by users in real conversations and generating a learning phrasebook based on them. This often leads to a decrease in users' motivation to learn and results in poor learning outcomes.

[1464] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for analyzing voice data and text data to recognize the user's emotions, means for adjusting the content of the phrase book based on the emotions, and means for notifying the user of the generated phrase book. This enables efficient foreign language learning according to the user's emotional state.

[1465] "User" refers to an individual who uses the system to learn.

[1466] "Means for recording conversations" refers to a device or system that records the user's voice in real time.

[1467] "Means for transmitting audio data to a server" refers to the technology or protocol for transferring recorded audio data to a server over a network.

[1468] "Natural language processing" refers to computer technology that analyzes voice and text data to understand and process their content.

[1469] "Means for extracting frequent expressions" refers to a technique for selecting specific words or phrases from the analyzed data based on their frequency.

[1470] "Means for translating into other languages" refers to a technology or system that translates the extracted expressions into a different language.

[1471] "Means for generating phrasebooks" refers to techniques for organizing translated expressions into a single collection for study.

[1472] "Means for recognizing user emotions" refers to technologies and systems that analyze voice data and text data to detect the user's emotional state.

[1473] "Means for adjusting phrase book content based on emotion" refers to a technique for changing or adjusting the content or difficulty of a phrase book in accordance with a detected emotional state.

[1474] "Means of notification" refers to the technology or system used to notify users of the generated phrasebook.

[1475] The present invention relates to a system that records user conversations in real time and analyzes and translates the data for foreign language learning. The system is designed to recognize the user's emotions and further enhance the learning experience.

[1476] Conversation recording and data storage

[1477] The device records the user's conversation in real time using a microphone built into a smartphone, smartwatch, smart glasses, etc. The recorded voice data is sent to a server as a file in an appropriate format such as WAV at regular intervals.

[1478] Examples:

[1479] A user says, "What a beautiful day today!"

[1480] The device records this conversation, saves the audio data with the file name "2023-10-01-1230.wav", and sends it to the server.

[1481] Analysis of audio data

[1482] The server converts the received voice data into text data using voice recognition technology. Specifically, it converts the voice data into text using a common voice recognition API (e.g., Google Cloud Speech-to-Text API). The converted text data is then analyzed using natural language processing (NLP) technology. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1483] Examples:

[1484] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API and receives the text data "What a beautiful day today!" as a result.

[1485] The server performs morphological analysis on the text data "The weather is so nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[1486] Emotion recognition

[1487] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. For example, it uses a common emotion analysis tool (e.g., IBM Watson Tone Analyzer) to detect emotions from the tone, pitch, and choice of words of the voice.

[1488] Examples:

[1489] The server detects the emotion of "joy" from the text data and tone of voice.

[1490] Extraction and translation of frequently occurring expressions

[1491] The server identifies frequently used words and phrases from the analysis results, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using common translation tools (e.g., Google Cloud Translation API).

[1492] Examples:

[1493] The server extracts the frequently occurring expression "nice weather" based on past analysis results.

[1494] The server translates "ii tenki" into "Good weather."

[1495] Phrasebook generation and provision

[1496] The translated frequently occurring expressions are organized in order of frequency and generated as a phrasebook. The server adjusts the contents of the phrasebook based on the user's recognized emotions. For example, if the server recognizes that the user is relaxed, more difficult phrases will be added. The generated phrasebook is provided to the user through a notification system. For example, this could be an in-app notification or an email notification.

[1497] Examples:

[1498] The server generates a collection of phrases containing "Good weather" based on the sentiment analysis results.

[1499] The server sends the generated phrasebook as an in-app notification.

[1500] Users can browse and learn from the phrasebook "Good weather" provided through the app.

[1501] Prompt Sentence Examples

[1502] If a user says "What a beautiful day today!" during a conversation, use the following prompt example:

[1503] "The user says, 'The weather is so nice today!' Based on this conversation, extract frequently occurring expressions, perform emotion recognition, and generate a phrasebook for foreign language learning based on the results."

[1504] This enables efficient foreign language learning according to the user's emotional state.

[1505] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1506] Step 1:

[1507] The device records the user's conversation in real time. Specific devices include smartphones, smart watches, and smart glasses. The input is the user's conversational voice, and the output is the recorded voice data.

[1508] Specific behavior:

[1509] A user says, "What a beautiful day today!"

[1510] The device records this conversation and saves it as audio data.

[1511] Step 2:

[1512] The device sends the recorded audio data to the server as a file at regular intervals. The audio data format is WAV, etc. The input is the recorded audio data, and the output is the file sent to the server.

[1513] Specific behavior:

[1514] The device saves the audio data with the file name "2023-10-01-1230.wav" and sends it to the server.

[1515] Step 3:

[1516] The server converts the received voice data into text data using speech recognition technology. Specifically, it uses the Google Cloud Speech-to-Text API, etc. The input is voice data, and the output is the converted text data.

[1517] Specific behavior:

[1518] The server uploads the "2023-10-01-1230.wav" file to the Google Cloud Speech-to-Text API, and receives the text data "What a beautiful day today!" as a result.

[1519] Step 4:

[1520] The server analyzes the converted text data using natural language processing (NLP) technology. As a result of the analysis, morphological analysis is performed to extract words, phrases, and sentence structure patterns. The input is text data, and the output is analyzed information (extracted words and phrases).

[1521] Specific behavior:

[1522] The server performs morphological analysis on the text data "The weather is very nice today!" and extracts words such as "today," "very," "nice," "weather," and "isn't it?"

[1523] Step 5:

[1524] The server uses the converted text and voice data to run an emotion engine to recognize the user's emotions. Specifically, emotions are detected from the tone, pitch, and word choice of the voice. The input is text and voice data, and the output is the recognized emotion.

[1525] Specific behavior:

[1526] The server detects the emotion of "joy" from the text data and tone of voice.

[1527] Step 6:

[1528] The server identifies frequently used words and phrases from the analysis results. Specifically, it calculates their frequency of occurrence. The input is the analyzed information (words and phrases), and the output is frequently used expressions.

[1529] Specific behavior:

[1530] The server extracts the frequently occurring expression "nice weather" based on the analysis results.

[1531] Step 7:

[1532] The server translates the identified frequent expressions into other languages ​​using a translation API. Specifically, it uses the Google Cloud Translation API. The input is the frequent expression, and the output is the translated expression.

[1533] Specific behavior:

[1534] The server translates "ii tenki" into "Good weather."

[1535] Step 8:

[1536] The server sorts the translated expressions by frequency of use to generate a phrasebook. The content of the phrasebook is adjusted based on the recognized emotions. The input is the translated expressions and emotional information, and the output is an adjusted phrasebook.

[1537] Specific behavior:

[1538] The server generates a list of phrases containing "Good weather" based on the results of sentiment analysis.

[1539] Step 9:

[1540] The server provides the generated phrasebook to the user through a notification system: the input is the generated phrasebook, and the output is the notified phrasebook.

[1541] Specific behavior:

[1542] The server sends the phrasebook as an in-app notification.

[1543] Users can browse and learn from a collection of phrases called "Good weather" provided through the app.

[1544] (Application example 2)

[1545] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1546] Conventional foreign language learning systems have the problem of being difficult to use efficiently because they do not take into account the user's emotional state or the context of real-life conversations. In particular, there is a need for a system that can analyze frequently used expressions in real time and provide appropriate learning content according to the user's emotions.

[1547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1548] In this invention, the server includes means for recording user conversations, means for transmitting the recorded voice data to the server, means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions, means for translating the extracted expressions into other languages, means for arranging the translated expressions in order of frequency of use and generating a phrasebook, means for displaying the phrasebook to the user via smart glasses, and means for recognizing the user's emotions from the analyzed data. This makes it possible to provide emotion-sensitive foreign language learning in real time based on the user's real-life conversation data.

[1549] "Means for recording user conversations" refers to devices or technologies for collecting voice data and storing user utterances.

[1550] The "means for transmitting recorded voice data to a server" refers to a device or method for transferring collected voice data to a server via a network.

[1551] "Means of analyzing voice data on a server using natural language processing and extracting frequently occurring expressions" refers to a technology that converts voice data into text data and identifies frequently used words and phrases.

[1552] The "means for translating extracted expressions into other languages" is a method for converting identified words and phrases into the foreign language that the user wishes to learn.

[1553] "Means for sorting translated expressions in order of frequency of use and generating a phrasebook" refers to a technology that sorts translated words and phrases in order of frequency and provides them as a single, comprehensive learning content.

[1554] The "means for displaying a phrasebook to a user via smart glasses" refers to a technique for displaying the generated phrasebook on the display of smart glasses worn by the user.

[1555] "Means for recognizing user emotions from analyzed data" refers to engines or technologies for identifying a user's emotional state based on voice data and text data.

[1556] The present invention provides a system for recording and analyzing a user's conversations and supporting foreign language learning based on the user's emotional state. The system includes: a means for recording the user's conversations; a means for transmitting the recorded voice data to a server; a means for analyzing the voice data on the server using natural language processing and extracting frequently occurring expressions; a means for translating the extracted expressions into other languages; a means for sorting the translated expressions in order of frequency of use and generating a phrasebook; a means for displaying the phrasebook to the user via smart glasses; and a means for recognizing the user's emotions from the analyzed data.

[1557] First, when a user wears smart glasses and engages in everyday conversation, the microphone built into the smart glasses collects voice data. This voice data is recorded in real time and sent to a server via a network. The voice data is then saved in an appropriate format (e.g., WAV format).

[1558] Next, on the server, the received voice data is converted into text data using speech recognition technology. Specifically, data conversion is performed using a speech recognition API such as the Google Speech-to-Text API. This converted text data is then analyzed using a natural language processing (NLP) engine, such as spaCy. This analysis involves morphological analysis to extract words, phrases, short sentences, and sentence structure patterns.

[1559] Next, based on the analyzed text and voice data, an emotion recognition engine, such as IBM Watson Tone Analyzer, is used to recognize the user's emotions. This engine detects emotions from the tone, pitch, and word choice of the voice. For example, if the user is speaking with an emotion of joy, it will be recognized as "joy."

[1560] The server then uses the analysis results to identify frequently used words and phrases, including calculating their frequency of occurrence, and translates the identified frequently used expressions into the foreign language the user wishes to learn using a translation API, such as the Google Cloud Translation API.

[1561] The translated frequently occurring expressions are organized in order of frequency and generated into a phrasebook. This phrasebook is sent in real time from the server to the smart glasses display and displayed to the user. This allows users to learn a foreign language naturally through everyday conversation.

[1562] As a specific example of use, if a user says, "What a beautiful day today!", this voice is recorded by the smart glasses and sent to the server. The speech recognition system on the server converts this voice data into text, obtaining the text data "What a beautiful day today!". Furthermore, the emotion recognition engine detects the emotion of "joy," extracts the frequently occurring expression "nice weather," and translates it into "Good weather." The translated phrases are organized in order of frequency and displayed on the smart glasses.

[1563] Through this process, efficient foreign language learning that responds to emotions can be realized based on the user's daily conversation data.

[1564] An example prompt might look like this:

[1565] "Your smart glasses are currently recording everyday conversations. Please explain the process of extracting common expressions from these conversations, translating them into English, and creating a phrasebook based on emotions."

[1566] As described above, the present invention is a system that provides emotion-responsive foreign language learning in real time based on conversation data from the user's real life.

[1567] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1568] Step 1:

[1569] The device records the user's conversation in real time. The device's built-in microphone collects audio data and saves it in a file format (e.g., WAV format). This outputs an audio file, which is raw data.

[1570] Step 2:

[1571] The device sends the recorded audio data to the server at regular intervals. The audio file is uploaded to the server via the network. The audio file received as input is transferred to the server, where it is saved.

[1572] Step 3:

[1573] The server converts the received voice data into text data using a speech recognition API (for example, Google Speech-to-Text API). It receives the voice file as input and performs data conversion processing to generate text data as output.

[1574] Step 4:

[1575] The server analyzes the generated text data using a natural language processing (NLP) engine (such as spaCy). During the analysis, morphological analysis is performed to extract words, phrases, short sentences, and sentence structure patterns. By receiving the text data as input and performing data analysis, the analysis results are obtained as output.

[1576] Step 5:

[1577] The server uses an emotion recognition engine (such as IBM Watson Tone Analyzer) to recognize the user's emotions based on the analyzed text data and voice data. It receives text data and voice parameters as input, performs emotion recognition processing, and generates emotion data as output.

[1578] Step 6:

[1579] The server identifies frequently used words and phrases from the analysis results, which includes calculating their frequency of occurrence. It receives the analysis results as input, extracts frequently occurring expressions, and outputs a list of frequently occurring expressions.

[1580] Step 7:

[1581] The server translates the extracted frequent expressions into other languages ​​using a translation API (e.g., Google Cloud Translation API). It receives a list of frequent expressions as input and performs translation processing, generating a list of translated expressions as output.

[1582] Step 8:

[1583] The server sorts the translated expressions in order of frequency of use and generates a phrasebook. It receives the list of translated expressions as input, performs sorting, and obtains a phrasebook as output.

[1584] Step 9:

[1585] The server sends the generated phrase book to the smart glasses display in real time and displays it to the user. The server receives the phrase book as input and performs communication processing, and the phrase book is displayed on the smart glasses as output.

[1586] This series of processes enables users to efficiently learn a foreign language through real-life conversations.

[1587] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1588] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1589] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1590] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1591] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1592] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1593] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1594] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1595] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1596] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1597] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1598] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1599] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1600] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1601] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1602] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1603] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1604] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1605] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1606] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1607] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1608] The following is further disclosed regarding the above embodiment.

[1609] (Claim 1)

[1610] means for recording user conversations;

[1611] means for transmitting the recorded voice data to a server;

[1612] A means for analyzing speech data on a server using natural language processing and extracting frequently occurring expressions;

[1613] a means for translating the extracted expressions into other languages;

[1614] A means of sorting translated expressions by frequency of use and generating phrasebooks

[1615] A system including:

[1616] (Claim 2)

[1617] means for notifying a user of the generated phrasebook;

[1618] 10. The system of claim 1.

[1619] (Claim 3)

[1620] The system further includes a means for recording and managing the frequency of use of the extracted expressions.

[1621] 10. The system of claim 1.

[1622] "Example 1"

[1623] (Claim 1)

[1624] means for recording user conversations;

[1625] means for transmitting the recorded voice data to a server;

[1626] A means for converting voice data into text data using natural language processing on a server;

[1627] A means for analyzing the converted text data and extracting words, phrases, short sentences, and sentence structure patterns;

[1628] A means for calculating the frequency of appearance of the extracted expressions and identifying frequently occurring expressions;

[1629] a means for translating the identified frequent expressions into other languages;

[1630] A means of sorting translated expressions by frequency of use and generating phrasebooks

[1631] A system including:

[1632] (Claim 2)

[1633] means for notifying a user of the generated phrasebook;

[1634] 10. The system of claim 1.

[1635] (Claim 3)

[1636] The system further includes a means for recording and managing the frequency of use of the extracted expressions.

[1637] 10. The system of claim 1.

[1638] "Application Example 1"

[1639] (Claim 1)

[1640] means for recording user conversations;

[1641] means for transmitting the recorded voice data to a server;

[1642] A means for analyzing speech data on a server using natural language processing and extracting frequently occurring expressions;

[1643] a means for translating the extracted expressions into other languages;

[1644] means for arranging the translated expressions in order of frequency of use to generate a phrasebook;

[1645] a means for recording a user's conversation in real time and storing the audio data in an appropriate format;

[1646] A means for analyzing text data obtained from the recorded voice data and extracting words and short sentences;

[1647] A means to translate the extracted frequently used expressions into foreign languages ​​by linking with a translation API via a server,

[1648] A way to provide generated phrasebooks to users via in-app or email notifications

[1649] A system including:

[1650] (Claim 2)

[1651] 10. The system of claim 1, further comprising: means for notifying a user of the generated phrasebook.

[1652] (Claim 3)

[1653] The system of claim 1, further comprising means for recording and managing the frequency of use of the extracted expressions.

[1654] "Example 2: Combining Emotion Engines"

[1655] (Claim 1)

[1656] means for recording user conversations;

[1657] means for transmitting the recorded voice data to a server;

[1658] A means for analyzing speech data on a server using natural language processing and extracting frequently occurring expressions;

[1659] a means for translating the extracted expressions into other languages;

[1660] means for arranging the translated expressions in order of frequency of use to generate a phrasebook;

[1661] means for analyzing the voice data and text data to recognize the user's emotions;

[1662] A way to tailor phrasebook content based on emotions

[1663] A system including:

[1664] (Claim 2)

[1665] 10. The system of claim 1, further comprising: means for notifying a user of the generated phrasebook.

[1666] (Claim 3)

[1667] The system of claim 1, further comprising means for recording and managing the frequency of use of the extracted expressions.

[1668] "Application example 2 when combining emotion engines"

[1669] (Claim 1)

[1670] means for recording user conversations;

[1671] means for transmitting the recorded voice data to a server;

[1672] A means for analyzing speech data on a server using natural language processing and extracting frequently occurring expressions;

[1673] a means for translating the extracted expressions into other languages;

[1674] means for arranging the translated expressions in order of frequency of use to generate a phrasebook;

[1675] means for displaying a phrasebook to a user via the smart glasses;

[1676] A means of recognizing user emotions from analyzed data

[1677] A system including:

[1678] (Claim 2)

[1679] means for notifying a user of the generated phrasebook;

[1680] 10. The system of claim 1.

[1681] (Claim 3)

[1682] The system further includes a means for recording and managing the frequency of use of the extracted expressions.

[1683] 10. The system of claim 1. [Explanation of symbols]

[1684] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for recording user conversations; means for transmitting the recorded voice data to a server; A means for analyzing speech data on a server using natural language processing and extracting frequently occurring expressions; a means for translating the extracted expressions into other languages; A means of sorting translated expressions by frequency of use and generating phrasebooks A system including:

2. means for notifying a user of the generated phrasebook; The system of claim 1 .

3. Further comprising means for recording and managing the frequency of use of the extracted expressions; The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A