system

The system addresses the inefficiency of conventional language learning by converting conversational content into text and audio, generating personalized materials, and utilizing emotion estimation to enhance learning effectiveness.

JP2026045553APending Publication Date: 2026-03-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Conventional techniques have not been effective enough in utilizing conversational content for language learning.

Method used

A system comprising a conversion unit, storage unit, and generation unit that converts conversational content into text and audio, analyzes the data, and generates personalized language learning materials tailored to the user's context, including emotion estimation and recognition of technical terms and slang.

Benefits of technology

Efficiently utilizes conversational content for language learning, providing personalized and context-aware language acquisition textbooks, enhancing learning effectiveness through visual and auditory aids, and real-time progress monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026045553000001_ABST
    Figure 2026045553000001_ABST
Patent Text Reader

Abstract

The system according to the embodiment aims to efficiently utilize conversation content for language learning. [Solution] A system according to an embodiment includes a conversion unit, a storage unit, and a generation unit. The conversion unit converts conversation content into text and audio. The storage unit stores the data converted by the conversion unit. The generation unit analyzes the data stored by the storage unit and generates text.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques have not been effective enough in utilizing conversational content for language learning, and there is room for improvement.

[0005] The system according to the embodiment aims to efficiently utilize conversation content for language learning. [Means for solving the problem]

[0006] The system according to the embodiment includes a conversion unit, a storage unit, and a generation unit. The conversion unit converts conversation content into text and audio. The storage unit stores the data converted by the conversion unit. The generation unit analyzes the data stored by the storage unit and generates text. [Effects of the Invention]

[0007] The system according to the embodiment can efficiently utilize conversation content for language learning. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION

[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0010] First, the terms used in the following description will be explained.

[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).

[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.

[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.

[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example 1) A language acquisition system according to an embodiment of the present invention uses a generative AI to convert everyday conversational content into text and audio in a selected language and save the data. The system can then use the saved data as a personalized language acquisition textbook. This allows users to convert the language of their own living environment, creating the ultimate in categorized learning materials. This allows users to acquire languages ​​regardless of the context, such as business or personal use. It also allows users to expand the language they use in their areas of expertise and interest. For example, English is not the only target language; it will be offered as the first language. Users can check other example sentences using the words used in the conversation, including expressions used in business and public settings. Furthermore, the system features a function that allows the generative AI to recreate scenes from conversations and turn them into videos. Selected words can be classified for review, and a function that allows users to view TOEIC and TOEFL reference scores based on their level of proficiency is also provided. A function is also provided that recommends and converts to text words, phrases, and example sentences to memorize. This allows the language acquisition system to provide personalized language acquisition textbooks.

[0029] A language learning system according to an embodiment includes a conversion unit, a storage unit, and a generation unit. The conversion unit converts conversational content into text and audio. Examples of conversational content include, but are not limited to, everyday conversation, business conversation, and conversations containing technical terms. The conversion unit converts the conversational content into text using, for example, speech recognition technology, and then converts it into audio using a text generation algorithm. The conversion unit can also analyze the conversational content and convert it into appropriate text and audio using a generation AI. For example, the conversion unit converts the conversational content into text using speech recognition technology, and the generation AI generates audio based on the text. The storage unit stores the data converted by the conversion unit. Examples of stored data include, but are not limited to, storing the data in a database or in a file format. The storage unit stores the converted data in, for example, a database, and searches and retrieves the data as needed. The storage unit can also store the data in a file format and back it up in external storage. The generation unit analyzes the data stored by the storage unit and generates text. The generated text includes, for example, sentence structure and a generation algorithm, but is not limited to, examples. The generator may analyze the stored data using natural language processing technology to generate appropriate text. The generator may also use a machine learning algorithm to learn from the stored data and generate more accurate text. This allows the language learning system according to the embodiment to provide a language learning text tailored to the user.

[0030] The language acquisition system includes an example sentence checking unit that checks other example sentences that use words that appear in the conversation. The example sentence checking unit checks other example sentences that use words that appear in the conversation. For example, the example sentence checking unit displays example sentences that include synonyms. The example sentence checking unit can also display usage examples in different contexts. For example, the example sentence checking unit displays usage examples in business situations and usage examples in everyday conversation. This allows the learning effect to be improved by checking other example sentences that use words that appear in the conversation.

[0031] The language acquisition system includes an expression confirmation unit that confirms expressions used in business or public use situations. The expression confirmation unit confirms expressions used in business or public use situations. For example, the expression confirmation unit displays expressions used in meetings or presentations. The expression confirmation unit can also display expressions used in speeches or presentations in public places. For example, the expression confirmation unit displays formal expressions used in business situations and casual expressions used in public situations. This allows the user to learn appropriate expressions by checking expressions used in business or public use situations.

[0032] The language acquisition system includes an animation unit that recreates scenes from conversations and creates animations. The animation unit recreates scenes from conversations and creates animations. For example, the animation unit recreates conversation scenes using animation. The animation unit can also recreate conversation scenes using live-action footage. For example, the animation unit recreates conversation scenes using simulations to visually enhance learning effects. In this way, by recreating and animating scenes from conversations, it is possible to visually enhance learning effects.

[0033] The language acquisition system includes a classification unit that classifies selected words for review. The classification unit classifies the selected words for review. For example, the classification unit prioritizes classification of frequently occurring words. The classification unit can also prioritize classification of highly important words. For example, the classification unit classifies words according to learning objectives, enabling efficient review. In this way, by classifying the selected words for review, efficient review can be achieved.

[0034] The language acquisition system includes a reference score display unit that displays a TOEIC or TOEFL reference score based on the level of proficiency. The reference score display unit displays a TOEIC or TOEFL reference score based on the level of proficiency. For example, the reference score display unit displays a reference score based on test results. The reference score display unit can also display a reference score based on learning progress. For example, the reference score display unit displays a reference score using a TOEIC or TOEFL score conversion method. This allows learning progress to be confirmed by displaying a TOEIC or TOEFL reference score based on the level of proficiency.

[0035] The language acquisition system includes a recommendation unit that recommends words, phrases, and example sentences to be memorized and converts them into text. The recommendation unit recommends words, phrases, and example sentences to be memorized and converts them into text. For example, the recommendation unit recommends frequently occurring words and phrases. The recommendation unit can also recommend important words and phrases. For example, the recommendation unit recommends words and phrases according to the learning objective and converts them into text. In this way, learning progresses efficiently by recommending words, phrases, and example sentences to be memorized and converting them into text.

[0036] The conversion unit can automatically recognize specific technical terms and slang and perform appropriate conversion when converting the content of a conversation. For example, when converting a conversation that includes business terms, the generation AI converts it into appropriate business terms. In addition, when converting a conversation that includes slang, the generation AI can also convert it into appropriate slang. For example, when converting a conversation that includes medical terms, the generation AI converts it into appropriate medical terms. This allows for accurate conversion by appropriately converting technical terms and slang. Recognition of technical terms and slang is performed using, for example, natural language processing technology or machine learning algorithms.

[0037] The conversion unit can optimize the conversion results based on the user's pronunciation and intonation. For example, if the user's pronunciation is unclear, the generation AI will take the context into account to make an appropriate conversion. In addition, if the user's intonation is different, the conversion unit can convert the generation AI to an appropriate intonation. For example, if the user has a strong accent, the conversion unit will convert to an appropriate accent. This allows for more natural conversion by taking the user's pronunciation and intonation into account. Pronunciation and intonation are evaluated using technologies such as speech waveform analysis and phoneme recognition.

[0038] When converting the content of a conversation, the conversion unit can perform appropriate conversion by taking into account the geographical and cultural background of the user. For example, if the user uses a dialect from a specific region, the generation AI converts it into an appropriate dialect. In addition, if the user has a specific cultural background, the conversion unit can also convert it into an appropriate cultural expression. For example, if the user uses a language from a different country, the conversion unit can convert it into an appropriate language. This allows for more appropriate conversion by taking into account the geographical and cultural background. Consideration of the geographical and cultural background is based on, for example, local customs and cultural practices.

[0039] When converting conversation content, the conversion unit can improve conversion accuracy by referring to the user's past conversation history. In the conversion unit, for example, the generation AI performs appropriate conversion based on words and phrases used by the user in the past. The conversion unit can also perform conversion by having the generation AI take into account appropriate context from the user's past conversation history. For example, the conversion unit analyzes the user's past conversation history, and the generation AI performs optimal conversion. In this way, by referring to the past conversation history, conversion accuracy is improved. Referring to the past conversation history is done based on, for example, conversations during a specific period or on a specific topic.

[0040] The storage unit can select a storage format based on the importance of the data when storing the data. For example, the storage unit stores important data in high resolution to preserve detailed information. The storage unit can also store normal data in standard resolution to achieve a balance. For example, the storage unit stores temporary data in low resolution to save capacity. This allows for efficient data management by selecting a storage format based on the importance of the data. The storage format can be selected based on criteria such as text format, audio format, or image format.

[0041] The storage unit can apply different storage algorithms depending on the data category when storing data. For example, the storage unit applies a storage algorithm with high security to business data. The storage unit can also apply a standard storage algorithm to private data. For example, the storage unit applies a simple storage algorithm to temporary data. This allows appropriate data storage by applying a storage algorithm depending on the data category. The storage algorithm is applied using techniques such as a compression algorithm or an encryption algorithm.

[0042] The storage unit can determine the storage priority based on the time of data submission when storing data. For example, the storage unit stores urgent data with priority to enable quick access. The storage unit can also store normal data with standard priority. For example, the storage unit stores temporary data with low priority to save capacity. In this way, determining the storage priority based on the time of data submission enables quick data access. The submission time is considered based on criteria such as the submission date and submission time.

[0043] The storage unit can adjust the order of storage based on the relevance of data when storing the data. For example, the storage unit can store highly relevant data preferentially to enable quick access. The storage unit can also store regular data in a standard order. For example, the storage unit can store less relevant data later to save space. In this way, adjusting the order of storage based on the relevance of data enables efficient data management. The relevance of data is evaluated based on criteria such as common tags or related topics.

[0044] When generating text, the generator can adjust the level of detail of the generated text based on the importance of the data. For example, the generator generates detailed text for important data to provide comprehensive information. The generator can also generate text with a standard level of detail for normal data. For example, the generator generates simple text for temporary data to save space. This allows for efficient text generation by adjusting the level of detail of the generated text based on the importance of the data. The adjustment of the level of detail of the generated text is performed based on criteria such as a detailed description or a concise summary.

[0045] The generation unit can apply different generation algorithms depending on the data category when generating text. For example, the generation unit applies a highly accurate generation algorithm to business data. The generation unit can also apply a standard generation algorithm to private data. For example, the generation unit applies a simple generation algorithm to temporary data. This makes it possible to generate appropriate text by applying a generation algorithm depending on the data category. The application of a generation algorithm is performed using techniques such as rule-based generation or machine learning-based generation.

[0046] When generating text, the generation unit can determine the generation priority based on the time of data submission. For example, the generation unit generates urgent data with priority to enable quick access. The generation unit can also generate normal data with standard priority. For example, the generation unit generates temporary data with low priority to save capacity. In this way, determining the generation priority based on the time of data submission enables quick text generation. The submission time is considered based on criteria such as the submission date and submission time.

[0047] The generator can adjust the order of generation based on the relevance of data when generating text. For example, the generator can generate highly relevant data preferentially to enable quick access. The generator can also generate normal data in a standard order. For example, the generator can generate less relevant data later to save space. In this way, adjusting the order of generation based on the relevance of data enables efficient text generation. The relevance of data is evaluated based on criteria such as common tags or related topics.

[0048] The example sentence checking unit can adjust the level of detail of the example sentence based on the importance of the word when checking the example sentence. For example, the example sentence checking unit provides detailed information for example sentences that include important words. The example sentence checking unit can also provide example sentences that include ordinary words with a standard level of detail. For example, the example sentence checking unit provides simple information for example sentences that include temporary words. This allows for efficient learning by adjusting the level of detail of the example sentence based on the importance of the word. The adjustment of the level of detail of the example sentence is performed based on criteria such as a detailed explanation or a concise summary.

[0049] When checking example sentences, the example sentence checking unit can determine the priority of example sentences based on the relevance of words. For example, the example sentence checking unit preferentially displays example sentences that include highly relevant words. The example sentence checking unit can also display example sentences that include words that have normal relevance with standard priority. For example, the example sentence checking unit displays example sentences that include less relevant words later. In this way, by determining the priority of example sentences based on the relevance of words, efficient learning becomes possible. The priority of example sentences is determined based on criteria such as the relevance and importance of words, for example.

[0050] The expression checking unit can adjust the level of detail of the expression based on the importance of the scene when checking the expression. For example, the expression checking unit provides detailed information for an expression including an important scene. The expression checking unit can also provide a standard level of detail for an expression including an ordinary scene. For example, the expression checking unit provides simple information for an expression including a temporary scene. This enables efficient learning by adjusting the level of detail of the expression based on the importance of the scene. The adjustment of the level of detail of the expression is performed based on criteria such as a detailed description or a concise summary.

[0051] The expression confirmation unit can determine the priority of expressions based on the relevance of scenes when confirming expressions. For example, the expression confirmation unit preferentially displays expressions including highly relevant scenes. The expression confirmation unit can also display expressions including scenes with normal relevance with standard priority. For example, the expression confirmation unit displays expressions including scenes with low relevance later. In this way, efficient learning is made possible by determining the priority of expressions based on the relevance of scenes. The priority of expressions is determined based on criteria such as the relevance and importance of scenes, for example.

[0052] The animation unit can adjust the level of detail of the video based on the importance of the conversation when creating the video. For example, the animation unit provides detailed information for videos that include important conversations. The animation unit can also provide a standard level of detail for videos that include normal conversations. For example, the animation unit provides simple information for videos that include temporary conversations. This allows for efficient learning by adjusting the level of detail of the video based on the importance of the conversation. The adjustment of the level of detail of the video is performed based on criteria such as a detailed explanation or a concise summary.

[0053] When creating a video, the video generation unit can apply different video reproduction algorithms depending on the category of the conversation. For example, the video generation unit applies a highly accurate reproduction algorithm to a video containing a business conversation. The video generation unit can also apply a standard reproduction algorithm to a video containing a private conversation. For example, the video generation unit applies a simple reproduction algorithm to a video containing a temporary conversation. This makes it possible to reproduce an appropriate video by applying a video reproduction algorithm depending on the category of the conversation. The video reproduction algorithm is applied using techniques such as rule-based generation or machine learning-based generation.

[0054] When creating videos, the video creator can determine the priority of the videos based on the time of submission of the conversation. For example, the video creator can prioritize urgent conversations to enable quick access. The video creator can also create videos of normal conversations with a standard priority. For example, the video creator can create videos of temporary conversations with a low priority to save capacity. In this way, by determining the priority of videos based on the time of submission of the conversation, quick video reproduction is possible. The submission time is considered based on criteria such as the submission date and submission time.

[0055] The animation unit can adjust the order of videos based on the relevance of the conversations when animating them. For example, the animation unit can prioritize animation of highly relevant conversations to enable quick access. The animation unit can also animate normal conversations in a standard order. For example, the animation unit can animate less relevant conversations later to save capacity. In this way, adjusting the order of videos based on the relevance of the conversations enables efficient video reproduction. The relevance of the conversations is evaluated based on criteria such as common tags or related topics, for example.

[0056] The classifier can adjust the level of detail of the classification based on the importance of the word during classification. For example, the classifier classifies important words in detail to provide comprehensive information. The classifier can also classify ordinary words with a standard level of detail. For example, the classifier can simply classify temporary words to save space. This allows for efficient learning by adjusting the level of detail of the classification based on the importance of the word. The adjustment of the level of detail of the classification is performed based on criteria such as a detailed description or a concise summary.

[0057] The classifier can determine the priority of classification based on the relevance of words during classification. For example, the classifier can prioritize highly relevant words to enable quick access. The classifier can also classify words with normal relevance with a standard priority. For example, the classifier can classify less relevant words later to save capacity. In this way, determining the priority of classification based on the relevance of words enables efficient learning. The priority of classification is determined based on criteria such as the relevance and importance of words, for example.

[0058] The reference point display unit can adjust the level of detail of the display based on the importance of the mastery level when displaying the reference points. For example, the reference point display unit provides detailed information for reference points including important mastery levels. The reference point display unit can also provide standard level of detail for reference points including normal mastery levels. For example, the reference point display unit provides simple information for reference points including temporary mastery levels. This allows for efficient learning by adjusting the level of detail of the display based on the importance of the mastery level. The adjustment of the level of detail of the display is performed based on criteria such as a detailed explanation or a concise summary.

[0059] The reference point display unit can determine the display priority based on the relevance of the mastery level when displaying the reference points. For example, the reference point display unit preferentially displays reference points including highly relevant mastery levels. The reference point display unit can also display reference points including normally relevant mastery levels with standard priority. For example, the reference point display unit displays reference points including less relevant mastery levels later. In this way, efficient learning is made possible by determining the display priority based on the relevance of the mastery level. The display priority is determined based on criteria such as the relevance and importance of the mastery level, for example.

[0060] When making a recommendation, the recommendation unit can adjust the level of detail of the recommendation based on the importance of the word or phrase. For example, the recommendation unit can recommend important words and phrases in detail to provide comprehensive information. The recommendation unit can also recommend ordinary words and phrases with a standard level of detail. For example, the recommendation unit can simply recommend temporary words and phrases to save space. This allows for efficient learning by adjusting the level of detail of the recommendation based on the importance of the word or phrase. The adjustment of the level of detail of the recommendation is performed based on criteria such as a detailed explanation or a concise summary, for example.

[0061] When making recommendations, the recommendation unit can apply different recommendation algorithms depending on the category of the word or phrase. For example, the recommendation unit applies a highly accurate recommendation algorithm to business terms. The recommendation unit can also apply a standard recommendation algorithm to private terms. For example, the recommendation unit applies a simple recommendation algorithm to temporary terms. This makes it possible to make appropriate recommendations by applying a recommendation algorithm depending on the category of the word or phrase. The recommendation algorithm is applied using technologies such as collaborative filtering and content-based filtering.

[0062] When making a recommendation, the recommendation unit can determine the priority of the recommendation based on the time of submission of the word or phrase. For example, the recommendation unit can prioritize urgent words and phrases to enable quick access. The recommendation unit can also recommend normal words and phrases with a standard priority. For example, the recommendation unit can recommend temporary words and phrases with a low priority to save capacity. This enables efficient learning by determining the priority of recommendations based on the time of submission of the word or phrase. The recommendation priority is determined based on criteria such as the time of submission and importance, for example.

[0063] When making recommendations, the recommendation unit can adjust the order of recommendations based on the relevance of words and phrases. For example, the recommendation unit can prioritize highly relevant words and phrases to enable quick access. The recommendation unit can also recommend words and phrases with normal relevance in a standard order. For example, the recommendation unit can recommend less relevant words and phrases later to save space. This allows for efficient learning by adjusting the order of recommendations based on the relevance of words and phrases. The adjustment of the order of recommendations is performed based on criteria such as relevance and importance, for example.

[0064] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0065] The language acquisition system may further include a learning style analysis unit that analyzes the user's learning style and suggests the optimal learning method. For example, if the user prefers visual learning, the learning style analysis unit may suggest learning materials that make extensive use of visual aids. If the user prefers auditory learning, the learning style analysis unit may suggest audio-based learning materials. Furthermore, if the user prefers hands-on learning, the learning system may suggest interactive simulations. This maximizes learning effectiveness by providing the optimal learning method according to the user's learning style.

[0066] The language acquisition system may further include a progress monitoring unit that monitors the user's learning progress in real time and provides appropriate feedback. For example, if the user is taking a long time to master a particular word or phrase, the progress monitoring unit may provide additional practice questions related to that word or phrase. If the user is making good progress, the progress monitoring unit may also suggest the next level of learning material. Furthermore, if the user loses motivation to learn, the progress monitoring unit may provide encouraging messages or rewards. This allows the system to support the user's learning progress in real time and improve learning effectiveness.

[0067] The language acquisition system may further include a learning plan creation unit that analyzes the user's learning history and creates an individual learning plan. For example, the learning plan creation unit may suggest what the user should study next based on what the user has previously studied. The learning plan may also be adjusted according to the user's learning pace. Furthermore, short-term and long-term learning plans may be created according to the user's goals. This allows for efficient learning by providing an individual learning plan based on the user's learning history.

[0068] The language acquisition system may further include an environment analysis unit that analyzes the user's learning environment and suggests the optimal learning environment. For example, if the user prefers to study in a quiet environment, the environment analysis unit may suggest a noise-canceling function. If the user prefers to study while moving, the environment analysis unit may suggest a portable learning device. If the user prefers to study in a group, the environment analysis unit may suggest an online learning group. This may improve learning effectiveness by providing the optimal learning method according to the user's learning environment.

[0069] The language learning system may further include a privacy protection unit that anonymizes the user's learning data and protects the privacy of the data. For example, the privacy protection unit may anonymize the user's personal information to reduce the risk of providing it to a third party. The privacy protection unit may also encrypt the user's learning data and store it securely. Furthermore, the system may manage the user's data access permissions and provide data only when necessary. This allows the learning data to be used effectively while protecting the user's privacy.

[0070] The language acquisition system may also include a trend prediction unit that analyzes the user's learning data and predicts learning trends. For example, the trend prediction unit predicts future learning trends based on what the user has learned in the past. It can also analyze the user's learning pace and interests and suggest what content to study next. Furthermore, it predicts short-term and long-term learning trends according to the user's goals. This allows for efficient learning by making trend predictions based on the user's learning data.

[0071] The processing flow of the first embodiment will be briefly explained below.

[0072] Step 1: The conversion unit converts the conversation content into text and audio. The conversation content can include everyday conversation, business conversation, and conversations containing technical terms. The conversion unit converts the conversation content into text using speech recognition technology, and then converts it into audio using a text generation algorithm. It can also use generation AI to analyze the conversation content and convert it into appropriate text and audio. For example, the conversation content can be converted into text using speech recognition technology, and then generation AI generates audio based on the text. Step 2: The storage unit stores the data converted by the conversion unit. The stored data can be stored in a database or in a file format. The storage unit stores the converted data in a database and searches and retrieves the data as needed. It can also store the data in a file format and back it up to external storage. Step 3: The generator analyzes the data stored by the storage unit and generates text. The generated text includes sentence structure, generation algorithms, etc. The generator uses natural language processing technology to analyze the stored data and generate appropriate text. It can also use machine learning algorithms to learn from the stored data and generate more accurate text.

[0073] (Example 2) A language acquisition system according to an embodiment of the present invention uses a generative AI to convert everyday conversational content into text and audio in a selected language and save the data. The system can then use the saved data as a personalized language acquisition textbook. This allows users to convert the language of their own living environment, creating the ultimate in categorized learning materials. This allows users to acquire languages ​​regardless of the context, such as business or personal use. It also allows users to expand the language they use in their areas of expertise and interest. For example, English is not the only target language; it will be offered as the first language. Users can check other example sentences using the words used in the conversation, including expressions used in business and public settings. Furthermore, the system features a function that allows the generative AI to recreate scenes from conversations and turn them into videos. Selected words can be classified for review, and a function that allows users to view TOEIC and TOEFL reference scores based on their level of proficiency is also provided. A function is also provided that recommends and converts to text words, phrases, and example sentences to memorize. This allows the language acquisition system to provide personalized language acquisition textbooks.

[0074] A language learning system according to an embodiment includes a conversion unit, a storage unit, and a generation unit. The conversion unit converts conversational content into text and audio. Examples of conversational content include, but are not limited to, everyday conversation, business conversation, and conversations containing technical terms. The conversion unit converts the conversational content into text using, for example, speech recognition technology, and then converts it into audio using a text generation algorithm. The conversion unit can also analyze the conversational content and convert it into appropriate text and audio using a generation AI. For example, the conversion unit converts the conversational content into text using speech recognition technology, and the generation AI generates audio based on the text. The storage unit stores the data converted by the conversion unit. Examples of stored data include, but are not limited to, storing the data in a database or in a file format. The storage unit stores the converted data in, for example, a database, and searches and retrieves the data as needed. The storage unit can also store the data in a file format and back it up in external storage. The generation unit analyzes the data stored by the storage unit and generates text. The generated text includes, for example, sentence structure and a generation algorithm, but is not limited to, examples. The generator may analyze the stored data using natural language processing technology to generate appropriate text. The generator may also use a machine learning algorithm to learn from the stored data and generate more accurate text. This allows the language learning system according to the embodiment to provide a language learning text tailored to the user.

[0075] The language acquisition system includes an example sentence checking unit that checks other example sentences that use words that appear in the conversation. The example sentence checking unit checks other example sentences that use words that appear in the conversation. For example, the example sentence checking unit displays example sentences that include synonyms. The example sentence checking unit can also display usage examples in different contexts. For example, the example sentence checking unit displays usage examples in business situations and usage examples in everyday conversation. This allows the learning effect to be improved by checking other example sentences that use words that appear in the conversation.

[0076] The language acquisition system includes an expression confirmation unit that confirms expressions used in business or public use situations. The expression confirmation unit confirms expressions used in business or public use situations. For example, the expression confirmation unit displays expressions used in meetings or presentations. The expression confirmation unit can also display expressions used in speeches or presentations in public places. For example, the expression confirmation unit displays formal expressions used in business situations and casual expressions used in public situations. This allows the user to learn appropriate expressions by checking expressions used in business or public use situations.

[0077] The language acquisition system includes an animation unit that recreates scenes from conversations and creates animations. The animation unit recreates scenes from conversations and creates animations. For example, the animation unit recreates conversation scenes using animation. The animation unit can also recreate conversation scenes using live-action footage. For example, the animation unit recreates conversation scenes using simulations to visually enhance learning effects. In this way, by recreating and animating scenes from conversations, it is possible to visually enhance learning effects.

[0078] The language acquisition system includes a classification unit that classifies selected words for review. The classification unit classifies the selected words for review. For example, the classification unit prioritizes classification of frequently occurring words. The classification unit can also prioritize classification of highly important words. For example, the classification unit classifies words according to learning objectives, enabling efficient review. In this way, by classifying the selected words for review, efficient review can be achieved.

[0079] The language acquisition system includes a reference score display unit that displays a TOEIC or TOEFL reference score based on the level of proficiency. The reference score display unit displays a TOEIC or TOEFL reference score based on the level of proficiency. For example, the reference score display unit displays a reference score based on test results. The reference score display unit can also display a reference score based on learning progress. For example, the reference score display unit displays a reference score using a TOEIC or TOEFL score conversion method. This allows learning progress to be confirmed by displaying a TOEIC or TOEFL reference score based on the level of proficiency.

[0080] The language acquisition system includes a recommendation unit that recommends words, phrases, and example sentences to be memorized and converts them into text. The recommendation unit recommends words, phrases, and example sentences to be memorized and converts them into text. For example, the recommendation unit recommends frequently occurring words and phrases. The recommendation unit can also recommend important words and phrases. For example, the recommendation unit recommends words and phrases according to the learning objective and converts them into text. In this way, learning progresses efficiently by recommending words, phrases, and example sentences to be memorized and converting them into text.

[0081] The conversion unit can estimate the user's emotions and adjust the conversion accuracy of the conversation content based on the estimated user emotions. For example, if the user is feeling stressed, the conversion unit causes the generation AI to perform a concise and easy-to-understand conversion. Furthermore, if the user is relaxed, the conversion unit can also cause the generation AI to perform a detailed conversion. For example, if the user is in a hurry, the conversion unit causes the generation AI to perform a quick conversion. This allows for more appropriate conversion by adjusting the conversion accuracy according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0082] The conversion unit can automatically recognize specific technical terms and slang and perform appropriate conversion when converting the content of a conversation. For example, when converting a conversation that includes business terms, the generation AI converts it into appropriate business terms. In addition, when converting a conversation that includes slang, the generation AI can also convert it into appropriate slang. For example, when converting a conversation that includes medical terms, the generation AI converts it into appropriate medical terms. This allows for accurate conversion by appropriately converting technical terms and slang. Recognition of technical terms and slang is performed using, for example, natural language processing technology or machine learning algorithms.

[0083] The conversion unit can optimize the conversion results based on the user's pronunciation and intonation. For example, if the user's pronunciation is unclear, the generation AI will take the context into account to make an appropriate conversion. In addition, if the user's intonation is different, the conversion unit can convert the generation AI to an appropriate intonation. For example, if the user has a strong accent, the conversion unit will convert to an appropriate accent. This allows for more natural conversion by taking the user's pronunciation and intonation into account. Pronunciation and intonation are evaluated using technologies such as speech waveform analysis and phoneme recognition.

[0084] The conversion unit can estimate the user's emotions and adjust the display method of the conversion results based on the estimated user emotions. For example, if the user is nervous, the generation AI can provide a simple, highly visible display method. Furthermore, if the user is relaxed, the conversion unit can also provide a display method that includes detailed information. For example, if the user is in a hurry, the conversion unit can provide a display method that focuses on the main points. This improves visibility by adjusting the display method according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0085] When converting the content of a conversation, the conversion unit can perform appropriate conversion by taking into account the geographical and cultural background of the user. For example, if the user uses a dialect from a specific region, the generation AI converts it into an appropriate dialect. In addition, if the user has a specific cultural background, the conversion unit can also convert it into an appropriate cultural expression. For example, if the user uses a language from a different country, the conversion unit can convert it into an appropriate language. This allows for more appropriate conversion by taking into account the geographical and cultural background. Consideration of the geographical and cultural background is based on, for example, local customs and cultural practices.

[0086] When converting conversation content, the conversion unit can improve conversion accuracy by referring to the user's past conversation history. In the conversion unit, for example, the generation AI performs appropriate conversion based on words and phrases used by the user in the past. The conversion unit can also perform conversion by having the generation AI take into account appropriate context from the user's past conversation history. For example, the conversion unit analyzes the user's past conversation history, and the generation AI performs optimal conversion. In this way, by referring to the past conversation history, conversion accuracy is improved. Referring to the past conversation history is done based on, for example, conversations during a specific period or on a specific topic.

[0087] The storage unit can estimate the user's emotions and determine the priority of stored data based on the estimated user emotions. For example, if the user is having an important conversation, the storage unit causes the generation AI to prioritize saving that data. In addition, if the user is relaxed, the storage unit can cause the generation AI to save data with normal priority. For example, if the user is in a hurry, the storage unit causes the generation AI to quickly save data. In this way, important data can be prioritized by determining the priority of stored data according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0088] The storage unit can select a storage format based on the importance of the data when storing the data. For example, the storage unit stores important data in high resolution to preserve detailed information. The storage unit can also store normal data in standard resolution to achieve a balance. For example, the storage unit stores temporary data in low resolution to save capacity. This allows for efficient data management by selecting a storage format based on the importance of the data. The storage format can be selected based on criteria such as text format, audio format, or image format.

[0089] The storage unit can apply different storage algorithms depending on the data category when storing data. For example, the storage unit applies a storage algorithm with high security to business data. The storage unit can also apply a standard storage algorithm to private data. For example, the storage unit applies a simple storage algorithm to temporary data. This allows appropriate data storage by applying a storage algorithm depending on the data category. The storage algorithm is applied using techniques such as a compression algorithm or an encryption algorithm.

[0090] The storage unit can estimate the user's emotions and adjust the access permissions for the stored data based on the estimated user emotions. For example, when the user saves important data, the generation AI in the storage unit sets high access permissions. In addition, when the user is relaxed, the generation AI in the storage unit can set normal access permissions. For example, when the user is in a hurry, the generation AI in the storage unit quickly sets access permissions. This improves data security by adjusting access permissions according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0091] The storage unit can determine the storage priority based on the time of data submission when storing data. For example, the storage unit stores urgent data with priority to enable quick access. The storage unit can also store normal data with standard priority. For example, the storage unit stores temporary data with low priority to save capacity. In this way, determining the storage priority based on the time of data submission enables quick data access. The submission time is considered based on criteria such as the submission date and submission time.

[0092] The storage unit can adjust the order of storage based on the relevance of data when storing the data. For example, the storage unit can store highly relevant data preferentially to enable quick access. The storage unit can also store regular data in a standard order. For example, the storage unit can store less relevant data later to save space. In this way, adjusting the order of storage based on the relevance of data enables efficient data management. The relevance of data is evaluated based on criteria such as common tags or related topics.

[0093] The generation unit can estimate the user's emotions and adjust the expression method of the generated text based on the estimated user's emotions. For example, if the user is relaxed, the generation AI generates text containing detailed expressions. Furthermore, if the user is in a hurry, the generation unit can generate text containing concise expressions. For example, if the user is excited, the generation unit generates text containing visually stimulating expressions. In this way, by adjusting the expression method of the text according to the user's emotions, more appropriate text is generated. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0094] When generating text, the generator can adjust the level of detail of the generated text based on the importance of the data. For example, the generator generates detailed text for important data to provide comprehensive information. The generator can also generate text with a standard level of detail for normal data. For example, the generator generates simple text for temporary data to save space. This allows for efficient text generation by adjusting the level of detail of the generated text based on the importance of the data. The adjustment of the level of detail of the generated text is performed based on criteria such as a detailed description or a concise summary.

[0095] The generation unit can apply different generation algorithms depending on the data category when generating text. For example, the generation unit applies a highly accurate generation algorithm to business data. The generation unit can also apply a standard generation algorithm to private data. For example, the generation unit applies a simple generation algorithm to temporary data. This makes it possible to generate appropriate text by applying a generation algorithm depending on the data category. The application of a generation algorithm is performed using techniques such as rule-based generation or machine learning-based generation.

[0096] The generation unit can estimate the user's emotions and adjust the length of the generated text based on the estimated user emotions. For example, if the user is in a hurry, the generation AI generates short, to-the-point text. Alternatively, if the user is relaxed, the generation unit can generate longer text with detailed explanations. For example, if the user is excited, the generation unit generates text with visually stimulating effects. This allows for the generation of more appropriate text by adjusting the length of the text according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0097] When generating text, the generation unit can determine the generation priority based on the time of data submission. For example, the generation unit generates urgent data with priority to enable quick access. The generation unit can also generate normal data with standard priority. For example, the generation unit generates temporary data with low priority to save capacity. In this way, determining the generation priority based on the time of data submission enables quick text generation. The submission time is considered based on criteria such as the submission date and submission time.

[0098] The generator can adjust the order of generation based on the relevance of data when generating text. For example, the generator can generate highly relevant data preferentially to enable quick access. The generator can also generate normal data in a standard order. For example, the generator can generate less relevant data later to save space. In this way, adjusting the order of generation based on the relevance of data enables efficient text generation. The relevance of data is evaluated based on criteria such as common tags or related topics.

[0099] The example sentence confirmation unit can estimate the user's emotions and adjust the display method of example sentences based on the estimated user's emotions. For example, if the user is nervous, the example sentence confirmation unit provides a simple, highly visible display method. Furthermore, if the user is relaxed, the example sentence confirmation unit can also provide a display method that includes detailed information. For example, if the user is in a hurry, the example sentence confirmation unit provides a display method that focuses on the main points. This improves visibility by adjusting the display method of example sentences according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0100] The example sentence checking unit can adjust the level of detail of the example sentence based on the importance of the word when checking the example sentence. For example, the example sentence checking unit provides detailed information for example sentences that include important words. The example sentence checking unit can also provide example sentences that include ordinary words with a standard level of detail. For example, the example sentence checking unit provides simple information for example sentences that include temporary words. This allows for efficient learning by adjusting the level of detail of the example sentence based on the importance of the word. The adjustment of the level of detail of the example sentence is performed based on criteria such as a detailed explanation or a concise summary.

[0101] The example sentence confirmation unit can estimate the user's emotions and adjust the display order of example sentences based on the estimated user emotions. For example, if the user is nervous, the example sentence confirmation unit causes the generation AI to prioritize displaying important example sentences. In addition, if the user is relaxed, the example sentence confirmation unit can also cause the generation AI to display detailed example sentences in an orderly manner. For example, if the user is in a hurry, the example sentence confirmation unit causes the generation AI to prioritize displaying example sentences that highlight the main points. This improves visibility by adjusting the display order of example sentences according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0102] When checking example sentences, the example sentence checking unit can determine the priority of example sentences based on the relevance of words. For example, the example sentence checking unit preferentially displays example sentences that include highly relevant words. The example sentence checking unit can also display example sentences that include words that have normal relevance with standard priority. For example, the example sentence checking unit displays example sentences that include less relevant words later. In this way, by determining the priority of example sentences based on the relevance of words, efficient learning becomes possible. The priority of example sentences is determined based on criteria such as the relevance and importance of words, for example.

[0103] The expression confirmation unit can estimate the user's emotion and adjust the display method of the expression based on the estimated user's emotion. For example, if the user is nervous, the expression confirmation unit provides a simple, highly visible display method. Furthermore, if the user is relaxed, the expression confirmation unit can also provide a display method including detailed information. For example, if the user is in a hurry, the expression confirmation unit provides a display method that focuses on the main points. In this way, visibility is improved by adjusting the display method of the expression according to the user's emotion. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis.

[0104] The expression checking unit can adjust the level of detail of the expression based on the importance of the scene when checking the expression. For example, the expression checking unit provides detailed information for an expression including an important scene. The expression checking unit can also provide a standard level of detail for an expression including an ordinary scene. For example, the expression checking unit provides simple information for an expression including a temporary scene. This enables efficient learning by adjusting the level of detail of the expression based on the importance of the scene. The adjustment of the level of detail of the expression is performed based on criteria such as a detailed description or a concise summary.

[0105] The expression confirmation unit can estimate the user's emotions and adjust the display order of expressions based on the estimated user emotions. For example, if the user is nervous, the expression confirmation unit causes the generation AI to prioritize displaying important expressions. In addition, if the user is relaxed, the expression confirmation unit can also cause the generation AI to display detailed expressions in an orderly manner. For example, if the user is in a hurry, the expression confirmation unit causes the generation AI to prioritize displaying expressions that highlight the main points. In this way, visibility is improved by adjusting the display order of expressions according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0106] The expression confirmation unit can determine the priority of expressions based on the relevance of scenes when confirming expressions. For example, the expression confirmation unit preferentially displays expressions including highly relevant scenes. The expression confirmation unit can also display expressions including scenes with normal relevance with standard priority. For example, the expression confirmation unit displays expressions including scenes with low relevance later. In this way, efficient learning is made possible by determining the priority of expressions based on the relevance of scenes. The priority of expressions is determined based on criteria such as the relevance and importance of scenes, for example.

[0107] The animation unit can estimate the user's emotions and adjust the video reproduction method based on the estimated user emotions. For example, if the user is relaxed, the animation unit generates a video in which the generation AI progresses at a leisurely pace. Furthermore, if the user is in a hurry, the animation unit can generate a video in which the generation AI emphasizes the shortest route. For example, if the user is excited, the animation unit generates a video with visually stimulating effects. This improves the visual learning effect by adjusting the video reproduction method according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0108] The animation unit can adjust the level of detail of the video based on the importance of the conversation when creating the video. For example, the animation unit provides detailed information for videos that include important conversations. The animation unit can also provide a standard level of detail for videos that include normal conversations. For example, the animation unit provides simple information for videos that include temporary conversations. This allows for efficient learning by adjusting the level of detail of the video based on the importance of the conversation. The adjustment of the level of detail of the video is performed based on criteria such as a detailed explanation or a concise summary.

[0109] When creating a video, the video generation unit can apply different video reproduction algorithms depending on the category of the conversation. For example, the video generation unit applies a highly accurate reproduction algorithm to a video containing a business conversation. The video generation unit can also apply a standard reproduction algorithm to a video containing a private conversation. For example, the video generation unit applies a simple reproduction algorithm to a video containing a temporary conversation. This makes it possible to reproduce an appropriate video by applying a video reproduction algorithm depending on the category of the conversation. The video reproduction algorithm is applied using techniques such as rule-based generation or machine learning-based generation.

[0110] The animation unit can estimate the user's emotions and adjust the video display method based on the estimated user emotions. For example, if the user is nervous, the generation AI of the animation unit can provide a simple, highly visible display method. Furthermore, if the user is relaxed, the animation unit can also provide a display method that includes detailed information. For example, if the user is in a hurry, the animation unit can provide a display method that focuses on the main points. This improves visibility by adjusting the video display method according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis. The generation AI can be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI.

[0111] When creating videos, the video creator can determine the priority of the videos based on the time of submission of the conversation. For example, the video creator can prioritize urgent conversations to enable quick access. The video creator can also create videos of normal conversations with a standard priority. For example, the video creator can create videos of temporary conversations with a low priority to save capacity. In this way, by determining the priority of videos based on the time of submission of the conversation, quick video reproduction is possible. The submission time is considered based on criteria such as the submission date and submission time.

[0112] The animation unit can adjust the order of videos based on the relevance of the conversations when animating them. For example, the animation unit can prioritize animation of highly relevant conversations to enable quick access. The animation unit can also animate normal conversations in a standard order. For example, the animation unit can animate less relevant conversations later to save capacity. In this way, adjusting the order of videos based on the relevance of the conversations enables efficient video reproduction. The relevance of the conversations is evaluated based on criteria such as common tags or related topics, for example.

[0113] The classification unit can estimate the user's emotions and adjust the word classification method based on the estimated user emotions. For example, if the user is relaxed, the generation AI can perform detailed classification. Alternatively, if the user is in a hurry, the classification unit can perform concise classification. For example, if the user is excited, the generation AI can perform visually stimulating classification. This allows for efficient learning by adjusting the word classification method according to the user's emotions. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis.

[0114] The classifier can adjust the level of detail of the classification based on the importance of the word during classification. For example, the classifier classifies important words in detail to provide comprehensive information. The classifier can also classify ordinary words with a standard level of detail. For example, the classifier can simply classify temporary words to save space. This allows for efficient learning by adjusting the level of detail of the classification based on the importance of the word. The adjustment of the level of detail of the classification is performed based on criteria such as a detailed description or a concise summary.

[0115] The classification unit can estimate the user's emotions and adjust the word classification order based on the estimated user emotions. For example, if the user is nervous, the generation AI will prioritize classifying important words. The classification unit can also classify detailed words in an orderly manner if the user is relaxed. For example, if the user is in a hurry, the classification unit will prioritize classifying words that capture the main points. This allows for efficient learning by adjusting the word classification order according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0116] The classifier can determine the priority of classification based on the relevance of words during classification. For example, the classifier can prioritize highly relevant words to enable quick access. The classifier can also classify words with normal relevance with a standard priority. For example, the classifier can classify less relevant words later to save capacity. In this way, determining the priority of classification based on the relevance of words enables efficient learning. The priority of classification is determined based on criteria such as the relevance and importance of words, for example.

[0117] The reference point display unit can estimate the user's emotion and adjust the display method of the reference points based on the estimated user's emotion. For example, when the user is nervous, the reference point display unit provides a simple, highly visible display method. Furthermore, when the user is relaxed, the reference point display unit can also provide a display method including detailed information. For example, when the user is in a hurry, the reference point display unit provides a display method that focuses on the main points. This improves visibility by adjusting the display method of the reference points according to the user's emotion. Emotion estimation is performed using technologies such as voice analysis, facial expression recognition, and text analysis.

[0118] The reference point display unit can adjust the level of detail of the display based on the importance of the mastery level when displaying the reference points. For example, the reference point display unit provides detailed information for reference points including important mastery levels. The reference point display unit can also provide standard level of detail for reference points including normal mastery levels. For example, the reference point display unit provides simple information for reference points including temporary mastery levels. This allows for efficient learning by adjusting the level of detail of the display based on the importance of the mastery level. The adjustment of the level of detail of the display is performed based on criteria such as a detailed explanation or a concise summary.

[0119] The reference point display unit can estimate the user's emotions and adjust the display order of reference points based on the estimated user emotions. For example, if the user is nervous, the generation AI can prioritize displaying important reference points. In addition, if the user is relaxed, the reference point display unit can also display detailed reference points in an orderly manner. For example, if the user is in a hurry, the generation AI can prioritize displaying reference points that highlight the key points. This improves visibility by adjusting the display order of reference points according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0120] The reference point display unit can determine the display priority based on the relevance of the mastery level when displaying the reference points. For example, the reference point display unit preferentially displays reference points including highly relevant mastery levels. The reference point display unit can also display reference points including normally relevant mastery levels with standard priority. For example, the reference point display unit displays reference points including less relevant mastery levels later. In this way, efficient learning is made possible by determining the display priority based on the relevance of the mastery level. The display priority is determined based on criteria such as the relevance and importance of the mastery level, for example.

[0121] The recommendation unit can estimate the user's emotions and adjust the recommendation method based on the estimated user emotions. For example, if the user is relaxed, the generation AI of the recommendation unit can make detailed recommendations. In addition, if the user is in a hurry, the generation AI of the recommendation unit can make concise recommendations. For example, if the user is excited, the generation AI of the recommendation unit can make visually stimulating recommendations. This enables efficient learning by adjusting the recommendation method according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0122] When making a recommendation, the recommendation unit can adjust the level of detail of the recommendation based on the importance of the word or phrase. For example, the recommendation unit can recommend important words and phrases in detail to provide comprehensive information. The recommendation unit can also recommend ordinary words and phrases with a standard level of detail. For example, the recommendation unit can simply recommend temporary words and phrases to save space. This allows for efficient learning by adjusting the level of detail of the recommendation based on the importance of the word or phrase. The adjustment of the level of detail of the recommendation is performed based on criteria such as a detailed explanation or a concise summary, for example.

[0123] When making recommendations, the recommendation unit can apply different recommendation algorithms depending on the category of the word or phrase. For example, the recommendation unit applies a highly accurate recommendation algorithm to business terms. The recommendation unit can also apply a standard recommendation algorithm to private terms. For example, the recommendation unit applies a simple recommendation algorithm to temporary terms. This makes it possible to make appropriate recommendations by applying a recommendation algorithm depending on the category of the word or phrase. The recommendation algorithm is applied using technologies such as collaborative filtering and content-based filtering.

[0124] The recommendation unit can estimate the user's emotions and adjust the display method of recommendations based on the estimated user emotions. For example, if the user is nervous, the generation AI of the recommendation unit can provide a simple, highly visible display method. In addition, if the user is relaxed, the generation AI can provide a display method that includes detailed information. For example, if the user is in a hurry, the recommendation unit can provide a display method that focuses on the main points. This improves visibility by adjusting the display method of recommendations according to the user's emotions. Emotions are estimated using technologies such as voice analysis, facial expression recognition, and text analysis.

[0125] When making a recommendation, the recommendation unit can determine the priority of the recommendation based on the time of submission of the word or phrase. For example, the recommendation unit can prioritize urgent words and phrases to enable quick access. The recommendation unit can also recommend normal words and phrases with a standard priority. For example, the recommendation unit can recommend temporary words and phrases with a low priority to save capacity. This enables efficient learning by determining the priority of recommendations based on the time of submission of the word or phrase. The recommendation priority is determined based on criteria such as the time of submission and importance, for example.

[0126] When making recommendations, the recommendation unit can adjust the order of recommendations based on the relevance of words and phrases. For example, the recommendation unit can prioritize highly relevant words and phrases to enable quick access. The recommendation unit can also recommend words and phrases with normal relevance in a standard order. For example, the recommendation unit can recommend less relevant words and phrases later to save space. This allows for efficient learning by adjusting the order of recommendations based on the relevance of words and phrases. The adjustment of the order of recommendations is performed based on criteria such as relevance and importance, for example. === Hard Collateral 1-1 === Each of the multiple elements, including the conversion unit, storage unit, generation unit, example sentence confirmation unit, expression confirmation unit, animation unit, classification unit, reference point display unit, and recommendation unit, is implemented, for example, by at least one of the smart device 14 and the data processing device 12. For example, the conversion unit is implemented by the processor 46 of the smart device 14 and converts the conversation content into text and audio. The storage unit stores the converted data in the database 24 of the data processing device 12. The generation unit analyzes the stored data using the specific processing unit 290 of the data processing device 12 and generates text. The example sentence confirmation unit checks other example sentences using words that appear in the conversation using the control unit 46A of the smart device 14. The expression confirmation unit checks expressions for business or public use scenarios using the control unit 46A of the smart device 14. The animation unit recreates and animates scenes from the conversation using the processor 46 of the smart device 14. The classification unit classifies words selected by the control unit 46A of the smart device 14 for review. The reference score display unit displays the TOEIC or TOEFL reference score based on the level of learning using the specific processing unit 290 of the data processing device 12. The recommendation unit uses the specific processing unit 290 of the data processing device 12 to recommend words, phrases, and example sentences that should be memorized and convert them into text. === Hard Collateral 1-2 === Each of the multiple elements, including the conversion unit, storage unit, generation unit, example sentence confirmation unit, expression confirmation unit, animation unit, classification unit, reference point display unit, and recommendation unit, is realized, for example, by at least one of the smart glasses 214 and the data processing device 12. For example, the conversion unit is realized by the processor 46 of the smart glasses 214 and converts the conversation content into text and audio. The storage unit stores the converted data in the database 24 of the data processing device 12. The generation unit analyzes the stored data using the specific processing unit 290 of the data processing device 12 and generates text. The example sentence confirmation unit checks other example sentences using words that appear in the conversation using the control unit 46A of the smart glasses 214. The expression confirmation unit checks expressions for business or public use scenarios using the control unit 46A of the smart glasses 214. The animation unit recreates and animates scenes from the conversation using the processor 46 of the smart glasses 214. The classification unit classifies words selected by the control unit 46A of the smart glasses 214 for review. The reference score display unit displays the TOEIC or TOEFL reference score based on the level of learning using the specific processing unit 290 of the data processing device 12. The recommendation unit uses the specific processing unit 290 of the data processing device 12 to recommend words, phrases, and example sentences that should be memorized and convert them into text. === Hard Collateral 1-3 === Each of the multiple elements, including the conversion unit, storage unit, generation unit, example sentence confirmation unit, expression confirmation unit, animation unit, classification unit, reference point display unit, and recommendation unit, is realized, for example, by at least one of the headset-type terminal 314 and the data processing device 12. For example, the conversion unit is realized by the processor 46 of the headset-type terminal 314 and converts the conversation content into text and audio. The storage unit stores the converted data in the database 24 of the data processing device 12. The generation unit analyzes the stored data using the specific processing unit 290 of the data processing device 12 and generates text. The example sentence confirmation unit checks other example sentences using words that appear in the conversation using the control unit 46A of the headset-type terminal 314. The expression confirmation unit checks expressions for business or public use scenarios using the control unit 46A of the headset-type terminal 314. The animation unit recreates and animates scenes from the conversation using the processor 46 of the headset-type terminal 314. The classification unit classifies words selected by the control unit 46A of the headset-type terminal 314 for review. The reference score display unit displays the TOEIC or TOEFL reference score based on the level of learning using the specific processing unit 290 of the data processing device 12. The recommendation unit uses the specific processing unit 290 of the data processing device 12 to recommend words, phrases, and example sentences that should be memorized and convert them into text. === Hard Collateral 1-4 === Each of the multiple elements, including the conversion unit, storage unit, generation unit, example sentence confirmation unit, expression confirmation unit, animation unit, classification unit, reference point display unit, and recommendation unit, is realized, for example, by at least one of the robot 414 and the data processing device 12. For example, the conversion unit is realized by the processor 46 of the robot 414 and converts the conversation content into text and audio. The storage unit stores the converted data in the database 24 of the data processing device 12. The generation unit analyzes the stored data using the specific processing unit 290 of the data processing device 12 and generates text. The example sentence confirmation unit checks other example sentences using words that appear in the conversation using the control unit 46A of the robot 414. The expression confirmation unit checks expressions in business or public usage scenarios using the control unit 46A of the robot 414. The animation unit recreates and animates scenes from the conversation using the processor 46 of the robot 414. The classification unit classifies words selected by the control unit 46A of the robot 414 for review. The reference score display unit displays the TOEIC or TOEFL reference score based on the level of learning using the specific processing unit 290 of the data processing device 12. The recommendation unit uses the specific processing unit 290 of the data processing device 12 to recommend words, phrases, and example sentences that should be memorized and convert them into text.

[0127] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0128] The language acquisition system may further include a learning style analysis unit that analyzes the user's learning style and suggests the optimal learning method. For example, if the user prefers visual learning, the learning style analysis unit may suggest learning materials that make extensive use of visual aids. If the user prefers auditory learning, the learning style analysis unit may suggest audio-based learning materials. Furthermore, if the user prefers hands-on learning, the learning system may suggest interactive simulations. This maximizes learning effectiveness by providing the optimal learning method according to the user's learning style.

[0129] The language acquisition system may further include a progress monitoring unit that monitors the user's learning progress in real time and provides appropriate feedback. For example, if the user is taking a long time to master a particular word or phrase, the progress monitoring unit may provide additional practice questions related to that word or phrase. If the user is making good progress, the progress monitoring unit may also suggest the next level of learning material. Furthermore, if the user loses motivation to learn, the progress monitoring unit may provide encouraging messages or rewards. This allows the system to support the user's learning progress in real time and improve learning effectiveness.

[0130] The language learning system may further include an emotion adaptation unit that estimates the user's emotions and adjusts the learning content based on the estimated user emotions. For example, if the user is stressed, the emotion adaptation unit may provide easy exercises that will help the user relax. If the user is excited, the emotion adaptation unit may provide challenging exercises. If the user is tired, the emotion adaptation unit may suggest a learning session that can be completed in a short time. This maximizes the learning effect by providing learning content that is tailored to the user's emotions.

[0131] The language acquisition system may further include a learning plan creation unit that analyzes the user's learning history and creates an individual learning plan. For example, the learning plan creation unit may suggest what the user should study next based on what the user has previously studied. The learning plan may also be adjusted according to the user's learning pace. Furthermore, short-term and long-term learning plans may be created according to the user's goals. This allows for efficient learning by providing an individual learning plan based on the user's learning history.

[0132] The language learning system may further include an emotion evaluation unit that estimates the user's emotions and evaluates the learning progress based on the estimated user emotions. For example, the emotion evaluation unit may highly evaluate the learning progress if the user has positive emotions. Alternatively, the emotion evaluation unit may carefully evaluate the learning progress if the user has negative emotions. Furthermore, the emotion evaluation unit may perform a standard evaluation if the user has neutral emotions. This allows the system to provide more appropriate feedback by evaluating the learning progress according to the user's emotions.

[0133] The language acquisition system may further include an environment analysis unit that analyzes the user's learning environment and suggests the optimal learning environment. For example, if the user prefers to study in a quiet environment, the environment analysis unit may suggest a noise-canceling function. If the user prefers to study while moving, the environment analysis unit may suggest a portable learning device. If the user prefers to study in a group, the environment analysis unit may suggest an online learning group. This may improve learning effectiveness by providing the optimal learning method according to the user's learning environment.

[0134] The language learning system may further include a motivation maintenance unit that estimates the user's emotions and maintains the user's motivation for learning based on the estimated emotions. For example, the motivation maintenance unit may provide encouraging messages or rewards if the user is losing motivation. It may also set challenging goals if the user is highly motivated. Furthermore, it may provide tasks of moderate difficulty if the user is moderately motivated. This allows the learning effect to be maximized by maintaining the user's motivation according to their emotions.

[0135] The language learning system may further include a privacy protection unit that anonymizes the user's learning data and protects the privacy of the data. For example, the privacy protection unit may anonymize the user's personal information to reduce the risk of providing it to a third party. The privacy protection unit may also encrypt the user's learning data and store it securely. Furthermore, the system may manage the user's data access permissions and provide data only when necessary. This allows the learning data to be used effectively while protecting the user's privacy.

[0136] The language learning system may further include a feedback adjustment unit that estimates the user's emotions and adjusts learning feedback based on the estimated user emotions. For example, the feedback adjustment unit may provide positive feedback when the user has positive emotions, or provide gentle feedback when the user has negative emotions, or provide standard feedback when the user has neutral emotions. This allows for improved learning effectiveness by providing feedback according to the user's emotions.

[0137] The language acquisition system may also include a trend prediction unit that analyzes the user's learning data and predicts learning trends. For example, the trend prediction unit predicts future learning trends based on what the user has learned in the past. It can also analyze the user's learning pace and interests and suggest what content to study next. Furthermore, it predicts short-term and long-term learning trends according to the user's goals. This allows for efficient learning by making trend predictions based on the user's learning data.

[0138] The processing flow of the second embodiment will be briefly explained below.

[0139] Step 1: The conversion unit converts the conversation content into text and audio. The conversation content can include everyday conversation, business conversation, and conversations containing technical terms. The conversion unit converts the conversation content into text using speech recognition technology, and then converts it into audio using a text generation algorithm. It can also use generation AI to analyze the conversation content and convert it into appropriate text and audio. For example, the conversation content can be converted into text using speech recognition technology, and then generation AI generates audio based on the text. Step 2: The storage unit stores the data converted by the conversion unit. The stored data can be stored in a database or in a file format. The storage unit stores the converted data in a database and searches and retrieves the data as needed. It can also store the data in a file format and back it up to external storage. Step 3: The generator analyzes the data stored by the storage unit and generates text. The generated text includes sentence structure, generation algorithms, etc. The generator uses natural language processing technology to analyze the stored data and generate appropriate text. It can also use machine learning algorithms to learn from the stored data and generate more accurate text.

[0140] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0141] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> Examples of the generative AI include a neural network (NN) and a neural network (NN). The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image (e.g., still image data or video data). The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in one or more data formats of voice data, text data, image data, etc. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and may perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processing of each of the above-mentioned parts is performed by an AI, the processing may be performed in part or entirely by the AI, but is not limited to these examples. The processing performed by an AI including the generative AI may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by an AI including the generative AI.

[0142] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0143] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0144] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0145] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0146] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0147] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0148] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0149] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0150] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0151] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0152] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0153] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0154] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0155] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0156] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0157] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt including an instruction, as well as inference data such as audio data indicating speech, text data indicating text, and image data indicating an image (e.g., still image data or video data). The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in one or more data formats, such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation models 58 include AIs other than the generation AI. Examples of AIs other than the generation AI include, but are not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and naive Bayes. These AIs can perform various types of processing, but are not limited to these examples. The AI ​​may also be an AI agent. When the processing of each of the above-described parts is performed by an AI, the processing may be performed in part or entirely by the AI, but is not limited to these examples. Processing performed by an AI, including the generation AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by an AI, including the generation AI.

[0158] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0159] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0160] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0161] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0162] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0163] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0164] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0165] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0166] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0167] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0168] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0169] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0170] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the identification processing unit 290 using these models.

[0171] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0172] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0173] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt including an instruction, as well as inference data such as audio data indicating speech, text data indicating text, and image data indicating an image (e.g., still image data or video data). The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in one or more data formats, such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation models 58 include AIs other than the generation AI. Examples of AIs other than the generation AI include, but are not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and naive Bayes. These AIs can perform various types of processing, but are not limited to these examples. The AI ​​may also be an AI agent. When the processing of each of the above-described parts is performed by an AI, the processing may be performed in part or entirely by the AI, but is not limited to these examples. Processing performed by an AI, including the generation AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by an AI, including the generation AI.

[0174] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0175] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0176] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0177] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0178] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0179] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0180] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0181] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0182] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0183] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0184] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0185] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0186] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0187] In the robot 414, the processor 46 performs the identification process. The storage 50 stores the identification program 60. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as the control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 also has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform the same process as the identification processing unit 290 using these models.

[0188] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0189] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0190] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt including an instruction, as well as inference data such as audio data indicating speech, text data indicating text, and image data indicating an image (e.g., still image data or video data). The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in one or more data formats, such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation models 58 include AIs other than the generation AI. Examples of AIs other than the generation AI include, but are not limited to, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and naive Bayes. These AIs can perform various types of processing, but are not limited to these examples. The AI ​​may also be an AI agent. When the processing of each of the above-described parts is performed by an AI, the processing may be performed in part or entirely by the AI, but is not limited to these examples. Processing performed by an AI, including the generation AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by an AI, including the generation AI.

[0191] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0192] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.

[0193] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0194] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0195] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0196] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0197] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.

[0198] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0199] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[0200] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.

[0201] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0202] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0203] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0204] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[0205] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0206] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0207] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0208] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.

[0209] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0210] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[0211] [Explanation of symbols]

[0212] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot

Claims

1. A conversion unit that converts the conversation content into text and voice; a storage unit for storing the data converted by the conversion unit; a generation unit that analyzes the data stored by the storage unit and generates text. A system characterized by:

2. Equipped with an example sentence checker that checks other example sentences using the words that appear in the conversation 2. The system of claim 1.

3. Equipped with an expression checker that checks expressions used in business or public situations 2. The system of claim 1.

4. Equipped with an animation section that recreates scenes from conversations and creates animations 2. The system of claim 1.

5. A classification unit is provided to classify selected words for review.

2. The system of claim 1.

6. Equipped with a reference score display that shows TOEIC or TOEFL reference scores based on the level of proficiency 2. The system of claim 1.

7. It has a recommendation section that recommends words, phrases, and example sentences to remember and converts them into text.

2. The system of claim 1.

8. The conversion unit Estimate the user's emotions and adjust the conversion accuracy of the conversation content based on the estimated user emotions.

2. The system of claim 1.

9. The conversion unit When converting conversation content, it automatically recognizes specific technical terms and slang and converts them appropriately.

2. The system of claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A