system

The system addresses the integration of word, grammar, and pronunciation learning by generating relevant sentences with specified accents and providing personalized feedback, enhancing learning efficiency and motivation.

JP2026068498APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Conventional methods for integrating words, grammar, and pronunciation learning in foreign languages lack efficient integration and fail to provide tailored learning environments and timely feedback, making it difficult for learners to solidify their knowledge and maintain motivation.

Method used

A system that uses natural language processing to generate relevant example sentences and speech with specified accents, records user pronunciation, and provides personalized feedback based on AI analysis and emotion recognition to enhance learning efficiency and motivation.

Benefits of technology

The system effectively integrates word, grammar, and pronunciation learning, providing timely and personalized feedback that enhances learning outcomes and maintains user motivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068498000001_ABST
    Figure 2026068498000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for generating related example sentences using natural language processing based on words entered by the user, A means for generating speech with a specified accent using speech generation technology according to the generated example sentences, A means of receiving recordings of pronunciation from users, analyzing the recording data, and generating feedback, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional foreign language learning, words, grammar, and pronunciation are learned separately, lacking efficient integration. Also, due to limited opportunities to interact with native speakers and receive feedback, it is difficult to solidify the learned content, and it is difficult to provide a learning environment tailored to individuals.

Means for Solving the Problems

[0005] This invention includes means for generating relevant example sentences using natural language processing based on words input by the user, and further generating speech with a specified accent using speech generation technology. Furthermore, by including means for recording the user's pronunciation and analyzing the recorded data to provide feedback, it provides a comprehensive and efficient language learning environment that meets individual learning needs.

[0006] A "user" refers to a person who uses the system to learn a foreign language.

[0007] A "word" refers to an individual vocabulary word in a particular language, and is a basic linguistic element that a user inputs as the target of learning.

[0008] "Natural language processing" is a technology that uses computers to process and understand human language, and is used for example sentence generation.

[0009] An "example sentence" refers to a sentence generated by natural language processing that contains specific words, and is used to enhance learning effectiveness.

[0010] "Speech generation technology" is a technology that generates speech from text information and provides the user with speech that corresponds to the selected accent.

[0011] "Accent" refers to a characteristic of pronunciation in spoken language, and is a style of pronunciation based on a particular region or culture.

[0012] "Recorded data" refers to audio data spoken by the user and recorded by the device.

[0013] "Feedback" refers to information that evaluates the user's pronunciation and indicates areas for improvement, and is provided to enhance learning effectiveness. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the numbered communication I / F (Interface) is an interface that includes a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention relates to a system for enabling users to learn foreign languages ​​more efficiently. The system operates on an infrastructure primarily consisting of the user, a terminal, and a server. Specifically, the learning process begins when the user inputs the words they wish to learn via the terminal's interface.

[0036] The server uses natural language processing to generate relevant example sentences based on words received from the user, and retrieves these from the learning database. These example sentences include the words entered by the user and have content and structure suitable for learning.

[0037] The server then uses speech generation technology to produce audio of example sentences, matching the accent and speed specified by the user. The generated audio data is sent to the user's device in real time. By listening to this, the user attempts to acquire near-native pronunciation.

[0038] Meanwhile, the terminal displays example sentences and audio data sent from the server on its user interface, providing an environment for the user to listen to and practice speaking. The user has the ability to practice speaking and record their pronunciation on the terminal.

[0039] Once recording is complete, the device sends the audio data to the server. The server analyzes the recording and generates feedback on the accuracy of pronunciation and areas for improvement. This feedback includes information that helps the user refine their pronunciation.

[0040] For example, if a user enters the word "bread," the server generates an example sentence such as "I would like some bread," creates an audio recording with the specified accent, and sends it to the user's device. The user listens to this audio and practices the pronunciation, recording their pronunciation. The server then provides feedback and guidelines for further practice. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user accesses the device and enters the word to be learned into the interface. The device prepares a request to send the entered word to the server.

[0044] Step 2:

[0045] The server parses word input requests received from the terminal and invokes a natural language processing engine to generate relevant example sentences. The generated example sentences are assembled by selecting appropriate content from the learning database.

[0046] Step 3:

[0047] Once an example sentence is generated, the server uses a speech generation module to produce audio with the specified accent and speed. The generated audio data is saved for subsequent processing.

[0048] Step 4:

[0049] The server sends the generated example sentences and audio data to the terminal. A notification indicating that the transmission is complete is recorded in the log.

[0050] Step 5:

[0051] The terminal displays the data received from the server and prepares for audio playback. The user interface is updated, and a button for the user to play the audio becomes available.

[0052] Step 6:

[0053] The user listens to the audio played on the device and practices pronouncing the provided example sentences. Afterwards, they press the "Start Recording" button on the device to record their pronunciation.

[0054] Step 7:

[0055] The device records the user's pronunciation and automatically sends the audio data to the server once recording is complete. This data is necessary to evaluate the quality of the user's pronunciation.

[0056] Step 8:

[0057] The server receives the recorded data and uses an AI analysis module to analyze the pronunciation. Based on the analysis results, it generates feedback that includes areas for improvement and additional instructional information.

[0058] Step 9:

[0059] The server sends the generated feedback to the terminal, which then displays it in the user interface. Based on this feedback, the user can then practice further.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In foreign language learning, there is a challenge in the lack of systems and methods that enable users to efficiently acquire native pronunciation. Furthermore, users often struggle to receive timely and appropriate feedback during pronunciation practice, hindering their progress. There is also a need for systems that manage the progress of individual learners and provide optimal learning methods based on that data.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user, means for generating speech with specified language characteristics using speech synthesis technology according to the generated sentences, and means for receiving voice recordings from the user, analyzing the recording data, evaluating the accuracy of pronunciation, and generating feedback. This enables the user to effectively practice foreign language pronunciation and receive real-time feedback, thereby improving learning efficiency.

[0065] A "user" is the entity that provides input to the system and advances the learning process.

[0066] A "word" is a word that a user inputs as the target of learning, and it serves as the basic information that the system uses to generate related sentences.

[0067] "Natural language processing" is a technology that allows computers to process and understand human language and generate related sentences.

[0068] A "text" is a sentence generated by natural language processing technology and used for learning.

[0069] "Speech synthesis technology" is a technology that converts text information into speech data and is used to generate speech according to user specifications.

[0070] "Speech" refers to audio data generated by speech synthesis technology, which users listen to to learn pronunciation.

[0071] "Recording" refers to the act of a user recording their own pronunciation, or the resulting audio data.

[0072] "Feedback" is information generated to evaluate the accuracy of the user's pronunciation and inform them of areas for improvement.

[0073] This invention is a system for users to efficiently learn foreign languages. The system consists of three main components: the user, the terminal, and the server. The specific operation of each component is described below.

[0074] The user enters words in the foreign language they want to learn into the terminal's interface. The entered words form the basis of information used throughout the system.

[0075] The terminal functions as a device for sending words entered by the user to the server. This terminal can be a computer or a smartphone. The terminal also plays a role in providing the user with information received from the server, both visually and audibly, to support the user's efficient learning.

[0076] The server uses a generative AI model to generate relevant sentences based on the received words. Natural language processing technology extracts appropriate sentences from a database, structuring content useful for learning. Specifically, the server generates audio data using speech synthesis technology based on the generated sentences. The accent and speaking speed of this audio are adjusted according to user specifications. The generated audio data is transmitted to the terminal in real time.

[0077] Users can listen to audio provided through the device and practice their pronunciation. The device also has a function to record the user's pronunciation and send it to the server.

[0078] The server analyzes the user's pronunciation using the received recording data and evaluates its accuracy. Based on the results, the system generates feedback to support the user's learning. The feedback is provided in an easy-to-understand format and includes suggestions for improvement and additional practice.

[0079] As a concrete example, if a user enters the word "bread," the system generates the sentence "I would like some bread," creates an audio recording with the specified accent, and sends it to the terminal. The following prompt is used in this process: "Describe a system that supports efficient foreign language learning by generating relevant example sentences based on words entered by the user and providing audio with the specified accent." The user listens to this audio, practices the pronunciation, and records their pronunciation. The server can then analyze the recorded pronunciation and provide specific feedback. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The user enters the word they want to learn into the device. The device sends this word as data to the server. The entered word becomes the basis for subsequent data generation processes on the server.

[0083] Step 2:

[0084] The server uses a generative AI model to generate related sentences based on the received words. It analyzes words using natural language processing techniques and extracts related sentences from a training database. The output is a training-appropriate sentence containing the input words.

[0085] Step 3:

[0086] The server processes the generated sentences using speech synthesis technology to produce speech based on the linguistic features specified by the user. This process creates audio data that the user can listen to and learn from. The output audio is immediately sent to the terminal.

[0087] Step 4:

[0088] The terminal displays and plays audio data and text received from the server on the user interface. The user prepares to practice pronunciation by listening to the audio. The terminal also provides functions such as play and pause.

[0089] Step 5:

[0090] The user listens to audio through the device and records their own pronunciation. The recording is done on the device, and the data is used to evaluate the user's learning. The recorded data is generated and saved.

[0091] Step 6:

[0092] The terminal sends the user's recorded data to the server. The server receives this data and analyzes it to evaluate the accuracy of pronunciation and areas for improvement. The evaluation results are generated as output.

[0093] Step 7:

[0094] The server generates helpful feedback for the user's learning based on the pronunciation evaluation results. This information includes specific guidelines to improve the user's pronunciation. The feedback is generated and immediately sent to the device.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] In foreign language learning, there is a need for a system that can effectively generate pronunciation practice materials tailored to individual learners and further evaluate their pronunciation. In particular, the challenge lies in enabling practice suited to each learner's pronunciation and rhythm, thereby supporting efficient learning.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes means for generating relevant sentences by natural language processing based on text information received from an information processing device, means for generating speech in a specified pronunciation style using speech synthesis technology according to the generated sentences, means for receiving records of pronunciation activities by users, analyzing the recorded information, and generating instructional information, and means for generating pronunciation practice materials in real time corresponding to word information input by users. This makes it possible to practice pronunciation according to the individual needs of learners.

[0100] An "information processing device" is a general term for electronic devices that have the functions of inputting, processing, and outputting data.

[0101] "Text information" refers to information expressed in the form of characters or strings of characters.

[0102] "Natural language processing" is a field of technology that enables machines to understand, interpret, and respond to human language.

[0103] A "relevant sentence" is an appropriate sentence generated based on the given words and context.

[0104] "Speech synthesis technology" is a technology for generating speech from text information.

[0105] "Specified pronunciation style" refers to the pronunciation style or format selected or set by the user.

[0106] "Record of pronunciation activity" refers to data that digitally saves what the user pronounces.

[0107] "Instructional information" refers to advice and feedback necessary for improving pronunciation and learning.

[0108] "Word information" refers to information about a specific word entered by the user.

[0109] "Pronunciation practice materials" are educational resources designed to train pronunciation.

[0110] "Real-time" refers to a situation or processing method where processing and responses occur immediately.

[0111] The system of this invention begins with the learner inputting specific word information. The user inputs the word to be learned into the interface using an information processing device such as a smartphone or smart glasses. This input information is sent to a server, which generates related sentences using natural language processing technology. A Python-based library (e.g., spaCy) is used for natural language processing.

[0112] The generated text is converted into audio data using speech synthesis technology, employing a specified pronunciation style. Google® Cloud Text-to-Speech API is used for this speech synthesis. Furthermore, this audio data is transmitted in real time to the user's information processing device, allowing the user to practice pronunciation.

[0113] The user's pronunciation activity is recorded by an information processing device, and the audio data is sent back to the server. The server analyzes the recorded data using the Google Cloud Speech-to-Text API and generates instructional information regarding the accuracy of the pronunciation. This instructional information is provided to the user, indicating the direction for further practice.

[0114] For example, if a user wants to learn the word "apple," the server generates a related sentence, "I eat an apple every day." The audio of this sentence is then generated using a specified pronunciation style and sent to the user's information processing device. The user listens to the audio, practices the pronunciation, and sends the recording back to the server. Based on the recording, the server provides feedback, indicating areas for improvement in pronunciation. In this way, learners can efficiently and personalizedly learn foreign language pronunciation.

[0115] Examples of prompts for a generative AI model include: "Based on the word entered by the user, generate natural-sounding example sentences containing that word. If possible, also consider the pronunciation of the example sentences."

[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0117] Step 1:

[0118] The user uses a smartphone or smart glasses to input the word information they want to learn. The entered word is sent as data from the device to the server.

[0119] Step 2:

[0120] The server performs natural language processing using a generative AI model based on the received word information. Specifically, it uses the natural language processing library spaCy to generate relevant sentences. In this process, the word information is the input data, and the generated sentences are the output data.

[0121] Step 3:

[0122] The server receives the generated text as text data and converts it into speech data using speech synthesis technology. During this process, the Google Cloud Text-to-Speech API is used to generate speech according to the specified pronunciation style. The output is an audio file.

[0123] Step 4:

[0124] The generated audio files are sent to the device in real time, and the user practices pronunciation by listening to them. The device provides the audio to the user using its audio playback function.

[0125] Step 5:

[0126] The user records their pronunciation using their device. The recorded audio data is sent from the device to the server. This audio file is the input data for the next step.

[0127] Step 6:

[0128] The server analyzes the received audio data using the Google Cloud Speech-to-Text API to evaluate the accuracy of the user's pronunciation. The resulting instructional information is the output data. This information includes an accuracy evaluation and feedback on areas for improvement.

[0129] Step 7:

[0130] The server generates feedback information which is sent to the terminal and notified to the user. The user can then use this guidance information to practice further. This feedback serves as a guideline for future learning.

[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0132] This invention relates to a system that generates relevant example sentences based on words entered by the user during foreign language learning, and presents the generated audio using speech generation technology. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and utilizes this information when providing feedback.

[0133] The user inputs a specific learning word via a terminal. The terminal sends the word to a server, which uses natural language processing to generate related example sentences. The generated example sentences are then spoken with the accent and speed specified by the user. The speech generation takes place on the server, and the generated audio files are sent to the terminal in real time.

[0134] The system has a function to record the user's pronunciation, and the recorded audio data is sent to a server. The server processes this audio data with an AI analysis module, evaluates the accuracy of the pronunciation, and generates feedback.

[0135] In addition, the emotion engine recognizes emotions from the user's facial expressions, tone of voice, and reaction speed during operation. The recognized emotion data is used to customize the content and format of subsequent feedback. For example, if the system detects that the user is feeling frustrated, it will provide more positive and encouraging feedback.

[0136] For example, if a user enters the word "apple" to learn it, the server will generate the example sentence "I like to eat an apple every day" and produce audio in the selected British English accent. The user listens to this audio, practices the pronunciation, and records it. The recorded audio is analyzed by the server, and feedback is provided, including areas for improvement in pronunciation. If the sentiment engine determines that the user is confused, the system will offer additional support information and encouraging comments.

[0137] This format allows users to receive a learning experience tailored to their individual emotional state, which is expected to improve learning efficiency.

[0138] The following describes the processing flow.

[0139] Step 1:

[0140] The user uses a device to input the words they want to learn into the interface. The device then constructs a request to send this input to the server.

[0141] Step 2:

[0142] The server activates a natural language processing engine based on the words received from the terminal and generates related example sentences. These example sentences are suitable for learning and are tailored to the user's proficiency level.

[0143] Step 3:

[0144] The server uses a speech generation module to convert the generated example sentences into speech. It customizes the speech with the accent and speed selected by the user and generates the speech data.

[0145] Step 4:

[0146] The server sends the generated example sentence data and customized audio files to the terminal. This transmission is done in real time, making them immediately available to the user.

[0147] Step 5:

[0148] The terminal displays the received data on the user interface and enables audio playback. The user plays the audio while reviewing the displayed example sentences.

[0149] Step 6:

[0150] The user listens to the audio played through the device and pronounces the given example sentences. Then, they activate the device's recording function to record their pronunciation.

[0151] Step 7:

[0152] The device records the user's pronunciation, and once recording is complete, it sends the audio data to the server. This audio data is used for analysis.

[0153] Step 8:

[0154] The server receives the recorded audio data and performs analysis using an AI analysis module. Based on the analysis results, it evaluates the accuracy of the pronunciation and generates feedback that includes specific areas for improvement.

[0155] Step 9:

[0156] The emotion engine analyzes the user's voice tone and device operation data to determine the user's emotional state. Based on the results, it adjusts the feedback content.

[0157] Step 10:

[0158] The server sends adjusted feedback to the device, which displays it in the user interface. The user can then review the displayed feedback and continue learning.

[0159] (Example 2)

[0160] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0161] Traditional foreign language learning systems have the problem of not maximizing learning effectiveness because they do not take into account the individual emotional state of the user. Furthermore, they have the challenge of not adequately addressing individual pronunciation improvement or maintaining motivation, as they only provide standardized feedback.

[0162] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0163] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user; means for generating speech with a specified accent and speed using speech synthesis technology according to the generated sentences; means for recording the user's pronunciation, analyzing the recording data, and generating feedback; and means for recognizing emotions from the user's facial expressions and tone of voice, and customizing the feedback based on that information. This makes it possible to provide the user with a more personalized learning experience, improve pronunciation effectively, and increase learning motivation.

[0164] A "user" refers to an individual who uses the system to learn a language.

[0165] "Words" refer to the vocabulary of a foreign language that the user inputs as the subject of learning.

[0166] "Natural language processing" refers to the technology that enables computers to understand and process human language.

[0167] "Sentence" refers to a meaningful text created by a generative AI model based on words entered by the user.

[0168] "Speech synthesis technology" refers to the technology of electronically generating and outputting speech.

[0169] "Accent" refers to the intonation and pitch specific to the pronunciation of language used in speech production.

[0170] "Speed" refers to the speed of audio playback, and it is a playback speed that can be adjusted according to the user's understanding.

[0171] "Recording" refers to the action of a user recording their own pronunciation as audio data.

[0172] "Analysis" refers to the process of analyzing recorded audio data and evaluating pronunciation.

[0173] "Feedback" refers to the information and advice provided to users based on the analysis results to improve their pronunciation.

[0174] "Facial expressions" refer to the expression of emotions conveyed through the user's facial movements.

[0175] "Voice tone" refers to the expression of emotion based on the pitch and volume of the user's voice.

[0176] "Emotion" refers to the psychological state recognized by the user during their interaction with the system.

[0177] "Customization" refers to the process of providing individualized feedback based on the user's emotions and learning progress.

[0178] This invention is a system for users to learn a language, which generates relevant sentences based on input foreign language words and provides audio using speech synthesis technology. Furthermore, it provides feedback tailored to the user's emotional state, enabling the personalization of the learning experience.

[0179] The user inputs the words they want to learn into the device. These words are sent to the server using communication technology. On the server, a generative AI model, integrated with natural language processing technology, is running and generates related sentences based on the input words. An example of a prompt message would be, "Generate a sentence using the following words."

[0180] Based on the generated sentence, the server uses speech synthesis software to produce an audio file with the specified accent and speed. This audio file is sent to the terminal in real time, and the user listens to it and practices pronunciation.

[0181] Users can record their pronunciation on their device, and this data is sent to a server for analysis. The server processes the recorded data with an AI analysis module and evaluates the accuracy of the pronunciation. Based on this evaluation, users are provided with feedback to improve their pronunciation.

[0182] In addition, the system incorporates an emotion recognition module that utilizes the device's camera and microphone to detect the user's facial expressions and tone of voice. This allows the system to analyze the user's emotions and adjust the content of the feedback accordingly. For example, if the user shows signs of confusion, the server will provide more positive and encouraging feedback to increase their motivation to learn.

[0183] In this way, users can efficiently learn a language by utilizing the voice and feedback provided by the system. Furthermore, because the learning process is individually optimized based on these operations, continuous evolution and development can be expected.

[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0185] Step 1:

[0186] The user enters the words they want to learn into the device. This input is done through the device's interface, and the words selected by the user are recorded as data on the device. The entered data is in the form of words.

[0187] Step 2:

[0188] The terminal sends word data entered by the user to the server. The data is sent to the server using a communication protocol, and the server receives it. The input is the user's words, and the output is the data transferred to the server.

[0189] Step 3:

[0190] The server uses a generative AI model to create related sentences based on the received words. A prompt (e.g., "Generate a sentence using the following words") is provided to the AI ​​model, and the resulting generated sentence is output. Natural language processing algorithms are applied for data processing.

[0191] Step 4:

[0192] The generated sentence is passed to the speech synthesis process on the server. The server generates an audio file with the specified accent and speed. To do this, it uses speech synthesis software to convert the sentence into audio data. The generated audio file is sent to the terminal as output.

[0193] Step 5:

[0194] The user listens to audio played from the device and practices their own pronunciation. During this practice, the user checks their pronunciation and operates the microphone to record their pronunciation. The output obtained from the audio playback is the user's pronunciation data.

[0195] Step 6:

[0196] The user records their pronunciation via their device and sends it to the server. The recorded data is saved in digital format and transferred to the server. The recorded data is input and ready for analysis on the server.

[0197] Step 7:

[0198] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of pronunciation. This evaluation is performed using a data analysis algorithm and forms the basis for generating feedback. The analysis results are output and become feedback data for the user.

[0199] Step 8:

[0200] During user interaction, the emotion engine analyzes facial expressions and voice tone to recognize emotions. Based on data collected via camera and microphone, emotion analysis is performed, and the recognized emotion data is output.

[0201] Step 9:

[0202] Based on the recognized emotion data, the server customizes the feedback. Combining the emotion data and pronunciation evaluation results, it generates the most appropriate feedback message for the user and sends it to the device. The customized feedback is then generated as output.

[0203] (Application Example 2)

[0204] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0205] Traditional foreign language learning systems often failed to adequately consider user emotions and merely provided uniform feedback, making it difficult to maintain learning motivation. Furthermore, the lack of interactive learning experiences in virtual spaces made it challenging to improve practical communication skills.

[0206] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0207] In this invention, the server includes means for generating relevant sentences using natural language processing based on information input from the user, means for generating sounds of a specified dialect using sound generation technology according to the generated sentences, means for receiving recordings of pronunciation by the user, analyzing the recorded data, and generating evaluations, and means for grasping emotions using a recognition engine and adjusting the evaluations. This enables a customized learning experience that takes into account the user's emotional state, making it possible to improve practical communication skills while increasing the user's motivation.

[0208] A "user" is an individual who uses the system to learn a foreign language, provides input information, and receives generated audio and evaluations.

[0209] "Information" refers to the words, phrases, and learning data that users input into the system during the learning process.

[0210] "Natural language processing" is a technology that analyzes language data received from users and generates appropriate sentences.

[0211] A "sentence" is an example sentence or related document generated by natural language processing to help the user learn.

[0212] "Sound generation technology" is a technology for converting generated text into speech in a specific dialect.

[0213] A "dialect" is a phonetic characteristic that differs depending on a particular region or culture, and it is the accent that a system selects when generating sounds.

[0214] "Records" refer to audio data that users input into the system for the purpose of correcting their pronunciation.

[0215] "Analysis" is a data processing procedure that evaluates recorded user pronunciation to identify accuracy and areas for improvement.

[0216] "Evaluation" refers to information used to provide feedback on the user's learning progress and pronunciation accuracy.

[0217] A "recognition engine" is a technology that identifies emotions from the user's voice and facial expressions and adjusts the system's response accordingly.

[0218] This system aims to improve the efficiency of foreign language learning by generating specific sentences based on words and phrases entered by the user, and then performing a series of processes to convert those sentences into speech in a specified dialect.

[0219] To enhance their virtual experience, users input information into the system through smart glasses. The device sends the user-specified word or phrase to the server. At this time, natural language processing technology is used to generate a meaningful sentence, and sound generation technology is used to create speech in the appropriate dialect of the input language.

[0220] The generated audio is delivered to the device in real time, and the user practices pronunciation while listening to it. During this process, the pronunciation is recorded on the device and sent back to the server. The server uses an AI analysis module (e.g., a generative AI model for natural language processing) to evaluate the pronunciation of this recorded data and generates feedback in a format that is easy for the user to accept.

[0221] Furthermore, the recognition engine grasps the user's emotions from their reactions and tone of voice, and provides optimized feedback accordingly. For example, when a user practices in a virtual store, it generates example sentences using the word "shopping cart," such as "I am looking for a fresh tomato in the market," to aid the user's repeated practice. In this case, an example of a prompt sentence input to the generation AI model would be, "Please enter the name of the product the user wants. Design a program that generates and speaks an appropriate example sentence."

[0222] This system personalizes the user's learning experience, maintains motivation through emotion-based feedback, and is expected to improve practical communication skills.

[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0224] Step 1:

[0225] The user inputs words and phrases into a device via smart glasses to learn a foreign language. The device receives the input information and sends it to a server. The input is the word the user wants to learn, and the output is the transfer of that word to the server.

[0226] Step 2:

[0227] The server analyzes the received words using natural language processing and generates related sentences. This process utilizes a generative AI model to output sentences with context related to the words. These generated sentences are intended to support the user's learning process.

[0228] Step 3:

[0229] Based on the generated sentence, the server uses sound generation technology to convert the sentence into speech. Here, the speech is adjusted according to the specified dialect and accent. The input is the generated sentence, and the output is dialect-adjusted audio data. The generated audio data is transmitted to the terminal in real time.

[0230] Step 4:

[0231] The device assists the user in practicing pronunciation while listening to audio. At this time, it is ready to record the user's pronunciation. The user records their voice, and the device returns the recorded data to the server. The input is the recorded audio, and the output is the transmission of data to the server.

[0232] Step 5:

[0233] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of the user's pronunciation. The evaluation process analyzes the input audio data and generates feedback on areas for improvement and successes in pronunciation. The output is feedback information that helps the user improve their learning.

[0234] Step 6:

[0235] The recognition engine analyzes the user's voice tone and reaction speed to understand their emotions. Based on this information, it adjusts the feedback content. The input is the user's voice and reaction data, and the output is customized feedback that corresponds to their emotions.

[0236] Step 7:

[0237] The server sends final feedback to the terminal, allowing the user to check their learning progress and identify areas for improvement. Using prompt sentences as an example, the system functions as a tool to enhance practical conversational skills by using a generative AI model to generate relevant example sentences when the user inputs a noun, then converting them into speech and playing them back.

[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0241] [Second Embodiment]

[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0254] This invention relates to a system for enabling users to learn foreign languages ​​more efficiently. The system operates on an infrastructure primarily consisting of the user, a terminal, and a server. Specifically, the learning process begins when the user inputs the words they wish to learn via the terminal's interface.

[0255] The server uses natural language processing to generate relevant example sentences based on words received from the user, and retrieves these from the learning database. These example sentences include the words entered by the user and have content and structure suitable for learning.

[0256] The server then uses speech generation technology to produce audio of example sentences, matching the accent and speed specified by the user. The generated audio data is sent to the user's device in real time. By listening to this, the user attempts to acquire near-native pronunciation.

[0257] Meanwhile, the terminal displays example sentences and audio data sent from the server on its user interface, providing an environment for the user to listen to and practice speaking. The user has the ability to practice speaking and record their pronunciation on the terminal.

[0258] Once recording is complete, the device sends the audio data to the server. The server analyzes the recording and generates feedback on the accuracy of pronunciation and areas for improvement. This feedback includes information that helps the user refine their pronunciation.

[0259] For example, if a user enters the word "bread," the server generates an example sentence such as "I would like some bread," creates an audio recording with the specified accent, and sends it to the user's device. The user listens to this audio and practices the pronunciation, recording their pronunciation. The server then provides feedback and guidelines for further practice. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0260] The following describes the processing flow.

[0261] Step 1:

[0262] The user accesses the device and enters the word to be learned into the interface. The device prepares a request to send the entered word to the server.

[0263] Step 2:

[0264] The server parses word input requests received from the terminal and invokes a natural language processing engine to generate relevant example sentences. The generated example sentences are assembled by selecting appropriate content from the learning database.

[0265] Step 3:

[0266] Once an example sentence is generated, the server uses a speech generation module to produce audio with the specified accent and speed. The generated audio data is saved for subsequent processing.

[0267] Step 4:

[0268] The server sends the generated example sentences and audio data to the terminal. A notification indicating that the transmission is complete is recorded in the log.

[0269] Step 5:

[0270] The terminal displays the data received from the server and prepares for audio playback. The user interface is updated, and a button for the user to play the audio becomes available.

[0271] Step 6:

[0272] The user listens to the audio played on the device and practices pronouncing the provided example sentences. Afterwards, they press the "Start Recording" button on the device to record their pronunciation.

[0273] Step 7:

[0274] The device records the user's pronunciation and automatically sends the audio data to the server once recording is complete. This data is necessary to evaluate the quality of the user's pronunciation.

[0275] Step 8:

[0276] The server receives the recorded data and uses an AI analysis module to analyze the pronunciation. Based on the analysis results, it generates feedback that includes areas for improvement and additional instructional information.

[0277] Step 9:

[0278] The server sends the generated feedback to the terminal, which then displays it in the user interface. Based on this feedback, the user can then practice further.

[0279] (Example 1)

[0280] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0281] In foreign language learning, there is a challenge in the lack of systems and methods that enable users to efficiently acquire native pronunciation. Furthermore, users often struggle to receive timely and appropriate feedback during pronunciation practice, hindering their progress. There is also a need for systems that manage the progress of individual learners and provide optimal learning methods based on that data.

[0282] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0283] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user, means for generating speech with specified language characteristics using speech synthesis technology according to the generated sentences, and means for receiving voice recordings from the user, analyzing the recording data, evaluating the accuracy of pronunciation, and generating feedback. This enables the user to effectively practice foreign language pronunciation and receive real-time feedback, thereby improving learning efficiency.

[0284] "User" refers to the entity that inputs to the system and advances the learning process.

[0285] "Word" refers to the word input by the user as the object of learning, and is the basic information for the system to generate related sentences.

[0286] "Natural language processing" refers to the technology by which a computer processes and understands human language and generates related sentences.

[0287] "Sentence" refers to the text generated by natural language processing technology and used for learning.

[0288] "Text-to-speech technology" refers to the technology that converts text information into voice data and is used to generate voice according to the user's specification.

[0289] "Voice" refers to the sound data generated by text-to-speech technology and used by the user to learn pronunciation by listening.

[0290] "Recording" refers to the act of the user recording their own pronunciation or the resulting voice data.

[0291] "Feedback" refers to the information generated to evaluate the accuracy of the user's pronunciation and inform the areas for improvement.

[0292] This invention is a system for the user to efficiently learn a foreign language. The system consists of three main components: the user, the terminal, and the server. The specific operations of each component are described below.

[0293] The user inputs the words of the foreign language to be learned into the interface of the terminal. The input words are the basis of the information used throughout the system.

[0294] The terminal functions as a device for sending words entered by the user to the server. This terminal can be a computer or a smartphone. The terminal also plays a role in providing the user with information received from the server, both visually and audibly, to support the user's efficient learning.

[0295] The server uses a generative AI model to generate relevant sentences based on the received words. Natural language processing technology extracts appropriate sentences from a database, structuring content useful for learning. Specifically, the server generates audio data using speech synthesis technology based on the generated sentences. The accent and speaking speed of this audio are adjusted according to user specifications. The generated audio data is transmitted to the terminal in real time.

[0296] Users can listen to audio provided through the device and practice their pronunciation. The device also has a function to record the user's pronunciation and send it to the server.

[0297] The server analyzes the user's pronunciation using the received recording data and evaluates its accuracy. Based on the results, the system generates feedback to support the user's learning. The feedback is provided in an easy-to-understand format and includes suggestions for improvement and additional practice.

[0298] As a concrete example, if a user enters the word "bread," the system generates the sentence "I would like some bread," creates an audio recording with the specified accent, and sends it to the terminal. The following prompt is used in this process: "Describe a system that supports efficient foreign language learning by generating relevant example sentences based on words entered by the user and providing audio with the specified accent." The user listens to this audio, practices the pronunciation, and records their pronunciation. The server can then analyze the recorded pronunciation and provide specific feedback. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0299] The flow of the specific process in Example 1 will be described using FIG. 11.

[0300] Step 1:

[0301] The user inputs the word to be learned into the terminal. The terminal transmits this word as data to the server. The input word becomes the information that forms the basis for the subsequent data generation process at the server.

[0302] Step 2:

[0303] Based on the received word, the server uses the generation AI model to generate related sentences. The words are analyzed using natural language processing technology, and related sentences are extracted from the learning database. As output, sentences suitable for learning that include the input word are generated.

[0304] Step 3:

[0305] The server applies the generated sentences to speech synthesis technology to generate speech based on the language features specified by the user. Through this process, speech data that the user can listen to and learn from is created. The output speech is immediately transmitted to the terminal.

[0306] Step 4:

[0307] The terminal displays and plays the speech data and sentences received from the server on the user interface. The user prepares to practice pronunciation by listening to the speech. The terminal also provides functions such as playback and pause.

[0308] Step 5:

[0309] The user listens to the speech through the terminal and records their own pronunciation. The recording is done on the terminal, and the data is used for the user's learning evaluation. Recording data is generated and saved.

[0310] Step 6:

[0311] The terminal sends the user's recorded data to the server. The server receives this data and analyzes it to evaluate the accuracy of pronunciation and areas for improvement. The evaluation results are generated as output.

[0312] Step 7:

[0313] The server generates helpful feedback for the user's learning based on the pronunciation evaluation results. This information includes specific guidelines to improve the user's pronunciation. The feedback is generated and immediately sent to the device.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] In foreign language learning, there is a need for a system that can effectively generate pronunciation practice materials tailored to individual learners and further evaluate their pronunciation. In particular, the challenge lies in enabling practice suited to each learner's pronunciation and rhythm, thereby supporting efficient learning.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes means for generating relevant sentences by natural language processing based on text information received from an information processing device, means for generating speech in a specified pronunciation style using speech synthesis technology according to the generated sentences, means for receiving records of pronunciation activities by users, analyzing the recorded information, and generating instructional information, and means for generating pronunciation practice materials in real time corresponding to word information input by users. This makes it possible to practice pronunciation according to the individual needs of learners.

[0319] An "information processing device" is a general term for electronic devices that have the functions of inputting, processing, and outputting data.

[0320] "Text information" refers to information expressed in the form of characters or strings of characters.

[0321] "Natural language processing" is a field of technology that enables machines to understand, interpret, and respond to human language.

[0322] A "relevant sentence" is an appropriate sentence generated based on the given words and context.

[0323] "Speech synthesis technology" is a technology for generating speech from text information.

[0324] "Specified pronunciation style" refers to the pronunciation style or format selected or set by the user.

[0325] "Record of pronunciation activity" refers to data that digitally saves what the user pronounces.

[0326] "Instructional information" refers to advice and feedback necessary for improving pronunciation and learning.

[0327] "Word information" refers to information about a specific word entered by the user.

[0328] "Pronunciation practice materials" are educational resources designed to train pronunciation.

[0329] "Real-time" refers to a situation or processing method where processing and responses occur immediately.

[0330] The system of this invention begins with the learner inputting specific word information. The user inputs the word to be learned into the interface using an information processing device such as a smartphone or smart glasses. This input information is sent to a server, which generates related sentences using natural language processing technology. A Python-based library (e.g., spaCy) is used for natural language processing.

[0331] The generated text is converted into audio data using a specified pronunciation style via speech synthesis technology. The Google Cloud Text-to-Speech API is used for speech synthesis. Furthermore, this audio data is sent in real time to the user's information processing device, allowing the user to practice pronunciation.

[0332] The user's pronunciation activity is recorded by an information processing device, and the audio data is sent back to the server. The server analyzes the recorded data using the Google Cloud Speech-to-Text API and generates instructional information regarding the accuracy of the pronunciation. This instructional information is provided to the user, indicating the direction for further practice.

[0333] For example, if a user wants to learn the word "apple," the server generates a related sentence, "I eat an apple every day." The audio of this sentence is then generated using a specified pronunciation style and sent to the user's information processing device. The user listens to the audio, practices the pronunciation, and sends the recording back to the server. Based on the recording, the server provides feedback, indicating areas for improvement in pronunciation. In this way, learners can efficiently and personalizedly learn foreign language pronunciation.

[0334] Examples of prompts for a generative AI model include: "Based on the word entered by the user, generate natural-sounding example sentences containing that word. If possible, also consider the pronunciation of the example sentences."

[0335] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0336] Step 1:

[0337] The user uses a smartphone or smart glasses to input the word information they want to learn. The entered word is sent as data from the device to the server.

[0338] Step 2:

[0339] The server performs natural language processing using a generative AI model based on the received word information. Specifically, it uses the natural language processing library spaCy to generate relevant sentences. In this process, the word information is the input data, and the generated sentences are the output data.

[0340] Step 3:

[0341] The server receives the generated text as text data and converts it into speech data using speech synthesis technology. During this process, the Google Cloud Text-to-Speech API is used to generate speech according to the specified pronunciation style. The output is an audio file.

[0342] Step 4:

[0343] The generated audio files are sent to the device in real time, and the user practices pronunciation by listening to them. The device provides the audio to the user using its audio playback function.

[0344] Step 5:

[0345] The user records their pronunciation using their device. The recorded audio data is sent from the device to the server. This audio file is the input data for the next step.

[0346] Step 6:

[0347] The server analyzes the received audio data using the Google Cloud Speech-to-Text API to evaluate the accuracy of the user's pronunciation. The resulting instructional information is the output data. This information includes an accuracy evaluation and feedback on areas for improvement.

[0348] Step 7:

[0349] The server generates feedback information which is sent to the terminal and notified to the user. The user can then use this guidance information to practice further. This feedback serves as a guideline for future learning.

[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0351] This invention relates to a system that generates relevant example sentences based on words entered by the user during foreign language learning, and presents the generated audio using speech generation technology. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and utilizes this information when providing feedback.

[0352] The user inputs a specific learning word via a terminal. The terminal sends the word to a server, which uses natural language processing to generate related example sentences. The generated example sentences are then spoken with the accent and speed specified by the user. The speech generation takes place on the server, and the generated audio files are sent to the terminal in real time.

[0353] The system has a function to record the user's pronunciation, and the recorded audio data is sent to a server. The server processes this audio data with an AI analysis module, evaluates the accuracy of the pronunciation, and generates feedback.

[0354] In addition, the emotion engine recognizes emotions from the user's facial expressions, tone of voice, and reaction speed during operation. The recognized emotion data is used to customize the content and format of subsequent feedback. For example, if the system detects that the user is feeling frustrated, it will provide more positive and encouraging feedback.

[0355] For example, if a user enters the word "apple" to learn it, the server will generate the example sentence "I like to eat an apple every day" and produce audio in the selected British English accent. The user listens to this audio, practices the pronunciation, and records it. The recorded audio is analyzed by the server, and feedback is provided, including areas for improvement in pronunciation. If the sentiment engine determines that the user is confused, the system will offer additional support information and encouraging comments.

[0356] This format allows users to receive a learning experience tailored to their individual emotional state, which is expected to improve learning efficiency.

[0357] The following describes the processing flow.

[0358] Step 1:

[0359] The user uses a device to input the words they want to learn into the interface. The device then constructs a request to send this input to the server.

[0360] Step 2:

[0361] The server activates a natural language processing engine based on the words received from the terminal and generates related example sentences. These example sentences are suitable for learning and are tailored to the user's proficiency level.

[0362] Step 3:

[0363] The server uses a speech generation module to convert the generated example sentences into speech. It customizes the speech with the accent and speed selected by the user and generates the speech data.

[0364] Step 4:

[0365] The server sends the generated example sentence data and customized audio files to the terminal. This transmission is done in real time, making them immediately available to the user.

[0366] Step 5:

[0367] The terminal displays the received data on the user interface and enables audio playback. The user plays the audio while reviewing the displayed example sentences.

[0368] Step 6:

[0369] The user listens to the audio played through the device and pronounces the given example sentences. Then, they activate the device's recording function to record their pronunciation.

[0370] Step 7:

[0371] The device records the user's pronunciation, and once recording is complete, it sends the audio data to the server. This audio data is used for analysis.

[0372] Step 8:

[0373] The server receives the recorded audio data and performs analysis using an AI analysis module. Based on the analysis results, it evaluates the accuracy of the pronunciation and generates feedback that includes specific areas for improvement.

[0374] Step 9:

[0375] The emotion engine analyzes the user's voice tone and device operation data to determine the user's emotional state. Based on the results, it adjusts the feedback content.

[0376] Step 10:

[0377] The server sends adjusted feedback to the device, which displays it in the user interface. The user can then review the displayed feedback and continue learning.

[0378] (Example 2)

[0379] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0380] Traditional foreign language learning systems have the problem of not maximizing learning effectiveness because they do not take into account the individual emotional state of the user. Furthermore, they have the challenge of not adequately addressing individual pronunciation improvement or maintaining motivation, as they only provide standardized feedback.

[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0382] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user; means for generating speech with a specified accent and speed using speech synthesis technology according to the generated sentences; means for recording the user's pronunciation, analyzing the recording data, and generating feedback; and means for recognizing emotions from the user's facial expressions and tone of voice, and customizing the feedback based on that information. This makes it possible to provide the user with a more personalized learning experience, improve pronunciation effectively, and increase learning motivation.

[0383] A "user" refers to an individual who uses the system to learn a language.

[0384] "Words" refer to the vocabulary of a foreign language that the user inputs as the subject of learning.

[0385] "Natural language processing" refers to the technology that enables computers to understand and process human language.

[0386] "Sentence" refers to a meaningful text created by a generative AI model based on words entered by the user.

[0387] "Speech synthesis technology" refers to the technology of electronically generating and outputting speech.

[0388] "Accent" refers to the intonation and pitch specific to the pronunciation of language used in speech production.

[0389] "Speed" refers to the speed of audio playback, and it is a playback speed that can be adjusted according to the user's understanding.

[0390] "Recording" refers to the action of a user recording their own pronunciation as audio data.

[0391] "Analysis" refers to the process of analyzing recorded audio data and evaluating pronunciation.

[0392] "Feedback" refers to the information and advice provided to users based on the analysis results to improve their pronunciation.

[0393] "Facial expressions" refer to the expression of emotions conveyed through the user's facial movements.

[0394] "Voice tone" refers to the expression of emotion based on the pitch and volume of the user's voice.

[0395] "Emotion" refers to the psychological state recognized by the user during their interaction with the system.

[0396] "Customization" refers to the process of providing individualized feedback based on the user's emotions and learning progress.

[0397] This invention is a system for users to learn a language, which generates relevant sentences based on input foreign language words and provides audio using speech synthesis technology. Furthermore, it provides feedback tailored to the user's emotional state, enabling the personalization of the learning experience.

[0398] The user inputs the words they want to learn into the device. These words are sent to the server using communication technology. On the server, a generative AI model, integrated with natural language processing technology, is running and generates relevant sentences based on the input words. An example of a prompt message would be, "Generate a sentence using the following words."

[0399] Based on the generated sentence, the server uses speech synthesis software to produce an audio file with the specified accent and speed. This audio file is sent to the terminal in real time, and the user listens to it and practices pronunciation.

[0400] Users can record their pronunciation on their device, and this data is sent to a server for analysis. The server processes the recorded data with an AI analysis module and evaluates the accuracy of the pronunciation. Based on this evaluation, users are provided with feedback to improve their pronunciation.

[0401] In addition, the system incorporates an emotion recognition module that utilizes the device's camera and microphone to detect the user's facial expressions and tone of voice. This allows the system to analyze the user's emotions and adjust the content of the feedback accordingly. For example, if the user shows signs of confusion, the server will provide more positive and encouraging feedback to increase their motivation to learn.

[0402] In this way, users can efficiently learn a language by utilizing the voice and feedback provided by the system. Furthermore, because the learning process is individually optimized based on these operations, continuous evolution and development can be expected.

[0403] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0404] Step 1:

[0405] The user enters the words they want to learn into the device. This input is done through the device's interface, and the words selected by the user are recorded as data on the device. The entered data is in the form of words.

[0406] Step 2:

[0407] The terminal sends word data entered by the user to the server. The data is sent to the server using a communication protocol, and the server receives it. The input is the user's words, and the output is the data transferred to the server.

[0408] Step 3:

[0409] The server uses a generative AI model to create related sentences based on the received words. A prompt (e.g., "Generate a sentence using the following words") is provided to the AI ​​model, and the resulting generated sentence is output. Natural language processing algorithms are applied for data processing.

[0410] Step 4:

[0411] The generated sentence is passed to the speech synthesis process on the server. The server generates an audio file with the specified accent and speed. To do this, it uses speech synthesis software to convert the sentence into audio data. The generated audio file is sent to the terminal as output.

[0412] Step 5:

[0413] The user listens to audio played from the device and practices their own pronunciation. During this practice, the user checks their pronunciation and operates the microphone to record their pronunciation. The output obtained from the audio playback is the user's pronunciation data.

[0414] Step 6:

[0415] The user records their pronunciation via their device and sends it to the server. The recorded data is saved in digital format and transferred to the server. The recorded data is input and ready for analysis on the server.

[0416] Step 7:

[0417] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of pronunciation. This evaluation is performed using a data analysis algorithm and forms the basis for generating feedback. The analysis results are output and become feedback data for the user.

[0418] Step 8:

[0419] During user interaction, the emotion engine analyzes facial expressions and voice tone to recognize emotions. Based on data collected via camera and microphone, emotion analysis is performed, and the recognized emotion data is output.

[0420] Step 9:

[0421] Based on the recognized emotion data, the server customizes the feedback. Combining the emotion data and pronunciation evaluation results, it generates the most appropriate feedback message for the user and sends it to the device. The customized feedback is then generated as output.

[0422] (Application Example 2)

[0423] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0424] Traditional foreign language learning systems often failed to adequately consider user emotions and merely provided uniform feedback, making it difficult to maintain learning motivation. Furthermore, the lack of interactive learning experiences in virtual spaces made it challenging to improve practical communication skills.

[0425] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0426] In this invention, the server includes means for generating relevant sentences using natural language processing based on information input from the user, means for generating sounds of a specified dialect using sound generation technology according to the generated sentences, means for receiving recordings of pronunciation by the user, analyzing the recorded data, and generating evaluations, and means for grasping emotions using a recognition engine and adjusting the evaluations. This enables a customized learning experience that takes into account the user's emotional state, making it possible to improve practical communication skills while increasing the user's motivation.

[0427] A "user" is an individual who uses the system to learn a foreign language, provides input information, and receives generated audio and evaluations.

[0428] "Information" refers to the words, phrases, and learning data that users input into the system during the learning process.

[0429] "Natural language processing" is a technology that analyzes language data received from users and generates appropriate sentences.

[0430] A "sentence" is an example sentence or related document generated by natural language processing to help the user learn.

[0431] "Sound generation technology" is a technology for converting generated text into speech in a specific dialect.

[0432] A "dialect" is a phonetic characteristic that differs depending on a particular region or culture, and it is the accent that a system selects when generating sounds.

[0433] "Records" refer to audio data that users input into the system for the purpose of correcting their pronunciation.

[0434] "Analysis" is a data processing procedure that evaluates recorded user pronunciation to identify accuracy and areas for improvement.

[0435] "Evaluation" refers to information used to provide feedback on the user's learning progress and pronunciation accuracy.

[0436] A "recognition engine" is a technology that identifies emotions from the user's voice and facial expressions and adjusts the system's response accordingly.

[0437] This system aims to improve the efficiency of foreign language learning by generating specific sentences based on words and phrases entered by the user, and then performing a series of processes to convert those sentences into speech in a specified dialect.

[0438] To enhance their virtual experience, users input information into the system through smart glasses. The device sends the user-specified word or phrase to the server. At this time, natural language processing technology is used to generate a meaningful sentence, and sound generation technology is used to create speech in the appropriate dialect of the input language.

[0439] The generated audio is delivered to the device in real time, and the user practices pronunciation while listening to it. During this process, the pronunciation is recorded on the device and sent back to the server. The server uses an AI analysis module (e.g., a generative AI model for natural language processing) to evaluate the pronunciation of this recorded data and generates feedback in a format that is easy for the user to accept.

[0440] Furthermore, the recognition engine grasps the user's emotions from their reactions and tone of voice, and provides optimized feedback accordingly. For example, when a user practices in a virtual store, it generates example sentences using the word "shopping cart," such as "I am looking for a fresh tomato in the market," to aid the user's repeated practice. In this case, an example of a prompt sentence input to the generation AI model would be, "Please enter the name of the product the user wants. Design a program that generates and speaks an appropriate example sentence."

[0441] This system personalizes the user's learning experience, maintains motivation through emotion-based feedback, and is expected to improve practical communication skills.

[0442] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0443] Step 1:

[0444] The user inputs words and phrases into a device via smart glasses to learn a foreign language. The device receives the input information and sends it to a server. The input is the word the user wants to learn, and the output is the transfer of that word to the server.

[0445] Step 2:

[0446] The server analyzes the received words using natural language processing and generates related sentences. This process utilizes a generative AI model to output sentences with context related to the words. These generated sentences are intended to support the user's learning process.

[0447] Step 3:

[0448] Based on the generated sentence, the server uses sound generation technology to convert the sentence into speech. Here, the speech is adjusted according to the specified dialect and accent. The input is the generated sentence, and the output is dialect-adjusted audio data. The generated audio data is transmitted to the terminal in real time.

[0449] Step 4:

[0450] The device assists the user in practicing pronunciation while listening to audio. At this time, it is ready to record the user's pronunciation. The user records their voice, and the device returns the recorded data to the server. The input is the recorded audio, and the output is the transmission of data to the server.

[0451] Step 5:

[0452] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of the user's pronunciation. The evaluation process analyzes the input audio data and generates feedback on areas for improvement and successes in pronunciation. The output is feedback information that helps the user improve their learning.

[0453] Step 6:

[0454] The recognition engine analyzes the user's voice tone and reaction speed to understand their emotions. Based on this information, it adjusts the feedback content. The input is the user's voice and reaction data, and the output is customized feedback that corresponds to their emotions.

[0455] Step 7:

[0456] The server sends final feedback to the terminal, allowing the user to check their learning progress and identify areas for improvement. Using prompt sentences as an example, the system functions as a tool to enhance practical conversational skills by using a generative AI model to generate relevant example sentences when the user inputs a noun, then converting them into speech and playing them back.

[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0460] [Third Embodiment]

[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0473] This invention relates to a system for enabling users to learn foreign languages ​​more efficiently. The system operates on an infrastructure primarily consisting of the user, a terminal, and a server. Specifically, the learning process begins when the user inputs the words they wish to learn via the terminal's interface.

[0474] The server uses natural language processing to generate relevant example sentences based on words received from the user, and retrieves these from the learning database. These example sentences include the words entered by the user and have content and structure suitable for learning.

[0475] The server then uses speech generation technology to produce audio of example sentences, matching the accent and speed specified by the user. The generated audio data is sent to the user's device in real time. By listening to this, the user attempts to acquire near-native pronunciation.

[0476] Meanwhile, the terminal displays example sentences and audio data sent from the server on its user interface, providing an environment for the user to listen to and practice speaking. The user has the ability to practice speaking and record their pronunciation on the terminal.

[0477] Once recording is complete, the device sends the audio data to the server. The server analyzes the recording and generates feedback on the accuracy of pronunciation and areas for improvement. This feedback includes information that helps the user refine their pronunciation.

[0478] For example, if a user enters the word "bread," the server generates an example sentence such as "I would like some bread," creates an audio recording with the specified accent, and sends it to the user's device. The user listens to this audio and practices the pronunciation, recording their pronunciation. The server then provides feedback and guidelines for further practice. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0479] The following describes the processing flow.

[0480] Step 1:

[0481] The user accesses the device and enters the word to be learned into the interface. The device prepares a request to send the entered word to the server.

[0482] Step 2:

[0483] The server parses word input requests received from the terminal and invokes a natural language processing engine to generate relevant example sentences. The generated example sentences are assembled by selecting appropriate content from the learning database.

[0484] Step 3:

[0485] Once an example sentence is generated, the server uses a speech generation module to produce audio with the specified accent and speed. The generated audio data is saved for subsequent processing.

[0486] Step 4:

[0487] The server sends the generated example sentences and audio data to the terminal. A notification indicating that the transmission is complete is recorded in the log.

[0488] Step 5:

[0489] The terminal displays the data received from the server and prepares for audio playback. The user interface is updated, and a button for the user to play the audio becomes available.

[0490] Step 6:

[0491] The user listens to the audio played on the device and practices pronouncing the provided example sentences. Afterwards, they press the "Start Recording" button on the device to record their pronunciation.

[0492] Step 7:

[0493] The device records the user's pronunciation and automatically sends the audio data to the server once recording is complete. This data is necessary to evaluate the quality of the user's pronunciation.

[0494] Step 8:

[0495] The server receives the recorded data and uses an AI analysis module to analyze the pronunciation. Based on the analysis results, it generates feedback that includes areas for improvement and additional instructional information.

[0496] Step 9:

[0497] The server sends the generated feedback to the terminal, which then displays it in the user interface. Based on this feedback, the user can then practice further.

[0498] (Example 1)

[0499] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0500] In foreign language learning, there is a challenge in the lack of systems and methods that enable users to efficiently acquire native pronunciation. Furthermore, users often struggle to receive timely and appropriate feedback during pronunciation practice, hindering their progress. There is also a need for systems that manage the progress of individual learners and provide optimal learning methods based on that data.

[0501] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0502] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user, means for generating speech with specified language characteristics using speech synthesis technology according to the generated sentences, and means for receiving voice recordings from the user, analyzing the recording data, evaluating the accuracy of pronunciation, and generating feedback. This enables the user to effectively practice foreign language pronunciation and receive real-time feedback, thereby improving learning efficiency.

[0503] A "user" is the entity that provides input to the system and advances the learning process.

[0504] A "word" is a word that a user inputs as the target of learning, and it serves as the basic information that the system uses to generate related sentences.

[0505] "Natural language processing" is a technology that allows computers to process and understand human language and generate related sentences.

[0506] A "text" is a sentence generated by natural language processing technology and used for learning.

[0507] "Speech synthesis technology" is a technology that converts text information into speech data and is used to generate speech according to user specifications.

[0508] "Speech" refers to audio data generated by speech synthesis technology, which users listen to to learn pronunciation.

[0509] "Recording" refers to the act of a user recording their own pronunciation, or the resulting audio data.

[0510] "Feedback" is information generated to evaluate the accuracy of the user's pronunciation and inform them of areas for improvement.

[0511] This invention is a system for users to efficiently learn foreign languages. The system consists of three main components: the user, the terminal, and the server. The specific operation of each component is described below.

[0512] The user enters words in the foreign language they want to learn into the terminal's interface. The entered words form the basis of information used throughout the system.

[0513] The terminal functions as a device for sending words entered by the user to the server. This terminal can be a computer or a smartphone. The terminal also plays a role in providing the user with information received from the server, both visually and audibly, to support the user's efficient learning.

[0514] The server uses a generative AI model to generate relevant sentences based on the received words. Natural language processing technology extracts appropriate sentences from a database, structuring content useful for learning. Specifically, the server generates audio data using speech synthesis technology based on the generated sentences. The accent and speaking speed of this audio are adjusted according to user specifications. The generated audio data is transmitted to the terminal in real time.

[0515] Users can listen to audio provided through the device and practice their pronunciation. The device also has a function to record the user's pronunciation and send it to the server.

[0516] The server analyzes the user's pronunciation using the received recording data and evaluates its accuracy. Based on the results, the system generates feedback to support the user's learning. The feedback is provided in an easy-to-understand format and includes suggestions for improvement and additional practice.

[0517] As a concrete example, if a user enters the word "bread," the system generates the sentence "I would like some bread," creates an audio recording with the specified accent, and sends it to the terminal. The following prompt is used in this process: "Describe a system that supports efficient foreign language learning by generating relevant example sentences based on words entered by the user and providing audio with the specified accent." The user listens to this audio, practices the pronunciation, and records their pronunciation. The server can then analyze the recorded pronunciation and provide specific feedback. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0518] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0519] Step 1:

[0520] The user enters the word they want to learn into the device. The device sends this word as data to the server. The entered word becomes the basis for subsequent data generation processes on the server.

[0521] Step 2:

[0522] The server uses a generative AI model to generate related sentences based on the received words. It analyzes words using natural language processing techniques and extracts related sentences from a training database. The output is a training-appropriate sentence containing the input words.

[0523] Step 3:

[0524] The server processes the generated sentences using speech synthesis technology to produce speech based on the linguistic features specified by the user. This process creates audio data that the user can listen to and learn from. The output audio is immediately sent to the terminal.

[0525] Step 4:

[0526] The terminal displays and plays audio data and text received from the server on the user interface. The user prepares to practice pronunciation by listening to the audio. The terminal also provides functions such as play and pause.

[0527] Step 5:

[0528] The user listens to audio through the device and records their own pronunciation. The recording is done on the device, and the data is used to evaluate the user's learning. The recorded data is generated and saved.

[0529] Step 6:

[0530] The terminal sends the user's recorded data to the server. The server receives this data and analyzes it to evaluate the accuracy of pronunciation and areas for improvement. The evaluation results are generated as output.

[0531] Step 7:

[0532] The server generates helpful feedback for the user's learning based on the pronunciation evaluation results. This information includes specific guidelines to improve the user's pronunciation. The feedback is generated and immediately sent to the device.

[0533] (Application Example 1)

[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0535] In foreign language learning, there is a need for a system that can effectively generate pronunciation practice materials tailored to individual learners and further evaluate their pronunciation. In particular, the challenge lies in enabling practice suited to each learner's pronunciation and rhythm, thereby supporting efficient learning.

[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0537] In this invention, the server includes means for generating relevant sentences by natural language processing based on text information received from an information processing device, means for generating speech in a specified pronunciation style using speech synthesis technology according to the generated sentences, means for receiving records of pronunciation activities by users, analyzing the recorded information, and generating instructional information, and means for generating pronunciation practice materials in real time corresponding to word information input by users. This makes it possible to practice pronunciation according to the individual needs of learners.

[0538] An "information processing device" is a general term for electronic devices that have the functions of inputting, processing, and outputting data.

[0539] "Text information" refers to information expressed in the form of characters or strings of characters.

[0540] "Natural language processing" is a field of technology that enables machines to understand, interpret, and respond to human language.

[0541] A "relevant sentence" is an appropriate sentence generated based on the given words and context.

[0542] "Speech synthesis technology" is a technology for generating speech from text information.

[0543] "Specified pronunciation style" refers to the pronunciation style or format selected or set by the user.

[0544] "Record of pronunciation activity" refers to data that digitally saves what the user pronounces.

[0545] "Instructional information" refers to advice and feedback necessary for improving pronunciation and learning.

[0546] "Word information" refers to information about a specific word entered by the user.

[0547] "Pronunciation practice materials" are educational resources designed to train pronunciation.

[0548] "Real-time" refers to a situation or processing method where processing and responses occur immediately.

[0549] The system of this invention begins with the learner inputting specific word information. The user inputs the word to be learned into the interface using an information processing device such as a smartphone or smart glasses. This input information is sent to a server, which generates related sentences using natural language processing technology. A Python-based library (e.g., spaCy) is used for natural language processing.

[0550] The generated text is converted into audio data using a specified pronunciation style via speech synthesis technology. The Google Cloud Text-to-Speech API is used for speech synthesis. Furthermore, this audio data is sent in real time to the user's information processing device, allowing the user to practice pronunciation.

[0551] The user's pronunciation activity is recorded by an information processing device, and the audio data is sent back to the server. The server analyzes the recorded data using the Google Cloud Speech-to-Text API and generates instructional information regarding the accuracy of the pronunciation. This instructional information is provided to the user, indicating the direction for further practice.

[0552] For example, if a user wants to learn the word "apple," the server generates a related sentence, "I eat an apple every day." The audio of this sentence is then generated using a specified pronunciation style and sent to the user's information processing device. The user listens to the audio, practices the pronunciation, and sends the recording back to the server. Based on the recording, the server provides feedback, indicating areas for improvement in pronunciation. In this way, learners can efficiently and personalizedly learn foreign language pronunciation.

[0553] Examples of prompts for a generative AI model include: "Based on the word entered by the user, generate natural-sounding example sentences containing that word. If possible, also consider the pronunciation of the example sentences."

[0554] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0555] Step 1:

[0556] The user uses a smartphone or smart glasses to input the word information they want to learn. The entered word is sent as data from the device to the server.

[0557] Step 2:

[0558] The server performs natural language processing using a generative AI model based on the received word information. Specifically, it uses the natural language processing library spaCy to generate relevant sentences. In this process, the word information is the input data, and the generated sentences are the output data.

[0559] Step 3:

[0560] The server receives the generated text as text data and converts it into speech data using speech synthesis technology. During this process, the Google Cloud Text-to-Speech API is used to generate speech according to the specified pronunciation style. The output is an audio file.

[0561] Step 4:

[0562] The generated audio files are sent to the device in real time, and the user practices pronunciation by listening to them. The device provides the audio to the user using its audio playback function.

[0563] Step 5:

[0564] The user records their pronunciation using their device. The recorded audio data is sent from the device to the server. This audio file is the input data for the next step.

[0565] Step 6:

[0566] The server analyzes the received audio data using the Google Cloud Speech-to-Text API to evaluate the accuracy of the user's pronunciation. The resulting instructional information is the output data. This information includes an accuracy evaluation and feedback on areas for improvement.

[0567] Step 7:

[0568] The server generates feedback information which is sent to the terminal and notified to the user. The user can then use this guidance information to practice further. This feedback serves as a guideline for future learning.

[0569] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0570] This invention relates to a system that generates relevant example sentences based on words entered by the user during foreign language learning, and presents the generated audio using speech generation technology. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and utilizes this information when providing feedback.

[0571] The user inputs a specific learning word via a terminal. The terminal sends the word to a server, which uses natural language processing to generate related example sentences. The generated example sentences are then spoken with the accent and speed specified by the user. The speech generation takes place on the server, and the generated audio files are sent to the terminal in real time.

[0572] The system has a function to record the user's pronunciation, and the recorded audio data is sent to a server. The server processes this audio data with an AI analysis module, evaluates the accuracy of the pronunciation, and generates feedback.

[0573] In addition, the emotion engine recognizes emotions from the user's facial expressions, tone of voice, and reaction speed during operation. The recognized emotion data is used to customize the content and format of subsequent feedback. For example, if the system detects that the user is feeling frustrated, it will provide more positive and encouraging feedback.

[0574] For example, if a user enters the word "apple" to learn it, the server will generate the example sentence "I like to eat an apple every day" and produce audio in the selected British English accent. The user listens to this audio, practices the pronunciation, and records it. The recorded audio is analyzed by the server, and feedback is provided, including areas for improvement in pronunciation. If the sentiment engine determines that the user is confused, the system will offer additional support information and encouraging comments.

[0575] This format allows users to receive a learning experience tailored to their individual emotional state, which is expected to improve learning efficiency.

[0576] The following describes the processing flow.

[0577] Step 1:

[0578] The user uses a device to input the words they want to learn into the interface. The device then constructs a request to send this input to the server.

[0579] Step 2:

[0580] The server activates a natural language processing engine based on the words received from the terminal and generates related example sentences. These example sentences are suitable for learning and are tailored to the user's proficiency level.

[0581] Step 3:

[0582] The server uses a speech generation module to convert the generated example sentences into speech. It customizes the speech with the accent and speed selected by the user and generates the speech data.

[0583] Step 4:

[0584] The server sends the generated example sentence data and customized audio files to the terminal. This transmission is done in real time, making them immediately available to the user.

[0585] Step 5:

[0586] The terminal displays the received data on the user interface and enables audio playback. The user plays the audio while reviewing the displayed example sentences.

[0587] Step 6:

[0588] The user listens to the audio played through the device and pronounces the given example sentences. Then, they activate the device's recording function to record their pronunciation.

[0589] Step 7:

[0590] The device records the user's pronunciation, and once recording is complete, it sends the audio data to the server. This audio data is used for analysis.

[0591] Step 8:

[0592] The server receives the recorded audio data and performs analysis using an AI analysis module. Based on the analysis results, it evaluates the accuracy of the pronunciation and generates feedback that includes specific areas for improvement.

[0593] Step 9:

[0594] The emotion engine analyzes the user's voice tone and device operation data to determine the user's emotional state. Based on the results, it adjusts the feedback content.

[0595] Step 10:

[0596] The server sends adjusted feedback to the device, which then displays it in the user interface. The user can review the displayed feedback and continue learning.

[0597] (Example 2)

[0598] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0599] Traditional foreign language learning systems have the problem of not maximizing learning effectiveness because they do not take into account the individual emotional state of the user. Furthermore, they have the challenge of not adequately addressing individual pronunciation improvement or maintaining motivation, as they only provide standardized feedback.

[0600] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0601] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user; means for generating speech with a specified accent and speed using speech synthesis technology according to the generated sentences; means for recording the user's pronunciation, analyzing the recording data, and generating feedback; and means for recognizing emotions from the user's facial expressions and tone of voice, and customizing the feedback based on that information. This makes it possible to provide the user with a more personalized learning experience, improve pronunciation effectively, and increase learning motivation.

[0602] A "user" refers to an individual who uses the system to learn a language.

[0603] "Words" refer to the vocabulary of a foreign language that the user inputs as the subject of learning.

[0604] "Natural language processing" refers to the technology that enables computers to understand and process human language.

[0605] "Sentence" refers to a meaningful text created by a generative AI model based on words entered by the user.

[0606] "Speech synthesis technology" refers to the technology of electronically generating and outputting speech.

[0607] "Accent" refers to the intonation and pitch specific to the pronunciation of language used in speech production.

[0608] "Speed" refers to the speed of audio playback, and it is a playback speed that can be adjusted according to the user's understanding.

[0609] "Recording" refers to the action of a user recording their own pronunciation as audio data.

[0610] "Analysis" refers to the process of analyzing recorded audio data and evaluating pronunciation.

[0611] "Feedback" refers to the information and advice provided to users based on the analysis results to improve their pronunciation.

[0612] "Facial expressions" refer to the expression of emotions conveyed through the user's facial movements.

[0613] "Voice tone" refers to the expression of emotion based on the pitch and volume of the user's voice.

[0614] "Emotion" refers to the psychological state recognized by the user during their interaction with the system.

[0615] "Customization" refers to the process of providing individualized feedback based on the user's emotions and learning progress.

[0616] This invention is a system for users to learn a language, which generates relevant sentences based on input foreign language words and provides audio using speech synthesis technology. Furthermore, it provides feedback tailored to the user's emotional state, enabling the personalization of the learning experience.

[0617] The user inputs the words they want to learn into the device. These words are sent to the server using communication technology. On the server, a generative AI model, integrated with natural language processing technology, is running and generates related sentences based on the input words. An example of a prompt message would be, "Generate a sentence using the following words."

[0618] Based on the generated sentence, the server uses speech synthesis software to produce an audio file with the specified accent and speed. This audio file is sent to the terminal in real time, and the user listens to it and practices pronunciation.

[0619] Users can record their pronunciation on their device, and this data is sent to a server for analysis. The server processes the recorded data with an AI analysis module and evaluates the accuracy of the pronunciation. Based on this evaluation, users are provided with feedback to improve their pronunciation.

[0620] In addition, the system incorporates an emotion recognition module that utilizes the device's camera and microphone to detect the user's facial expressions and tone of voice. This allows the system to analyze the user's emotions and adjust the content of the feedback accordingly. For example, if the user shows signs of confusion, the server will provide more positive and encouraging feedback to increase their motivation to learn.

[0621] In this way, users can efficiently learn a language by utilizing the voice and feedback provided by the system. Furthermore, because the learning process is individually optimized based on these operations, continuous evolution and development can be expected.

[0622] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0623] Step 1:

[0624] The user enters the words they want to learn into the device. This input is done through the device's interface, and the words selected by the user are recorded as data on the device. The entered data is in the form of words.

[0625] Step 2:

[0626] The terminal sends word data entered by the user to the server. The data is sent to the server using a communication protocol, and the server receives it. The input is the user's words, and the output is the data transferred to the server.

[0627] Step 3:

[0628] The server uses a generative AI model to create related sentences based on the received words. A prompt (e.g., "Generate a sentence using the following words") is provided to the AI ​​model, and the resulting generated sentence is output. Natural language processing algorithms are applied for data processing.

[0629] Step 4:

[0630] The generated sentence is passed to the speech synthesis process on the server. The server generates an audio file with the specified accent and speed. To do this, it uses speech synthesis software to convert the sentence into audio data. The generated audio file is sent to the terminal as output.

[0631] Step 5:

[0632] The user listens to audio played from the device and practices their own pronunciation. During this practice, the user checks their pronunciation and operates the microphone to record their pronunciation. The output obtained from the audio playback is the user's pronunciation data.

[0633] Step 6:

[0634] The user records their pronunciation via their device and sends it to the server. The recorded data is saved in digital format and transferred to the server. The recorded data is input and ready for analysis on the server.

[0635] Step 7:

[0636] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of pronunciation. This evaluation is performed using a data analysis algorithm and forms the basis for generating feedback. The analysis results are output and become feedback data for the user.

[0637] Step 8:

[0638] During user interaction, the emotion engine analyzes facial expressions and voice tone to recognize emotions. Based on data collected via camera and microphone, emotion analysis is performed, and the recognized emotion data is output.

[0639] Step 9:

[0640] Based on the recognized emotion data, the server customizes the feedback. Combining the emotion data and pronunciation evaluation results, it generates the most appropriate feedback message for the user and sends it to the device. The customized feedback is then generated as output.

[0641] (Application Example 2)

[0642] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0643] Traditional foreign language learning systems often failed to adequately consider user emotions and merely provided uniform feedback, making it difficult to maintain learning motivation. Furthermore, the lack of interactive learning experiences in virtual spaces made it challenging to improve practical communication skills.

[0644] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0645] In this invention, the server includes means for generating relevant sentences using natural language processing based on information input from the user, means for generating sounds of a specified dialect using sound generation technology according to the generated sentences, means for receiving recordings of pronunciation by the user, analyzing the recorded data, and generating evaluations, and means for grasping emotions using a recognition engine and adjusting the evaluations. This enables a customized learning experience that takes into account the user's emotional state, making it possible to improve practical communication skills while increasing the user's motivation.

[0646] A "user" is an individual who uses the system to learn a foreign language, provides input information, and receives generated audio and evaluations.

[0647] "Information" refers to the words, phrases, and learning data that users input into the system during the learning process.

[0648] "Natural language processing" is a technology that analyzes language data received from users and generates appropriate sentences.

[0649] A "sentence" is an example sentence or related document generated by natural language processing to help the user learn.

[0650] "Sound generation technology" is a technology for converting generated text into speech in a specific dialect.

[0651] A "dialect" is a phonetic characteristic that differs depending on a particular region or culture, and it is the accent that a system selects when generating sounds.

[0652] "Records" refer to audio data that users input into the system for the purpose of correcting their pronunciation.

[0653] "Analysis" is a data processing procedure that evaluates recorded user pronunciation to identify accuracy and areas for improvement.

[0654] "Evaluation" refers to information used to provide feedback on the user's learning progress and pronunciation accuracy.

[0655] A "recognition engine" is a technology that identifies emotions from the user's voice and facial expressions and adjusts the system's response accordingly.

[0656] This system aims to improve the efficiency of foreign language learning by generating specific sentences based on words and phrases entered by the user, and then performing a series of processes to convert those sentences into speech in a specified dialect.

[0657] To enhance their virtual experience, users input information into the system through smart glasses. The device sends the user-specified word or phrase to the server. At this time, natural language processing technology is used to generate a meaningful sentence, and sound generation technology is used to create speech in the appropriate dialect of the input language.

[0658] The generated audio is delivered to the device in real time, and the user practices pronunciation while listening to it. During this process, the pronunciation is recorded on the device and sent back to the server. The server uses an AI analysis module (e.g., a generative AI model for natural language processing) to evaluate the pronunciation of this recorded data and generates feedback in a format that is easy for the user to accept.

[0659] Furthermore, the recognition engine grasps the user's emotions from their reactions and tone of voice, and provides optimized feedback accordingly. For example, when a user practices in a virtual store, it generates example sentences using the word "shopping cart," such as "I am looking for a fresh tomato in the market," to aid the user's repeated practice. In this case, an example of a prompt sentence input to the generation AI model would be, "Please enter the name of the product the user wants. Design a program that generates and speaks an appropriate example sentence."

[0660] This system personalizes the user's learning experience, maintains motivation through emotion-based feedback, and is expected to improve practical communication skills.

[0661] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0662] Step 1:

[0663] The user inputs words and phrases into a device via smart glasses to learn a foreign language. The device receives the input information and sends it to a server. The input is the word the user wants to learn, and the output is the transfer of that word to the server.

[0664] Step 2:

[0665] The server analyzes the received words using natural language processing and generates related sentences. This process utilizes a generative AI model to output sentences with context related to the words. These generated sentences are intended to support the user's learning process.

[0666] Step 3:

[0667] Based on the generated sentence, the server uses sound generation technology to convert the sentence into speech. Here, the speech is adjusted according to the specified dialect and accent. The input is the generated sentence, and the output is dialect-adjusted audio data. The generated audio data is transmitted to the terminal in real time.

[0668] Step 4:

[0669] The device assists the user in practicing pronunciation while listening to audio. At this time, it is ready to record the user's pronunciation. The user records their voice, and the device returns the recorded data to the server. The input is the recorded audio, and the output is the transmission of data to the server.

[0670] Step 5:

[0671] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of the user's pronunciation. The evaluation process analyzes the input audio data and generates feedback on areas for improvement and successes in pronunciation. The output is feedback information that helps the user improve their learning.

[0672] Step 6:

[0673] The recognition engine analyzes the user's voice tone and reaction speed to understand their emotions. Based on this information, it adjusts the feedback content. The input is the user's voice and reaction data, and the output is customized feedback that corresponds to their emotions.

[0674] Step 7:

[0675] The server sends final feedback to the terminal, allowing the user to check their learning progress and identify areas for improvement. Using prompt sentences as an example, the system functions as a tool to enhance practical conversational skills by using a generative AI model to generate relevant example sentences when the user inputs a noun, then converting them into speech and playing them back.

[0676] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0678] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0679] [Fourth Embodiment]

[0680] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0681] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0682] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0683] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0684] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0685] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0686] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0687] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0688] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0691] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0693] This invention relates to a system for enabling users to learn foreign languages ​​more efficiently. The system operates on an infrastructure primarily consisting of the user, a terminal, and a server. Specifically, the learning process begins when the user inputs the words they wish to learn via the terminal's interface.

[0694] The server uses natural language processing to generate relevant example sentences based on words received from the user, and retrieves these from the learning database. These example sentences include the words entered by the user and have content and structure suitable for learning.

[0695] The server then uses speech generation technology to produce audio of example sentences, matching the accent and speed specified by the user. The generated audio data is sent to the user's device in real time. By listening to this, the user attempts to acquire near-native pronunciation.

[0696] Meanwhile, the terminal displays example sentences and audio data sent from the server on its user interface, providing an environment for the user to listen to and practice speaking. The user has the ability to practice speaking and record their pronunciation on the terminal.

[0697] Once recording is complete, the device sends the audio data to the server. The server analyzes the recording and generates feedback on the accuracy of pronunciation and areas for improvement. This feedback includes information that helps the user refine their pronunciation.

[0698] For example, if a user enters the word "bread," the server generates an example sentence such as "I would like some bread," creates an audio recording with the specified accent, and sends it to the user's device. The user listens to this audio and practices the pronunciation, recording their pronunciation. The server then provides feedback and guidelines for further practice. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0699] The following describes the processing flow.

[0700] Step 1:

[0701] The user accesses the device and enters the word to be learned into the interface. The device prepares a request to send the entered word to the server.

[0702] Step 2:

[0703] The server parses word input requests received from the terminal and invokes a natural language processing engine to generate relevant example sentences. The generated example sentences are assembled by selecting appropriate content from the learning database.

[0704] Step 3:

[0705] Once an example sentence is generated, the server uses a speech generation module to produce audio with the specified accent and speed. The generated audio data is saved for subsequent processing.

[0706] Step 4:

[0707] The server sends the generated example sentences and audio data to the terminal. A notification indicating that the transmission is complete is recorded in the log.

[0708] Step 5:

[0709] The terminal displays the data received from the server and prepares for audio playback. The user interface is updated, and a button for the user to play the audio becomes available.

[0710] Step 6:

[0711] The user listens to the audio played on the device and practices pronouncing the provided example sentences. Afterwards, they press the "Start Recording" button on the device to record their pronunciation.

[0712] Step 7:

[0713] The device records the user's pronunciation and automatically sends the audio data to the server once recording is complete. This data is necessary to evaluate the quality of the user's pronunciation.

[0714] Step 8:

[0715] The server receives the recorded data and uses an AI analysis module to analyze the pronunciation. Based on the analysis results, it generates feedback that includes areas for improvement and additional instructional information.

[0716] Step 9:

[0717] The server sends the generated feedback to the terminal, which then displays it in the user interface. Based on this feedback, the user can then practice further.

[0718] (Example 1)

[0719] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0720] In foreign language learning, there is a challenge in the lack of systems and methods that enable users to efficiently acquire native pronunciation. Furthermore, users often struggle to receive timely and appropriate feedback during pronunciation practice, hindering their progress. There is also a need for systems that manage the progress of individual learners and provide optimal learning methods based on that data.

[0721] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0722] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user, means for generating speech with specified language characteristics using speech synthesis technology according to the generated sentences, and means for receiving voice recordings from the user, analyzing the recording data, evaluating the accuracy of pronunciation, and generating feedback. This enables the user to effectively practice foreign language pronunciation and receive real-time feedback, thereby improving learning efficiency.

[0723] A "user" is the entity that provides input to the system and advances the learning process.

[0724] A "word" is a word that a user inputs as the target of learning, and it serves as the basic information that the system uses to generate related sentences.

[0725] "Natural language processing" is a technology that allows computers to process and understand human language and generate related sentences.

[0726] A "text" is a sentence generated by natural language processing technology and used for learning.

[0727] "Speech synthesis technology" is a technology that converts text information into speech data and is used to generate speech according to user specifications.

[0728] "Speech" refers to audio data generated by speech synthesis technology, which users listen to to learn pronunciation.

[0729] "Recording" refers to the act of a user recording their own pronunciation, or the resulting audio data.

[0730] "Feedback" is information generated to evaluate the accuracy of the user's pronunciation and inform them of areas for improvement.

[0731] This invention is a system for users to efficiently learn foreign languages. The system consists of three main components: the user, the terminal, and the server. The specific operation of each component is described below.

[0732] The user enters words in the foreign language they want to learn into the terminal's interface. The entered words form the basis of information used throughout the system.

[0733] The terminal functions as a device for sending words entered by the user to the server. This terminal can be a computer or a smartphone. The terminal also plays a role in providing the user with information received from the server, both visually and audibly, to support the user's efficient learning.

[0734] The server uses a generative AI model to generate relevant sentences based on the received words. Natural language processing technology extracts appropriate sentences from a database, structuring content useful for learning. Specifically, the server generates audio data using speech synthesis technology based on the generated sentences. The accent and speaking speed of this audio are adjusted according to user specifications. The generated audio data is transmitted to the terminal in real time.

[0735] Users can listen to audio provided through the device and practice their pronunciation. The device also has a function to record the user's pronunciation and send it to the server.

[0736] The server analyzes the user's pronunciation using the received recording data and evaluates its accuracy. Based on the results, the system generates feedback to support the user's learning. The feedback is provided in an easy-to-understand format and includes suggestions for improvement and additional practice.

[0737] As a concrete example, if a user enters the word "bread," the system generates the sentence "I would like some bread," creates an audio recording with the specified accent, and sends it to the terminal. The following prompt is used in this process: "Describe a system that supports efficient foreign language learning by generating relevant example sentences based on words entered by the user and providing audio with the specified accent." The user listens to this audio, practices the pronunciation, and records their pronunciation. The server can then analyze the recorded pronunciation and provide specific feedback. This allows users to efficiently learn foreign language pronunciation and grammar at home.

[0738] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0739] Step 1:

[0740] The user enters the word they want to learn into the device. The device sends this word as data to the server. The entered word becomes the basis for subsequent data generation processes on the server.

[0741] Step 2:

[0742] The server uses a generative AI model to generate related sentences based on the received words. It analyzes words using natural language processing techniques and extracts related sentences from a training database. The output is a training-appropriate sentence containing the input words.

[0743] Step 3:

[0744] The server processes the generated sentences using speech synthesis technology to produce speech based on the linguistic features specified by the user. This process creates audio data that the user can listen to and learn from. The output audio is immediately sent to the terminal.

[0745] Step 4:

[0746] The terminal displays and plays audio data and text received from the server on the user interface. The user prepares to practice pronunciation by listening to the audio. The terminal also provides functions such as play and pause.

[0747] Step 5:

[0748] The user listens to audio through the device and records their own pronunciation. The recording is done on the device, and the data is used to evaluate the user's learning. The recorded data is generated and saved.

[0749] Step 6:

[0750] The terminal sends the user's recorded data to the server. The server receives this data and analyzes it to evaluate the accuracy of pronunciation and areas for improvement. The evaluation results are generated as output.

[0751] Step 7:

[0752] The server generates helpful feedback for the user's learning based on the pronunciation evaluation results. This information includes specific guidelines to improve the user's pronunciation. The feedback is generated and immediately sent to the device.

[0753] (Application Example 1)

[0754] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0755] In foreign language learning, there is a need for a system that can effectively generate pronunciation practice materials tailored to individual learners and further evaluate their pronunciation. In particular, the challenge lies in enabling practice suited to each learner's pronunciation and rhythm, thereby supporting efficient learning.

[0756] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0757] In this invention, the server includes means for generating relevant sentences by natural language processing based on text information received from an information processing device, means for generating speech in a specified pronunciation style using speech synthesis technology according to the generated sentences, means for receiving records of pronunciation activities by users, analyzing the recorded information, and generating instructional information, and means for generating pronunciation practice materials in real time corresponding to word information input by users. This makes it possible to practice pronunciation according to the individual needs of learners.

[0758] An "information processing device" is a general term for electronic devices that have the functions of inputting, processing, and outputting data.

[0759] "Text information" refers to information expressed in the form of characters or strings of characters.

[0760] "Natural language processing" is a field of technology that enables machines to understand, interpret, and respond to human language.

[0761] A "relevant sentence" is an appropriate sentence generated based on the given words and context.

[0762] "Speech synthesis technology" is a technology for generating speech from text information.

[0763] "Specified pronunciation style" refers to the pronunciation style or format selected or set by the user.

[0764] "Record of pronunciation activity" refers to data that digitally saves what the user pronounces.

[0765] "Instructional information" refers to advice and feedback necessary for improving pronunciation and learning.

[0766] "Word information" refers to information about a specific word entered by the user.

[0767] "Pronunciation practice materials" are educational resources designed to train pronunciation.

[0768] "Real-time" refers to a situation or processing method where processing and responses occur immediately.

[0769] The system of this invention begins with the learner inputting specific word information. The user inputs the word to be learned into the interface using an information processing device such as a smartphone or smart glasses. This input information is sent to a server, which generates related sentences using natural language processing technology. A Python-based library (e.g., spaCy) is used for natural language processing.

[0770] The generated text is converted into audio data using a specified pronunciation style via speech synthesis technology. The Google Cloud Text-to-Speech API is used for speech synthesis. Furthermore, this audio data is sent in real time to the user's information processing device, allowing the user to practice pronunciation.

[0771] The user's pronunciation activity is recorded by an information processing device, and the audio data is sent back to the server. The server analyzes the recorded data using the Google Cloud Speech-to-Text API and generates instructional information regarding the accuracy of the pronunciation. This instructional information is provided to the user, indicating the direction for further practice.

[0772] For example, if a user wants to learn the word "apple," the server generates a related sentence, "I eat an apple every day." The audio of this sentence is then generated using a specified pronunciation style and sent to the user's information processing device. The user listens to the audio, practices the pronunciation, and sends the recording back to the server. Based on the recording, the server provides feedback, indicating areas for improvement in pronunciation. In this way, learners can efficiently and personalizedly learn foreign language pronunciation.

[0773] Examples of prompts for a generative AI model include: "Based on the word entered by the user, generate natural-sounding example sentences containing that word. If possible, also consider the pronunciation of the example sentences."

[0774] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0775] Step 1:

[0776] The user uses a smartphone or smart glasses to input the word information they want to learn. The entered word is sent as data from the device to the server.

[0777] Step 2:

[0778] The server performs natural language processing using a generative AI model based on the received word information. Specifically, it uses the natural language processing library spaCy to generate relevant sentences. In this process, the word information is the input data, and the generated sentences are the output data.

[0779] Step 3:

[0780] The server receives the generated text as text data and converts it into speech data using speech synthesis technology. During this process, the Google Cloud Text-to-Speech API is used to generate speech according to the specified pronunciation style. The output is an audio file.

[0781] Step 4:

[0782] The generated audio files are sent to the device in real time, and the user practices pronunciation by listening to them. The device provides the audio to the user using its audio playback function.

[0783] Step 5:

[0784] The user records their pronunciation using their device. The recorded audio data is sent from the device to the server. This audio file is the input data for the next step.

[0785] Step 6:

[0786] The server analyzes the received audio data using the Google Cloud Speech-to-Text API to evaluate the accuracy of the user's pronunciation. The resulting instructional information is the output data. This information includes an accuracy evaluation and feedback on areas for improvement.

[0787] Step 7:

[0788] The server generates feedback information which is sent to the terminal and notified to the user. The user can then use this guidance information to practice further. This feedback serves as a guideline for future learning.

[0789] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0790] This invention relates to a system that generates relevant example sentences based on words entered by the user during foreign language learning, and presents the generated audio using speech generation technology. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions and utilizes this information when providing feedback.

[0791] The user inputs a specific learning word via a terminal. The terminal sends the word to a server, which uses natural language processing to generate related example sentences. The generated example sentences are then spoken with the accent and speed specified by the user. The speech generation takes place on the server, and the generated audio files are sent to the terminal in real time.

[0792] The system has a function to record the user's pronunciation, and the recorded audio data is sent to a server. The server processes this audio data with an AI analysis module, evaluates the accuracy of the pronunciation, and generates feedback.

[0793] In addition, the emotion engine recognizes emotions from the user's facial expressions, tone of voice, and reaction speed during operation. The recognized emotion data is used to customize the content and format of subsequent feedback. For example, if the system detects that the user is feeling frustrated, it will provide more positive and encouraging feedback.

[0794] For example, if a user enters the word "apple" to learn it, the server will generate the example sentence "I like to eat an apple every day" and produce audio in the selected British English accent. The user listens to this audio, practices the pronunciation, and records it. The recorded audio is analyzed by the server, and feedback is provided, including areas for improvement in pronunciation. If the sentiment engine determines that the user is confused, the system will offer additional support information and encouraging comments.

[0795] This format allows users to receive a learning experience tailored to their individual emotional state, which is expected to improve learning efficiency.

[0796] The following describes the processing flow.

[0797] Step 1:

[0798] The user uses a device to input the words they want to learn into the interface. The device then constructs a request to send this input to the server.

[0799] Step 2:

[0800] The server activates a natural language processing engine based on the words received from the terminal and generates related example sentences. These example sentences are suitable for learning and are tailored to the user's proficiency level.

[0801] Step 3:

[0802] The server uses a speech generation module to convert the generated example sentences into speech. It customizes the speech with the accent and speed selected by the user and generates the speech data.

[0803] Step 4:

[0804] The server sends the generated example sentence data and customized audio files to the terminal. This transmission is done in real time, making them immediately available to the user.

[0805] Step 5:

[0806] The terminal displays the received data on the user interface and enables audio playback. The user plays the audio while reviewing the displayed example sentences.

[0807] Step 6:

[0808] The user listens to the audio played through the device and pronounces the given example sentences. Then, they activate the device's recording function to record their pronunciation.

[0809] Step 7:

[0810] The device records the user's pronunciation, and once recording is complete, it sends the audio data to the server. This audio data is used for analysis.

[0811] Step 8:

[0812] The server receives the recorded audio data and performs analysis using an AI analysis module. Based on the analysis results, it evaluates the accuracy of the pronunciation and generates feedback that includes specific areas for improvement.

[0813] Step 9:

[0814] The emotion engine analyzes the user's voice tone and device operation data to determine the user's emotional state. Based on the results, it adjusts the feedback content.

[0815] Step 10:

[0816] The server sends adjusted feedback to the device, which then displays it in the user interface. The user can review the displayed feedback and continue learning.

[0817] (Example 2)

[0818] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0819] Traditional foreign language learning systems have the problem of not maximizing learning effectiveness because they do not take into account the individual emotional state of the user. Furthermore, they have the challenge of not adequately addressing individual pronunciation improvement or maintaining motivation, relying only on standardized feedback.

[0820] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0821] In this invention, the server includes means for generating relevant sentences using natural language processing based on words input by the user; means for generating speech with a specified accent and speed using speech synthesis technology according to the generated sentences; means for recording the user's pronunciation, analyzing the recording data, and generating feedback; and means for recognizing emotions from the user's facial expressions and tone of voice, and customizing the feedback based on that information. This makes it possible to provide the user with a more personalized learning experience, improve pronunciation effectively, and increase learning motivation.

[0822] A "user" refers to an individual who uses the system to learn a language.

[0823] "Words" refer to the vocabulary of a foreign language that the user inputs as the subject of learning.

[0824] "Natural language processing" refers to the technology that enables computers to understand and process human language.

[0825] "Sentence" refers to a meaningful text created by a generative AI model based on words entered by the user.

[0826] "Speech synthesis technology" refers to the technology of electronically generating and outputting speech.

[0827] "Accent" refers to the intonation and pitch specific to the pronunciation of language used in speech production.

[0828] "Speed" refers to the speed of audio playback, and it is a playback speed that can be adjusted according to the user's understanding.

[0829] "Recording" refers to the action of a user recording their own pronunciation as audio data.

[0830] "Analysis" refers to the process of analyzing recorded audio data and evaluating pronunciation.

[0831] "Feedback" refers to the information and advice provided to users based on the analysis results to improve their pronunciation.

[0832] "Facial expressions" refer to the expression of emotions conveyed through the user's facial movements.

[0833] "Voice tone" refers to the expression of emotion based on the pitch and volume of the user's voice.

[0834] "Emotion" refers to the psychological state recognized by the user during their interaction with the system.

[0835] "Customization" refers to the process of providing individualized feedback based on the user's emotions and learning progress.

[0836] This invention is a system for users to learn a language, which generates relevant sentences based on input foreign language words and provides audio using speech synthesis technology. Furthermore, it provides feedback tailored to the user's emotional state, enabling the personalization of the learning experience.

[0837] The user inputs the words they want to learn into the device. These words are sent to the server using communication technology. On the server, a generative AI model, integrated with natural language processing technology, is running and generates related sentences based on the input words. An example of a prompt message would be, "Generate a sentence using the following words."

[0838] Based on the generated sentence, the server uses speech synthesis software to produce an audio file with the specified accent and speed. This audio file is sent to the terminal in real time, and the user listens to it and practices pronunciation.

[0839] Users can record their pronunciation on their device, and this data is sent to a server for analysis. The server processes the recorded data with an AI analysis module and evaluates the accuracy of the pronunciation. Based on this evaluation, users are provided with feedback to improve their pronunciation.

[0840] In addition, the system incorporates an emotion recognition module that utilizes the device's camera and microphone to detect the user's facial expressions and tone of voice. This allows the system to analyze the user's emotions and adjust the content of the feedback accordingly. For example, if the user shows signs of confusion, the server will provide more positive and encouraging feedback to increase their motivation to learn.

[0841] In this way, users can efficiently learn a language by utilizing the voice and feedback provided by the system. Furthermore, because the learning process is individually optimized based on these operations, continuous evolution and development can be expected.

[0842] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0843] Step 1:

[0844] The user enters the words they want to learn into the device. This input is done through the device's interface, and the words selected by the user are recorded as data on the device. The entered data is in the form of words.

[0845] Step 2:

[0846] The terminal sends word data entered by the user to the server. The data is sent to the server using a communication protocol, and the server receives it. The input is the user's words, and the output is the data transferred to the server.

[0847] Step 3:

[0848] The server uses a generative AI model to create related sentences based on the received words. A prompt (e.g., "Generate a sentence using the following words") is provided to the AI ​​model, and the resulting generated sentence is output. Natural language processing algorithms are applied for data processing.

[0849] Step 4:

[0850] The generated sentence is passed to the speech synthesis process on the server. The server generates an audio file with the specified accent and speed. To do this, it uses speech synthesis software to convert the sentence into audio data. The generated audio file is sent to the terminal as output.

[0851] Step 5:

[0852] The user listens to audio played from the device and practices their own pronunciation. During this practice, the user checks their pronunciation and operates the microphone to record their pronunciation. The output obtained from the audio playback is the user's pronunciation data.

[0853] Step 6:

[0854] The user records their pronunciation via their device and sends it to the server. The recorded data is saved in digital format and transferred to the server. The recorded data is input and ready for analysis on the server.

[0855] Step 7:

[0856] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of pronunciation. This evaluation is performed using a data analysis algorithm and forms the basis for generating feedback. The analysis results are output and become feedback data for the user.

[0857] Step 8:

[0858] During user interaction, the emotion engine analyzes facial expressions and voice tone to recognize emotions. Based on data collected via camera and microphone, emotion analysis is performed, and the recognized emotion data is output.

[0859] Step 9:

[0860] Based on the recognized emotion data, the server customizes the feedback. Combining the emotion data and pronunciation evaluation results, it generates the most appropriate feedback message for the user and sends it to the device. The customized feedback is then generated as output.

[0861] (Application Example 2)

[0862] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0863] Traditional foreign language learning systems often failed to adequately consider user emotions and merely provided uniform feedback, making it difficult to maintain learning motivation. Furthermore, the lack of interactive learning experiences in virtual spaces made it challenging to improve practical communication skills.

[0864] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0865] In this invention, the server includes means for generating relevant sentences using natural language processing based on information input from the user, means for generating sounds of a specified dialect using sound generation technology according to the generated sentences, means for receiving recordings of pronunciation by the user, analyzing the recorded data, and generating evaluations, and means for grasping emotions using a recognition engine and adjusting the evaluations. This enables a customized learning experience that takes into account the user's emotional state, making it possible to improve practical communication skills while increasing the user's motivation.

[0866] A "user" is an individual who uses the system to learn a foreign language, provides input information, and receives generated audio and evaluations.

[0867] "Information" refers to the words, phrases, and learning data that users input into the system during the learning process.

[0868] "Natural language processing" is a technology that analyzes language data received from users and generates appropriate sentences.

[0869] A "sentence" is an example sentence or related document generated by natural language processing to help the user learn.

[0870] "Sound generation technology" is a technology for converting generated text into speech in a specific dialect.

[0871] A "dialect" is a phonetic characteristic that differs depending on a particular region or culture, and it is the accent that a system selects when generating sounds.

[0872] "Records" refer to audio data that users input into the system for the purpose of correcting their pronunciation.

[0873] "Analysis" is a data processing procedure that evaluates recorded user pronunciation to identify accuracy and areas for improvement.

[0874] "Evaluation" refers to information used to provide feedback on the user's learning progress and pronunciation accuracy.

[0875] A "recognition engine" is a technology that identifies emotions from the user's voice and facial expressions and adjusts the system's response accordingly.

[0876] This system aims to improve the efficiency of foreign language learning by generating specific sentences based on words and phrases entered by the user, and then performing a series of processes to convert those sentences into speech in a specified dialect.

[0877] To enhance their virtual experience, users input information into the system through smart glasses. The device sends the user-specified word or phrase to the server. At this time, natural language processing technology is used to generate a meaningful sentence, and sound generation technology is used to create speech in the appropriate dialect of the input language.

[0878] The generated audio is delivered to the device in real time, and the user practices pronunciation while listening to it. During this process, the pronunciation is recorded on the device and sent back to the server. The server uses an AI analysis module (e.g., a generative AI model for natural language processing) to evaluate the pronunciation of this recorded data and generates feedback in a format that is easy for the user to accept.

[0879] Furthermore, the recognition engine grasps the user's emotions from their reactions and tone of voice, and provides optimized feedback accordingly. For example, when a user practices in a virtual store, it generates example sentences using the word "shopping cart," such as "I am looking for a fresh tomato in the market," to aid the user's repeated practice. In this case, an example of a prompt sentence input to the generation AI model would be, "Please enter the name of the product the user wants. Design a program that generates and speaks an appropriate example sentence."

[0880] This system personalizes the user's learning experience, maintains motivation through emotion-based feedback, and is expected to improve practical communication skills.

[0881] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0882] Step 1:

[0883] The user inputs words and phrases into a device via smart glasses to learn a foreign language. The device receives the input information and sends it to a server. The input is the word the user wants to learn, and the output is the transfer of that word to the server.

[0884] Step 2:

[0885] The server analyzes the received words using natural language processing and generates related sentences. This process utilizes a generative AI model to output sentences with context related to the words. These generated sentences are intended to support the user's learning process.

[0886] Step 3:

[0887] Based on the generated sentence, the server uses sound generation technology to convert the sentence into speech. Here, the speech is adjusted according to the specified dialect and accent. The input is the generated sentence, and the output is dialect-adjusted audio data. The generated audio data is transmitted to the terminal in real time.

[0888] Step 4:

[0889] The device assists the user in practicing pronunciation while listening to audio. At this time, it is ready to record the user's pronunciation. The user records their voice, and the device returns the recorded data to the server. The input is the recorded audio, and the output is the transmission of data to the server.

[0890] Step 5:

[0891] The server analyzes the recorded audio data using an AI analysis module to evaluate the accuracy of the user's pronunciation. The evaluation process analyzes the input audio data and generates feedback on areas for improvement and successes in pronunciation. The output is feedback information that helps the user improve their learning.

[0892] Step 6:

[0893] The recognition engine analyzes the user's voice tone and reaction speed to understand their emotions. Based on this information, it adjusts the feedback content. The input is the user's voice and reaction data, and the output is customized feedback that corresponds to their emotions.

[0894] Step 7:

[0895] The server sends final feedback to the terminal, allowing the user to check their learning progress and identify areas for improvement. Using prompt sentences as an example, the system functions as a tool to enhance practical conversational skills by using a generative AI model to generate relevant example sentences when the user inputs a noun, then converting them into speech and playing them back.

[0896] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0897] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0898] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0899] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0900] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0901] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0902] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0903] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0904] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0905] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0906] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0907] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0908] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0909] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0910] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0911] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0912] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0913] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0914] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0915] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0916] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0917] The following is further disclosed regarding the embodiments described above.

[0918] (Claim 1)

[0919] A means for generating related example sentences using natural language processing based on words entered by the user,

[0920] A means for generating speech with a specified accent using speech generation technology according to the generated example sentences,

[0921] A means of receiving recordings of pronunciation from users, analyzing the recording data, and generating feedback,

[0922] A system that includes this.

[0923] (Claim 2)

[0924] The system according to claim 1, which customizes the voice according to the accent and speaking speed selected by the user.

[0925] (Claim 3)

[0926] The system according to claim 1, comprising a function for recording the user's language learning history and tracking their progress.

[0927] "Example 1"

[0928] (Claim 1)

[0929] A means for generating related sentences using natural language processing based on words entered by the user,

[0930] A means for generating speech with specified language characteristics using speech synthesis technology in response to a generated sentence,

[0931] A means of receiving voice recordings from users, analyzing the recording data, evaluating the accuracy of pronunciation, and generating feedback,

[0932] A means of using an interface to present information useful for learning to users in real time,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, which adjusts the sound according to the pronunciation characteristics and speed selected by the user.

[0936] (Claim 3)

[0937] The system according to claim 1, comprising a function to record the user's language learning history and track and analyze the progress of learning.

[0938] "Application Example 1"

[0939] (Claim 1)

[0940] A means for generating related texts using natural language processing based on text information received from an information processing device,

[0941] A means for generating speech with a specified pronunciation style using speech synthesis technology in response to the generated text,

[0942] A means of receiving records of pronunciation activities by users, analyzing that recorded information, and generating instructional information,

[0943] A means for generating pronunciation practice materials in real time corresponding to word information entered by the user,

[0944] A system that includes this.

[0945] (Claim 2)

[0946] The system according to claim 1, which adjusts the voice according to the pronunciation style and speaking speed selected by the user.

[0947] (Claim 3)

[0948] The system according to claim 1, comprising a function for recording the user's language learning activity history and analyzing their progress.

[0949] "Example 2 of combining an emotion engine"

[0950] (Claim 1)

[0951] A means for generating related sentences using natural language processing based on words input by the user,

[0952] A means for generating speech with a specified accent and speed using speech synthesis technology in response to a generated sentence,

[0953] A means for recording user pronunciation, analyzing the recorded data, and generating feedback,

[0954] A means of recognizing emotions from the user's facial expressions and tone of voice, and customizing feedback based on that information,

[0955] A system that includes this.

[0956] (Claim 2)

[0957] The system according to claim 1, which adjusts the voice according to the accent and speaking speed selected by the user.

[0958] (Claim 3)

[0959] The system according to claim 1, comprising a function for recording the user's language learning history and tracking their progress.

[0960] "Application example 2 when combining with an emotional engine"

[0961] (Claim 1)

[0962] A means for generating relevant sentences using natural language processing based on information input by the user,

[0963] A means for generating sounds of a specified dialect using sound generation technology in response to a generated sentence,

[0964] A means for receiving pronunciation recordings from users, analyzing that recording data, and generating evaluations,

[0965] A means of understanding and adjusting emotions using a recognition engine,

[0966] A system that includes this.

[0967] (Claim 2)

[0968] The system according to claim 1, which allows the user to set the sound according to the dialect and speech rate selected by the user.

[0969] (Claim 3)

[0970] The system according to claim 1, which records a user's language learning history, tracks their progress, and presents information including practical examples in a virtual space. [Explanation of Symbols]

[0971] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for generating related example sentences using natural language processing based on words entered by the user, A means for generating speech with a specified accent using speech generation technology according to the generated example sentences, A means of receiving recordings of pronunciation from users, analyzing the recording data, and generating feedback, A system that includes this.

2. The system according to claim 1, which customizes the voice according to the accent and speaking speed selected by the user.

3. The system according to claim 1, further comprising a function for recording the user's language learning history and tracking their progress.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A