system
The system addresses the challenges of traditional English learning by integrating audio-visual content and feedback, enabling intuitive practice and accurate pronunciation evaluation.
Patent Information
- Application Number
- JP2024138619
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
Smart Images

Figure 2026036104000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Traditional English learning methods generally require a process that requires Japanese translation, which makes it difficult to intuitively acquire English. Furthermore, monotonous learning steps make it difficult to continue learning, making it difficult to efficiently improve listening and speaking skills. Furthermore, when it comes to pronunciation practice, learners have difficulty determining whether their pronunciation is correct, which makes it difficult to receive accurate feedback. [Means for solving the problem]
[0005] The present invention provides a system that integrates the following: means for acquiring English language learning material data; means for generating audio data based on the acquired English language learning material data; means for generating visual content corresponding to the generated audio data; means for associating and saving the audio data and visual content; means for converting the audio data into katakana characters; means for analyzing the audio data and generating pitch bars; means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal; means for recording a learner's pronunciation; means for transmitting the pronunciation data to a server; means for analyzing and evaluating the pronunciation data; and means for displaying pronunciation evaluation results. This system allows learners to intuitively understand English through sounds and pictures and practice correct pronunciation using katakana transcriptions and pitch bars. Furthermore, by providing feedback based on the pronunciation evaluation results, learners can receive assistance in improving their pronunciation.
[0006] "English language learning material data" refers to data on learning materials for the purpose of learning English, and includes sentences, vocabulary, conversation examples, test questions, and the like.
[0007] "Audio data" refers to audio information corresponding to English language teaching material data, and includes generated audio and recorded audio.
[0008] "Visual content" refers to visual content that corresponds to audio data, and includes still images, illustrations, animated videos, and the like.
[0009] A "katakana character string" is a character string that expresses English pronunciation in katakana and is composed of Japanese katakana characters.
[0010] The "pitch bar" is a bar that visually indicates the pitch and timing of each syllable in the audio data, and is used on a karaoke scoring screen, for example.
[0011] "Device" refers to an electronic device used by a learner, such as a computer, smartphone, or tablet, that has the ability to display audio data and visual content.
[0012] The "pronunciation evaluation results" are obtained by analyzing the speech data produced by the learner and showing the accuracy and quality of the pronunciation as numerical values and feedback.
[0013] A "server" is a central processing unit that acquires English language teaching material data, generates audio data and visual content, and stores and analyzes data, and refers to a computer on a network that exchanges data with terminals.
[0014] "Feedback" refers to evaluations and advice provided to learners based on the results of pronunciation assessment, and includes information that helps learners improve their own pronunciation. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The system for implementing the present invention mainly consists of a server and a terminal. The server is responsible for multiple functions, including acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and pronunciation evaluation. Meanwhile, the terminal provides an interface for learners and is responsible for displaying audio data, visual content, katakana strings, and pitch bars, as well as recording and transmitting audio.
[0037] Main server processing
[0038] 1. Acquiring English teaching material data
[0039] The server retrieves English learning material data from the database, including sentences, vocabulary, conversation examples, test questions, etc.
[0040] 2. Generating audio data
[0041] The server uses a generative AI to generate voice data based on the acquired English learning material data, and uses a speech synthesis engine to create natural pronunciation and intonation.
[0042] 3. Visual content generation
[0043] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[0044] 4. Katakana String Generation
[0045] The server converts the audio data into katakana characters, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[0046] 5. Generating interval bars
[0047] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0048] 6. Data transmission
[0049] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0050] Main processing of the device
[0051] 1. Receiving and integrating data
[0052] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the learner.
[0053] 2. Display the learning interface
[0054] The device displays an interface that allows learners to view visual content, katakana strings, and pitch bars while listening to the audio.
[0055] 3. Recording pronunciation practice
[0056] The device provides a function to record pronunciation when the learner pronounces along with the audio data.
[0057] 4. Sending pronunciation data
[0058] The device sends the recorded pronunciation data to the server.
[0059] Pronunciation assessment and feedback
[0060] 1. Analysis of pronunciation evaluation
[0061] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0062] 2. Generating evaluation results
[0063] The server uses a pronunciation analysis tool to generate an evaluation result and transmits the result to the terminal.
[0064] 3. Providing Feedback
[0065] The device displays the evaluation results received from the server to the learner and provides detailed feedback on which areas need improvement.
[0066] Specific examples
[0067] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[0068] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0069] The server generates an animated video of the scene "a man opening a door."
[0070] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[0071] This data is sent to the terminal, which then integrates it and displays it to the learner.
[0072] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[0073] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[0074] The device then displays this feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[0075] This allows learners to learn English interactively and intuitively, encouraging continuous learning.
[0076] The processing flow will be explained below.
[0077] Understood. The process flow is explained in detail below.
[0078] Processing on the server
[0079] Step 1:
[0080] The server retrieves English learning material data from the database, such as TOEIC listening questions.
[0081] Step 2:
[0082] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with correct pronunciation and intonation.
[0083] Step 3:
[0084] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, using generative AI and multimedia tools.
[0085] Step 4:
[0086] The server converts the audio data into katakana strings, using a speech analysis tool to accurately represent the English pronunciation in katakana.
[0087] Step 5:
[0088] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0089] Step 6:
[0090] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0091] Processing on the device
[0092] Step 7:
[0093] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[0094] Step 8:
[0095] The device displays the integrated data to the learner and provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[0096] Step 9:
[0097] As learners practice pronunciation following the audio guide, the device records their pronunciation.
[0098] Step 10:
[0099] The device sends the recorded pronunciation data to the server.
[0100] Processing on the server (continued)
[0101] Step 11:
[0102] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0103] Step 12:
[0104] The server uses a pronunciation analysis tool to generate a pronunciation evaluation result and transmits the result to the terminal.
[0105] Processing on the terminal (continued)
[0106] Step 13:
[0107] The device receives pronunciation evaluation results from the server and displays them to the learner, providing detailed feedback on areas that need improvement based on the evaluation results.
[0108] This allows learners to improve their English listening and speaking skills in an interactive and intuitive way.
[0109] Example 1
[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0111] A major challenge in modern English learning is the lack of interactive and effective learning materials. In particular, the lack of speech recognition and feedback makes it difficult for learners to self-evaluate the accuracy of their pronunciation. There is also a lack of visually appealing content to keep learners engaged.
[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0113] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal, means for recording a learner's pronunciation, means for transmitting the recorded pronunciation data to the server, means for analyzing the pronunciation data and evaluating pronunciation, means for displaying the pronunciation evaluation results, means for integrating data received from the server and displaying a learning interface, and means for allowing a learner to view the visual content, katakana characters, and pitch bars while listening to audio within the learning interface. This allows learners to study English in an interactive and visually appealing environment, have their pronunciation accuracy evaluated, and receive effective feedback.
[0114] "English language learning material data" refers to information such as sentences, vocabulary, conversation examples, and test questions that learners use to study English.
[0115] "Audio data" refers to audio files with natural pronunciation and intonation that are generated based on the acquired English teaching material data.
[0116] "Visual content" refers to content that learners can refer to visually, such as illustrations or animated videos that correspond to audio data.
[0117] A "katakana string" is a representation of English pronunciation analyzed from audio data written in Japanese katakana characters.
[0118] The "pitch bar" is a karaoke-style bar that visually displays the pitch and timing of each syllable in the audio data.
[0119] A "terminal" is a device used by a learner (e.g., a PC, tablet, smartphone, etc.).
[0120] The "server" is a computer system responsible for acquiring English language teaching material data, generating audio data and visual content, analyzing katakana character strings, generating pitch bars, and managing and transmitting this data.
[0121] The "learning interface" is a user interface that allows learners to view visual content, katakana character strings, and pitch bars while listening to audio on their device.
[0122] "Pronunciation data" refers to audio data recorded by learners during pronunciation practice.
[0123] "Pronunciation evaluation" is the process of analyzing recorded pronunciation data and evaluating its accuracy.
[0124] "Feedback" is information that indicates to the learner which parts have been pronounced correctly and which parts need improvement, based on the pronunciation evaluation results.
[0125] The system for implementing the present invention is composed of a server and a terminal. The specific hardware and software and how they process data or perform calculations will be described below.
[0126] Main server processing
[0127] 1. Acquiring English teaching material data
[0128] When the server receives the request, it connects to the database and retrieves the English learning material data. The database used is a relational database management system (RDBMS) such as MySQL (registered trademark). For example, an SQL query such as SELECT FROM learning material WHERE ID = ? is used.
[0129] 2. Generating audio data
[0130] The server generates audio data using a speech synthesis engine based on the acquired English learning material data. The speech synthesis engine used is the Google (registered trademark) Cloud Text-to-Speech API. The audio data is generated in MP3 or WAV format with natural pronunciation and intonation.
[0131] 3. Visual content generation
[0132] The server uses the generative AI model to generate visual content (such as illustrations and animated videos) corresponding to the English learning material data. For example, the generative AI model DALL-E generates an illustration based on the prompt "The man is opening a door," and creates an animated video using Adobe After Effects.
[0133] 4. Katakana String Generation
[0134] The server uses a speech analysis tool to convert the voice data into a katakana string, for example, Google Cloud Speech-to-Text, which converts the voice into text and then converts the resulting text according to the katakana character conversion rules.
[0135] 5. Generating interval bars
[0136] The server analyzes the audio data to identify the pitch and timing of each syllable. This analysis is performed using the audio analysis tool Praat, which extracts pitch information and then converts it into pitch bars using the visual tool FFmpeg.
[0137] 6. Data transmission
[0138] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal as an HTTP response.
[0139] Main processing of the device
[0140] 1. Receiving and integrating data
[0141] The device receives the audio data, visual content, katakana characters, and pitch bars sent from the server, and the received data is integrated and prepared for display to the learner.
[0142] 2. Display the learning interface
[0143] The device uses HTML5 and JavaScript (registered trademark) to display visual content, katakana characters, and pitch bars while the learner is playing the audio. The interface is designed to be intuitive for learners to operate.
[0144] 3. Recording pronunciation practice
[0145] The device provides the ability to record pronunciation as the learner speaks along with the audio data, using standard recording features such as the browser's MediaRecorder API.
[0146] 4. Sending pronunciation data
[0147] The device sends the recorded pronunciation data to the server as an HTTP POST request, which the server then processes and analyzes.
[0148] Pronunciation assessment and feedback
[0149] 1. Analysis of pronunciation evaluation
[0150] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, using voice analysis tools such as Google Cloud Speech-to-Text.
[0151] 2. Generating evaluation results
[0152] The server uses an evaluation tool to evaluate the accuracy of the pronunciation and generate feedback information, including which parts were pronounced correctly and which parts need improvement.
[0153] 3. Providing Feedback
[0154] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and charts.
[0155] Specific examples
[0156] For example, the processing flow based on the TOEIC listening question "The man is opening a door." is as follows:
[0157] 1. The server retrieves the English teaching material data "The man is opening a door." from the database.
[0158] 2. The server uses Google Cloud Text-to-Speech to generate audio data for "The man is opening a door."
[0159] 3. The server uses DALL-E to generate an illustration of the scene of a man opening a door, and then creates an animated video using Adobe After Effects.
[0160] 4. The server uses Google Cloud Speech-to-Text to convert the audio into the katakana string "The man is opening a door."
[0161] 5. The server uses Praat to analyze pitch and timing, and creates karaoke-style pitch bars with FFmpeg.
[0162] 6. The server sends this data to the terminal.
[0163] 7. The device integrates the received data and displays it to the learner.
[0164] 8. When the user says "The Man is Opening a Door," the audio is recorded and sent to the server.
[0165] 9. The server analyzes the voice data and generates an evaluation result (e.g., 90 points).
[0166] 10. The device displays the assessment results to the learner, showing which parts of their pronunciation are correct and which parts need improvement.
[0167] Prompt Sentence Examples
[0168] The prompt sentence for generating learning content based on the English sentence "The boy is eating an apple" is as follows:
[0169] Generate learning content based on the following English sentence: "The boy is eating an apple." Generate the following content:
[0170] 1. Audio data
[0171] 2. Katakana string
[0172] 3. Karaoke-style pitch bar
[0173] 4. Animated video of the scene
[0174] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0175] Step 1:
[0176] The server receives the request and connects to the database to retrieve English learning material data. This data includes sentences, vocabulary, conversation examples, test questions, etc. Specifically, it executes an SQL query (e.g., SELECT FROM learning material WHERE ID = ?). The input is the request and learning material ID, and the output is the corresponding English learning material data.
[0177] Step 2:
[0178] The server generates audio data based on the acquired English language learning material data. Using the Google Cloud Text-to-Speech API, it sends a text request and receives audio data. The input is the English language learning material data, and the output is the generated audio data (MP3 or WAV format).
[0179] Step 3:
[0180] The server generates visual content using a generative AI model. DALL-E is used to generate an illustration based on the prompt "The man is opening a door," and an animation video is created using Adobe After Effects. The input is the prompt and the AI model, and the output is the visual content (illustration and animation video).
[0181] Step 4:
[0182] The server uses a speech analysis tool to convert the voice data into katakana strings. Google Cloud Speech-to-Text is used to convert the voice data into text, which is then further converted according to katakana conversion rules. The input is voice data, and the output is katakana strings.
[0183] Step 5:
[0184] The server analyzes the audio data and generates pitch bars. It uses Praat to analyze the pitch and timing of the audio, and then converts it into visual pitch bars using FFmpeg. The input is the audio data, and the output is pitch bars.
[0185] Step 6:
[0186] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. It packages these data as an HTTP response and delivers it to the terminal. The input is the set of generated data, and the output is the transmission packet.
[0187] Step 7:
[0188] The device analyzes and integrates the data received from the server and displays it on the learning interface. This uses HTML5 and JavaScript to provide a user interface that learners can operate intuitively. The input is the data received from the server, and the output is the screen display of the learning interface.
[0189] Step 8:
[0190] When the user speaks in sync with the audio data, the device records the pronunciation. It uses the browser's MediaRecorder API to capture microphone input. The input is audio data, and the output is the recorded data.
[0191] Step 9:
[0192] The device sends the recorded pronunciation data to the server as an HTTP POST request. The server receives the pronunciation data and proceeds to the next processing step. The input is the recorded data, and the output is an HTTP request.
[0193] Step 10:
[0194] The server analyzes the received pronunciation data and performs pronunciation evaluation. It extracts speech features using Google Cloud Speech-to-Text and executes the evaluation algorithm. The input is the pronunciation data, and the output is the pronunciation evaluation result.
[0195] Step 11:
[0196] The server generates an evaluation result and feedback information, which includes the correct pronunciation and areas that need improvement. The input is the output of the evaluation algorithm, and the output is the detailed evaluation result and feedback information.
[0197] Step 12:
[0198] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and diagrams. The input is the evaluation results, and the output is the feedback displayed on the learning interface.
[0199] (Application example 1)
[0200] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0201] Conventional English learning systems and product description systems often do not provide sufficient support for learners and customers when practicing pronunciation or receiving feedback. In particular, in physical stores, there is a lack of interactive systems that allow non-native speakers to obtain accurate information and practice pronunciation appropriately when purchasing products. This results in reduced learning and purchasing efficiency for learners and customers, and reduced satisfaction in physical stores.
[0202] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0203] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for acquiring product information, means for generating English audio guidance based on the acquired product information, means for generating visual content corresponding to the generated audio guidance, means for scanning QR codes (registered trademark), means for providing guidance information based on the product information scanned from the QR code, and means for evaluating the pronunciation when reading product descriptions. This enables non-native speakers to efficiently acquire product information in physical stores and learn accurate pronunciation through pronunciation practice.
[0204] "English language learning material data" refers to data that includes information such as sentences, vocabulary, conversation examples, and test questions for the purpose of learning English.
[0205] "Audio data" refers to data that includes information on voices synthesized based on English language teaching material data and guidance information.
[0206] "Visual content" refers to information that includes visual content such as illustrations and animated videos that correspond to the audio data.
[0207] A "katakana character string" is a character string that expresses the pronunciation of audio data in katakana.
[0208] A "pitch bar" is a visual bar that indicates pitch and timing based on audio data.
[0209] "Terminal" refers to a device that a user uses as an interface, including smartphones, tablets, etc.
[0210] "Pronunciation data" refers to data containing voice information generated when a user records their voice.
[0211] A "server" is a computer system that processes and stores various data online.
[0212] "Product information" is data that includes detailed product information, such as product descriptions, usage instructions, and characteristics.
[0213] "Voice guidance" is data that includes voice explanations generated based on product information.
[0214] A "QR code" is a two-dimensional barcode used to encode data, including product information.
[0215] "Guide information" refers to information including detailed product descriptions and usage instructions that can be provided by scanning the QR code.
[0216] "Reading aloud" is the act of reading aloud a specified text.
[0217] The system for implementing this invention comprises a server and a terminal. The server is responsible for various data processing and generation, and the terminal provides an interface to the user. Specific embodiments will be described below.
[0218] Server-side processing
[0219] 1. Acquiring English teaching material data
[0220] The server retrieves English learning material data from the database, which includes various learning materials such as sentences, vocabulary, conversation examples, and test questions.
[0221] 2. Generating audio data
[0222] The server uses a speech synthesis engine to generate voice data with natural pronunciation and intonation based on the acquired English learning material data, and uses a generative AI model to achieve more natural pronunciation.
[0223] 3. Visual Content Generation
[0224] The server generates visual content (such as illustrations or animated videos) corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[0225] 4. Data Association and Storage
[0226] The generated audio data, visual content, katakana character strings, and pitch bars are associated and saved.
[0227] 5. Katakana String Generation
[0228] The server converts the audio data into katakana strings, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[0229] 6. Generating interval bars
[0230] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0231] 7. Data transmission
[0232] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0233] 8. Analysis and Evaluation of Pronunciation Data
[0234] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0235] Terminal side processing
[0236] 1. Receiving and integrating data
[0237] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the user.
[0238] 2. Display the learning interface
[0239] The device displays an interface that allows the user to view visual content, katakana characters, and pitch bars while playing the audio.
[0240] 3. Recording pronunciation practice
[0241] The terminal provides a function for recording pronunciation when the user pronounces along with the audio data.
[0242] 4. Sending pronunciation data
[0243] The device sends the recorded pronunciation data to the server.
[0244] 5. Displaying pronunciation evaluation results
[0245] The device displays the evaluation results received from the server to the user, providing detailed feedback on which areas need improvement.
[0246] Specific examples
[0247] For example, when a user scans a QR code printed with "12345" in a physical store, the product description is played in English. The user repeats the description and records their own pronunciation, which is then sent to the server. The server analyzes the user's pronunciation and provides evaluation feedback, which is then displayed to the user on the device.
[0248] Hardware and software used
[0249] Hardware
[0250] Server (high performance computer)
[0251] Device (smartphone, tablet, etc.)
[0252] software
[0253] Database Management Systems
[0254] Speech synthesis engine
[0255] Generative AI Models
[0256] Audio analysis tools
[0257] Prompt Sentence Examples
[0258] "You scan the QR code to hear a detailed product description in English, then record yourself repeating the description and receive feedback in the app."
[0259] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0260] Step 1:
[0261] Acquisition of English teaching material data
[0262] The server retrieves English learning material data from the database. At this time, it receives a request containing information about the learning material selected by the user, and retrieves the target learning material from the database based on that request. The retrieved data includes sentences, vocabulary, conversation examples, test questions, etc.
[0263] Input: User's request to select learning materials
[0264] Output: Acquired English teaching material data
[0265] Step 2:
[0266] Generate audio data
[0267] The server generates voice data using a generative AI model and a speech synthesis engine based on the acquired English learning material data. Specifically, the learning material text is sent as input to the speech synthesis engine, which generates voice data with natural pronunciation and intonation.
[0268] Input: English teaching material data
[0269] Output: Generated audio data
[0270] Step 3:
[0271] Visual content generation
[0272] The server generates visual content (illustrations and animated videos) that correspond to the generated audio data, using generative AI and multimedia tools to create visual materials that match the audio data.
[0273] Input: Audio data
[0274] Output: The generated visual content
[0275] Step 4:
[0276] Data association and storage
[0277] The server associates the generated audio data, visual content, katakana character strings, and pitch bars and stores them in a unified manner. At this time, the association information for each data is also saved.
[0278] Input: Audio data, visual content, katakana strings, pitch bars
[0279] Output: A set of associated data
[0280] Step 5:
[0281] Katakana string generation
[0282] The server converts the voice data into a katakana string. Using a voice analysis tool, the pronunciation of the voice data is analyzed and converted into katakana notation.
[0283] Input: Audio data
[0284] Output: Katakana string
[0285] Step 6:
[0286] Generating interval bars
[0287] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates pitch bars using a voice analysis tool.
[0288] Input: Audio data
[0289] Output: Generated interval bars
[0290] Step 7:
[0291] Sending data
[0292] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device using protocols such as HTTP or WebSocket.
[0293] Input: Audio data, visual content, katakana strings, pitch bars
[0294] Output: Data sent to the terminal
[0295] Step 8:
[0296] Receiving and integrating data
[0297] The device receives and integrates the data sent from the server, providing a viewable interface for the user.
[0298] Input: Send data
[0299] Output: Unified learning interface
[0300] Step 9:
[0301] View the learning interface
[0302] The device displays an interface that allows users to view visual content, katakana characters, and pitch bars while listening to the audio, allowing them to begin learning.
[0303] Input: Integrated dataset
[0304] Output: The displayed learning interface
[0305] Step 10:
[0306] Pronunciation practice recording
[0307] Users can use the device's recording function to practice pronunciation along with the audio data, and the recorded pronunciation data is saved on the device.
[0308] Input: User pronunciation
[0309] Output: Recorded pronunciation data
[0310] Step 11:
[0311] Sending pronunciation data
[0312] The device sends the recorded pronunciation data to the server using protocols such as HTTP or WebSocket.
[0313] Input: Recorded pronunciation data
[0314] Output: Data sent to the server
[0315] Step 12:
[0316] Analysis and evaluation of pronunciation data
[0317] The server analyzes the received pronunciation data, evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, and calculates an evaluation score using a pronunciation evaluation algorithm.
[0318] Input: Received pronunciation data
[0319] Output: Evaluation score
[0320] Step 13:
[0321] Displaying pronunciation evaluation results
[0322] The terminal displays the pronunciation evaluation results sent from the server to the user, allowing the user to check the accuracy of their own pronunciation and understand which parts need improvement.
[0323] Input: Rating score
[0324] Output: Display of pronunciation evaluation results
[0325] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0326] The system for implementing the present invention consists of three main components: a server, a terminal, and an emotion engine. The server acquires English learning material data, generates audio data, generates visual content, generates katakana strings, generates pitch bars, and evaluates pronunciation. The terminal functions as a learner interface and is responsible for audio playback, visual content display, pronunciation recording, and data transmission. The emotion engine recognizes the user's emotions and adjusts learning content and feedback based on them.
[0327] Processing on the server
[0328] 1. Acquiring English teaching material data
[0329] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[0330] 2. Generating audio data
[0331] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create audio files with natural pronunciation and intonation.
[0332] 3. Visual content generation
[0333] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create content that is visually easy to understand.
[0334] 4. Katakana String Generation
[0335] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[0336] 5. Generating interval bars
[0337] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0338] 6. Data transmission
[0339] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0340] Processing on the device
[0341] 1. Receiving and integrating data
[0342] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[0343] 2. Display the learning interface
[0344] The terminal displays the integrated data to the learner, providing an interface that allows them to refer to visual content, katakana characters, and pitch bars while playing the audio.
[0345] 3. Use of Emotion Engine
[0346] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, for example, by analyzing the user's emotions in real time using a camera or microphone.
[0347] 4. Recording pronunciation practice
[0348] The device records the learner's pronunciation as they practice it following the audio guide.
[0349] 5. Sending pronunciation data
[0350] The device sends the recorded pronunciation data to the server.
[0351] Processing on the server (continued)
[0352] 6. Analysis of pronunciation evaluation
[0353] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0354] 7. Generating evaluation results
[0355] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[0356] 8. Generate feedback
[0357] The server takes into account the emotional data from the emotion engine and generates optimal feedback content and tone.
[0358] Processing on the terminal (continued)
[0359] 9. Viewing assessment results and providing feedback
[0360] The device receives the evaluation results from the server and displays them to the learner. Based on the evaluation results, detailed feedback is provided on areas that need improvement. The content and tone of the feedback is adjusted appropriately based on emotion recognition data from the emotion engine.
[0361] Specific examples
[0362] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[0363] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0364] The server generates an animated video of the scene "a man opening a door."
[0365] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[0366] This data is sent to the terminal, which then integrates it and displays it to the learner.
[0367] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[0368] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[0369] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[0370] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[0371] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[0372] This allows users to self-evaluate their learning and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[0373] The processing flow will be explained below.
[0374] Understood. The process flow is explained in detail below.
[0375] Processing on the server
[0376] Step 1:
[0377] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[0378] Step 2:
[0379] The server uses a generative AI to generate voice data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with natural pronunciation and intonation.
[0380] Step 3:
[0381] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, created using generative AI and multimedia tools.
[0382] Step 4:
[0383] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[0384] Step 5:
[0385] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0386] Step 6:
[0387] The server transmits the generated audio data, visual content, katakana character strings, and pitch bars to the terminal.
[0388] Processing on the device
[0389] Step 7:
[0390] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[0391] Step 8:
[0392] The terminal displays the integrated data to the learner and serves as a learning interface, allowing the learner to view audio playback, visual content, katakana strings, and pitch bars.
[0393] Step 9:
[0394] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, and captures emotion data in real time using a camera and microphone.
[0395] Step 10:
[0396] As learners practice pronunciation following the audio guide, their pronunciation is recorded.
[0397] Step 11:
[0398] The device transmits the recorded pronunciation data and emotion data to the server.
[0399] Processing on the server (continued)
[0400] Step 12:
[0401] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0402] Step 13:
[0403] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[0404] Step 14:
[0405] The server takes into account the emotional data and generates feedback content and tone based on the learner's emotions.
[0406] Processing on the terminal (continued)
[0407] Step 15:
[0408] The device receives the evaluation results and feedback from the server and displays them to the learner. Based on the evaluation results and emotion-based feedback, the learner is shown which areas need improvement.
[0409] Specific examples
[0410] For example, in the case of the TOEIC listening question "The man is opening a door," the process proceeds as follows:
[0411] Step 1:
[0412] The server retrieves the relevant statement from the database.
[0413] Step 2:
[0414] The server uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0415] Step 3:
[0416] The server generates an animated video of the scene "a man opening a door."
[0417] Step 4:
[0418] The server converts the speech into the katakana string "The man is opening a door."
[0419] Step 5:
[0420] The server analyzes the audio data, identifies the pitch and timing of each syllable, and creates pitch bars.
[0421] Step 6:
[0422] These data are transmitted to the terminal.
[0423] Step 7:
[0424] The terminal consolidates the received data and displays it to the learner.
[0425] Step 8:
[0426] It provides an interface that simultaneously displays audio playback, visual content, katakana characters, and pitch bars.
[0427] Step 9:
[0428] The device uses an emotion engine to acquire emotional data from the learner's facial expressions and voice, and analyzes emotions in real time using a camera and microphone.
[0429] Step 10:
[0430] When the learner says "The man is opening a door," the device records it.
[0431] Step 11:
[0432] The device sends the recorded data and emotion data to the server.
[0433] Step 12:
[0434] The server analyzes the recording and evaluates the accuracy of the pronunciation.
[0435] Step 13:
[0436] The server generates an evaluation result and transmits it to the terminal.
[0437] Step 14:
[0438] The server adjusts the content and tone of the feedback based on the emotional data, generates the feedback, and sends it to the device.
[0439] Step 15:
[0440] The device displays the assessment results and feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[0441] This allows learners to efficiently improve their English listening and speaking skills while receiving flexible feedback that reflects their emotions.
[0442] Example 2
[0443] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0444] In conventional English learning systems, evaluation and feedback of learners' pronunciation is mechanical, making it difficult to respond flexibly to the learner's emotions and learning progress. As a result, there are problems such as a decrease in learner motivation and an impediment to effective learning. The purpose of this invention is to improve learning efficiency and motivation by recognizing learners' emotions and providing feedback accordingly.
[0445] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0446] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for recognizing the learner's emotions, and means for adjusting the content and tone of the feedback based on the emotions. This makes it possible to provide flexible and effective feedback based on the learner's emotions and the accuracy of their pronunciation.
[0447] "English language learning data" is a general term for various data used in English learning, such as English listening questions, vocabulary lists, and example sentences.
[0448] "Audio data" refers to audio files with natural pronunciation and intonation based on English language learning material data generated by the server.
[0449] "Visual content" refers to content that is easy to understand visually, such as illustrations or animated videos that correspond to audio data.
[0450] A "katakana character string" is a character string in which audio data is written in Japanese katakana.
[0451] "Pitch bar" refers to a bar that visually displays the pitch and timing of each syllable in the analyzed audio data.
[0452] "Terminal" refers to a device used by a learner that plays audio, displays visual content, records pronunciation, transmits data, etc.
[0453] "Learner" refers to a user who uses this system to learn English.
[0454] An "emotion engine" is a software component that recognizes a user's emotions and adjusts learning content and feedback based on those emotions.
[0455] "Pronunciation data" refers to audio data recorded when a learner practices pronunciation.
[0456] The "pronunciation evaluation result" indicates an evaluation of the accuracy of pronunciation obtained by the server by analyzing the pronunciation data.
[0457] "Feedback" refers to information or instructions provided to learners about their learning progress and areas for improvement.
[0458] The system for implementing this invention consists of three main components: a server, a terminal, and an emotion engine. The server is responsible for acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and evaluating pronunciation. The terminal functions as a learner interface, playing audio, displaying visual content, recording pronunciation, and transmitting data. The emotion engine also recognizes the user's emotions and adjusts learning content and feedback based on those emotions.
[0459] Processing on the server
[0460] 1. Acquiring English teaching material data
[0461] The server queries and retrieves English learning material data from a database (e.g., PostgreSQL), including TOEIC listening questions, vocabulary lists, and example sentences.
[0462] 2. Generating audio data
[0463] Based on the acquired English learning material data, the server uses a speech synthesis engine (e.g., Google Cloud Text-to-Speech API) to generate voice data with natural pronunciation and intonation.
[0464] 3. Visual content generation
[0465] The server generates visual content (e.g., illustrations or animated videos) corresponding to the generated audio data using Adobe Animate or image generation technology. For example, it generates an animated video corresponding to the sentence "The man is opening a door."
[0466] 4. Katakana String Generation
[0467] The server uses a speech analysis tool (e.g., the Julius speech engine) to convert the speech data into a katakana string, resulting in the katakana transcription "The Man is Opening a Door."
[0468] 5. Generating interval bars
[0469] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars using tools such as praat.
[0470] 6. Data transmission
[0471] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device as a JSON-formatted package.
[0472] Processing on the device
[0473] 1. Receiving and integrating data
[0474] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[0475] 2. Display the learning interface
[0476] The device uses the integrated data to display to the learner, playing audio and displaying an interface that provides visual content, katakana characters, and pitch bars.
[0477] 3. Use of Emotion Engine
[0478] The device uses an emotion engine (for example, Microsoft® Azure® Emotion API) to recognize emotions from the learner's facial expressions and voice. It also analyzes the user's emotions in real time using a camera and microphone.
[0479] 4. Recording pronunciation practice
[0480] The device records the learner's pronunciation as they practice it following the audio guide.
[0481] 5. Sending pronunciation data
[0482] The device sends the recorded pronunciation data to the server via a REST API.
[0483] Processing on the server (continued)
[0484] 1. Analysis of pronunciation evaluation
[0485] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The analysis is performed using the Julius speech engine and praat.
[0486] 2. Generating evaluation results
[0487] The server converts the evaluation results obtained from the pronunciation analysis into a score and sends it to the terminal, generating feedback such as "90 points."
[0488] 3. Generate feedback
[0489] The server takes into account the emotional data and generates optimal feedback content and tone.
[0490] Processing on the terminal (continued)
[0491] 1. Viewing assessment results and providing feedback
[0492] The device displays the evaluation results received from the server to the learner, provides detailed feedback on areas that need improvement based on the evaluation results, and adjusts the tone of the feedback based on emotional data.
[0493] Specific examples
[0494] For example, in the case of the TOEIC listening question "The man is opening a door," the processing proceeds as follows:
[0495] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0496] The server generates an animated video of the scene "a man opening a door."
[0497] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[0498] This data is sent to the terminal, which then integrates it and displays it to the learner.
[0499] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[0500] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback, such as a score of 90.
[0501] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[0502] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[0503] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[0504] Example prompt: "Based on the TOEIC listening question 'The man is opening a door,' please generate natural-sounding audio data, visual content, katakana characters, and pitch bars and send them to the device."
[0505] This allows users to self-evaluate and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[0506] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0507] Step 1:
[0508] Acquisition of English teaching material data
[0509] The server retrieves English language learning material data from the database. The input is the query information stored on the server, and the output is the retrieved English language learning material data. Specifically, the server executes an SQL query to retrieve data such as TOEIC listening questions, vocabulary lists, and example sentences.
[0510] Step 2:
[0511] Generate audio data
[0512] The server uses a speech synthesis engine to generate voice data based on the acquired English language learning material data. The input is the acquired English language learning material data, and the output is the generated voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to generate natural-sounding voice from text data such as "The man is opening a door."
[0513] Step 3:
[0514] Visual content generation
[0515] The server generates visual content corresponding to the audio data. The input is the generated audio data, and the output is the generated visual content. Specifically, the server uses Adobe Animate to create an animated video of the scene "a man opening a door."
[0516] Step 4:
[0517] Katakana string generation
[0518] The server converts the speech data into a katakana string. The input is the generated speech data, and the output is a katakana string. Specifically, the server uses the Julius speech engine to convert the speech data into the katakana notation "The man is opening a door."
[0519] Step 5:
[0520] Generating interval bars
[0521] The server analyzes the audio data and generates pitch bars. The input is the generated audio data, and the output is pitch bars. Specifically, the server uses praat to analyze the pitch and timing of each syllable and generate karaoke-style pitch bars.
[0522] Step 6:
[0523] Sending data
[0524] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device. The input is all the generated data, and the output is a data package. Specifically, the server compiles this data into JSON format and sends it to the device using the HTTP protocol.
[0525] Step 7:
[0526] Receiving and integrating data
[0527] The device receives and integrates the audio data, visual content, katakana text, and pitch bars sent from the server. The input is the data package sent from the server, and the output is the integrated learning content. Specifically, the device parses the JSON-formatted data and displays the audio data and visual content on a single screen.
[0528] Step 8:
[0529] View the learning interface
[0530] The device uses the integrated data to provide an interface to be displayed to the learner. The input is the integrated learning content, and the output is the displayed learning interface. Specifically, the device plays audio while synchronously drawing an interface on the screen that displays visual content, katakana characters, and pitch bars.
[0531] Step 9:
[0532] Using the Emotion Engine
[0533] The device recognizes emotions from the learner's facial expressions and voice. The input is the learner's facial and voice data, and the output is emotional data. Specifically, the device uses Microsoft Azure's Emotion API to analyze the learner's emotions in real time from facial expressions captured by the camera and voice data collected by the microphone.
[0534] Step 10:
[0535] Pronunciation practice recording
[0536] The device records the learner's pronunciation as they practice pronunciation following the audio guide. The input is the learner's pronunciation, and the output is the recorded pronunciation data. Specifically, the device uses a built-in microphone to record the learner's pronunciation into an audio file.
[0537] Step 11:
[0538] Sending pronunciation data
[0539] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the pronunciation data sent to the server. Specifically, the device uses a REST API to upload the audio file to the server.
[0540] Step 12:
[0541] Pronunciation evaluation analysis
[0542] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the recorded pronunciation data, the katakana string, and the pitch bar, and the output is the pronunciation evaluation result. Specifically, the server uses the Julius speech engine and praat to calculate the degree of match for each phoneme and convert it into a score.
[0543] Step 13:
[0544] Generating evaluation results
[0545] The server generates pronunciation evaluation results and sends them to the terminal. The input is the pronunciation evaluation results, and the output is the evaluation results sent from the server to the terminal. Specifically, the server summarizes the evaluation results and sends them to the terminal using the HTTP protocol.
[0546] Step 14:
[0547] Viewing assessment results and providing feedback
[0548] The device receives the evaluation results from the server, displays them to the learner, and provides feedback. The input is the evaluation results and emotional data, and the output is the displayed feedback. Specifically, the device highlights areas that need improvement based on the evaluation results and provides flexible feedback based on the emotional data.
[0549] (Application example 2)
[0550] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0551] Conventional English learning systems are unable to provide feedback that reflects learners' emotions, limiting their ability to support effective learning. Furthermore, the feedback provided for English pronunciation evaluation and correction is fixed, making it difficult to respond individually to learners' emotions and reactions. As a result, issues such as a decline in learning motivation and a delay in learning progress have arisen.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting to katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to the terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for acquiring emotion data based on the evaluation results, and means for displaying the pronunciation evaluation results and the emotion data. This enables flexible and individual feedback based on the learner's emotions.
[0553] "English language learning material data" refers to data on learning materials used in learning English, and includes vocabulary lists, sample sentences, listening questions, and the like.
[0554] "Speech data" refers to acoustic information generated based on acquired English language learning material data, and includes English pronunciation and intonation.
[0555] "Visual content" includes visual information such as illustrations and animated videos generated in response to audio data.
[0556] A "katakana character string" is a character string obtained by converting the generated voice data into Japanese katakana notation.
[0557] A "pitch bar" is a bar that visually represents the pitch and timing of each syllable by analyzing audio data.
[0558] A "terminal" is a device that receives the generated audio data, visual content, katakana character strings, and pitch bars, and provides them to a learner.
[0559] "Pronunciation data" refers to data that is recorded by a learner pronouncing words and sent to the server.
[0560] The "pronunciation evaluation result" is the result of analyzing the recorded pronunciation data and evaluating its accuracy.
[0561] "Emotion data" is information about emotions obtained from the learner's facial expressions, voice, etc.
[0562] "Feedback" refers to guidance and advice provided to learners based on pronunciation evaluation results and emotional data.
[0563] The system for implementing this invention mainly consists of three main components: a server, a terminal, and an emotion engine.
[0564] Server Processing
[0565] The server implements the following functions:
[0566] 1. Acquisition of English learning material data: The server acquires English learning material data from the database. The acquired data includes listening questions, vocabulary lists, sample sentences, etc.
[0567] 2. Audio data generation: The server generates audio data using a speech synthesis engine based on the acquired English learning material data. Using this generative AI model, an audio file with natural pronunciation and intonation is created.
[0568] 3. Visual content generation: The server generates visual content such as illustrations and animated videos corresponding to the audio data. Generative AI and multimedia tools are used to create content that is easy to understand visually.
[0569] 4. Katakana string generation: The server converts the audio data into Katakana strings, using a speech analysis tool to represent the English pronunciation in Japanese Katakana.
[0570] 5. Generating pitch bars: The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0571] 6. Data transmission: The server transmits the generated audio data, visual content, katakana string, and pitch bar to the terminal.
[0572] Terminal handling
[0573] The device implements the following functions:
[0574] 1. Data reception and integration: The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[0575] 2. Display of learning interface: The device displays the integrated data to the learner. It provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[0576] 3. Recording pronunciation practice: The device records the learner's pronunciation as they practice pronunciation following the audio guide.
[0577] 4. Sending pronunciation data: The device sends the recorded pronunciation data to the server.
[0578] Emotion engine processing
[0579] The emotion engine implements the following functions:
[0580] 1. Emotion Recognition: The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice. For example, the device analyzes the user's emotions in real time using a camera or microphone.
[0581] 2. Sending emotion data: The device sends the recognized emotion data to the server.
[0582] Reprocessing on the server
[0583] The server implements the following functions:
[0584] 1. Pronunciation evaluation analysis: The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0585] 2. Generating evaluation results: The server generates evaluation results using the pronunciation analysis tool and sends the evaluation results to the terminal.
[0586] 3. Feedback Generation: The server takes into account the emotional data from the emotion engine and generates the optimal feedback content and tone.
[0587] Reprocessing on the device
[0588] The device implements the following functions:
[0589] 1. Displaying assessment results and providing feedback: The device displays the assessment results received from the server to the learner. Based on the assessment results, detailed feedback is provided on which areas need improvement. Appropriate feedback content and tone are provided based on emotion recognition data from the emotion engine.
[0590] Specific examples
[0591] For example, consider a step in which a user tries on a new jacket in a virtual fitting room. The device recognizes the user's emotions from their facial expressions and sends the data to the server. The server generates feedback based on the emotion data, such as "Why don't you try on a jacket in a lighter color?" and sends it to the device. The device then provides the feedback to the user visually and audibly.
[0592] The following is an example of a prompt sentence to be used in the specific example:
[0593] Analyze an image of a user trying on a new jacket in front of a mirror. From their facial expression, they appear to be somewhat satisfied, but would like further style suggestions. Generate feedback content about the color and design of the jacket.
[0594] This significantly improves the user experience in the virtual fitting room and enables efficient and effective feedback.
[0595] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0596] Step 1:
[0597] The server retrieves English learning material data from the database. This data includes listening questions, vocabulary lists, sample sentences, etc. The input is learning material information in the learning material database, and the output is the retrieved information for each learning material data.
[0598] Step 2:
[0599] The server generates audio data based on the acquired English learning material data, using a generative AI model to create an audio file with natural pronunciation and intonation. The input is the learning material data, and the output is the generated audio data.
[0600] Step 3:
[0601] The server generates visual content corresponding to the generated audio data. Here, generative AI and multimedia tools are used to create visually easy-to-understand content such as illustrations and animated videos. The input is audio data, and the output is visual content.
[0602] Step 4:
[0603] The server converts the generated voice data into a katakana string. The server uses a voice analysis tool to convert the voice data into Japanese katakana. The input is the voice data, and the output is a katakana string.
[0604] Step 5:
[0605] The server analyzes the audio data and generates pitch bars, which are a visual representation of the pitch and timing of each syllable. The input is the audio data, and the output is pitch bars.
[0606] Step 6:
[0607] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. The input is each generated data, and the output is the transmission of that data to the terminal.
[0608] Step 7:
[0609] The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server. The input is the multiple data sent from the server, and the output is the integrated training data.
[0610] Step 8:
[0611] The terminal displays the integrated data to the learner. It provides an interface that allows the learner to refer to visual content, katakana characters, and pitch bars while playing the audio. The input is the integrated data, and the output is the display of the learning interface.
[0612] Step 9:
[0613] The device records the user's pronunciation as they practice their pronunciation following the audio guide. The input is the user's pronunciation, and the output is the recorded pronunciation data.
[0614] Step 10:
[0615] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the completion of transmission to the server.
[0616] Step 11:
[0617] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the pronunciation data and evaluation criteria (katakana string and pitch bar), and the output is the evaluation result.
[0618] Step 12:
[0619] The server generates optimal feedback content and tone based on the pronunciation evaluation results and emotional data, and sends it to the device. At this time, the server inputs the previously defined prompt sentence into the generative AI model to generate feedback. The input is the evaluation results and emotional data, and the output is the feedback content.
[0620] Step 13:
[0621] The terminal displays the evaluation results and feedback received from the server to the learner, indicating which parts were pronounced correctly and which parts need improvement. The input is the evaluation results and feedback content, and the output is the display and feedback provided to the learner.
[0622] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0623] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0624] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0625] [Second embodiment]
[0626] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0627] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0628] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0629] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0630] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0631] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0632] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0633] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0634] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0635] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0636] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0637] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0638] The system for implementing the present invention mainly consists of a server and a terminal. The server is responsible for multiple functions, including acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and pronunciation evaluation. Meanwhile, the terminal provides an interface for learners and is responsible for displaying audio data, visual content, katakana strings, and pitch bars, as well as recording and transmitting audio.
[0639] Main server processing
[0640] 1. Acquiring English teaching material data
[0641] The server retrieves English learning material data from the database, including sentences, vocabulary, conversation examples, test questions, etc.
[0642] 2. Generating audio data
[0643] The server uses a generative AI to generate voice data based on the acquired English learning material data, and uses a speech synthesis engine to create natural pronunciation and intonation.
[0644] 3. Visual content generation
[0645] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[0646] 4. Katakana String Generation
[0647] The server converts the audio data into katakana characters, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[0648] 5. Generating interval bars
[0649] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0650] 6. Data transmission
[0651] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0652] Main processing of the device
[0653] 1. Receiving and integrating data
[0654] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the learner.
[0655] 2. Display the learning interface
[0656] The device displays an interface that allows learners to view visual content, katakana strings, and pitch bars while listening to the audio.
[0657] 3. Recording pronunciation practice
[0658] The device provides a function to record pronunciation when the learner pronounces along with the audio data.
[0659] 4. Sending pronunciation data
[0660] The device sends the recorded pronunciation data to the server.
[0661] Pronunciation assessment and feedback
[0662] 1. Analysis of pronunciation evaluation
[0663] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0664] 2. Generating evaluation results
[0665] The server uses a pronunciation analysis tool to generate an evaluation result and transmits the result to the terminal.
[0666] 3. Providing Feedback
[0667] The device displays the evaluation results received from the server to the learner and provides detailed feedback on which areas need improvement.
[0668] Specific examples
[0669] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[0670] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0671] The server generates an animated video of the scene "a man opening a door."
[0672] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[0673] This data is sent to the terminal, which then integrates it and displays it to the learner.
[0674] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[0675] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[0676] The device then displays this feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[0677] This allows learners to learn English interactively and intuitively, encouraging continuous learning.
[0678] The processing flow will be explained below.
[0679] Understood. The process flow is explained in detail below.
[0680] Processing on the server
[0681] Step 1:
[0682] The server retrieves English learning material data from the database, such as TOEIC listening questions.
[0683] Step 2:
[0684] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with correct pronunciation and intonation.
[0685] Step 3:
[0686] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, using generative AI and multimedia tools.
[0687] Step 4:
[0688] The server converts the audio data into katakana strings, using a speech analysis tool to accurately represent the English pronunciation in katakana.
[0689] Step 5:
[0690] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0691] Step 6:
[0692] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0693] Processing on the device
[0694] Step 7:
[0695] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[0696] Step 8:
[0697] The device displays the integrated data to the learner and provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[0698] Step 9:
[0699] As learners practice pronunciation following the audio guide, the device records their pronunciation.
[0700] Step 10:
[0701] The device sends the recorded pronunciation data to the server.
[0702] Processing on the server (continued)
[0703] Step 11:
[0704] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0705] Step 12:
[0706] The server uses a pronunciation analysis tool to generate a pronunciation evaluation result and transmits the result to the terminal.
[0707] Processing on the terminal (continued)
[0708] Step 13:
[0709] The device receives pronunciation evaluation results from the server and displays them to the learner, providing detailed feedback on areas that need improvement based on the evaluation results.
[0710] This allows learners to improve their English listening and speaking skills in an interactive and intuitive way.
[0711] Example 1
[0712] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0713] A major challenge in modern English learning is the lack of interactive and effective learning materials. In particular, the lack of speech recognition and feedback makes it difficult for learners to self-evaluate the accuracy of their pronunciation. There is also a lack of visually appealing content to keep learners engaged.
[0714] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0715] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal, means for recording a learner's pronunciation, means for transmitting the recorded pronunciation data to the server, means for analyzing the pronunciation data and evaluating pronunciation, means for displaying the pronunciation evaluation results, means for integrating data received from the server and displaying a learning interface, and means for allowing a learner to view the visual content, katakana characters, and pitch bars while listening to audio within the learning interface. This allows learners to study English in an interactive and visually appealing environment, have their pronunciation accuracy evaluated, and receive effective feedback.
[0716] "English language learning material data" refers to information such as sentences, vocabulary, conversation examples, and test questions that learners use to study English.
[0717] "Audio data" refers to audio files with natural pronunciation and intonation that are generated based on the acquired English teaching material data.
[0718] "Visual content" refers to content that learners can refer to visually, such as illustrations or animated videos that correspond to audio data.
[0719] A "katakana string" is a representation of English pronunciation analyzed from audio data written in Japanese katakana characters.
[0720] The "pitch bar" is a karaoke-style bar that visually displays the pitch and timing of each syllable in the audio data.
[0721] A "terminal" is a device used by a learner (e.g., a PC, tablet, smartphone, etc.).
[0722] The "server" is a computer system responsible for acquiring English language teaching material data, generating audio data and visual content, analyzing katakana character strings, generating pitch bars, and managing and transmitting this data.
[0723] The "learning interface" is a user interface that allows learners to view visual content, katakana character strings, and pitch bars while listening to audio on their device.
[0724] "Pronunciation data" refers to audio data recorded by learners during pronunciation practice.
[0725] "Pronunciation evaluation" is the process of analyzing recorded pronunciation data and evaluating its accuracy.
[0726] "Feedback" is information that indicates to the learner which parts have been pronounced correctly and which parts need improvement, based on the pronunciation evaluation results.
[0727] The system for implementing the present invention is composed of a server and a terminal. The specific hardware and software and how they process data or perform calculations will be described below.
[0728] Main server processing
[0729] 1. Acquiring English teaching material data
[0730] When the server receives the request, it connects to the database to retrieve the English learning material data. The database used is a relational database management system (RDBMS) such as MySQL. For example, an SQL query such as SELECT FROM learning material WHERE ID = ? is used.
[0731] 2. Generating audio data
[0732] The server generates audio data using a speech synthesis engine based on the acquired English learning material data. The speech synthesis engine used is the Google Cloud Text-to-Speech API. The audio data is generated in MP3 or WAV format with natural pronunciation and intonation.
[0733] 3. Visual content generation
[0734] The server uses the generative AI model to generate visual content (such as illustrations and animated videos) corresponding to the English learning material data. For example, the generative AI model DALL-E generates an illustration based on the prompt "The man is opening a door," and creates an animated video using Adobe After Effects.
[0735] 4. Katakana String Generation
[0736] The server uses a speech analysis tool to convert the voice data into a katakana string, for example, Google Cloud Speech-to-Text, which converts the voice into text and then converts the resulting text according to the katakana character conversion rules.
[0737] 5. Generating interval bars
[0738] The server analyzes the audio data to identify the pitch and timing of each syllable. This analysis is performed using the audio analysis tool Praat, which extracts pitch information and then converts it into pitch bars using the visual tool FFmpeg.
[0739] 6. Data transmission
[0740] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal as an HTTP response.
[0741] Main processing of the device
[0742] 1. Receiving and integrating data
[0743] The device receives the audio data, visual content, katakana characters, and pitch bars sent from the server, and the received data is integrated and prepared for display to the learner.
[0744] 2. Display the learning interface
[0745] The device uses HTML5 and JavaScript to display visual content, katakana characters, and pitch bars while the learner is playing the audio. The interface is designed to be intuitive for learners to use.
[0746] 3. Recording pronunciation practice
[0747] The device provides the ability to record pronunciation as the learner speaks along with the audio data, using standard recording features such as the browser's MediaRecorder API.
[0748] 4. Sending pronunciation data
[0749] The device sends the recorded pronunciation data to the server as an HTTP POST request, which the server then processes and analyzes.
[0750] Pronunciation assessment and feedback
[0751] 1. Analysis of pronunciation evaluation
[0752] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, using voice analysis tools such as Google Cloud Speech-to-Text.
[0753] 2. Generating evaluation results
[0754] The server uses an evaluation tool to evaluate the accuracy of the pronunciation and generate feedback information, including which parts were pronounced correctly and which parts need improvement.
[0755] 3. Providing Feedback
[0756] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and charts.
[0757] Specific examples
[0758] For example, the processing flow based on the TOEIC listening question "The man is opening a door." is as follows:
[0759] 1. The server retrieves the English teaching material data "The man is opening a door." from the database.
[0760] 2. The server uses Google Cloud Text-to-Speech to generate audio data for "The man is opening a door."
[0761] 3. The server uses DALL-E to generate an illustration of the scene of a man opening a door, and then creates an animated video using Adobe After Effects.
[0762] 4. The server uses Google Cloud Speech-to-Text to convert the audio into the katakana string "The man is opening a door."
[0763] 5. The server uses Praat to analyze pitch and timing, and creates karaoke-style pitch bars with FFmpeg.
[0764] 6. The server sends this data to the terminal.
[0765] 7. The device integrates the received data and displays it to the learner.
[0766] 8. When the user says "The Man is Opening a Door," the audio is recorded and sent to the server.
[0767] 9. The server analyzes the voice data and generates an evaluation result (e.g., 90 points).
[0768] 10. The device displays the assessment results to the learner, showing which parts of their pronunciation are correct and which parts need improvement.
[0769] Prompt Sentence Examples
[0770] The prompt sentence for generating learning content based on the English sentence "The boy is eating an apple" is as follows:
[0771] Generate learning content based on the following English sentence: "The boy is eating an apple." Generate the following content:
[0772] 1. Audio data
[0773] 2. Katakana string
[0774] 3. Karaoke-style pitch bar
[0775] 4. Animated video of the scene
[0776] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0777] Step 1:
[0778] The server receives the request and connects to the database to retrieve English learning material data. This data includes sentences, vocabulary, conversation examples, test questions, etc. Specifically, it executes an SQL query (e.g., SELECT FROM learning material WHERE ID = ?). The input is the request and learning material ID, and the output is the corresponding English learning material data.
[0779] Step 2:
[0780] The server generates audio data based on the acquired English language learning material data. Using the Google Cloud Text-to-Speech API, it sends a text request and receives audio data. The input is the English language learning material data, and the output is the generated audio data (MP3 or WAV format).
[0781] Step 3:
[0782] The server generates visual content using a generative AI model. DALL-E is used to generate an illustration based on the prompt "The man is opening a door," and an animation video is created using Adobe After Effects. The input is the prompt and the AI model, and the output is the visual content (illustration and animation video).
[0783] Step 4:
[0784] The server uses a speech analysis tool to convert the voice data into katakana strings. Google Cloud Speech-to-Text is used to convert the voice data into text, which is then further converted according to katakana conversion rules. The input is voice data, and the output is katakana strings.
[0785] Step 5:
[0786] The server analyzes the audio data and generates pitch bars. It uses Praat to analyze the pitch and timing of the audio, and then converts it into visual pitch bars using FFmpeg. The input is the audio data, and the output is pitch bars.
[0787] Step 6:
[0788] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. It packages these data as an HTTP response and delivers it to the terminal. The input is the set of generated data, and the output is the transmission packet.
[0789] Step 7:
[0790] The device analyzes and integrates the data received from the server and displays it on the learning interface. This uses HTML5 and JavaScript to provide a user interface that learners can operate intuitively. The input is the data received from the server, and the output is the screen display of the learning interface.
[0791] Step 8:
[0792] When the user speaks in sync with the audio data, the device records the pronunciation. It uses the browser's MediaRecorder API to capture microphone input. The input is audio data, and the output is the recorded data.
[0793] Step 9:
[0794] The device sends the recorded pronunciation data to the server as an HTTP POST request. The server receives the pronunciation data and proceeds to the next processing step. The input is the recorded data, and the output is an HTTP request.
[0795] Step 10:
[0796] The server analyzes the received pronunciation data and performs pronunciation evaluation. It extracts speech features using Google Cloud Speech-to-Text and executes the evaluation algorithm. The input is the pronunciation data, and the output is the pronunciation evaluation result.
[0797] Step 11:
[0798] The server generates an evaluation result and feedback information, which includes the correct pronunciation and areas that need improvement. The input is the output of the evaluation algorithm, and the output is the detailed evaluation result and feedback information.
[0799] Step 12:
[0800] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and diagrams. The input is the evaluation results, and the output is the feedback displayed on the learning interface.
[0801] (Application example 1)
[0802] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0803] Conventional English learning systems and product description systems often do not provide sufficient support for learners and customers when practicing pronunciation or receiving feedback. In particular, in physical stores, there is a lack of interactive systems that allow non-native speakers to obtain accurate information and practice pronunciation appropriately when purchasing products. This results in reduced learning and purchasing efficiency for learners and customers, and reduced satisfaction in physical stores.
[0804] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0805] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for acquiring product information, means for generating English audio guidance based on the acquired product information, means for generating visual content corresponding to the generated audio guidance, means for scanning QR codes, means for providing guidance information based on the product information scanned from the QR codes, and means for evaluating the pronunciation of product descriptions when they are read aloud. This enables non-native speakers to efficiently acquire product information in physical stores and learn accurate pronunciation through pronunciation practice.
[0806] "English language learning material data" refers to data that includes information such as sentences, vocabulary, conversation examples, and test questions for the purpose of learning English.
[0807] "Audio data" refers to data that includes information on voices synthesized based on English language teaching material data and guidance information.
[0808] "Visual content" refers to information that includes visual content such as illustrations and animated videos that correspond to the audio data.
[0809] A "katakana character string" is a character string that expresses the pronunciation of audio data in katakana.
[0810] A "pitch bar" is a visual bar that indicates pitch and timing based on audio data.
[0811] "Terminal" refers to a device that a user uses as an interface, including smartphones, tablets, etc.
[0812] "Pronunciation data" refers to data containing voice information generated when a user records their voice.
[0813] A "server" is a computer system that processes and stores various data online.
[0814] "Product information" is data that includes detailed product information, such as product descriptions, usage instructions, and characteristics.
[0815] "Voice guidance" is data that includes voice explanations generated based on product information.
[0816] A "QR code" is a two-dimensional barcode used to encode data, including product information.
[0817] "Guide information" refers to information including detailed product descriptions and usage instructions that can be provided by scanning the QR code.
[0818] "Reading aloud" is the act of reading aloud a specified text.
[0819] The system for implementing this invention comprises a server and a terminal. The server is responsible for various data processing and generation, and the terminal provides an interface to the user. Specific embodiments will be described below.
[0820] Server-side processing
[0821] 1. Acquiring English teaching material data
[0822] The server retrieves English learning material data from the database, which includes various learning materials such as sentences, vocabulary, conversation examples, and test questions.
[0823] 2. Generating audio data
[0824] The server uses a speech synthesis engine to generate voice data with natural pronunciation and intonation based on the acquired English learning material data, and uses a generative AI model to achieve more natural pronunciation.
[0825] 3. Visual Content Generation
[0826] The server generates visual content (such as illustrations or animated videos) corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[0827] 4. Data Association and Storage
[0828] The generated audio data, visual content, katakana character strings, and pitch bars are associated and saved.
[0829] 5. Katakana String Generation
[0830] The server converts the audio data into katakana strings, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[0831] 6. Generating interval bars
[0832] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0833] 7. Data transmission
[0834] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0835] 8. Analysis and Evaluation of Pronunciation Data
[0836] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0837] Terminal side processing
[0838] 1. Receiving and integrating data
[0839] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the user.
[0840] 2. Display the learning interface
[0841] The device displays an interface that allows the user to view visual content, katakana characters, and pitch bars while playing the audio.
[0842] 3. Recording pronunciation practice
[0843] The terminal provides a function for recording pronunciation when the user pronounces along with the audio data.
[0844] 4. Sending pronunciation data
[0845] The device sends the recorded pronunciation data to the server.
[0846] 5. Displaying pronunciation evaluation results
[0847] The device displays the evaluation results received from the server to the user, providing detailed feedback on which areas need improvement.
[0848] Specific examples
[0849] For example, when a user scans a QR code printed with "12345" in a physical store, the product description is played in English. The user repeats the description and records their own pronunciation, which is then sent to the server. The server analyzes the user's pronunciation and provides evaluation feedback, which is then displayed to the user on the device.
[0850] Hardware and software used
[0851] Hardware
[0852] Server (high performance computer)
[0853] Device (smartphone, tablet, etc.)
[0854] software
[0855] Database Management Systems
[0856] Speech synthesis engine
[0857] Generative AI Models
[0858] Audio analysis tools
[0859] Prompt Sentence Examples
[0860] "You scan the QR code to hear a detailed product description in English, then record yourself repeating the description and receive feedback in the app."
[0861] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0862] Step 1:
[0863] Acquisition of English teaching material data
[0864] The server retrieves English learning material data from the database. At this time, it receives a request containing information about the learning material selected by the user, and retrieves the target learning material from the database based on that request. The retrieved data includes sentences, vocabulary, conversation examples, test questions, etc.
[0865] Input: User's request to select learning materials
[0866] Output: Acquired English teaching material data
[0867] Step 2:
[0868] Generate audio data
[0869] The server generates voice data using a generative AI model and a speech synthesis engine based on the acquired English learning material data. Specifically, the learning material text is sent as input to the speech synthesis engine, which generates voice data with natural pronunciation and intonation.
[0870] Input: English teaching material data
[0871] Output: Generated audio data
[0872] Step 3:
[0873] Visual content generation
[0874] The server generates visual content (illustrations and animated videos) that correspond to the generated audio data, using generative AI and multimedia tools to create visual materials that match the audio data.
[0875] Input: Audio data
[0876] Output: The generated visual content
[0877] Step 4:
[0878] Data association and storage
[0879] The server associates the generated audio data, visual content, katakana character strings, and pitch bars and stores them in a unified manner. At this time, the association information for each data is also saved.
[0880] Input: Audio data, visual content, katakana strings, pitch bars
[0881] Output: A set of associated data
[0882] Step 5:
[0883] Katakana string generation
[0884] The server converts the voice data into a katakana string. Using a voice analysis tool, the pronunciation of the voice data is analyzed and converted into katakana notation.
[0885] Input: Audio data
[0886] Output: Katakana string
[0887] Step 6:
[0888] Generating interval bars
[0889] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates pitch bars using a voice analysis tool.
[0890] Input: Audio data
[0891] Output: Generated interval bars
[0892] Step 7:
[0893] Sending data
[0894] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device using protocols such as HTTP or WebSocket.
[0895] Input: Audio data, visual content, katakana strings, pitch bars
[0896] Output: Data sent to the terminal
[0897] Step 8:
[0898] Receiving and integrating data
[0899] The device receives and integrates the data sent from the server, providing a viewable interface for the user.
[0900] Input: Send data
[0901] Output: Unified learning interface
[0902] Step 9:
[0903] View the learning interface
[0904] The device displays an interface that allows users to view visual content, katakana characters, and pitch bars while listening to the audio, allowing them to begin learning.
[0905] Input: Integrated dataset
[0906] Output: The displayed learning interface
[0907] Step 10:
[0908] Pronunciation practice recording
[0909] Users can use the device's recording function to practice pronunciation along with the audio data, and the recorded pronunciation data is saved on the device.
[0910] Input: User pronunciation
[0911] Output: Recorded pronunciation data
[0912] Step 11:
[0913] Sending pronunciation data
[0914] The device sends the recorded pronunciation data to the server using protocols such as HTTP or WebSocket.
[0915] Input: Recorded pronunciation data
[0916] Output: Data sent to the server
[0917] Step 12:
[0918] Analysis and evaluation of pronunciation data
[0919] The server analyzes the received pronunciation data, evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, and calculates an evaluation score using a pronunciation evaluation algorithm.
[0920] Input: Received pronunciation data
[0921] Output: Evaluation score
[0922] Step 13:
[0923] Displaying pronunciation evaluation results
[0924] The terminal displays the pronunciation evaluation results sent from the server to the user, allowing the user to check the accuracy of their own pronunciation and understand which parts need improvement.
[0925] Input: Rating score
[0926] Output: Display of pronunciation evaluation results
[0927] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0928] The system for implementing the present invention consists of three main components: a server, a terminal, and an emotion engine. The server acquires English learning material data, generates audio data, generates visual content, generates katakana strings, generates pitch bars, and evaluates pronunciation. The terminal functions as a learner interface and is responsible for audio playback, visual content display, pronunciation recording, and data transmission. The emotion engine recognizes the user's emotions and adjusts learning content and feedback based on them.
[0929] Processing on the server
[0930] 1. Acquiring English teaching material data
[0931] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[0932] 2. Generating audio data
[0933] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create audio files with natural pronunciation and intonation.
[0934] 3. Visual content generation
[0935] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create content that is visually easy to understand.
[0936] 4. Katakana String Generation
[0937] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[0938] 5. Generating interval bars
[0939] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0940] 6. Data transmission
[0941] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[0942] Processing on the device
[0943] 1. Receiving and integrating data
[0944] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[0945] 2. Display the learning interface
[0946] The terminal displays the integrated data to the learner, providing an interface that allows them to refer to visual content, katakana characters, and pitch bars while playing the audio.
[0947] 3. Use of Emotion Engine
[0948] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, for example, by analyzing the user's emotions in real time using a camera or microphone.
[0949] 4. Recording pronunciation practice
[0950] The device records the learner's pronunciation as they practice it following the audio guide.
[0951] 5. Sending pronunciation data
[0952] The device sends the recorded pronunciation data to the server.
[0953] Processing on the server (continued)
[0954] 6. Analysis of pronunciation evaluation
[0955] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[0956] 7. Generating evaluation results
[0957] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[0958] 8. Generate feedback
[0959] The server takes into account the emotional data from the emotion engine and generates optimal feedback content and tone.
[0960] Processing on the terminal (continued)
[0961] 9. Viewing assessment results and providing feedback
[0962] The device receives the evaluation results from the server and displays them to the learner. Based on the evaluation results, detailed feedback is provided on areas that need improvement. The content and tone of the feedback is adjusted appropriately based on emotion recognition data from the emotion engine.
[0963] Specific examples
[0964] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[0965] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[0966] The server generates an animated video of the scene "a man opening a door."
[0967] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[0968] This data is sent to the terminal, which then integrates it and displays it to the learner.
[0969] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[0970] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[0971] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[0972] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[0973] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[0974] This allows users to self-evaluate their learning and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[0975] The processing flow will be explained below.
[0976] Understood. The process flow is explained in detail below.
[0977] Processing on the server
[0978] Step 1:
[0979] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[0980] Step 2:
[0981] The server uses a generative AI to generate voice data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with natural pronunciation and intonation.
[0982] Step 3:
[0983] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, created using generative AI and multimedia tools.
[0984] Step 4:
[0985] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[0986] Step 5:
[0987] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[0988] Step 6:
[0989] The server transmits the generated audio data, visual content, katakana character strings, and pitch bars to the terminal.
[0990] Processing on the device
[0991] Step 7:
[0992] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[0993] Step 8:
[0994] The terminal displays the integrated data to the learner and serves as a learning interface, allowing the learner to view audio playback, visual content, katakana strings, and pitch bars.
[0995] Step 9:
[0996] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, and captures emotion data in real time using a camera and microphone.
[0997] Step 10:
[0998] As learners practice pronunciation following the audio guide, their pronunciation is recorded.
[0999] Step 11:
[1000] The device transmits the recorded pronunciation data and emotion data to the server.
[1001] Processing on the server (continued)
[1002] Step 12:
[1003] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1004] Step 13:
[1005] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[1006] Step 14:
[1007] The server takes into account the emotional data and generates feedback content and tone based on the learner's emotions.
[1008] Processing on the terminal (continued)
[1009] Step 15:
[1010] The device receives the evaluation results and feedback from the server and displays them to the learner. Based on the evaluation results and emotion-based feedback, the learner is shown which areas need improvement.
[1011] Specific examples
[1012] For example, in the case of the TOEIC listening question "The man is opening a door," the process proceeds as follows:
[1013] Step 1:
[1014] The server retrieves the relevant statement from the database.
[1015] Step 2:
[1016] The server uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1017] Step 3:
[1018] The server generates an animated video of the scene "a man opening a door."
[1019] Step 4:
[1020] The server converts the speech into the katakana string "The man is opening a door."
[1021] Step 5:
[1022] The server analyzes the audio data, identifies the pitch and timing of each syllable, and creates pitch bars.
[1023] Step 6:
[1024] These data are transmitted to the terminal.
[1025] Step 7:
[1026] The terminal consolidates the received data and displays it to the learner.
[1027] Step 8:
[1028] It provides an interface that simultaneously displays audio playback, visual content, katakana characters, and pitch bars.
[1029] Step 9:
[1030] The device uses an emotion engine to acquire emotional data from the learner's facial expressions and voice, and analyzes emotions in real time using a camera and microphone.
[1031] Step 10:
[1032] When the learner says "The man is opening a door," the device records it.
[1033] Step 11:
[1034] The device sends the recorded data and emotion data to the server.
[1035] Step 12:
[1036] The server analyzes the recording and evaluates the accuracy of the pronunciation.
[1037] Step 13:
[1038] The server generates an evaluation result and transmits it to the terminal.
[1039] Step 14:
[1040] The server adjusts the content and tone of the feedback based on the emotional data, generates the feedback, and sends it to the device.
[1041] Step 15:
[1042] The device displays the assessment results and feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[1043] This allows learners to efficiently improve their English listening and speaking skills while receiving flexible feedback that reflects their emotions.
[1044] Example 2
[1045] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1046] In conventional English learning systems, evaluation and feedback of learners' pronunciation is mechanical, making it difficult to respond flexibly to the learner's emotions and learning progress. As a result, there are problems such as a decrease in learner motivation and an impediment to effective learning. The purpose of this invention is to improve learning efficiency and motivation by recognizing learners' emotions and providing feedback accordingly.
[1047] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1048] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for recognizing the learner's emotions, and means for adjusting the content and tone of the feedback based on the emotions. This makes it possible to provide flexible and effective feedback based on the learner's emotions and the accuracy of their pronunciation.
[1049] "English language learning data" is a general term for various data used in English learning, such as English listening questions, vocabulary lists, and example sentences.
[1050] "Audio data" refers to audio files with natural pronunciation and intonation based on English language learning material data generated by the server.
[1051] "Visual content" refers to content that is easy to understand visually, such as illustrations or animated videos that correspond to audio data.
[1052] A "katakana character string" is a character string in which audio data is written in Japanese katakana.
[1053] "Pitch bar" refers to a bar that visually displays the pitch and timing of each syllable in the analyzed audio data.
[1054] "Terminal" refers to a device used by a learner that plays audio, displays visual content, records pronunciation, transmits data, etc.
[1055] "Learner" refers to a user who uses this system to learn English.
[1056] An "emotion engine" is a software component that recognizes a user's emotions and adjusts learning content and feedback based on those emotions.
[1057] "Pronunciation data" refers to audio data recorded when a learner practices pronunciation.
[1058] The "pronunciation evaluation result" indicates an evaluation of the accuracy of pronunciation obtained by the server by analyzing the pronunciation data.
[1059] "Feedback" refers to information or instructions provided to learners about their learning progress and areas for improvement.
[1060] The system for implementing this invention consists of three main components: a server, a terminal, and an emotion engine. The server is responsible for acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and evaluating pronunciation. The terminal functions as a learner interface, playing audio, displaying visual content, recording pronunciation, and transmitting data. The emotion engine also recognizes the user's emotions and adjusts learning content and feedback based on those emotions.
[1061] Processing on the server
[1062] 1. Acquiring English teaching material data
[1063] The server queries and retrieves English learning material data from a database (e.g., PostgreSQL), including TOEIC listening questions, vocabulary lists, and example sentences.
[1064] 2. Generating audio data
[1065] Based on the acquired English learning material data, the server uses a speech synthesis engine (e.g., Google Cloud Text-to-Speech API) to generate voice data with natural pronunciation and intonation.
[1066] 3. Visual content generation
[1067] The server generates visual content (e.g., illustrations or animated videos) corresponding to the generated audio data using Adobe Animate or image generation technology. For example, it generates an animated video corresponding to the sentence "The man is opening a door."
[1068] 4. Katakana String Generation
[1069] The server uses a speech analysis tool (e.g., the Julius speech engine) to convert the speech data into a katakana string, resulting in the katakana transcription "The Man is Opening a Door."
[1070] 5. Generating interval bars
[1071] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars using tools such as praat.
[1072] 6. Data transmission
[1073] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device as a JSON-formatted package.
[1074] Processing on the device
[1075] 1. Receiving and integrating data
[1076] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[1077] 2. Display the learning interface
[1078] The device uses the integrated data to display to the learner, playing audio and displaying an interface that provides visual content, katakana characters, and pitch bars.
[1079] 3. Use of Emotion Engine
[1080] The device uses an emotion engine (for example, Microsoft Azure's Emotion API) to recognize emotions from the learner's facial expressions and voice. It also analyzes the user's emotions in real time using a camera and microphone.
[1081] 4. Recording pronunciation practice
[1082] The device records the learner's pronunciation as they practice it following the audio guide.
[1083] 5. Sending pronunciation data
[1084] The device sends the recorded pronunciation data to the server via a REST API.
[1085] Processing on the server (continued)
[1086] 1. Analysis of pronunciation evaluation
[1087] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The analysis is performed using the Julius speech engine and praat.
[1088] 2. Generating evaluation results
[1089] The server converts the evaluation results obtained from the pronunciation analysis into a score and sends it to the terminal, generating feedback such as "90 points."
[1090] 3. Generate feedback
[1091] The server takes into account the emotional data and generates optimal feedback content and tone.
[1092] Processing on the terminal (continued)
[1093] 1. Viewing assessment results and providing feedback
[1094] The device displays the evaluation results received from the server to the learner, provides detailed feedback on areas that need improvement based on the evaluation results, and adjusts the tone of the feedback based on emotional data.
[1095] Specific examples
[1096] For example, in the case of the TOEIC listening question "The man is opening a door," the processing proceeds as follows:
[1097] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1098] The server generates an animated video of the scene "a man opening a door."
[1099] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[1100] This data is sent to the terminal, which then integrates it and displays it to the learner.
[1101] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[1102] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback, such as a score of 90.
[1103] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[1104] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[1105] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[1106] Example prompt: "Based on the TOEIC listening question 'The man is opening a door,' please generate natural-sounding audio data, visual content, katakana characters, and pitch bars and send them to the device."
[1107] This allows users to self-evaluate and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[1108] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1109] Step 1:
[1110] Acquisition of English teaching material data
[1111] The server retrieves English language learning material data from the database. The input is the query information stored on the server, and the output is the retrieved English language learning material data. Specifically, the server executes an SQL query to retrieve data such as TOEIC listening questions, vocabulary lists, and example sentences.
[1112] Step 2:
[1113] Generate audio data
[1114] The server uses a speech synthesis engine to generate voice data based on the acquired English language learning material data. The input is the acquired English language learning material data, and the output is the generated voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to generate natural-sounding voice from text data such as "The man is opening a door."
[1115] Step 3:
[1116] Visual content generation
[1117] The server generates visual content corresponding to the audio data. The input is the generated audio data, and the output is the generated visual content. Specifically, the server uses Adobe Animate to create an animated video of the scene "a man opening a door."
[1118] Step 4:
[1119] Katakana string generation
[1120] The server converts the speech data into a katakana string. The input is the generated speech data, and the output is a katakana string. Specifically, the server uses the Julius speech engine to convert the speech data into the katakana notation "The man is opening a door."
[1121] Step 5:
[1122] Generating interval bars
[1123] The server analyzes the audio data and generates pitch bars. The input is the generated audio data, and the output is pitch bars. Specifically, the server uses praat to analyze the pitch and timing of each syllable and generate karaoke-style pitch bars.
[1124] Step 6:
[1125] Sending data
[1126] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device. The input is all the generated data, and the output is a data package. Specifically, the server compiles this data into JSON format and sends it to the device using the HTTP protocol.
[1127] Step 7:
[1128] Receiving and integrating data
[1129] The device receives and integrates the audio data, visual content, katakana text, and pitch bars sent from the server. The input is the data package sent from the server, and the output is the integrated learning content. Specifically, the device parses the JSON-formatted data and displays the audio data and visual content on a single screen.
[1130] Step 8:
[1131] View the learning interface
[1132] The device uses the integrated data to provide an interface to be displayed to the learner. The input is the integrated learning content, and the output is the displayed learning interface. Specifically, the device plays audio while synchronously drawing an interface on the screen that displays visual content, katakana characters, and pitch bars.
[1133] Step 9:
[1134] Using the Emotion Engine
[1135] The device recognizes emotions from the learner's facial expressions and voice. The input is the learner's facial and voice data, and the output is emotional data. Specifically, the device uses Microsoft Azure's Emotion API to analyze the learner's emotions in real time from facial expressions captured by the camera and voice data collected by the microphone.
[1136] Step 10:
[1137] Pronunciation practice recording
[1138] The device records the learner's pronunciation as they practice pronunciation following the audio guide. The input is the learner's pronunciation, and the output is the recorded pronunciation data. Specifically, the device uses a built-in microphone to record the learner's pronunciation into an audio file.
[1139] Step 11:
[1140] Sending pronunciation data
[1141] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the pronunciation data sent to the server. Specifically, the device uses a REST API to upload the audio file to the server.
[1142] Step 12:
[1143] Pronunciation evaluation analysis
[1144] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the recorded pronunciation data, the katakana string, and the pitch bar, and the output is the pronunciation evaluation result. Specifically, the server uses the Julius speech engine and praat to calculate the degree of match for each phoneme and convert it into a score.
[1145] Step 13:
[1146] Generating evaluation results
[1147] The server generates pronunciation evaluation results and sends them to the terminal. The input is the pronunciation evaluation results, and the output is the evaluation results sent from the server to the terminal. Specifically, the server summarizes the evaluation results and sends them to the terminal using the HTTP protocol.
[1148] Step 14:
[1149] Viewing assessment results and providing feedback
[1150] The device receives the evaluation results from the server, displays them to the learner, and provides feedback. The input is the evaluation results and emotional data, and the output is the displayed feedback. Specifically, the device highlights areas that need improvement based on the evaluation results and provides flexible feedback based on the emotional data.
[1151] (Application example 2)
[1152] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1153] Conventional English learning systems are unable to provide feedback that reflects learners' emotions, limiting their ability to support effective learning. Furthermore, the feedback provided for English pronunciation evaluation and correction is fixed, making it difficult to respond individually to learners' emotions and reactions. As a result, issues such as a decline in learning motivation and a delay in learning progress have arisen.
[1154] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting to katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to the terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for acquiring emotion data based on the evaluation results, and means for displaying the pronunciation evaluation results and the emotion data. This enables flexible and individual feedback based on the learner's emotions.
[1155] "English language learning material data" refers to data on learning materials used in learning English, and includes vocabulary lists, sample sentences, listening questions, and the like.
[1156] "Speech data" refers to acoustic information generated based on acquired English language learning material data, and includes English pronunciation and intonation.
[1157] "Visual content" includes visual information such as illustrations and animated videos generated in response to audio data.
[1158] A "katakana character string" is a character string obtained by converting the generated voice data into Japanese katakana notation.
[1159] A "pitch bar" is a bar that visually represents the pitch and timing of each syllable by analyzing audio data.
[1160] A "terminal" is a device that receives the generated audio data, visual content, katakana character strings, and pitch bars, and provides them to a learner.
[1161] "Pronunciation data" refers to data that is recorded by a learner pronouncing words and sent to the server.
[1162] The "pronunciation evaluation result" is the result of analyzing the recorded pronunciation data and evaluating its accuracy.
[1163] "Emotion data" is information about emotions obtained from the learner's facial expressions, voice, etc.
[1164] "Feedback" refers to guidance and advice provided to learners based on pronunciation evaluation results and emotional data.
[1165] The system for implementing this invention mainly consists of three main components: a server, a terminal, and an emotion engine.
[1166] Server Processing
[1167] The server implements the following functions:
[1168] 1. Acquisition of English learning material data: The server acquires English learning material data from the database. The acquired data includes listening questions, vocabulary lists, sample sentences, etc.
[1169] 2. Audio data generation: The server generates audio data using a speech synthesis engine based on the acquired English learning material data. Using this generative AI model, an audio file with natural pronunciation and intonation is created.
[1170] 3. Visual content generation: The server generates visual content such as illustrations and animated videos corresponding to the audio data. Generative AI and multimedia tools are used to create content that is easy to understand visually.
[1171] 4. Katakana string generation: The server converts the audio data into Katakana strings, using a speech analysis tool to represent the English pronunciation in Japanese Katakana.
[1172] 5. Generating pitch bars: The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1173] 6. Data transmission: The server transmits the generated audio data, visual content, katakana string, and pitch bar to the terminal.
[1174] Terminal handling
[1175] The device implements the following functions:
[1176] 1. Data reception and integration: The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[1177] 2. Display of learning interface: The device displays the integrated data to the learner. It provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[1178] 3. Recording pronunciation practice: The device records the learner's pronunciation as they practice pronunciation following the audio guide.
[1179] 4. Sending pronunciation data: The device sends the recorded pronunciation data to the server.
[1180] Emotion engine processing
[1181] The emotion engine implements the following functions:
[1182] 1. Emotion Recognition: The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice. For example, the device analyzes the user's emotions in real time using a camera or microphone.
[1183] 2. Sending emotion data: The device sends the recognized emotion data to the server.
[1184] Reprocessing on the server
[1185] The server implements the following functions:
[1186] 1. Pronunciation evaluation analysis: The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1187] 2. Generating evaluation results: The server generates evaluation results using the pronunciation analysis tool and sends the evaluation results to the terminal.
[1188] 3. Feedback Generation: The server takes into account the emotional data from the emotion engine and generates the optimal feedback content and tone.
[1189] Reprocessing on the device
[1190] The device implements the following functions:
[1191] 1. Displaying assessment results and providing feedback: The device displays the assessment results received from the server to the learner. Based on the assessment results, detailed feedback is provided on which areas need improvement. Appropriate feedback content and tone are provided based on emotion recognition data from the emotion engine.
[1192] Specific examples
[1193] For example, consider a step in which a user tries on a new jacket in a virtual fitting room. The device recognizes the user's emotions from their facial expressions and sends the data to the server. The server generates feedback based on the emotion data, such as "Why don't you try on a jacket in a lighter color?" and sends it to the device. The device then provides the feedback to the user visually and audibly.
[1194] The following is an example of a prompt sentence to be used in the specific example:
[1195] Analyze an image of a user trying on a new jacket in front of a mirror. From their facial expression, they appear to be somewhat satisfied, but would like further style suggestions. Generate feedback content about the color and design of the jacket.
[1196] This significantly improves the user experience in the virtual fitting room and enables efficient and effective feedback.
[1197] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1198] Step 1:
[1199] The server retrieves English learning material data from the database. This data includes listening questions, vocabulary lists, sample sentences, etc. The input is learning material information in the learning material database, and the output is the retrieved information for each learning material data.
[1200] Step 2:
[1201] The server generates audio data based on the acquired English learning material data, using a generative AI model to create an audio file with natural pronunciation and intonation. The input is the learning material data, and the output is the generated audio data.
[1202] Step 3:
[1203] The server generates visual content corresponding to the generated audio data. Here, generative AI and multimedia tools are used to create visually easy-to-understand content such as illustrations and animated videos. The input is audio data, and the output is visual content.
[1204] Step 4:
[1205] The server converts the generated voice data into a katakana string. The server uses a voice analysis tool to convert the voice data into Japanese katakana. The input is the voice data, and the output is a katakana string.
[1206] Step 5:
[1207] The server analyzes the audio data and generates pitch bars, which are a visual representation of the pitch and timing of each syllable. The input is the audio data, and the output is pitch bars.
[1208] Step 6:
[1209] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. The input is each generated data, and the output is the transmission of that data to the terminal.
[1210] Step 7:
[1211] The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server. The input is the multiple data sent from the server, and the output is the integrated training data.
[1212] Step 8:
[1213] The terminal displays the integrated data to the learner. It provides an interface that allows the learner to refer to visual content, katakana characters, and pitch bars while playing the audio. The input is the integrated data, and the output is the display of the learning interface.
[1214] Step 9:
[1215] The device records the user's pronunciation as they practice their pronunciation following the audio guide. The input is the user's pronunciation, and the output is the recorded pronunciation data.
[1216] Step 10:
[1217] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the completion of transmission to the server.
[1218] Step 11:
[1219] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the pronunciation data and evaluation criteria (katakana string and pitch bar), and the output is the evaluation result.
[1220] Step 12:
[1221] The server generates optimal feedback content and tone based on the pronunciation evaluation results and emotional data, and sends it to the device. At this time, the server inputs the previously defined prompt sentence into the generative AI model to generate feedback. The input is the evaluation results and emotional data, and the output is the feedback content.
[1222] Step 13:
[1223] The terminal displays the evaluation results and feedback received from the server to the learner, indicating which parts were pronounced correctly and which parts need improvement. The input is the evaluation results and feedback content, and the output is the display and feedback provided to the learner.
[1224] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1225] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1226] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1227] [Third embodiment]
[1228] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1229] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1230] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1231] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1232] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1234] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1235] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1236] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1237] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1238] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1239] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1240] The system for implementing the present invention mainly consists of a server and a terminal. The server is responsible for multiple functions, including acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and pronunciation evaluation. Meanwhile, the terminal provides an interface for learners and is responsible for displaying audio data, visual content, katakana strings, and pitch bars, as well as recording and transmitting audio.
[1241] Main server processing
[1242] 1. Acquiring English teaching material data
[1243] The server retrieves English learning material data from the database, including sentences, vocabulary, conversation examples, test questions, etc.
[1244] 2. Generating audio data
[1245] The server uses a generative AI to generate voice data based on the acquired English learning material data, and uses a speech synthesis engine to create natural pronunciation and intonation.
[1246] 3. Visual content generation
[1247] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[1248] 4. Katakana String Generation
[1249] The server converts the audio data into katakana characters, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[1250] 5. Generating interval bars
[1251] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1252] 6. Data transmission
[1253] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1254] Main processing of the device
[1255] 1. Receiving and integrating data
[1256] The terminal integrates the audio data, visual content, katakana strings, and pitch bars received from the server and provides a viewable interface for the learner.
[1257] 2. Display the learning interface
[1258] The device displays an interface that allows learners to view visual content, katakana strings, and pitch bars while listening to the audio.
[1259] 3. Recording pronunciation practice
[1260] The device provides a function to record pronunciation when the learner pronounces along with the audio data.
[1261] 4. Sending pronunciation data
[1262] The device sends the recorded pronunciation data to the server.
[1263] Pronunciation assessment and feedback
[1264] 1. Analysis of pronunciation evaluation
[1265] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1266] 2. Generating evaluation results
[1267] The server uses a pronunciation analysis tool to generate an evaluation result and transmits the result to the terminal.
[1268] 3. Providing Feedback
[1269] The device displays the evaluation results received from the server to the learner and provides detailed feedback on which areas need improvement.
[1270] Specific examples
[1271] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[1272] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1273] The server generates an animated video of the scene "a man opening a door."
[1274] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[1275] This data is sent to the terminal, which then integrates it and displays it to the learner.
[1276] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[1277] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[1278] The device then displays this feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[1279] This allows learners to learn English interactively and intuitively, encouraging continuous learning.
[1280] The processing flow will be explained below.
[1281] Understood. The process flow is explained in detail below.
[1282] Processing on the server
[1283] Step 1:
[1284] The server retrieves English learning material data from the database, such as TOEIC listening questions.
[1285] Step 2:
[1286] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with correct pronunciation and intonation.
[1287] Step 3:
[1288] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, using generative AI and multimedia tools.
[1289] Step 4:
[1290] The server converts the audio data into katakana strings, using a speech analysis tool to accurately represent the English pronunciation in katakana.
[1291] Step 5:
[1292] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1293] Step 6:
[1294] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1295] Processing on the device
[1296] Step 7:
[1297] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[1298] Step 8:
[1299] The device displays the integrated data to the learner and provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[1300] Step 9:
[1301] As learners practice pronunciation following the audio guide, the device records their pronunciation.
[1302] Step 10:
[1303] The device sends the recorded pronunciation data to the server.
[1304] Processing on the server (continued)
[1305] Step 11:
[1306] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1307] Step 12:
[1308] The server uses a pronunciation analysis tool to generate a pronunciation evaluation result and transmits the result to the terminal.
[1309] Processing on the terminal (continued)
[1310] Step 13:
[1311] The device receives pronunciation evaluation results from the server and displays them to the learner, providing detailed feedback on areas that need improvement based on the evaluation results.
[1312] This allows learners to improve their English listening and speaking skills in an interactive and intuitive way.
[1313] Example 1
[1314] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1315] A major challenge in modern English learning is the lack of interactive and effective learning materials. In particular, the lack of speech recognition and feedback makes it difficult for learners to self-evaluate the accuracy of their pronunciation. There is also a lack of visually appealing content to keep learners engaged.
[1316] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1317] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal, means for recording a learner's pronunciation, means for transmitting the recorded pronunciation data to the server, means for analyzing the pronunciation data and evaluating pronunciation, means for displaying the pronunciation evaluation results, means for integrating data received from the server and displaying a learning interface, and means for allowing a learner to view the visual content, katakana characters, and pitch bars while listening to audio within the learning interface. This allows learners to study English in an interactive and visually appealing environment, have their pronunciation accuracy evaluated, and receive effective feedback.
[1318] "English language learning material data" refers to information such as sentences, vocabulary, conversation examples, and test questions that learners use to study English.
[1319] "Audio data" refers to audio files with natural pronunciation and intonation that are generated based on the acquired English teaching material data.
[1320] "Visual content" refers to content that learners can refer to visually, such as illustrations or animated videos that correspond to audio data.
[1321] A "katakana string" is a representation of English pronunciation analyzed from audio data written in Japanese katakana characters.
[1322] The "pitch bar" is a karaoke-style bar that visually displays the pitch and timing of each syllable in the audio data.
[1323] A "terminal" is a device used by a learner (e.g., a PC, tablet, smartphone, etc.).
[1324] The "server" is a computer system responsible for acquiring English language teaching material data, generating audio data and visual content, analyzing katakana character strings, generating pitch bars, and managing and transmitting this data.
[1325] The "learning interface" is a user interface that allows learners to view visual content, katakana character strings, and pitch bars while listening to audio on their device.
[1326] "Pronunciation data" refers to audio data recorded by learners during pronunciation practice.
[1327] "Pronunciation evaluation" is the process of analyzing recorded pronunciation data and evaluating its accuracy.
[1328] "Feedback" is information that indicates to the learner which parts have been pronounced correctly and which parts need improvement, based on the pronunciation evaluation results.
[1329] The system for implementing the present invention is composed of a server and a terminal. The specific hardware and software and how they process data or perform calculations will be described below.
[1330] Main server processing
[1331] 1. Acquiring English teaching material data
[1332] When the server receives the request, it connects to the database to retrieve the English learning material data. The database used is a relational database management system (RDBMS) such as MySQL. For example, an SQL query such as SELECT FROM learning material WHERE ID = ? is used.
[1333] 2. Generating audio data
[1334] The server generates audio data using a speech synthesis engine based on the acquired English learning material data. The speech synthesis engine used is the Google Cloud Text-to-Speech API. The audio data is generated in MP3 or WAV format with natural pronunciation and intonation.
[1335] 3. Visual content generation
[1336] The server uses a generative AI model to generate visual content (such as illustrations and animated videos) corresponding to the English learning material data. For example, the generative AI model DALL-E generates an illustration based on the prompt "The man is opening a door," and creates an animated video using Adobe After Effects.
[1337] 4. Katakana String Generation
[1338] The server uses a speech analysis tool to convert the voice data into a katakana string, for example, Google Cloud Speech-to-Text is used to convert the voice into text, and the resulting text is then converted according to the katakana character conversion rules.
[1339] 5. Generating interval bars
[1340] The server analyzes the audio data to identify the pitch and timing of each syllable. This analysis is performed using the audio analysis tool Praat, which extracts pitch information and then converts it into pitch bars using the visual tool FFmpeg.
[1341] 6. Data transmission
[1342] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal as an HTTP response.
[1343] Main processing of the device
[1344] 1. Receiving and integrating data
[1345] The device receives the audio data, visual content, katakana characters, and pitch bars sent from the server, and the received data is integrated and prepared for display to the learner.
[1346] 2. Display the learning interface
[1347] The device uses HTML5 and JavaScript to display visual content, katakana characters, and pitch bars while the learner is playing the audio. The interface is designed to be intuitive for learners to use.
[1348] 3. Recording pronunciation practice
[1349] The device provides the ability to record pronunciation as the learner speaks along with the audio data, using standard recording features such as the browser's MediaRecorder API.
[1350] 4. Sending pronunciation data
[1351] The device sends the recorded pronunciation data to the server as an HTTP POST request, which the server then processes and analyzes.
[1352] Pronunciation assessment and feedback
[1353] 1. Analysis of pronunciation evaluation
[1354] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, using voice analysis tools such as Google Cloud Speech-to-Text.
[1355] 2. Generating evaluation results
[1356] The server uses an evaluation tool to evaluate the accuracy of the pronunciation and generate feedback information, including which parts were pronounced correctly and which parts need improvement.
[1357] 3. Providing Feedback
[1358] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and charts.
[1359] Specific examples
[1360] For example, the processing flow based on the TOEIC listening question "The man is opening a door." is as follows:
[1361] 1. The server retrieves the English teaching material data "The man is opening a door." from the database.
[1362] 2. The server uses Google Cloud Text-to-Speech to generate audio data for "The man is opening a door."
[1363] 3. The server uses DALL-E to generate an illustration of the scene of a man opening a door, and then creates an animated video using Adobe After Effects.
[1364] 4. The server uses Google Cloud Speech-to-Text to convert the audio into the katakana string "The man is opening a door."
[1365] 5. The server uses Praat to analyze pitch and timing, and creates karaoke-style pitch bars with FFmpeg.
[1366] 6. The server sends this data to the terminal.
[1367] 7. The device integrates the received data and displays it to the learner.
[1368] 8. When the user says "The Man is Opening a Door," the audio is recorded and sent to the server.
[1369] 9. The server analyzes the voice data and generates an evaluation result (e.g., 90 points).
[1370] 10. The device displays the assessment results to the learner, showing which parts of their pronunciation are correct and which parts need improvement.
[1371] Prompt Sentence Examples
[1372] The prompt sentence for generating learning content based on the English sentence "The boy is eating an apple" is as follows:
[1373] Generate learning content based on the following sentence: "The boy is eating an apple." Generate the following content:
[1374] 1. Audio data
[1375] 2. Katakana string
[1376] 3. Karaoke-style pitch bar
[1377] 4. Animated video of the scene
[1378] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] The server receives the request and connects to the database to retrieve English learning material data. This data includes sentences, vocabulary, conversation examples, test questions, etc. Specifically, it executes an SQL query (e.g., SELECT FROM learning material WHERE ID = ?). The input is the request and learning material ID, and the output is the corresponding English learning material data.
[1381] Step 2:
[1382] The server generates audio data based on the acquired English language learning material data. Using the Google Cloud Text-to-Speech API, it sends a text request and receives audio data. The input is the English language learning material data, and the output is the generated audio data (MP3 or WAV format).
[1383] Step 3:
[1384] The server generates visual content using a generative AI model. DALL-E is used to generate an illustration based on the prompt "The man is opening a door," and an animation video is created using Adobe After Effects. The input is the prompt and the AI model, and the output is the visual content (illustration and animation video).
[1385] Step 4:
[1386] The server uses a speech analysis tool to convert the voice data into katakana strings. Google Cloud Speech-to-Text is used to convert the voice data into text, which is then further converted according to katakana conversion rules. The input is voice data, and the output is katakana strings.
[1387] Step 5:
[1388] The server analyzes audio data and generates pitch bars. It uses Praat to analyze the pitch and timing of the audio, and then converts it into visual pitch bars using FFmpeg. The input is audio data, and the output is pitch bars.
[1389] Step 6:
[1390] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. It packages these data as an HTTP response and delivers it to the terminal. The input is the set of generated data, and the output is the transmission packet.
[1391] Step 7:
[1392] The device analyzes and integrates the data received from the server and displays it on the learning interface. This uses HTML5 and JavaScript to provide a user interface that learners can operate intuitively. The input is the data received from the server, and the output is the screen display of the learning interface.
[1393] Step 8:
[1394] When the user speaks in sync with the audio data, the device records the pronunciation. It uses the browser's MediaRecorder API to capture microphone input. The input is audio data, and the output is the recorded data.
[1395] Step 9:
[1396] The device sends the recorded pronunciation data to the server as an HTTP POST request. The server receives the pronunciation data and proceeds to the next processing step. The input is the recorded data, and the output is an HTTP request.
[1397] Step 10:
[1398] The server analyzes the received pronunciation data and performs pronunciation evaluation. It extracts speech features using Google Cloud Speech-to-Text and executes the evaluation algorithm. The input is the pronunciation data, and the output is the pronunciation evaluation result.
[1399] Step 11:
[1400] The server generates an evaluation result and feedback information, which includes the correct pronunciation and areas that need improvement. The input is the output of the evaluation algorithm, and the output is the detailed evaluation result and feedback information.
[1401] Step 12:
[1402] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and diagrams. The input is the evaluation results, and the output is the feedback displayed on the learning interface.
[1403] (Application example 1)
[1404] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1405] Conventional English learning systems and product description systems often do not provide sufficient support for learners and customers when practicing pronunciation or receiving feedback. In particular, in physical stores, there is a lack of interactive systems that allow non-native speakers to obtain accurate information and practice pronunciation appropriately when purchasing products. This results in reduced learning and purchasing efficiency for learners and customers, and reduced satisfaction in physical stores.
[1406] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1407] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for acquiring product information, means for generating English audio guidance based on the acquired product information, means for generating visual content corresponding to the generated audio guidance, means for scanning QR codes, means for providing guidance information based on the product information scanned from the QR codes, and means for evaluating the pronunciation of product descriptions when they are read aloud. This enables non-native speakers to efficiently acquire product information in physical stores and learn accurate pronunciation through pronunciation practice.
[1408] "English language learning material data" refers to data that includes information such as sentences, vocabulary, conversation examples, and test questions for the purpose of learning English.
[1409] "Audio data" refers to data that includes information on voices synthesized based on English language teaching material data and guidance information.
[1410] "Visual content" refers to information that includes visual content such as illustrations and animated videos that correspond to the audio data.
[1411] A "katakana character string" is a character string that expresses the pronunciation of audio data in katakana.
[1412] A "pitch bar" is a visual bar that indicates pitch and timing based on audio data.
[1413] "Terminal" refers to a device that a user uses as an interface, including smartphones, tablets, etc.
[1414] "Pronunciation data" refers to data containing voice information generated when a user records their voice.
[1415] A "server" is a computer system that processes and stores various data online.
[1416] "Product information" is data that includes detailed product information, such as product descriptions, usage instructions, and characteristics.
[1417] "Voice guidance" is data that includes voice explanations generated based on product information.
[1418] A "QR code" is a two-dimensional barcode used to encode data, including product information.
[1419] "Guide information" refers to information including detailed product descriptions and usage instructions that can be provided by scanning the QR code.
[1420] "Reading aloud" is the act of reading aloud a specified text.
[1421] The system for implementing this invention comprises a server and a terminal. The server is responsible for various data processing and generation, and the terminal provides an interface to the user. Specific embodiments will be described below.
[1422] Server-side processing
[1423] 1. Acquiring English teaching material data
[1424] The server retrieves English learning material data from the database, which includes various learning materials such as sentences, vocabulary, conversation examples, and test questions.
[1425] 2. Generating audio data
[1426] The server uses a speech synthesis engine to generate voice data with natural pronunciation and intonation based on the acquired English learning material data, and uses a generative AI model to achieve more natural pronunciation.
[1427] 3. Visual Content Generation
[1428] The server generates visual content (such as illustrations or animated videos) corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[1429] 4. Data Association and Storage
[1430] The generated audio data, visual content, katakana character strings, and pitch bars are associated and saved.
[1431] 5. Katakana String Generation
[1432] The server converts the audio data into katakana strings, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[1433] 6. Generating interval bars
[1434] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1435] 7. Data transmission
[1436] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1437] 8. Analysis and Evaluation of Pronunciation Data
[1438] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1439] Terminal side processing
[1440] 1. Receiving and integrating data
[1441] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the user.
[1442] 2. Display the learning interface
[1443] The device displays an interface that allows the user to view visual content, katakana characters, and pitch bars while playing the audio.
[1444] 3. Recording pronunciation practice
[1445] The terminal provides a function for recording pronunciation when the user pronounces along with the audio data.
[1446] 4. Sending pronunciation data
[1447] The device sends the recorded pronunciation data to the server.
[1448] 5. Displaying pronunciation evaluation results
[1449] The device displays the evaluation results received from the server to the user, providing detailed feedback on which areas need improvement.
[1450] Specific examples
[1451] For example, when a user scans a QR code printed with "12345" in a physical store, the product description is played in English. The user repeats the description and records their own pronunciation, which is then sent to the server. The server analyzes the user's pronunciation and provides evaluation feedback, which is then displayed to the user on the device.
[1452] Hardware and software used
[1453] Hardware
[1454] Server (high performance computer)
[1455] Device (smartphone, tablet, etc.)
[1456] software
[1457] Database Management Systems
[1458] Speech synthesis engine
[1459] Generative AI Models
[1460] Audio analysis tools
[1461] Prompt Sentence Examples
[1462] "You scan the QR code to hear a detailed product description in English, then record yourself repeating the description and receive feedback in the app."
[1463] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1464] Step 1:
[1465] Acquisition of English teaching material data
[1466] The server retrieves English learning material data from the database. At this time, it receives a request containing information about the learning material selected by the user, and retrieves the target learning material from the database based on that request. The retrieved data includes sentences, vocabulary, conversation examples, test questions, etc.
[1467] Input: User's request to select learning materials
[1468] Output: Acquired English teaching material data
[1469] Step 2:
[1470] Generating audio data
[1471] The server generates voice data using a generative AI model and a speech synthesis engine based on the acquired English learning material data. Specifically, the learning material text is sent as input to the speech synthesis engine, which generates voice data with natural pronunciation and intonation.
[1472] Input: English teaching material data
[1473] Output: Generated audio data
[1474] Step 3:
[1475] Visual content generation
[1476] The server generates visual content (illustrations and animated videos) that correspond to the generated audio data, using generative AI and multimedia tools to create visual materials that match the audio data.
[1477] Input: Audio data
[1478] Output: The generated visual content
[1479] Step 4:
[1480] Data association and storage
[1481] The server associates the generated audio data, visual content, katakana character strings, and pitch bars and stores them in a unified manner. At this time, the association information for each data is also saved.
[1482] Input: Audio data, visual content, katakana strings, pitch bars
[1483] Output: A set of associated data
[1484] Step 5:
[1485] Katakana string generation
[1486] The server converts the voice data into a katakana string. Using a voice analysis tool, the pronunciation of the voice data is analyzed and converted into katakana notation.
[1487] Input: Audio data
[1488] Output: Katakana string
[1489] Step 6:
[1490] Generating interval bars
[1491] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates pitch bars using a voice analysis tool.
[1492] Input: Audio data
[1493] Output: Generated interval bars
[1494] Step 7:
[1495] Sending data
[1496] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device using protocols such as HTTP or WebSocket.
[1497] Input: Audio data, visual content, katakana strings, pitch bars
[1498] Output: Data sent to the terminal
[1499] Step 8:
[1500] Receiving and integrating data
[1501] The device receives and integrates the data sent from the server, providing a viewable interface for the user.
[1502] Input: Send data
[1503] Output: Unified learning interface
[1504] Step 9:
[1505] View the learning interface
[1506] The device displays an interface that allows users to view visual content, katakana characters, and pitch bars while listening to the audio, allowing them to begin learning.
[1507] Input: Integrated dataset
[1508] Output: The displayed learning interface
[1509] Step 10:
[1510] Pronunciation practice recording
[1511] Users can use the device's recording function to practice pronunciation along with the audio data, and the recorded pronunciation data is saved on the device.
[1512] Input: User pronunciation
[1513] Output: Recorded pronunciation data
[1514] Step 11:
[1515] Sending pronunciation data
[1516] The device sends the recorded pronunciation data to the server using protocols such as HTTP or WebSocket.
[1517] Input: Recorded pronunciation data
[1518] Output: Data sent to the server
[1519] Step 12:
[1520] Analysis and evaluation of pronunciation data
[1521] The server analyzes the received pronunciation data, evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, and calculates an evaluation score using a pronunciation evaluation algorithm.
[1522] Input: Received pronunciation data
[1523] Output: Evaluation score
[1524] Step 13:
[1525] Displaying pronunciation evaluation results
[1526] The terminal displays the pronunciation evaluation results sent from the server to the user, allowing the user to check the accuracy of their own pronunciation and understand which parts need improvement.
[1527] Input: Rating score
[1528] Output: Display of pronunciation evaluation results
[1529] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1530] The system for implementing the present invention consists of three main components: a server, a terminal, and an emotion engine. The server acquires English learning material data, generates audio data, generates visual content, generates katakana strings, generates pitch bars, and evaluates pronunciation. The terminal functions as a learner interface and is responsible for audio playback, visual content display, pronunciation recording, and data transmission. The emotion engine recognizes the user's emotions and adjusts learning content and feedback based on them.
[1531] Processing on the server
[1532] 1. Acquiring English teaching material data
[1533] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[1534] 2. Generating audio data
[1535] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create audio files with natural pronunciation and intonation.
[1536] 3. Visual content generation
[1537] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create content that is visually easy to understand.
[1538] 4. Katakana String Generation
[1539] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[1540] 5. Generating interval bars
[1541] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1542] 6. Data transmission
[1543] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1544] Processing on the device
[1545] 1. Receiving and integrating data
[1546] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[1547] 2. Display the learning interface
[1548] The terminal displays the integrated data to the learner, providing an interface that allows them to refer to visual content, katakana characters, and pitch bars while playing the audio.
[1549] 3. Use of Emotion Engine
[1550] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, for example, by analyzing the user's emotions in real time using a camera or microphone.
[1551] 4. Recording pronunciation practice
[1552] The device records the learner's pronunciation as they practice it following the audio guide.
[1553] 5. Sending pronunciation data
[1554] The device sends the recorded pronunciation data to the server.
[1555] Processing on the server (continued)
[1556] 6. Analysis of pronunciation evaluation
[1557] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1558] 7. Generating evaluation results
[1559] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[1560] 8. Generate feedback
[1561] The server takes into account the emotional data from the emotion engine and generates optimal feedback content and tone.
[1562] Processing on the terminal (continued)
[1563] 9. Viewing assessment results and providing feedback
[1564] The device receives the evaluation results from the server and displays them to the learner. Based on the evaluation results, detailed feedback is provided on areas that need improvement. The content and tone of the feedback is adjusted appropriately based on emotion recognition data from the emotion engine.
[1565] Specific examples
[1566] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[1567] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1568] The server generates an animated video of the scene "a man opening a door."
[1569] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[1570] This data is sent to the terminal, which then integrates it and displays it to the learner.
[1571] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[1572] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[1573] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[1574] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[1575] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[1576] This allows users to self-evaluate their learning and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[1577] The processing flow will be explained below.
[1578] Understood. The process flow is explained in detail below.
[1579] Processing on the server
[1580] Step 1:
[1581] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[1582] Step 2:
[1583] The server uses a generative AI to generate voice data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with natural pronunciation and intonation.
[1584] Step 3:
[1585] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, created using generative AI and multimedia tools.
[1586] Step 4:
[1587] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[1588] Step 5:
[1589] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1590] Step 6:
[1591] The server transmits the generated audio data, visual content, katakana character strings, and pitch bars to the terminal.
[1592] Processing on the device
[1593] Step 7:
[1594] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[1595] Step 8:
[1596] The terminal displays the integrated data to the learner and serves as a learning interface, allowing the learner to view audio playback, visual content, katakana strings, and pitch bars.
[1597] Step 9:
[1598] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, and captures emotion data in real time using a camera and microphone.
[1599] Step 10:
[1600] As learners practice pronunciation following the audio guide, their pronunciation is recorded.
[1601] Step 11:
[1602] The device transmits the recorded pronunciation data and emotion data to the server.
[1603] Processing on the server (continued)
[1604] Step 12:
[1605] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1606] Step 13:
[1607] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[1608] Step 14:
[1609] The server takes into account the emotional data and generates feedback content and tone based on the learner's emotions.
[1610] Processing on the terminal (continued)
[1611] Step 15:
[1612] The device receives the evaluation results and feedback from the server and displays them to the learner. Based on the evaluation results and emotion-based feedback, the learner is shown which areas need improvement.
[1613] Specific examples
[1614] For example, in the case of the TOEIC listening question "The man is opening a door," the process proceeds as follows:
[1615] Step 1:
[1616] The server retrieves the relevant statement from the database.
[1617] Step 2:
[1618] The server uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1619] Step 3:
[1620] The server generates an animated video of the scene "a man opening a door."
[1621] Step 4:
[1622] The server converts the speech into the katakana string "The man is opening a door."
[1623] Step 5:
[1624] The server analyzes the audio data, identifies the pitch and timing of each syllable, and creates pitch bars.
[1625] Step 6:
[1626] These data are transmitted to the terminal.
[1627] Step 7:
[1628] The terminal consolidates the received data and displays it to the learner.
[1629] Step 8:
[1630] It provides an interface that simultaneously displays audio playback, visual content, katakana characters, and pitch bars.
[1631] Step 9:
[1632] The device uses an emotion engine to acquire emotional data from the learner's facial expressions and voice, and analyzes emotions in real time using a camera and microphone.
[1633] Step 10:
[1634] When the learner says "The man is opening a door," the device records it.
[1635] Step 11:
[1636] The device sends the recorded data and emotion data to the server.
[1637] Step 12:
[1638] The server analyzes the recording and evaluates the accuracy of the pronunciation.
[1639] Step 13:
[1640] The server generates an evaluation result and transmits it to the terminal.
[1641] Step 14:
[1642] The server adjusts the content and tone of the feedback based on the emotional data, generates the feedback, and sends it to the device.
[1643] Step 15:
[1644] The device displays the assessment results and feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[1645] This allows learners to efficiently improve their English listening and speaking skills while receiving flexible feedback that reflects their emotions.
[1646] Example 2
[1647] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1648] In conventional English learning systems, evaluation and feedback of learners' pronunciation is mechanical, making it difficult to respond flexibly to the learner's emotions and learning progress. As a result, there are problems such as a decrease in learner motivation and an impediment to effective learning. The purpose of this invention is to improve learning efficiency and motivation by recognizing learners' emotions and providing feedback accordingly.
[1649] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1650] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for recognizing the learner's emotions, and means for adjusting the content and tone of the feedback based on the emotions. This makes it possible to provide flexible and effective feedback based on the learner's emotions and the accuracy of their pronunciation.
[1651] "English language learning data" is a general term for various data used in English learning, such as English listening questions, vocabulary lists, and example sentences.
[1652] "Audio data" refers to audio files with natural pronunciation and intonation based on English language learning material data generated by the server.
[1653] "Visual content" refers to content that is easy to understand visually, such as illustrations or animated videos that correspond to audio data.
[1654] A "katakana character string" is a character string in which audio data is written in Japanese katakana.
[1655] "Pitch bar" refers to a bar that visually displays the pitch and timing of each syllable in the analyzed audio data.
[1656] "Terminal" refers to a device used by a learner that plays audio, displays visual content, records pronunciation, transmits data, etc.
[1657] "Learner" refers to a user who uses this system to learn English.
[1658] An "emotion engine" is a software component that recognizes a user's emotions and adjusts learning content and feedback based on those emotions.
[1659] "Pronunciation data" refers to audio data recorded when a learner practices pronunciation.
[1660] The "pronunciation evaluation result" indicates an evaluation of the accuracy of pronunciation obtained by the server by analyzing the pronunciation data.
[1661] "Feedback" refers to information or instructions provided to learners about their learning progress and areas for improvement.
[1662] The system for implementing this invention consists of three main components: a server, a terminal, and an emotion engine. The server is responsible for acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and evaluating pronunciation. The terminal functions as a learner interface, playing audio, displaying visual content, recording pronunciation, and transmitting data. The emotion engine also recognizes the user's emotions and adjusts learning content and feedback based on those emotions.
[1663] Processing on the server
[1664] 1. Acquiring English teaching material data
[1665] The server queries and retrieves English learning material data from a database (e.g., PostgreSQL), including TOEIC listening questions, vocabulary lists, and example sentences.
[1666] 2. Generating audio data
[1667] Based on the acquired English learning material data, the server uses a speech synthesis engine (e.g., Google Cloud Text-to-Speech API) to generate voice data with natural pronunciation and intonation.
[1668] 3. Visual content generation
[1669] The server generates visual content (e.g., illustrations or animated videos) corresponding to the generated audio data using Adobe Animate or image generation technology. For example, it generates an animated video corresponding to the sentence "The man is opening a door."
[1670] 4. Katakana String Generation
[1671] The server uses a speech analysis tool (e.g., the Julius speech engine) to convert the speech data into a katakana string, resulting in the katakana transcription "The Man is Opening a Door."
[1672] 5. Generating interval bars
[1673] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars using tools such as praat.
[1674] 6. Data transmission
[1675] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device as a JSON-formatted package.
[1676] Processing on the device
[1677] 1. Receiving and integrating data
[1678] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[1679] 2. Display the learning interface
[1680] The device uses the integrated data to display to the learner, playing audio and displaying an interface that provides visual content, katakana characters, and pitch bars.
[1681] 3. Use of Emotion Engine
[1682] The device uses an emotion engine (for example, Microsoft Azure's Emotion API) to recognize emotions from the learner's facial expressions and voice. It also analyzes the user's emotions in real time using a camera and microphone.
[1683] 4. Recording pronunciation practice
[1684] The device records the learner's pronunciation as they practice it following the audio guide.
[1685] 5. Sending pronunciation data
[1686] The device sends the recorded pronunciation data to the server via a REST API.
[1687] Processing on the server (continued)
[1688] 1. Analysis of pronunciation evaluation
[1689] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The analysis is performed using the Julius speech engine and praat.
[1690] 2. Generating evaluation results
[1691] The server converts the evaluation results obtained from the pronunciation analysis into a score and sends it to the terminal, generating feedback such as "90 points."
[1692] 3. Generate feedback
[1693] The server takes into account the emotional data and generates optimal feedback content and tone.
[1694] Processing on the terminal (continued)
[1695] 1. Viewing assessment results and providing feedback
[1696] The device displays the evaluation results received from the server to the learner, provides detailed feedback on areas that need improvement based on the evaluation results, and adjusts the tone of the feedback based on emotional data.
[1697] Specific examples
[1698] For example, in the case of the TOEIC listening question "The man is opening a door," the processing proceeds as follows:
[1699] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1700] The server generates an animated video of the scene "a man opening a door."
[1701] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[1702] This data is sent to the terminal, which then integrates it and displays it to the learner.
[1703] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[1704] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback, such as a score of 90.
[1705] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[1706] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[1707] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[1708] Example prompt: "Based on the TOEIC listening question 'The man is opening a door,' please generate natural-sounding audio data, visual content, katakana characters, and pitch bars and send them to the device."
[1709] This allows users to self-evaluate and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[1710] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1711] Step 1:
[1712] Acquisition of English teaching material data
[1713] The server retrieves English language learning material data from the database. The input is the query information stored on the server, and the output is the retrieved English language learning material data. Specifically, the server executes an SQL query to retrieve data such as TOEIC listening questions, vocabulary lists, and example sentences.
[1714] Step 2:
[1715] Generating audio data
[1716] The server uses a speech synthesis engine to generate voice data based on the acquired English language learning material data. The input is the acquired English language learning material data, and the output is the generated voice data. Specifically, the server uses the Google Cloud Text-to-Speech API to generate natural-sounding voice from text data such as "The man is opening a door."
[1717] Step 3:
[1718] Visual content generation
[1719] The server generates visual content corresponding to the audio data. The input is the generated audio data, and the output is the generated visual content. Specifically, the server uses Adobe Animate to create an animated video of the scene "a man opening a door."
[1720] Step 4:
[1721] Katakana string generation
[1722] The server converts the speech data into a katakana string. The input is the generated speech data, and the output is a katakana string. Specifically, the server uses the Julius speech engine to convert the speech data into the katakana notation "The man is opening a door."
[1723] Step 5:
[1724] Generating interval bars
[1725] The server analyzes the audio data and generates pitch bars. The input is the generated audio data, and the output is pitch bars. Specifically, the server uses praat to analyze the pitch and timing of each syllable and generate karaoke-style pitch bars.
[1726] Step 6:
[1727] Sending data
[1728] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device. The input is all the generated data, and the output is a data package. Specifically, the server compiles this data into JSON format and sends it to the device using the HTTP protocol.
[1729] Step 7:
[1730] Receiving and integrating data
[1731] The device receives and integrates the audio data, visual content, katakana text, and pitch bars sent from the server. The input is the data package sent from the server, and the output is the integrated learning content. Specifically, the device parses the JSON-formatted data and displays the audio data and visual content on a single screen.
[1732] Step 8:
[1733] View the learning interface
[1734] The device uses the integrated data to provide an interface to be displayed to the learner. The input is the integrated learning content, and the output is the displayed learning interface. Specifically, the device plays audio while synchronously drawing an interface on the screen that displays visual content, katakana characters, and pitch bars.
[1735] Step 9:
[1736] Using the Emotion Engine
[1737] The device recognizes emotions from the learner's facial expressions and voice. The input is the learner's facial and voice data, and the output is emotional data. Specifically, the device uses Microsoft Azure's Emotion API to analyze the learner's emotions in real time from facial expressions captured by the camera and voice data collected by the microphone.
[1738] Step 10:
[1739] Pronunciation practice recording
[1740] The device records the learner's pronunciation as they practice pronunciation following the audio guide. The input is the learner's pronunciation, and the output is the recorded pronunciation data. Specifically, the device uses a built-in microphone to record the learner's pronunciation into an audio file.
[1741] Step 11:
[1742] Sending pronunciation data
[1743] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the pronunciation data sent to the server. Specifically, the device uses a REST API to upload the audio file to the server.
[1744] Step 12:
[1745] Pronunciation evaluation analysis
[1746] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the recorded pronunciation data, the katakana string, and the pitch bar, and the output is the pronunciation evaluation result. Specifically, the server uses the Julius speech engine and praat to calculate the degree of match for each phoneme and convert it into a score.
[1747] Step 13:
[1748] Generating evaluation results
[1749] The server generates pronunciation evaluation results and sends them to the terminal. The input is the pronunciation evaluation results, and the output is the evaluation results sent from the server to the terminal. Specifically, the server summarizes the evaluation results concisely and sends them to the terminal using the HTTP protocol.
[1750] Step 14:
[1751] Viewing assessment results and providing feedback
[1752] The device receives the evaluation results from the server, displays them to the learner, and provides feedback. The input is the evaluation results and emotional data, and the output is the displayed feedback. Specifically, the device highlights areas that need improvement based on the evaluation results and provides flexible feedback based on the emotional data.
[1753] (Application example 2)
[1754] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1755] Conventional English learning systems are unable to provide feedback that reflects learners' emotions, limiting their ability to support effective learning. Furthermore, the feedback provided for English pronunciation evaluation and correction is fixed, making it difficult to respond individually to learners' emotions and reactions. As a result, issues such as a decline in learning motivation and a delay in learning progress have arisen.
[1756] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting to katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to the terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for acquiring emotion data based on the evaluation results, and means for displaying the pronunciation evaluation results and the emotion data. This enables flexible and individual feedback based on the learner's emotions.
[1757] "English language learning material data" refers to data on learning materials used in learning English, and includes vocabulary lists, sample sentences, listening questions, and the like.
[1758] "Speech data" refers to acoustic information generated based on acquired English language learning material data, and includes English pronunciation and intonation.
[1759] "Visual content" includes visual information such as illustrations and animated videos generated in response to audio data.
[1760] A "katakana character string" is a character string obtained by converting the generated voice data into Japanese katakana notation.
[1761] A "pitch bar" is a bar that visually represents the pitch and timing of each syllable by analyzing audio data.
[1762] A "terminal" is a device that receives the generated audio data, visual content, katakana character strings, and pitch bars, and provides them to a learner.
[1763] "Pronunciation data" refers to data that is recorded by a learner pronouncing words and sent to the server.
[1764] The "pronunciation evaluation result" is the result of analyzing the recorded pronunciation data and evaluating its accuracy.
[1765] "Emotion data" is information about emotions obtained from the learner's facial expressions, voice, etc.
[1766] "Feedback" refers to guidance and advice provided to learners based on pronunciation evaluation results and emotional data.
[1767] The system for implementing this invention mainly consists of three main components: a server, a terminal, and an emotion engine.
[1768] Server Processing
[1769] The server implements the following functions:
[1770] 1. Acquisition of English learning material data: The server acquires English learning material data from the database. The acquired data includes listening questions, vocabulary lists, sample sentences, etc.
[1771] 2. Audio data generation: The server generates audio data using a speech synthesis engine based on the acquired English learning material data. Using this generative AI model, an audio file with natural pronunciation and intonation is created.
[1772] 3. Visual content generation: The server generates visual content such as illustrations and animated videos corresponding to the audio data. Generative AI and multimedia tools are used to create content that is easy to understand visually.
[1773] 4. Katakana string generation: The server converts the audio data into Katakana strings, using a speech analysis tool to represent the English pronunciation in Japanese Katakana.
[1774] 5. Generating pitch bars: The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1775] 6. Data transmission: The server transmits the generated audio data, visual content, katakana string, and pitch bar to the terminal.
[1776] Terminal handling
[1777] The device implements the following functions:
[1778] 1. Data reception and integration: The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[1779] 2. Display of learning interface: The device displays the integrated data to the learner. It provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[1780] 3. Recording pronunciation practice: The device records the learner's pronunciation as they practice pronunciation following the audio guide.
[1781] 4. Sending pronunciation data: The device sends the recorded pronunciation data to the server.
[1782] Emotion engine processing
[1783] The emotion engine implements the following functions:
[1784] 1. Emotion Recognition: The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice. For example, the device analyzes the user's emotions in real time using a camera or microphone.
[1785] 2. Sending emotion data: The device sends the recognized emotion data to the server.
[1786] Reprocessing on the server
[1787] The server implements the following functions:
[1788] 1. Pronunciation evaluation analysis: The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1789] 2. Generating evaluation results: The server generates evaluation results using the pronunciation analysis tool and sends the evaluation results to the terminal.
[1790] 3. Feedback Generation: The server takes into account the emotional data from the emotion engine and generates the optimal feedback content and tone.
[1791] Reprocessing on the device
[1792] The device implements the following functions:
[1793] 1. Displaying assessment results and providing feedback: The device displays the assessment results received from the server to the learner. Based on the assessment results, detailed feedback is provided on which areas need improvement. Appropriate feedback content and tone are provided based on emotion recognition data from the emotion engine.
[1794] Specific examples
[1795] For example, consider a step in which a user tries on a new jacket in a virtual fitting room. The device recognizes the user's emotions from their facial expressions and sends the data to the server. The server generates feedback based on the emotion data, such as "Why don't you try on a jacket in a lighter color?" and sends it to the device. The device then provides the feedback to the user visually and audibly.
[1796] The following is an example of a prompt sentence to be used in the specific example:
[1797] Analyze an image of a user trying on a new jacket in front of a mirror. From their facial expression, they appear to be somewhat satisfied, but would like further style suggestions. Generate feedback content about the color and design of the jacket.
[1798] This significantly improves the user experience in the virtual fitting room and enables efficient and effective feedback.
[1799] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1800] Step 1:
[1801] The server retrieves English learning material data from the database. This data includes listening questions, vocabulary lists, sample sentences, etc. The input is learning material information in the learning material database, and the output is the retrieved information for each learning material data.
[1802] Step 2:
[1803] The server generates audio data based on the acquired English learning material data, using a generative AI model to create an audio file with natural pronunciation and intonation. The input is the learning material data, and the output is the generated audio data.
[1804] Step 3:
[1805] The server generates visual content corresponding to the generated audio data. Here, generative AI and multimedia tools are used to create visually easy-to-understand content such as illustrations and animated videos. The input is audio data, and the output is visual content.
[1806] Step 4:
[1807] The server converts the generated voice data into a katakana string. The server uses a voice analysis tool to convert the voice data into Japanese katakana. The input is the voice data, and the output is a katakana string.
[1808] Step 5:
[1809] The server analyzes the audio data and generates pitch bars, which are a visual representation of the pitch and timing of each syllable. The input is the audio data, and the output is pitch bars.
[1810] Step 6:
[1811] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. The input is each generated data, and the output is the transmission of that data to the terminal.
[1812] Step 7:
[1813] The device receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server. The input is the multiple data sent from the server, and the output is the integrated training data.
[1814] Step 8:
[1815] The terminal displays the integrated data to the learner. It provides an interface that allows the learner to refer to visual content, katakana characters, and pitch bars while playing the audio. The input is the integrated data, and the output is the display of the learning interface.
[1816] Step 9:
[1817] The device records the user's pronunciation as they practice their pronunciation following the audio guide. The input is the user's pronunciation, and the output is the recorded pronunciation data.
[1818] Step 10:
[1819] The device sends the recorded pronunciation data to the server. The input is the recorded pronunciation data, and the output is the completion of transmission to the server.
[1820] Step 11:
[1821] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar. The input is the pronunciation data and evaluation criteria (katakana string and pitch bar), and the output is the evaluation result.
[1822] Step 12:
[1823] The server generates optimal feedback content and tone based on the pronunciation evaluation results and emotional data, and sends it to the device. At this time, the server inputs the previously defined prompt sentence into the generative AI model to generate feedback. The input is the evaluation results and emotional data, and the output is the feedback content.
[1824] Step 13:
[1825] The terminal displays the evaluation results and feedback received from the server to the learner, indicating which parts were pronounced correctly and which parts need improvement. The input is the evaluation results and feedback content, and the output is the display and feedback provided to the learner.
[1826] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1827] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1828] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1829] [Fourth embodiment]
[1830] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1831] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1832] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1833] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1834] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1835] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1836] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1837] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1838] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1839] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1840] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1841] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1842] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1843] The system for implementing the present invention mainly consists of a server and a terminal. The server is responsible for multiple functions, including acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and pronunciation evaluation. Meanwhile, the terminal provides an interface for learners and is responsible for displaying audio data, visual content, katakana strings, and pitch bars, as well as recording and transmitting audio.
[1844] Main server processing
[1845] 1. Acquiring English teaching material data
[1846] The server retrieves English learning material data from the database, including sentences, vocabulary, conversation examples, test questions, etc.
[1847] 2. Generating audio data
[1848] The server uses a generative AI to generate voice data based on the acquired English learning material data, and uses a speech synthesis engine to create natural pronunciation and intonation.
[1849] 3. Visual content generation
[1850] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[1851] 4. Katakana String Generation
[1852] The server converts the audio data into katakana characters, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[1853] 5. Generating interval bars
[1854] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1855] 6. Data transmission
[1856] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1857] Main processing of the device
[1858] 1. Receiving and integrating data
[1859] The terminal integrates the audio data, visual content, katakana strings, and pitch bars received from the server and provides a viewable interface for the learner.
[1860] 2. Display the learning interface
[1861] The device displays an interface that allows learners to view visual content, katakana strings, and pitch bars while listening to the audio.
[1862] 3. Recording pronunciation practice
[1863] The device provides a function to record pronunciation when the learner pronounces along with the audio data.
[1864] 4. Sending pronunciation data
[1865] The device sends the recorded pronunciation data to the server.
[1866] Pronunciation assessment and feedback
[1867] 1. Analysis of pronunciation evaluation
[1868] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1869] 2. Generating evaluation results
[1870] The server uses a pronunciation analysis tool to generate an evaluation result and transmits the result to the terminal.
[1871] 3. Providing Feedback
[1872] The device displays the evaluation results received from the server to the learner and provides detailed feedback on which areas need improvement.
[1873] Specific examples
[1874] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[1875] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[1876] The server generates an animated video of the scene "a man opening a door."
[1877] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[1878] This data is sent to the terminal, which then integrates it and displays it to the learner.
[1879] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[1880] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[1881] The device then displays this feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[1882] This allows learners to learn English interactively and intuitively, encouraging continuous learning.
[1883] The processing flow will be explained below.
[1884] Understood. The process flow is explained in detail below.
[1885] Processing on the server
[1886] Step 1:
[1887] The server retrieves English learning material data from the database, such as TOEIC listening questions.
[1888] Step 2:
[1889] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with correct pronunciation and intonation.
[1890] Step 3:
[1891] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, using generative AI and multimedia tools.
[1892] Step 4:
[1893] The server converts the audio data into katakana strings, using a speech analysis tool to accurately represent the English pronunciation in katakana.
[1894] Step 5:
[1895] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[1896] Step 6:
[1897] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[1898] Processing on the device
[1899] Step 7:
[1900] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[1901] Step 8:
[1902] The device displays the integrated data to the learner and provides an interface where the learner can refer to visual content, katakana characters, and pitch bars while listening to the audio.
[1903] Step 9:
[1904] As learners practice pronunciation following the audio guide, the device records their pronunciation.
[1905] Step 10:
[1906] The device sends the recorded pronunciation data to the server.
[1907] Processing on the server (continued)
[1908] Step 11:
[1909] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[1910] Step 12:
[1911] The server uses a pronunciation analysis tool to generate a pronunciation evaluation result and transmits the result to the terminal.
[1912] Processing on the terminal (continued)
[1913] Step 13:
[1914] The device receives pronunciation evaluation results from the server and displays them to the learner, providing detailed feedback on areas that need improvement based on the evaluation results.
[1915] This allows learners to improve their English listening and speaking skills in an interactive and intuitive way.
[1916] Example 1
[1917] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1918] A major challenge in modern English learning is the lack of interactive and effective learning materials. In particular, the lack of speech recognition and feedback makes it difficult for learners to self-evaluate the accuracy of their pronunciation. There is also a lack of visually appealing content to keep learners engaged.
[1919] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1920] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, visual content, katakana characters, and pitch bars to a terminal, means for recording a learner's pronunciation, means for transmitting the recorded pronunciation data to the server, means for analyzing the pronunciation data and evaluating pronunciation, means for displaying the pronunciation evaluation results, means for integrating data received from the server and displaying a learning interface, and means for allowing a learner to view the visual content, katakana characters, and pitch bars while listening to audio within the learning interface. This allows learners to study English in an interactive and visually appealing environment, have their pronunciation accuracy evaluated, and receive effective feedback.
[1921] "English language learning material data" refers to information such as sentences, vocabulary, conversation examples, and test questions that learners use to study English.
[1922] "Audio data" refers to audio files with natural pronunciation and intonation that are generated based on the acquired English teaching material data.
[1923] "Visual content" refers to content that learners can refer to visually, such as illustrations or animated videos that correspond to audio data.
[1924] A "katakana string" is a representation of English pronunciation analyzed from audio data written in Japanese katakana characters.
[1925] The "pitch bar" is a karaoke-style bar that visually displays the pitch and timing of each syllable in the audio data.
[1926] A "terminal" is a device used by a learner (e.g., a PC, tablet, smartphone, etc.).
[1927] The "server" is a computer system responsible for acquiring English language teaching material data, generating audio data and visual content, analyzing katakana character strings, generating pitch bars, and managing and transmitting this data.
[1928] The "learning interface" is a user interface that allows learners to view visual content, katakana character strings, and pitch bars while listening to audio on their device.
[1929] "Pronunciation data" refers to audio data recorded by learners during pronunciation practice.
[1930] "Pronunciation evaluation" is the process of analyzing recorded pronunciation data and evaluating its accuracy.
[1931] "Feedback" is information that indicates to the learner which parts have been pronounced correctly and which parts need improvement, based on the pronunciation evaluation results.
[1932] The system for implementing the present invention is composed of a server and a terminal. The specific hardware and software and how they process data or perform calculations will be described below.
[1933] Main server processing
[1934] 1. Acquiring English teaching material data
[1935] When the server receives the request, it connects to the database to retrieve the English learning material data. The database used is a relational database management system (RDBMS) such as MySQL. For example, an SQL query such as SELECT FROM learning material WHERE ID = ? is used.
[1936] 2. Generating audio data
[1937] The server generates audio data using a speech synthesis engine based on the acquired English learning material data. The speech synthesis engine used is the Google Cloud Text-to-Speech API. The audio data is generated in MP3 or WAV format with natural pronunciation and intonation.
[1938] 3. Visual content generation
[1939] The server uses a generative AI model to generate visual content (such as illustrations and animated videos) corresponding to the English learning material data. For example, the generative AI model DALL-E generates an illustration based on the prompt "The man is opening a door," and creates an animated video using Adobe After Effects.
[1940] 4. Katakana String Generation
[1941] The server uses a speech analysis tool to convert the voice data into a katakana string, for example, Google Cloud Speech-to-Text is used to convert the voice into text, and the resulting text is then converted according to the katakana character conversion rules.
[1942] 5. Generating interval bars
[1943] The server analyzes the audio data to identify the pitch and timing of each syllable. This analysis is performed using the audio analysis tool Praat, which extracts pitch information and then converts it into pitch bars using the visual tool FFmpeg.
[1944] 6. Data transmission
[1945] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal as an HTTP response.
[1946] Main processing of the device
[1947] 1. Receiving and integrating data
[1948] The device receives the audio data, visual content, katakana characters, and pitch bars sent from the server, and the received data is integrated and prepared for display to the learner.
[1949] 2. Display the learning interface
[1950] The device uses HTML5 and JavaScript to display visual content, katakana characters, and pitch bars while the learner is playing the audio. The interface is designed to be intuitive for learners to use.
[1951] 3. Recording pronunciation practice
[1952] The device provides the ability to record pronunciation as the learner speaks along with the audio data, using standard recording features such as the browser's MediaRecorder API.
[1953] 4. Sending pronunciation data
[1954] The device sends the recorded pronunciation data to the server as an HTTP POST request, which the server then processes and analyzes.
[1955] Pronunciation assessment and feedback
[1956] 1. Analysis of pronunciation evaluation
[1957] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, using voice analysis tools such as Google Cloud Speech-to-Text.
[1958] 2. Generating evaluation results
[1959] The server uses an evaluation tool to evaluate the accuracy of the pronunciation and generate feedback information, including which parts were pronounced correctly and which parts need improvement.
[1960] 3. Providing Feedback
[1961] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and charts.
[1962] Specific examples
[1963] For example, the processing flow based on the TOEIC listening question "The man is opening a door." is as follows:
[1964] 1. The server retrieves the English teaching material data "The man is opening a door." from the database.
[1965] 2. The server uses Google Cloud Text-to-Speech to generate audio data for "The man is opening a door."
[1966] 3. The server uses DALL-E to generate an illustration of the scene of a man opening a door, and then creates an animated video using Adobe After Effects.
[1967] 4. The server uses Google Cloud Speech-to-Text to convert the audio into the katakana string "The man is opening a door."
[1968] 5. The server uses Praat to analyze pitch and timing, and creates karaoke-style pitch bars with FFmpeg.
[1969] 6. The server sends this data to the terminal.
[1970] 7. The device integrates the received data and displays it to the learner.
[1971] 8. When the user says "The Man is Opening a Door," the audio is recorded and sent to the server.
[1972] 9. The server analyzes the voice data and generates an evaluation result (e.g., 90 points).
[1973] 10. The device displays the assessment results to the learner, showing which parts of their pronunciation are correct and which parts need improvement.
[1974] Prompt Sentence Examples
[1975] The prompt sentence for generating learning content based on the English sentence "The boy is eating an apple" is as follows:
[1976] Generate learning content based on the following sentence: "The boy is eating an apple." Generate the following content:
[1977] 1. Audio data
[1978] 2. Katakana string
[1979] 3. Karaoke-style pitch bar
[1980] 4. Animated video of the scene
[1981] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1982] Step 1:
[1983] The server receives the request and connects to the database to retrieve English learning material data. This data includes sentences, vocabulary, conversation examples, test questions, etc. Specifically, it executes an SQL query (e.g., SELECT FROM learning material WHERE ID = ?). The input is the request and learning material ID, and the output is the corresponding English learning material data.
[1984] Step 2:
[1985] The server generates audio data based on the acquired English language learning material data. Using the Google Cloud Text-to-Speech API, it sends a text request and receives audio data. The input is the English language learning material data, and the output is the generated audio data (MP3 or WAV format).
[1986] Step 3:
[1987] The server generates visual content using a generative AI model. DALL-E is used to generate an illustration based on the prompt "The man is opening a door," and an animation video is created using Adobe After Effects. The input is the prompt and the AI model, and the output is the visual content (illustration and animation video).
[1988] Step 4:
[1989] The server uses a speech analysis tool to convert the voice data into katakana strings. Google Cloud Speech-to-Text is used to convert the voice data into text, which is then further converted according to katakana conversion rules. The input is voice data, and the output is katakana strings.
[1990] Step 5:
[1991] The server analyzes audio data and generates pitch bars. It uses Praat to analyze the pitch and timing of the audio, and then converts it into visual pitch bars using FFmpeg. The input is audio data, and the output is pitch bars.
[1992] Step 6:
[1993] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the terminal. It packages these data as an HTTP response and delivers it to the terminal. The input is the set of generated data, and the output is the transmission packet.
[1994] Step 7:
[1995] The device analyzes and integrates the data received from the server and displays it on the learning interface. This uses HTML5 and JavaScript to provide a user interface that learners can operate intuitively. The input is the data received from the server, and the output is the screen display of the learning interface.
[1996] Step 8:
[1997] When the user speaks in sync with the audio data, the device records the pronunciation. It uses the browser's MediaRecorder API to capture microphone input. The input is audio data, and the output is the recorded data.
[1998] Step 9:
[1999] The device sends the recorded pronunciation data to the server as an HTTP POST request. The server receives the pronunciation data and proceeds to the next processing step. The input is the recorded data, and the output is an HTTP request.
[2000] Step 10:
[2001] The server analyzes the received pronunciation data and performs pronunciation evaluation. It extracts speech features using Google Cloud Speech-to-Text and executes the evaluation algorithm. The input is the pronunciation data, and the output is the pronunciation evaluation result.
[2002] Step 11:
[2003] The server generates an evaluation result and feedback information, which includes the correct pronunciation and areas that need improvement. The input is the output of the evaluation algorithm, and the output is the detailed evaluation result and feedback information.
[2004] Step 12:
[2005] The device receives the evaluation results from the server and displays them to the learner, specifically indicating which areas need improvement. The evaluation results are displayed visually in an easy-to-understand manner using colors and diagrams. The input is the evaluation results, and the output is the feedback displayed on the learning interface.
[2006] (Application example 1)
[2007] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2008] Conventional English learning systems and product description systems often do not provide sufficient support for learners and customers when practicing pronunciation or receiving feedback. In particular, in physical stores, there is a lack of interactive systems that allow non-native speakers to obtain accurate information and practice pronunciation appropriately when purchasing products. This results in reduced learning and purchasing efficiency for learners and customers, and reduced satisfaction in physical stores.
[2009] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2010] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for acquiring product information, means for generating English audio guidance based on the acquired product information, means for generating visual content corresponding to the generated audio guidance, means for scanning QR codes, means for providing guidance information based on the product information scanned from the QR codes, and means for evaluating the pronunciation of product descriptions when they are read aloud. This enables non-native speakers to efficiently acquire product information in physical stores and learn accurate pronunciation through pronunciation practice.
[2011] "English language learning material data" refers to data that includes information such as sentences, vocabulary, conversation examples, and test questions for the purpose of learning English.
[2012] "Audio data" refers to data that includes information on voices synthesized based on English language teaching material data and guidance information.
[2013] "Visual content" refers to information that includes visual content such as illustrations and animated videos that correspond to the audio data.
[2014] A "katakana character string" is a character string that expresses the pronunciation of audio data in katakana.
[2015] A "pitch bar" is a visual bar that indicates pitch and timing based on audio data.
[2016] "Terminal" refers to a device that a user uses as an interface, including smartphones, tablets, etc.
[2017] "Pronunciation data" refers to data containing voice information generated when a user records their voice.
[2018] A "server" is a computer system that processes and stores various data online.
[2019] "Product information" is data that includes detailed product information, such as product descriptions, usage instructions, and characteristics.
[2020] "Voice guidance" is data that includes voice explanations generated based on product information.
[2021] A "QR code" is a two-dimensional barcode used to encode data, including product information.
[2022] "Guide information" refers to information including detailed product descriptions and usage instructions that can be provided by scanning the QR code.
[2023] "Reading aloud" is the act of reading aloud a specified text.
[2024] The system for implementing this invention comprises a server and a terminal. The server is responsible for various data processing and generation, and the terminal provides an interface to the user. Specific embodiments will be described below.
[2025] Server-side processing
[2026] 1. Acquiring English teaching material data
[2027] The server retrieves English learning material data from the database, which includes various learning materials such as sentences, vocabulary, conversation examples, and test questions.
[2028] 2. Generating audio data
[2029] The server uses a speech synthesis engine to generate voice data with natural pronunciation and intonation based on the acquired English learning material data, and uses a generative AI model to achieve more natural pronunciation.
[2030] 3. Visual Content Generation
[2031] The server generates visual content (such as illustrations or animated videos) corresponding to the audio data, utilizing generative AI and multimedia tools to create appropriate visual materials.
[2032] 4. Data Association and Storage
[2033] The generated audio data, visual content, katakana character strings, and pitch bars are associated and saved.
[2034] 5. Katakana String Generation
[2035] The server converts the audio data into katakana strings, and uses a speech analysis tool to accurately represent the English pronunciation in katakana.
[2036] 6. Generating interval bars
[2037] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[2038] 7. Data transmission
[2039] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[2040] 8. Analysis and Evaluation of Pronunciation Data
[2041] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[2042] Terminal side processing
[2043] 1. Receiving and integrating data
[2044] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server and provides a viewable interface for the user.
[2045] 2. Display the learning interface
[2046] The device displays an interface that allows the user to view visual content, katakana characters, and pitch bars while playing the audio.
[2047] 3. Recording pronunciation practice
[2048] The terminal provides a function for recording pronunciation when the user pronounces along with the audio data.
[2049] 4. Sending pronunciation data
[2050] The device sends the recorded pronunciation data to the server.
[2051] 5. Displaying pronunciation evaluation results
[2052] The device displays the evaluation results received from the server to the user, providing detailed feedback on which areas need improvement.
[2053] Specific examples
[2054] For example, when a user scans a QR code printed with "12345" in a physical store, the product description is played in English. The user repeats the description and records their own pronunciation, which is then sent to the server. The server analyzes the user's pronunciation and provides evaluation feedback, which is then displayed to the user on the device.
[2055] Hardware and software used
[2056] Hardware
[2057] Server (high performance computer)
[2058] Device (smartphone, tablet, etc.)
[2059] software
[2060] Database Management Systems
[2061] Speech synthesis engine
[2062] Generative AI Models
[2063] Audio analysis tools
[2064] Prompt Sentence Examples
[2065] "You scan the QR code to hear a detailed product description in English, then record yourself repeating the description and receive feedback in the app."
[2066] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2067] Step 1:
[2068] Acquisition of English teaching material data
[2069] The server retrieves English learning material data from the database. At this time, it receives a request containing information about the learning material selected by the user, and retrieves the target learning material from the database based on that request. The retrieved data includes sentences, vocabulary, conversation examples, test questions, etc.
[2070] Input: User's request to select learning materials
[2071] Output: Acquired English teaching material data
[2072] Step 2:
[2073] Generating audio data
[2074] The server generates voice data using a generative AI model and a speech synthesis engine based on the acquired English learning material data. Specifically, the learning material text is sent as input to the speech synthesis engine, which generates voice data with natural pronunciation and intonation.
[2075] Input: English teaching material data
[2076] Output: Generated audio data
[2077] Step 3:
[2078] Visual content generation
[2079] The server generates visual content (illustrations and animated videos) that correspond to the generated audio data, using generative AI and multimedia tools to create visual materials that match the audio data.
[2080] Input: Audio data
[2081] Output: The generated visual content
[2082] Step 4:
[2083] Data association and storage
[2084] The server associates the generated audio data, visual content, katakana character strings, and pitch bars and stores them in a unified manner. At this time, the association information for each data is also saved.
[2085] Input: Audio data, visual content, katakana strings, pitch bars
[2086] Output: A set of associated data
[2087] Step 5:
[2088] Katakana string generation
[2089] The server converts the voice data into a katakana string. Using a voice analysis tool, the pronunciation of the voice data is analyzed and converted into katakana notation.
[2090] Input: Audio data
[2091] Output: Katakana string
[2092] Step 6:
[2093] Generating interval bars
[2094] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates pitch bars using a voice analysis tool.
[2095] Input: Audio data
[2096] Output: Generated interval bars
[2097] Step 7:
[2098] Sending data
[2099] The server sends the generated audio data, visual content, katakana strings, and pitch bars to the device using protocols such as HTTP or WebSocket.
[2100] Input: Audio data, visual content, katakana strings, pitch bars
[2101] Output: Data sent to the terminal
[2102] Step 8:
[2103] Receiving and integrating data
[2104] The device receives and integrates the data sent from the server, providing a viewable interface for the user.
[2105] Input: Send data
[2106] Output: Unified learning interface
[2107] Step 9:
[2108] View the learning interface
[2109] The device displays an interface that allows users to view visual content, katakana characters, and pitch bars while listening to the audio, allowing them to begin learning.
[2110] Input: Integrated dataset
[2111] Output: The displayed learning interface
[2112] Step 10:
[2113] Pronunciation practice recording
[2114] Users can use the device's recording function to practice pronunciation along with the audio data, and the recorded pronunciation data is saved on the device.
[2115] Input: User pronunciation
[2116] Output: Recorded pronunciation data
[2117] Step 11:
[2118] Sending pronunciation data
[2119] The device sends the recorded pronunciation data to the server using protocols such as HTTP or WebSocket.
[2120] Input: Recorded pronunciation data
[2121] Output: Data sent to the server
[2122] Step 12:
[2123] Analysis and evaluation of pronunciation data
[2124] The server analyzes the received pronunciation data, evaluates the accuracy of the pronunciation based on the katakana string and pitch bar, and calculates an evaluation score using a pronunciation evaluation algorithm.
[2125] Input: Received pronunciation data
[2126] Output: Evaluation score
[2127] Step 13:
[2128] Displaying pronunciation evaluation results
[2129] The terminal displays the pronunciation evaluation results sent from the server to the user, allowing the user to check the accuracy of their own pronunciation and understand which parts need improvement.
[2130] Input: Rating score
[2131] Output: Display of pronunciation evaluation results
[2132] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2133] The system for implementing the present invention consists of three main components: a server, a terminal, and an emotion engine. The server acquires English learning material data, generates audio data, generates visual content, generates katakana strings, generates pitch bars, and evaluates pronunciation. The terminal functions as a learner interface and is responsible for audio playback, visual content display, pronunciation recording, and data transmission. The emotion engine recognizes the user's emotions and adjusts learning content and feedback based on them.
[2134] Processing on the server
[2135] 1. Acquiring English teaching material data
[2136] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[2137] 2. Generating audio data
[2138] The server uses a generative AI to generate audio data based on the acquired English learning material data, and a speech synthesis engine to create audio files with natural pronunciation and intonation.
[2139] 3. Visual content generation
[2140] The server generates visual content such as illustrations and animated videos corresponding to the audio data, utilizing generative AI and multimedia tools to create content that is visually easy to understand.
[2141] 4. Katakana String Generation
[2142] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[2143] 5. Generating interval bars
[2144] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[2145] 6. Data transmission
[2146] The server transmits the generated audio data, visual content, katakana characters, and pitch bars to the terminal.
[2147] Processing on the device
[2148] 1. Receiving and integrating data
[2149] The terminal receives and integrates the audio data, visual content, katakana characters, and pitch bars sent from the server.
[2150] 2. Display the learning interface
[2151] The terminal displays the integrated data to the learner, providing an interface that allows them to refer to visual content, katakana characters, and pitch bars while playing the audio.
[2152] 3. Use of Emotion Engine
[2153] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, for example, by analyzing the user's emotions in real time using a camera or microphone.
[2154] 4. Recording pronunciation practice
[2155] The device records the learner's pronunciation as they practice it following the audio guide.
[2156] 5. Sending pronunciation data
[2157] The device sends the recorded pronunciation data to the server.
[2158] Processing on the server (continued)
[2159] 6. Analysis of pronunciation evaluation
[2160] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[2161] 7. Generating evaluation results
[2162] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[2163] 8. Generate feedback
[2164] The server takes into account the emotional data from the emotion engine and generates optimal feedback content and tone.
[2165] Processing on the terminal (continued)
[2166] 9. Viewing assessment results and providing feedback
[2167] The device receives the evaluation results from the server and displays them to the learner. Based on the evaluation results, detailed feedback is provided on areas that need improvement. The content and tone of the feedback is adjusted appropriately based on emotion recognition data from the emotion engine.
[2168] Specific examples
[2169] For example, if we take the TOEIC listening question "The man is opening a door," the process would proceed as follows:
[2170] The server retrieves the relevant sentence and uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[2171] The server generates an animated video of the scene "a man opening a door."
[2172] The server converts the audio into the katakana string "The Man is Opening a Door," and then analyzes the pitch and timing of each syllable to create pitch bars.
[2173] This data is sent to the terminal, which then integrates it and displays it to the learner.
[2174] When a user says "The man is opening a door," the audio is recorded and sent to a server.
[2175] The server analyzes the audio data, evaluates the accuracy of the pronunciation, and generates feedback on a scale of 90 points.
[2176] The device recognizes the user's emotions using an emotion engine and transmits the emotion data to the server.
[2177] The server generates optimal feedback content and tones based on the emotional data and sends them back to the device.
[2178] The device displays the evaluation results and feedback to the user, showing which parts were pronounced correctly and which parts need improvement.
[2179] This allows users to self-evaluate their learning and receive flexible feedback based on their emotions, enabling them to efficiently and effectively improve their English listening and speaking skills.
[2180] The processing flow will be explained below.
[2181] Understood. The process flow is explained in detail below.
[2182] Processing on the server
[2183] Step 1:
[2184] The server retrieves English learning material data from the database, including TOEIC listening questions, vocabulary lists, and sample sentences.
[2185] Step 2:
[2186] The server uses a generative AI to generate voice data based on the acquired English learning material data, and a speech synthesis engine to create an audio file with natural pronunciation and intonation.
[2187] Step 3:
[2188] The server generates visual content corresponding to the generated audio data, including illustrations and animated videos, created using generative AI and multimedia tools.
[2189] Step 4:
[2190] The server converts the audio data into a katakana string, and uses a speech analysis tool to represent the English pronunciation in Japanese katakana.
[2191] Step 5:
[2192] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style pitch bars.
[2193] Step 6:
[2194] The server transmits the generated audio data, visual content, katakana character strings, and pitch bars to the terminal.
[2195] Processing on the device
[2196] Step 7:
[2197] The terminal integrates the audio data, visual content, katakana characters, and pitch bars received from the server.
[2198] Step 8:
[2199] The terminal displays the integrated data to the learner and serves as a learning interface, allowing the learner to view audio playback, visual content, katakana strings, and pitch bars.
[2200] Step 9:
[2201] The device uses an emotion engine to recognize emotions from the learner's facial expressions and voice, and captures emotion data in real time using a camera and microphone.
[2202] Step 10:
[2203] As learners practice pronunciation following the audio guide, their pronunciation is recorded.
[2204] Step 11:
[2205] The device transmits the recorded pronunciation data and emotion data to the server.
[2206] Processing on the server (continued)
[2207] Step 12:
[2208] The server analyzes the received pronunciation data and evaluates the accuracy of the pronunciation based on the katakana string and pitch bar.
[2209] Step 13:
[2210] The server generates an evaluation result using a pronunciation analysis tool and transmits the evaluation result to the terminal.
[2211] Step 14:
[2212] The server takes into account the emotional data and generates feedback content and tone based on the learner's emotions.
[2213] Processing on the terminal (continued)
[2214] Step 15:
[2215] The device receives the evaluation results and feedback from the server and displays them to the learner. Based on the evaluation results and emotion-based feedback, the learner is shown which areas need improvement.
[2216] Specific examples
[2217] For example, in the case of the TOEIC listening question "The man is opening a door," the process proceeds as follows:
[2218] Step 1:
[2219] The server retrieves the relevant statement from the database.
[2220] Step 2:
[2221] The server uses a speech synthesis engine to generate audio data saying "The man is opening a door."
[2222] Step 3:
[2223] The server generates an animated video of the scene "a man opening a door."
[2224] Step 4:
[2225] The server converts the speech into the katakana string "The man is opening a door."
[2226] Step 5:
[2227] The server analyzes the audio data, identifies the pitch and timing of each syllable, and creates pitch bars.
[2228] Step 6:
[2229] These data are transmitted to the terminal.
[2230] Step 7:
[2231] The terminal consolidates the received data and displays it to the learner.
[2232] Step 8:
[2233] It provides an interface that simultaneously displays audio playback, visual content, katakana characters, and pitch bars.
[2234] Step 9:
[2235] The device uses an emotion engine to acquire emotional data from the learner's facial expressions and voice, and analyzes emotions in real time using a camera and microphone.
[2236] Step 10:
[2237] When the learner says "The man is opening a door," the device records it.
[2238] Step 11:
[2239] The device sends the recorded data and emotion data to the server.
[2240] Step 12:
[2241] The server analyzes the recording and evaluates the accuracy of the pronunciation.
[2242] Step 13:
[2243] The server generates an evaluation result and transmits it to the terminal.
[2244] Step 14:
[2245] The server adjusts the content and tone of the feedback based on the emotional data, generates the feedback, and sends it to the device.
[2246] Step 15:
[2247] The device displays the assessment results and feedback to the learner, showing which parts were pronounced correctly and which parts need improvement.
[2248] This allows learners to efficiently improve their English listening and speaking skills while receiving flexible feedback that reflects their emotions.
[2249] Example 2
[2250] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2251] In conventional English learning systems, evaluation and feedback of learners' pronunciation is mechanical, making it difficult to respond flexibly to the learner's emotions and learning progress. As a result, there are problems such as a decrease in learner motivation and an impediment to effective learning. The purpose of this invention is to improve learning efficiency and motivation by recognizing learners' emotions and providing feedback accordingly.
[2252] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2253] In this invention, the server includes means for acquiring English language learning material data, means for generating audio data based on the acquired English language learning material data, means for generating visual content corresponding to the generated audio data, means for associating and saving the audio data and the visual content, means for converting the audio data into katakana characters, means for analyzing the audio data and generating pitch bars, means for transmitting the audio data, the visual content, the katakana characters, and the pitch bars to a terminal, means for recording the learner's pronunciation, means for transmitting the pronunciation data to the server, means for analyzing and evaluating the pronunciation data, means for displaying the pronunciation evaluation results, means for recognizing the learner's emotions, and means for adjusting the content and tone of the feedback based on the emotions. This makes it possible to provide flexible and effective feedback based on the learner's emotions and the accuracy of their pronunciation.
[2254] "English language learning data" is a general term for various data used in English learning, such as English listening questions, vocabulary lists, and example sentences.
[2255] "Audio data" refers to audio files with natural pronunciation and intonation based on English language learning material data generated by the server.
[2256] "Visual content" refers to content that is easy to understand visually, such as illustrations or animated videos that correspond to audio data.
[2257] A "katakana character string" is a character string in which audio data is written in Japanese katakana.
[2258] "Pitch bar" refers to a bar that visually displays the pitch and timing of each syllable in the analyzed audio data.
[2259] "Terminal" refers to a device used by a learner that plays audio, displays visual content, records pronunciation, transmits data, etc.
[2260] "Learner" refers to a user who uses this system to learn English.
[2261] An "emotion engine" is a software component that recognizes a user's emotions and adjusts learning content and feedback based on those emotions.
[2262] "Pronunciation data" refers to audio data recorded when a learner practices pronunciation.
[2263] The "pronunciation evaluation result" indicates an evaluation of the accuracy of pronunciation obtained by the server by analyzing the pronunciation data.
[2264] "Feedback" refers to information or instructions provided to learners about their learning progress and areas for improvement.
[2265] The system for implementing this invention consists of three main components: a server, a terminal, and an emotion engine. The server is responsible for acquiring English learning material data, generating audio data, generating visual content, generating katakana strings, generating pitch bars, and evaluating pronunciation. The terminal functions as a learner interface, playing audio, displaying visual content, recording pronunciation, and transmitting data. The emotion engine also recognizes the user's emotions and adjusts learning content and feedback based on those emotions.
[2266] Processing on the server
[2267] 1. Acquiring English teaching material data
[2268] The server queries and retrieves English learning material data from a database (e.g., PostgreSQL), including TOEIC listening questions, vocabulary lists, and example sentences.
[2269] 2. Generating audio data
[2270] Based on the acquired English learning material data, the server uses a speech synthesis engine (e.g., Google Cloud Text-to-Speech API) to generate voice data with natural pronunciation and intonation.
[2271] 3. Visual content generation
[2272] The server generates visual content (e.g., illustrations or animated videos) corresponding to the generated audio data using Adobe Animate or image generation technology. For example, it generates an animated video corresponding to the sentence "The man is opening a door."
[2273] 4. Katakana String Generation
[2274] The server uses a speech analysis tool (e.g., the Julius speech engine) to convert the speech data into a katakana string, resulting in the katakana transcription "The Man is Opening a Door."
[2275] 5. Generating interval bars
[2276] The server analyzes the audio data, identifies the pitch and timing of each syllable, and generates karaoke-style p...
Claims
1. A means of obtaining English teaching material data; means for generating audio data based on the acquired English language teaching material data; means for generating visual content corresponding to the generated audio data; means for storing the audio data and the visual content in association with each other; A means for converting the text into Katakana characters; means for analyzing audio data to generate pitch bars; A means for transmitting audio data, visual content, katakana character strings, and pitch bars to a terminal; a means of recording the learner's pronunciation; means for transmitting pronunciation data to a server; means for analyzing and evaluating the pronunciation data; a means for displaying the pronunciation evaluation result; A system including:
2. The system of claim 1 further comprising means for providing feedback based on the pronunciation evaluation results.
3. The system of claim 1 , wherein the generated visual content is an animated video.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A