Information processing device and information processing method
The RAG system effectively addresses the challenge of providing adaptive pronunciation feedback by analyzing user pronunciation, determining accuracy scores, and generating tailored content, enhancing learning experiences through personalized and dynamic feedback.
Patent Information
- Application Number
- PCT/JP2024/023214
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies for improving pronunciation lack effective methods to provide personalized and adaptive content that accurately assess and address pronunciation errors, leading to inefficient learning experiences.
An information processing device and method that utilizes a Retrieval-Augmented Generation (RAG) system to analyze user pronunciation, determine accuracy scores, and generate tailored content, including text, images, and audio feedback based on pronunciation analysis, adapting feedback based on user improvement.
Provides personalized and effective content for improving pronunciation by objectively evaluating user progress and dynamically adjusting feedback to enhance learning outcomes.
Smart Images

Figure JP2024023214_02012026_PF_FP_ABST
Abstract
Description
Information processing device and information processing method
[0001] The present disclosure relates to an information processing device and an information processing method.
[0002] Patent Literature 1 discloses a device that provides feedback materials to help users improve their pronunciation. This device analyzes the frequency spectrum of the user's voice data to identify formant frequencies and derives the error from pre-stored formant frequency reference values. The device also displays the error from the formant frequency reference values as a graphic, and also displays advice materials to help users improve their pronunciation.
[0003] JP 2013-88552 A
[0004] For example, a technology for providing content to a user is known as a technology used in language learning, and in such a technology, it is required to provide effective content for improving pronunciation.
[0005] The present disclosure aims to provide effective content for improving pronunciation.
[0006] An information processing device according to one aspect of the present disclosure includes a determination unit that determines first advice information relating to advice on one pronunciation based on the accuracy of the one pronunciation based on first pronunciation information relating to the one pronunciation, and a generation unit that generates first generation information for generating first content corresponding to the one pronunciation based on the first advice information, wherein the determination unit determines second advice information relating to advice on a different pronunciation based on both the accuracy of the one pronunciation based on the first pronunciation information and the accuracy of a different pronunciation based on second pronunciation information relating to a different pronunciation different from the one pronunciation that is obtained after the generation unit generates the first generation information, and the generation unit generates second generation information for generating second content corresponding to the different pronunciation based on the second advice information.
[0007] An information processing method according to another aspect of the present disclosure is executed by a processor, and includes the steps of: determining first advice information relating to advice on one pronunciation based on the accuracy of one pronunciation based on first pronunciation information relating to the one pronunciation; generating first generation information for generating first content corresponding to the one pronunciation based on the first advice information; determining second advice information relating to advice on another pronunciation based on both the accuracy of the one pronunciation based on the first pronunciation information and the accuracy of another pronunciation based on second pronunciation information relating to another pronunciation that is different from the one pronunciation and that is obtained after the first generation information is generated; and generating second generation information for generating second content corresponding to the another pronunciation based on the second advice information.
[0008] According to the present disclosure, effective content for improving pronunciation can be provided.
[0009] 1 is a block diagram showing a system including a RAG system according to the present embodiment; FIG. 2 is a flowchart showing an operation method of the RAG system according to the present embodiment; FIG. 3 is a diagram showing an example of determining the accuracy of a pronunciation; FIG. 4 is a diagram showing an example of determining first advice information; FIG. 5 is a diagram showing an example of generating first generation information; FIG. 6 is a diagram showing an example of outputting first content; FIG. 7 is a flowchart showing an example of second content output processing in FIG. 2; FIG. 8 is a diagram showing an example of determining second advice information; FIG. 9 is a diagram showing an example of generating second generation information; and FIG. 10 is a block diagram showing the hardware configuration of a RAG system according to the present embodiment.
[0010] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.
[0011] FIG. 1 is a block diagram showing a system 10 including a RAG system 1 according to the present embodiment. The system 10 provides a user with a service for language learning, for example. The service includes, for example, displaying foreign language text, acquiring the user's voice, and outputting content for the user's language learning based on the acquired user's voice. The content is, for example, learning content for improving the user's pronunciation. The content includes, for example, text, images (including video), and audio. The system 10 includes a RAG (Retrieval-Augmented Generation) system 1 (information processing device), a terminal 2, a database 3, and a server device 4.
[0012] The RAG system 1 outputs content. The RAG system 1 is configured to be able to communicate with each of the terminal 2, database 3, and server device 4 via a network. The network includes, for example, a wireless communication network and a fixed communication network. The RAG system 1, for example, acquires pronunciation information related to the user's pronunciation. The RAG system 1 outputs content to the terminal 2 based on the acquired pronunciation information. The RAG system 1, for example, causes the terminal 2 to output content.
[0013] The terminal 2 is a device used by a user who wishes to receive a language learning service. The terminal 2 may be, for example, a personal computer, a smartphone, a tablet terminal, a game console, or the like. The type of the terminal 2 is not particularly limited. Note that while only one terminal 2 is illustrated in FIG. 1, the system 10 may include any number of terminals 2, such as two or more.
[0014] The terminal 2 includes, for example, an input unit and an output unit. The input unit receives, for example, voice input from the user. If the terminal 2 is a smartphone, the input unit is a microphone mounted on the terminal 2. The output unit outputs content output from the RAG system 1 to the user, for example. If the terminal 2 is a smartphone, the output unit is a display and speaker mounted on the terminal 2.
[0015] The database 3 stores predetermined information. For example, the database 3 stores sentences in multiple foreign languages, reference pronunciations indicating the correct pronunciation for each of the sentences in multiple foreign languages, and content. The reference pronunciations are pronunciations that serve as a reference for evaluating the accuracy of a user's pronunciation. The reference pronunciations include, for example, the pronunciations of each word included in the sentences in the foreign languages. The database 3 is stored, for example, in a device external to the RAG system 1.
[0016] The server device 4 is a device that enables the provision of content using a generative AI model 41. The server device 4 stores the generative AI model 41. The server device 4 stores, for example, one type of generative AI model 41. The generative AI model 41 is a model that can generate content in response to input of generation information (described below, for example, a prompt) including input information, according to any one or a combination of instructions, context, questions, and output formats indicated by the generation information, and return the generated content as response information.
[0017] The generative AI model 41 may be, for example, an interactive AI that includes a large-scale language model (LLM) and a user interface (UI) for interacting with the user, enabling text or voice chat with the user. Examples of such generative AI models 41 include ChatGPT, GPT (registered trademark)-3.5, GPT-4V, PaLM2, etc. The generative AI model 41 may be, for example, an AI model specialized for images (e.g., Audio2Face, DALL-E2, etc.). The generative AI model 41 may be, for example, an AI model specialized for voice (e.g., VALL-E, etc.).
[0018] The generative AI model 41 may be stored in the server device 4, or may be stored in another device connected to the server device 4 via a network so that information can be exchanged with the user via the server device 4. The generative AI model 41 may be stored in, for example, the RAG system 1. Note that although only one server device 4 is illustrated in FIG. 1, the system 10 may include multiple server devices 4.
[0019] Next, a description will be given of the functional configuration of the RAG system 1. The RAG system 1 includes an acquisition unit 11, a determination unit 12, a decision unit 13, a generation unit 14, and an output unit 15 as its functional configuration.
[0020] The acquisition unit 11 acquires pronunciation information relating to pronunciation. The pronunciation information is, for example, a voice including pronunciation. The acquisition unit 11 may acquire, for example, a waveform of the user's voice.
[0021] The determination unit 12 determines the accuracy of the pronunciation based on the speech acquired by the acquisition unit 11. The determination unit 12 calculates, for example, a score as the accuracy of the pronunciation. The score may be expressed, for example, by an integer greater than or equal to 0 and less than or equal to 100. The determination unit 12 may calculate the pronunciation score using known means.
[0022] The determination unit 12 may, for example, acquire the reference pronunciation from the database 3. The determination unit 12 may calculate a pronunciation score based on the acquired reference pronunciation and the user's voice acquired by the acquisition unit 11. For example, the determination unit 12 compares the user's pronunciation for each word included in the user's voice with the reference pronunciation stored in the database 3. The determination unit 12 determines, for each word included in the user's voice, whether the word is pronounced according to the reference pronunciation. For example, the determination unit 12 may calculate the pronunciation score by multiplying the ratio of the number of words pronounced according to the reference pronunciation to the number of words included in a sentence in a foreign language by a certain coefficient (for example, 100). The coefficient may, for example, be set in advance.
[0023] However, the method by which the determination unit 12 determines the accuracy of pronunciation is not limited to the method described above. In the above example, the determination unit 12 determines, for each word included in the user's speech, whether the word is pronounced according to the reference pronunciation. However, the determination unit 12 may determine, for each syllable included in the user's speech, whether the syllable is pronounced according to the reference pronunciation. The determination unit 12 may calculate, as the pronunciation score, a value obtained by multiplying the ratio of the number of syllables pronounced according to the reference pronunciation to the number of syllables included in the foreign language sentence by a certain coefficient (e.g., 100).
[0024] The determination unit 13 determines advice information relating to pronunciation advice based on the accuracy of pronunciation based on the pronunciation information acquired by the acquisition unit 11. The determination unit 13 determines the advice information based on, for example, the pronunciation score calculated by the determination unit 12. The advice information is, for example, an instruction to generate advice necessary to improve the user's pronunciation. Specific examples of advice information will be described later.
[0025] The generation unit 14 generates generation information for generating content according to the pronunciation, based on the advice information determined by the determination unit 13. The generation information is, for example, a prompt to be input to the generation AI model 41. Specific examples of the generation information will be described later.
[0026] The output unit 15 inputs the generation information generated by the generation unit 14 to the generation AI model 41. The output unit 15 acquires the content output from the generation AI model 41. The output unit 15 outputs the acquired content. The output unit 15, for example, outputs the content to the terminal 2. The output unit 15, for example, outputs the content to the output unit of the terminal 2.
[0027] Next, an operation method (including an example of an information processing method) of the RAG system 1 according to this embodiment will be described. FIG. 2 is a flowchart showing an example of the operation method of the RAG system 1. First, the output unit 15 displays a foreign language sentence on the terminal 2 (step S1). In step S1, the output unit 15 acquires the foreign language sentence from the database 3. The output unit 15 outputs the acquired foreign language sentence to the terminal 2. Then, the output unit 15 displays the foreign language sentence on the output unit of the terminal 2.
[0028] Next, as shown in Figure 3, the user, for example, reads out a sentence in a foreign language displayed on terminal 2 and inputs audio containing the first pronunciation (hereinafter referred to as "first audio") via the input unit of terminal 2.
[0029] Next, the acquisition unit 11 acquires pronunciation information related to the first pronunciation (hereinafter referred to as "first pronunciation information") (step S2). The first pronunciation information is, for example, a first speech including the first pronunciation. That is, in step S2, the acquisition unit 11 acquires, for example, the first speech as the first pronunciation information.
[0030] 3, the determination unit 12 determines the accuracy of the first pronunciation based on the first speech acquired in step S2 (step S3). In step S3, the determination unit 12 calculates, for example, a pronunciation score (hereinafter referred to as a "first score") as the accuracy of the first pronunciation.
[0031] In the example of FIG. 3 , the determination unit 12 acquires the reference pronunciation from the database 3. The determination unit 12 compares the acquired reference pronunciation with the user's pronunciation included in the first speech acquired in step S2. Then, the determination unit 12 determines that, among the words "He," "sold," "the," and "sheet" included in the foreign language sentence, "He," "sold," and "the" are pronounced according to the reference pronunciation, and "sheet" is not pronounced according to the reference pronunciation. The determination unit 12 calculates, as the first score, 75, which is a value obtained by multiplying the ratio of the number of words pronounced according to the reference pronunciation (e.g., 3) to the number of words included in the foreign language sentence (e.g., 4) by a coefficient (e.g., 100).
[0032] 4, the determination unit 13 determines advice information relating to advice on one pronunciation (hereinafter referred to as "first advice information") based on the accuracy of the one pronunciation based on the first pronunciation information (step S4). In step S4, the determination unit 13 determines the first advice information based on, for example, the first score calculated in step S3.
[0033] For example, if the score calculated in step S3 is equal to or greater than a certain threshold, it is possible that the user's pronunciation is generally correct and that the pronunciation can be sufficiently improved by providing a text explanation of the correct pronunciation. In this case, the determination unit 13 determines, for example, an instruction to generate a text explanation as the advice information. On the other hand, if the score calculated in step S3 is less than a certain threshold, it is possible that the user's pronunciation contains many inaccuracies and that the pronunciation cannot be improved by a text explanation. It is therefore considered reasonable to encourage the user to improve their pronunciation by outputting an image showing the correct way to use their mouths to the user. In this case, the determination unit 13 determines, for example, an instruction to generate an image explanation as the advice information.
[0034] In the above example, the determination unit 13 determines the advice information based on rules. However, the determination unit 13 may determine the advice information using a machine learning model generated in advance. The machine learning model may be generated, for example, using data including the user's voice as an explanatory variable and data including a pronunciation score difference before and after content is output to the user as a target variable. The machine learning model may be stored in, for example, the RAG system 1. The determination unit 13 inputs the user's voice into the machine learning model to obtain the pronunciation score difference for each piece of content. The determination unit 13 may determine, as the advice information, an instruction to generate an explanation using the content with the largest obtained score difference.
[0035] 5, the generator 14 generates generation information (hereinafter referred to as "first generation information") for generating content (hereinafter referred to as "first content") corresponding to one pronunciation based on the first advice information generated in step S4 (step S5). In the example of FIG. 5, the generation information includes a role, a task, a condition, and input information.
[0036] The role is the part that specifies the position from which advice is generated. The task is the part that specifies the required task. The condition is the part that specifies the constraints required to execute the task. The input information is the information that is the premise for executing the task.
[0037] The generating unit 14 generates, for example, first generation information including a preset role. For example, the generating unit 14 generates first generation information including a role of "foreign language teacher."
[0038] The generation unit 14 generates first generation information including, for example, a preset task. For example, the generation unit 14 generates first generation information including, as a task, an instruction to "generate advice for correct pronunciation."
[0039] The generation unit 14 determines the conditions to be included in the first generation information, for example, based on the first advice information. The generation unit 14 generates second generation information including the determined conditions. For example, when the first advice information includes an instruction to generate a text explanation, the generation unit 14 generates first generation information including, as a condition, an instruction to provide advice on pronunciation using text. For example, the generation unit 14 generates first generation information including, as a condition, an instruction to "generate text providing advice on how to pronounce parts of the sentence that differ from the student's pronunciation and the correct pronunciation to achieve the correct pronunciation."
[0040] The generation unit 14 generates first generation information that includes, as input information, the foreign language sentence acquired in step S1, the reference pronunciation acquired in step S3, the user's pronunciation contained in the first voice acquired in step S2, and the first score calculated in step S3.
[0041] Next, as shown in FIG. 6 , the output unit 15 outputs the first content (step S6). In step S6, the output unit 15 transmits the first generation information generated in step S5 to the server device 4 and inputs the first generation information to the generation AI model 41. The generation AI model 41 generates the first content using the first generation information as input. The generation AI model 41 generates the first content based on, for example, the role, task, conditions, and input information included in the first generation information. The generation AI model 41 generates, for example, a sentence that provides advice for improving the user's pronunciation that differs from the reference pronunciation. The generation AI model 41 generates the generated sentence as the first content.
[0042] In step S6, the server device 4 outputs the first content generated by the generative AI model 41 to the output unit 15. The output unit 15 acquires the first content output from the generative AI model 41. The output unit 15 outputs the acquired first content. The output unit 15, for example, outputs the first content to the terminal 2. The output unit 15, for example, causes the output unit of the terminal 2 to output the first content.
[0043] Next, the RAG system 1 determines whether the first score calculated in step S3 is equal to or greater than a first threshold (step S7). The first threshold is, for example, set in advance. If the first score is equal to or greater than the first threshold, it is considered that the user's pronunciation is generally correct and there is little need to encourage improvement of pronunciation. Therefore, if it is determined that the first score is equal to or greater than the first threshold (YES in step S7), the RAG system 1 ends the series of operations.
[0044] Conversely, if the first score is less than the first threshold, it is considered that there are many inaccuracies in the user's pronunciation and that there is a strong need to encourage improvement in pronunciation. Therefore, if it is determined that the first score is less than the first threshold (NO in step S7), the RAG system 1 executes a second content output process (step S8). The second content output process is a process of outputting content (hereinafter referred to as "second content") based on both the first pronunciation information and pronunciation information relating to another pronunciation different from the first pronunciation.
[0045] Next, the second content output process will be described in detail. Fig. 7 is a flowchart showing an example of the second content output process shown in Fig. 2. In the second content output process, first, the output unit 15 displays a sentence in a foreign language (step S11). In step S11, the output unit 15 executes the same process as in step S1. For example, the output unit 15 causes the output unit of the terminal 2 to display a sentence similar to the sentence in a foreign language displayed in step S1.
[0046] Next, the user reads out the foreign language sentence displayed on terminal 2, for example, and inputs a voice including a different pronunciation (hereinafter referred to as "second voice") via the input unit of terminal 2. At this time, the user reads out the foreign language sentence after perceiving the first content output via terminal 2, for example. The user reads out the foreign language sentence after attempting to improve their own pronunciation based on the advice provided by the first content.
[0047] Next, after the generation unit 14 generates the first generation information, the acquisition unit 11 acquires pronunciation information relating to a different pronunciation (hereinafter referred to as "second pronunciation information") (step S12). The second pronunciation information is, for example, a second voice including a different pronunciation. That is, in step S12, the acquisition unit 11 acquires, for example, the second voice as the second pronunciation information. In step S12, the acquisition unit 11 executes the same process as in step S2.
[0048] Next, the determination unit 12 determines the accuracy of the different pronunciation based on the second speech acquired in step S12 (step S13). In step S13, the determination unit 12 calculates, for example, a pronunciation score (hereinafter referred to as the "second score") as the accuracy of the different pronunciation. In step S13, the determination unit 12 executes the same process as in step S3.
[0049] 8, the determination unit 13 calculates a score difference by subtracting the first score calculated in step S3 from the second score calculated in step S13 (step S14). The score difference indicates, for example, the degree to which the user's pronunciation has improved as a result of the first content being output to the user. A larger score difference indicates a greater improvement in the user's pronunciation as a result of the output of the first content, and a smaller score difference indicates a less significant improvement in the user's pronunciation.
[0050] Next, the determination unit 13 determines whether the score difference calculated in step S14 is equal to or greater than a second threshold (a certain threshold) (step S15). The second threshold is, for example, set in advance. If it is determined that the score difference is less than the second threshold (NO in step S15), the determination unit 13 identifies content that is not included in the first content (step S16). In the example of FIG. 6, the first content is a sentence that provides advice for improving pronunciation that differs from the reference pronunciation. Therefore, the determination unit 13 identifies content other than the sentence (e.g., images and audio). If it is determined that the score difference is equal to or greater than the second threshold (YES in step S15), the determination unit 13 executes step S17.
[0051] Next, the determiner 13 determines advice information relating to advice on another pronunciation (hereinafter referred to as "second advice information") based on both the accuracy of the one pronunciation based on the first pronunciation information acquired in step S2 and the accuracy of the other pronunciation based on the second pronunciation information acquired in step S12 (step S17). In step S17, the determiner 13 determines the second advice information based on, for example, the score difference calculated in step S14. In step S17, if the calculated score difference is less than a second threshold, for example, the determiner 13 determines, as the second advice information, an instruction to generate second content including content not included in the first content identified in step S16.
[0052] In the example of Fig. 8, the determination unit 13 calculates a value of 4 as the score difference in step S14. Furthermore, if the second threshold is 10, the determination unit 13 determines in step S15 that the score difference (4) is less than the second threshold (10). Next, in step S16, the determination unit 13 identifies images and audio as content not included in the first content. For example, the determination unit 13 identifies a video of a mouth shape as content not included in the first content. For example, the determination unit 13 determines an instruction to generate an explanation using a video of a mouth shape as the second advice information.
[0053] 9, the generation unit 14 generates generation information for generating second content corresponding to a different pronunciation (hereinafter referred to as "second generation information") based on the second advice information generated in step S17 (step S18). In step S18, the generation unit 14 executes the same process as in step S5.
[0054] The generation unit 14 may generate the second generation information in a manner similar to the manner in which the first generation information is generated. In the example of Fig. 9, the generation unit 14 generates the second generation information including, for example, a role of "foreign language teacher." The generation unit 14 generates the second generation information including, for example, an instruction to "generate advice for correct pronunciation" as a task.
[0055] The generation unit 14 determines the conditions to be included in the second generation information, for example, based on the second advice information. The generation unit 14 generates the second generation information including the determined conditions. For example, if the second advice information includes an instruction to generate an explanation using a video of mouth shapes, the generation unit 14 generates second generation information including, as a condition, an instruction to provide advice on pronunciation using the video of mouth shapes. For example, the generation unit 14 generates second generation information including, as a condition, an instruction to "display a video of the mouth shapes of someone who is pronouncing correctly for parts of the student's pronunciation that are not pronounced correctly."
[0056] The generation unit 14 generates second generation information that includes, as input information, the foreign language sentence acquired in step S11, the reference pronunciation acquired in step S3, the user's pronunciation contained in the second voice acquired in step S2, and the second score calculated in step S3.
[0057] 10, the output unit 15 outputs the second content (step S19). In step S19, the output unit 15 inputs the second generation information generated in step S18 to the generation AI model 41. The generation AI model 41 generates the second content based on, for example, the role, task, condition, and input information included in the second generation information. The generation AI model 41, for example, obtains from the database 3 a video of the mouth shape of a person who is pronouncing words correctly. The generation AI model 41 generates the obtained video as the second content.
[0058] In step S19, the server device 4 outputs the second content generated by the generative AI model 41 to the output unit 15. The output unit 15 acquires the second content output from the generative AI model 41. The output unit 15 outputs the acquired second content. The output unit 15, for example, outputs the second content to the terminal 2. The output unit 15, for example, causes the output unit of the terminal 2 to output the second content. After the above processing, the RAG system 1 completes the series of operations.
[0059] Next, the effects of the RAG system 1 will be described. For example, the RAG system 1 can evaluate the degree of improvement in pronunciation based on one pronunciation before the first content is output to the user and another pronunciation after the first content is output. Furthermore, second generation information for generating second content corresponding to another pronunciation is generated based on the degree of improvement in pronunciation, so that second content corresponding to the degree of improvement in pronunciation can be output to the user. Therefore, effective content for improving pronunciation can be provided.
[0060] Furthermore, the method of operating the RAG system 1 described above has the same effects as the RAG system 1.
[0061] The first pronunciation information is a first voice including one pronunciation, and the second pronunciation information is a second voice including another pronunciation. The RAG system 1 further includes a determination unit 12 that determines the accuracy of the one pronunciation based on the first voice. The determination unit 12 determines the accuracy of the other pronunciation based on the second voice. In this case, for example, a separate device for determining the accuracy of pronunciation is not required, and the user can receive advice on pronunciation by inputting voice. Therefore, advice on pronunciation can be easily received.
[0062] The determination unit 12 calculates a first score as the accuracy of one pronunciation and a second score as the accuracy of another pronunciation. In this case, the accuracy of the speech can be evaluated numerically, making the determination of the accuracy of the speech more objective.
[0063] The determination unit 13 determines the second advice information based on the score difference obtained by subtracting the first score from the second score. In this case, by calculating the score difference as the degree of improvement in pronunciation, the degree of improvement can be expressed numerically. Because the second advice information is determined based on the score difference expressed numerically, it is possible to reduce the possibility that the process for determining the second advice information will become complicated.
[0064] When the score difference is less than the second threshold, the determiner 13 determines, as the second advice information, an instruction to generate second content including content not included in the first content. When the score difference is less than the second threshold, it can be said that the degree of improvement in pronunciation before and after the first content is output to the user is small. In this case, the first content may not have been effective for the user in improving their pronunciation. When the degree of improvement in pronunciation is small, the determiner 13 determines, as the second advice information, an instruction to generate second content including content not included in the first content, so that, for example, second content including content other than the content that was not effective for the user can be generated. Therefore, more effective content for improving pronunciation can be provided.
[0065] The RAG system 1 further includes an output unit 15 that inputs the first generation information to the generation AI model 41 and outputs the first content output from the generation AI model 41, and the output unit 15 inputs the second generation information to the generation AI model 41 and outputs the second content output from the generation AI model 41. In this case, content for improving the user's pronunciation can be provided.
[0066] Next, a modified example of the RAG system 1 will be described.
[0067] (1) In the above example, the determination unit 13 determines, as the advice information, an instruction to generate text and video as content. However, the determination unit 13 may determine, as the advice information, an instruction to generate audio as content. For example, the determination unit 13 may determine, as the first advice information, an instruction to generate first content including a third voice, which is a voice obtained by correcting one pronunciation included in the first voice. In this case, the generation unit 14 generates, for example, first generation information based on the instruction to generate first content including the third voice, the instruction including, as a condition, an instruction to "play audio in which an incorrect pronunciation of the user's pronunciation has been corrected." The generation unit 14 generates generation information that includes, as input information, the first voice and a reference pronunciation pre-stored in the database 3.
[0068] The output unit 15 inputs the first generation information generated by the generation unit 14 to the generation AI model 41. The generation AI model 41 generates a third speech based on the conditions included in the first generation information. The generation AI model 41 uses known means to correct incorrectly pronounced portions of the user's pronunciation included in the first speech, thereby generating the third speech. The server device 4 outputs the third speech generated by the generation AI model 41 to the output unit 15. The output unit 15 outputs the third speech output by the generation AI model 41 to the terminal 2. The output unit 15 outputs the third speech via the output unit of the terminal 2. In this case, the generation AI model 41 may be an AI model specialized for speech.
[0069] The determination unit 13 may determine, as the second advice information, an instruction to generate a second content including a fourth voice, which is a voice obtained by correcting another pronunciation included in the second voice, by performing a process similar to the above-mentioned process for determining the first advice information.
[0070] (2) In the above example, the server device 4 stored one type of generative AI model 41. However, the server device 4 may store multiple types of generative AI models 41. The multiple types of generative AI models 41 may include, for example, the interactive AI described above, an AI model specialized for images, and an AI model specialized for voice. The output unit 15 may determine the generative AI model 41 to which the generative information is to be input based on the content of the instruction determined as the advice information.
[0071] When the advice information includes an instruction to generate a text explanation, the output unit 15 determines an interactive AI as the generation AI model 41 to which the generation information is to be input. When the advice information includes an instruction to generate an image explanation, the output unit 15 determines an AI model specialized in images as the generation AI model 41 to which the generation information is to be input. When the advice information includes an instruction to generate content including audio, the output unit 15 determines an AI model specialized in audio as the generation AI model 41 to which the generation information is to be input. The output unit 15 inputs the generation information to the determined generation AI model 41.
[0072] (3) In the above example, the RAG system 1 includes the determination unit 12, and the acquisition unit 11 acquires speech including pronunciation as pronunciation information. However, the RAG system 1 does not need to include the determination unit 12. In this case, the acquisition unit 11 may acquire the accuracy of pronunciation as pronunciation information. For example, a user may input speech to a determination device external to the RAG system 1. The determination device may determine the accuracy of pronunciation included in the speech based on the input speech. The determination device may input the determined accuracy of pronunciation to the RAG system 1 as pronunciation information. The acquisition unit 11 may acquire the pronunciation information input by the determination device.
[0073] (4) In the above example, the determination unit 12 calculated a score as the accuracy of pronunciation. However, the determination unit 12 may not necessarily calculate a score, as long as it determines the accuracy of pronunciation. For example, the determination unit 12 may determine whether the pronunciation is accurate as the accuracy of pronunciation. For example, the determination unit 12 may acquire determination information indicating whether the pronunciation is accurate. The determination unit 13 may determine advice information based on the determination information acquired by the determination unit 12.
[0074] (5) In the above example, when the score difference is less than the second threshold, the determiner 13 determines, as the second advice information, an instruction to generate the second content including content not included in the first content. However, when the score difference is less than the second threshold, the determiner 13 may determine, as the second advice information, an instruction to generate the second content including content included in the first content.
[0075] (6) In the above example, the database 3 is stored in a device external to the RAG system 1. However, the database 3 may also be stored in the RAG system 1.
[0076] (7) In the above example, the system 10 included the RAG system 1, the terminal 2, the database 3, and the server device 4. An example was described in which each functional unit (the acquisition unit 11, the determination unit 12, the decision unit 13, the generation unit 14, and the output unit 15) was realized by processing in the RAG system 1. However, each functional unit may be realized by processing in the terminal 2. In this case, the system 10 does not need to include the RAG system 1. Also, if the database 3 is stored in the terminal 2, the system 10 does not need to include the database 3. The generative AI model 41 may be stored in the terminal 2, for example. In this case, the system 10 does not need to include the server device 4.
[0077] (8) In the above example, the RAG system 1 includes the acquisition unit 11. However, the RAG system 1 does not necessarily need to include the acquisition unit 11.
[0078] The information processing device and information processing method of the present disclosure have the following configuration.
[0079] [1] An information processing device comprising: a determination unit that determines first advice information related to advice on a first pronunciation based on accuracy of the first pronunciation based on first pronunciation information related to the first pronunciation; and a generation unit that generates first generation information for generating first content corresponding to the first pronunciation based on the first advice information, wherein the determination unit determines second advice information related to advice on the different pronunciation based on both the accuracy of the first pronunciation based on the first pronunciation information and the accuracy of the different pronunciation based on second pronunciation information related to the different pronunciation that is acquired after the generation unit generates the first generation information and is different from the first pronunciation, and the generation unit generates second generation information for generating second content corresponding to the different pronunciation based on the second advice information.
[0080] [2] The information processing device described in [1], wherein the first pronunciation information is a first voice including the one pronunciation, and the second pronunciation information is a second voice including the other pronunciation, and the information processing device further includes a judgment unit that judges the accuracy of the one pronunciation based on the first voice, and the judgment unit judges the accuracy of the other pronunciation based on the second voice.
[0081] [3] The information processing device according to [2], wherein the determination unit calculates a first score as the accuracy of the one pronunciation, and calculates a second score as the accuracy of the other pronunciation.
[0082] [4] The information processing device according to [3], wherein the determination unit determines the second advice information based on a score difference obtained by subtracting the first score from the second score.
[0083] [5] The information processing device according to [4], wherein, when the score difference is less than a threshold, the determination unit determines, as the second advice information, an instruction to generate the second content including content not included in the first content.
[0084] [6] An information processing device according to any one of [1] to [5], further comprising an output unit that inputs the first generation information into a generative AI model and outputs the first content output from the generative AI model, wherein the output unit inputs the second generation information into the generative AI model and outputs the second content output from the generative AI model.
[0085] [7] An information processing device according to any one of claims [1] to [6], wherein the first pronunciation information is a first voice including the one pronunciation, the second pronunciation information is a second voice including the other pronunciation, and the determination unit determines, as the first advice information, an instruction to generate the first content including a third voice that is a voice obtained by modifying the one pronunciation included in the first voice, and determines, as the second advice information, an instruction to generate the second content including a fourth voice that is a voice obtained by modifying the other voice included in the second voice.
[0086] [8] An information processing method executed by a processor, comprising: a step of determining first advice information relating to advice on one pronunciation based on accuracy of the one pronunciation based on first pronunciation information relating to the one pronunciation; a step of generating first generation information for generating first content corresponding to the one pronunciation based on the first advice information; a step of determining second advice information relating to advice on the other pronunciation based on both the accuracy of the one pronunciation based on the first pronunciation information and the accuracy of the other pronunciation based on second pronunciation information relating to the other pronunciation different from the one pronunciation, which is obtained after the first generation information is generated; and a step of generating second generation information for generating second content corresponding to the other pronunciation based on the second advice information.
[0087] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.
[0088] Functions include, but are not limited to, judgment, determination, discrimination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0089] 11 is a diagram showing an example of the hardware configuration of the RAG system 1 according to this embodiment. The RAG system 1 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0090] In the following description, the term "device" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the RAG system 1 may be configured to include one or more of the devices shown in the figure, or may be configured to exclude some of the devices.
[0091] Each function in the RAG system 1 is realized by loading specified software (programs) onto hardware such as a processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via a communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0092] The processor 1001, for example, runs an operating system to control the entire computer. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, at least one of the functional units of the RAG system 1 described above may be realized by the processor 1001.
[0093] The processor 1001 also reads programs (program code), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The program may be a program that causes a computer to execute at least some of the operations described in the above-described embodiments. For example, at least one of the functional units of the RAG system 1 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0094] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing the accompanying determination method according to one embodiment of the present disclosure.
[0095] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0096] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, at least one of the functional units of the RAG system 1 described above may be realized by the communication device 1004. The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated.
[0097] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that accepts input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0098] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0099] The RAG system 1 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0100] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0101] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0102] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0103] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0104] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0105] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0106] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0107] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0108] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0109] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0110] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0111] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0112] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.
[0113] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
[0114] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0115] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0116] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0117] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0118] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0119] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0120] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0121] 1...RAG system (information processing device), 11...acquisition unit, 12...judgment unit, 13...decision unit, 14...generation unit, 15...output unit, 41...generated AI model
Claims
1. An information processing device comprising: a determination unit that determines first advice information relating to advice on a first pronunciation based on the accuracy of the first pronunciation based on first pronunciation information relating to the first pronunciation; and a generation unit that generates first generation information for generating first content corresponding to the first pronunciation based on the first advice information, wherein the determination unit determines second advice information relating to advice on the different pronunciation based on both the accuracy of the first pronunciation based on the first pronunciation information and the accuracy of the different pronunciation based on second pronunciation information relating to a different pronunciation different from the first pronunciation that is obtained after the generation unit generates the first generation information, and the generation unit generates second generation information for generating second content corresponding to the different pronunciation based on the second advice information.
2. The information processing device of claim 1, wherein the first pronunciation information is a first voice including the one pronunciation, and the second pronunciation information is a second voice including the other pronunciation, and the information processing device further comprises a judgment unit that judges the accuracy of the one pronunciation based on the first voice, and the judgment unit judges the accuracy of the other pronunciation based on the second voice.
3. The information processing device according to claim 2, wherein the determination unit calculates a first score as the accuracy of the one pronunciation, and calculates a second score as the accuracy of the other pronunciation.
4. The information processing device according to claim 3, wherein the determination unit determines the second advice information based on a score difference obtained by subtracting the first score from the second score.
5. The information processing device according to claim 4, wherein the determination unit determines, when the score difference is less than a certain threshold, that an instruction to generate the second content including content not included in the first content is the second advice information.
6. An information processing device as described in claim 1, further comprising an output unit that inputs the first generation information into a generative AI model and outputs the first content output from the generative AI model, wherein the output unit inputs the second generation information into the generative AI model and outputs the second content output from the generative AI model.
7. The information processing device of claim 1, wherein the first pronunciation information is a first voice including the one pronunciation, the second pronunciation information is a second voice including the other pronunciation, and the determination unit determines, as the first advice information, an instruction to generate the first content including a third voice that is a modified version of the one pronunciation included in the first voice, and determines, as the second advice information, an instruction to generate the second content including a fourth voice that is a modified version of the other voice included in the second voice.
8. An information processing method executed by a processor, comprising: a step of determining first advice information relating to advice on one pronunciation based on accuracy of the one pronunciation based on first pronunciation information relating to the one pronunciation; a step of generating first generation information for generating first content corresponding to the one pronunciation based on the first advice information; a step of determining second advice information relating to advice on the other pronunciation based on both the accuracy of the one pronunciation based on the first pronunciation information and the accuracy of the other pronunciation based on second pronunciation information relating to a different pronunciation that is different from the one pronunciation and that is obtained after the first generation information is generated; and a step of generating second generation information for generating second content corresponding to the other pronunciation based on the second advice information.
Citation Information
Patent Citations
Language learning system
JP2006139162A
Information processing system and information processor
JP2019087276A
Information processing system, information processing server, information processing program, and information providing method
WO2015107748A1