Information processing method, program, and information processing device
The method addresses the challenge of inappropriate intervention in language learning systems by detecting stumbling blocks and providing tailored feedback based on mastery and confidence levels, improving language learning outcomes.
Patent Information
- Application Number
- PCT/JP2025/023778
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-22
AI Technical Summary
Existing dialogue systems for language learning lack the ability to appropriately determine whether intervention in a learner's speech is necessary, particularly in identifying stumbling blocks and providing targeted feedback.
An information processing method that utilizes a server to detect stumbling blocks in a learner's speech, estimate the degree of mastery and confidence for each learning item, and generate appropriate intervention or additional question data based on these factors to support language learning.
Enables effective and targeted feedback to learners by determining the necessity of intervention and providing tailored responses, thereby enhancing language learning efficiency.
Smart Images

Figure JP2025023778_22012026_PF_FP_ABST
Abstract
Description
Information processing method, program, and information processing device
[0001] The present invention relates to an information processing method, a program, and an information processing device.
[0002] Various dialogue systems for language learning have been proposed. For example, Patent Literature 1 discloses an information processing method in which a question at one of a plurality of levels is output (uttered), an answer to the question is received from an interlocutor (learner), the level is gradually increased, and questions are output one after another, and when a breakdown such as a lack of understanding of the content or disfluency is detected from the answer, the level of the question at the time the breakdown was detected is determined to be the level of the interlocutor.
[0003] Japanese Patent Application Laid-Open No. 2023-142373
[0004] In one aspect, an object of the present invention is to provide an information processing method etc. that can appropriately determine whether or not it is necessary to intervene in a learner's speech.
[0005] The information processing method involves acquiring speech data in a language being learned by a learner, detecting stumbling blocks from the speech data, and, if a stumbling block is detected, estimating from the speech data the degree of mastery of the learning item corresponding to the stumbling block and a confidence level indicating the likelihood of that mastery, and having the computer execute a process to determine whether or not it is necessary to intervene in the learner's speech based on the confidence level.
[0006] In one aspect, it is possible to appropriately determine whether or not it is necessary to intervene in the learner's speech.
[0007] It is a diagram showing an example of the configuration of an interactive system. It is a block diagram showing an example of the configuration of a server. It is a diagram showing an example of the record layout of a learner DB and a learning item DB. It is a diagram showing an overview of an embodiment. It is a flowchart showing an example of a processing procedure executed by a server.
[0008] The present invention will be described in detail below with reference to the drawings showing embodiments thereof. (Embodiment) Fig. 1 is a diagram showing an example of the configuration of a dialogue system. In this embodiment, a dialogue system will be described that engages in a dialogue with a learner in a learning language, and that, when a predetermined stumbling point is detected in the learner's speech, determines whether or not intervention is necessary in the learner's speech and provides the learner with feedback according to the determination result. The dialogue system includes an information processing device 1 and a terminal 2. Each device is communicatively connected via a network N such as the Internet.
[0009] The information processing device 1 is an information processing device capable of various information processing and information transmission and reception, such as a server computer or a personal computer. In this embodiment, the information processing device 1 is assumed to be a server computer, and for simplicity, will be referred to as server 1 below. The server 1 conducts a dialogue with the learner in a learning language (e.g., English) in a question-and-answer format. Specifically, the server 1 outputs question data (voice) to the learner according to a predetermined scenario and accepts input of speech data (voice) from the learner as a response to the question data.
[0010] In this embodiment, the dialogue with the learner is conducted in a question and answer format, but the dialogue format is not limited to the question and answer format.
[0011] In this embodiment, the server 1 detects stumbling blocks in the learner's speech data and provides appropriate feedback (e.g., corrections to the stumbling blocks). Specifically, as described below, for each predetermined learning item, the server 1 counts the number of times the learning item was used correctly (the number of times the learning item was spoken without stumbling) and the number of times the learning item was not used correctly (the number of times the learning item was stumbling). Based on the counted number, the server 1 estimates the degree of mastery of the learning item corresponding to the stumbling block and a confidence factor indicating the likelihood of that mastery. The server 1 then determines whether or not intervention in the learner's speech is necessary based on the confidence factor, and if it determines that intervention is necessary, generates and outputs intervening utterance data for intervening in the learner's speech.
[0012] The terminal 2 is a terminal device used by the learner, such as a personal computer, a smartphone, a tablet terminal, etc. The server 1 outputs question data to the learner via the terminal 2 and accepts input of utterance data as a response to the question data.
[0013] 2 is a block diagram showing an example configuration of the server 1. The server 1 includes a control unit 11, a main memory unit 12, a communication unit 13, and an auxiliary memory unit 14. The control unit 11 has one or more arithmetic processing devices such as a central processing unit (CPU), a micro-processing unit (MPU), or a graphics processing unit (GPU), and performs various information processing, control processing, and the like by reading and executing a program P stored in the auxiliary memory unit 14. The main memory unit 12 is a temporary storage area such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), and temporarily stores data necessary for the control unit 11 to execute arithmetic processing. The communication unit 13 is a communication module for performing communication-related processing and transmits and receives information to and from the outside.
[0014] The auxiliary storage unit 14 is a non-volatile storage area such as a large-capacity memory or a hard disk, and stores a program P (program product) and other data necessary for the control unit 11 to execute processing. The auxiliary storage unit 14 also stores a learner DB 141 and a learning item DB 142. The learner DB 141 is a database that stores information about each learner. The learning item DB 142 is a database that stores pre-defined learning items and intervention methods (templates of intervention utterance data) for intervening with a learner regarding the learning items.
[0015] The auxiliary storage unit 14 may be an external storage device connected to the server 1. The server 1 may be a multi-computer consisting of multiple computers, or may be a virtual machine virtually constructed by software.
[0016] Furthermore, in this embodiment, the server 1 is not limited to the above configuration, and may include, for example, an input unit that accepts operation input, a display unit that displays images, etc. Furthermore, the server 1 may be provided with a reading unit that reads a portable storage medium 1a such as a CD (Compact Disk)-ROM or a DVD (Digital Versatile Disc)-ROM, and may read and execute the program P from the portable storage medium 1a.
[0017] FIG. 3 is a diagram showing an example of the record layout of the learner DB 141 and the learning item DB 142. As shown in FIG.
[0018] The learner DB 141 includes a learner ID column, a proficiency column, a learning goal column, a learning item column, an n column, and an N column. The learner ID column stores a learner ID for identifying each learner. The proficiency column, the learning goal column, and the learning item column store the learner's proficiency in the language being studied, learning goal, and learning item, respectively, in association with the learner ID. The proficiency column is data that expresses the learner's language ability using discrete and / or continuous values. In this embodiment, the CEFR (Common European Framework of Reference for Languages) level, which is evaluated on a six-level scale of A1, A2, B1, B2, C1, and C2, is used. In the CEFR, A1 is the lowest level and C2 is the highest level. The n column and the N column correspond to the learner ID and the learning item, respectively, and store the number of times n that the target learning item was used correctly, and the total number of times (number of samples) N that the learning item was used correctly n and the number of times m that the learning item was used incorrectly.
[0019] The learning item DB 142 includes a learning item sequence, a proficiency sequence, a confidence / establishment sequence, and an intervention template sequence. The learning item sequence stores learning items. The proficiency sequence, the confidence / establishment sequence, and the intervention template sequence each store, in association with a learning item, a proficiency, a confidence, and an establishment (defined as "high" if above a threshold and "low" if below the threshold), as well as a template for intervention utterance data.
[0020] 4 is a diagram showing an outline of the embodiment, and the outline of the embodiment will be described with reference to FIG.
[0021] As already mentioned, the server 1 communicates with the learner in a question-and-answer format. Specifically, the server 1 outputs question data to the terminal 2 according to a predetermined scenario and plays back audio. The server 1 accepts input of speech data (audio) from the learner as a response to the question data.
[0022] The server 1 performs speech recognition on the input speech data, converts the speech data from speech to text, and detects stumbling points from the speech data (text and / or speech).
[0023] A stumbling block refers to an undesirable part of an utterance, such as a grammatical error, a long silence, a pronunciation error, an inconsistent discourse structure (the order of sentences or clauses, the use of conjunctions), etc. In this embodiment, as an example, grammatical errors are detected as stumbling blocks from utterance data (text).
[0024] For example, the server 1 detects stumbling points using a machine learning model as an example of a stumbling point detection means. Specifically, the server 1 detects stumbling points using GECToR (Grammatical Error Correction: Tag, Not Rewrite). GECToR is a model that identifies stumbling points (grammatical errors) in an input sentence as a sequence labeling problem and outputs a tag indicating the type of stumbling point along with a sentence corrected for the stumbling point. For example, if the sentence "I go to the store yesterday." is input, the server 1 outputs the correction result "I went to the store yesterday." and a tag "$VERB_FORM_VB_VBD" (to change to past tense) indicating the operation of correcting "go" to "went."
[0025] In this way, when utterance data is input, the server 1 detects the stumbling block using a machine learning model (correction model) that outputs data that corrects the stumbling block in the utterance data, i.e., corrected utterance data, and a tag that indicates the type of stumbling block. For example, if the utterance data acquired from the terminal 2 differs from the corrected utterance data output from the model, the server 1 detects the difference as the stumbling block.
[0026] It should be noted that the server 1 only needs to be able to appropriately detect stumbling points from the speech data, and the means for detecting them is not limited to the machine learning model (GECToR) described above.
[0027] Next, when the server 1 detects a stumbling block from the speech data, it determines whether or not it is necessary to intervene in the learner's speech based on the detection result of the stumbling block, etc. Specifically, the server 1 identifies one or more learning items included in the speech data, estimates the degree of understanding and confidence of each learning item from the detection result of the stumbling block and the identification result of the learning item, and determines the need for intervention based on the confidence.
[0028] First, the server 1 identifies the learning items included in the speech data. Learning items are key points for a learner to learn the language being studied, and a finite number of learning items are defined in the learning item DB 142 (and the learner DB 141). For example, if the language being studied is English, learning items related to grammar include whether verb conjugations are used correctly, whether the word order and position of each word are correct, etc.
[0029] For example, the server 1 identifies learning items from the corrected utterance data corrected by the above machine learning model. Specifically, the server 1 performs morphological analysis on the corrected utterance data and identifies the part of speech of each morpheme (e.g., word) that makes up the utterance data. The server 1 identifies learning items included in the utterance data based on the part-of-speech identification result. For example, if the corrected utterance data is "I went to the store yesterday," the server 1 can obtain the following information: went: verb (past tense of go) sentence structure: SVM (subject, verb, modifier) to: preposition yesterday: adverb. The server 1 searches the learning item DB 142 using this result as a query and identifies that the utterance sentence is composed of the following learning items: "Conjugation form: past tense" "Word order / position: SVM sentence structure" "Word order / position: preposition selection" "Word order / position: adverb position"
[0030] Based on the results of detecting stumbling points and identifying learning items, the server 1 counts, for each learning item, the number of times the learning item was used correctly and the number of times the learning item was used incorrectly.The server 1 then estimates the degree of mastery of the learning item corresponding to the stumbling point and a confidence factor indicating the likelihood of that mastery based on the number of times the learning item was used correctly and the number of times the learning item was used incorrectly.
[0031] Specifically, the server 1 estimates the degree of fixation according to the following formula (1).
[0032]
[0033] R represents learner u's degree of mastery of learning item i, n represents the number of times learner u was able to use learning item i correctly, m represents the number of times learner u failed to use learning item i correctly, and N represents the sum of n and m (the number of times learning item i was observed). In this way, mastery represents the percentage of times the target learning item was used correctly out of the number N of times the target learning item was observed during a conversation, including the number of times learner u stumbled (the number of times grammatical errors were made).
[0034] Since learners are constantly growing, N may not be counted from the time they start using the system, but may instead focus on the most recent phenomena and set an upper limit on the value of N, such as the most recent M observations (e.g., M = 10).
[0035] For example, if the learner's speech data is "I go to the store yesterday," and the corrected speech data is "I went to the store yesterday," no stumbling blocks are detected for the learning items "word order / position: SVM sentence pattern," "word order / position: preposition selection," and "word order / position: adverb position," so the number of times n each learning item was used correctly is incremented by 1. In contrast, a stumbling block is detected for the learning item "conjugation form: past tense," so the number of times m this learning item was used incorrectly is incremented by 1. In this way, the server 1 sequentially detects stumbling blocks from the learner's speech data to identify the learning items, and counts the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly.
[0036] For each learner and learning item, the server 1 stores (saves) in the learner DB 141 the number of times n the learning item was used correctly, and the total value N of the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly (see FIG. 3). The server 1 reads each parameter n and N from the learner DB 141 and increments the parameters n and N each time a learning item is observed. The server 1 updates the parameters n and N stored in the learner DB 141, for example, after the dialogue ends.
[0037] The server 1 calculates (estimates) the degree of mastery based on the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly according to formula (1). The server 1 classifies the calculated degree of mastery into two values based on a threshold according to formula (2) below. Note that the classification does not necessarily have to be binary.
[0038]
[0039] T R is the threshold value of the degree of mastery of the learning item i of the learner u. "0" indicates that the learning item has not been mastered, and "1" indicates that the learning item has been mastered. For example, the server 1 uses the threshold value T R is set to 0.7, and if the degree of retention is 0.7 or more, it is determined to be "mastered," and if it is less than 0.7, it is determined to be "not mastered."
[0040] This threshold may vary depending on the learner's learning goal. For example, if the learner's learning goal is to acquire everyday English conversation skills, the server 1 may consider some errors to be acceptable and set the threshold at 0.6. Alternatively, the server 1 may set the threshold at 0.8 for a learner whose learning goal is to study abroad or acquire business English conversation skills. Also, different thresholds may be set for each learning item.
[0041] The server 1 estimates the confidence level, along with the retention level, as a measure of how reliable the retention level is. For example, the confidence level is defined as the sum of the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly, i.e., the number of samples N of the learning item. For example, the server 1 classifies the confidence level into two values based on the number of samples N, according to the following formula (3). Note that the classification does not necessarily have to be binary.
[0042]
[0043] T C is the threshold of the confidence level of the learning item i of the learner u. "0" indicates a low confidence level, and "1" indicates a high confidence level. For example, the server 1 uses the threshold T C is set to 10 (M or less), and if N is 10 or more, the determination can be considered to be based on a sufficient number of samples, and therefore the certainty is determined to be high. On the other hand, if N is less than 10, the server 1 determines that the number of samples is insufficient, and therefore the certainty is low.
[0044] In this way, the server 1 estimates the degree of mastery and the degree of confidence of the learning item based on the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly. Note that this estimation method is one example, and the present embodiment is not limited to this.
[0045] For example, when utterance data is input, the server 1 may construct a machine learning model (prediction model) that outputs the degree of mastery of each learning item included in the utterance data, and use the model to estimate the degree of mastery and confidence. Specifically, the server 1 generates a model that outputs a probability value between 0 and 1 representing the degree of mastery for training utterance data (or stumbling points) using training data labeled with the correct value for the degree of mastery of the learning item. The server 1 outputs a probability value by inputting the learner's utterance data into the model. For example, the server 1 determines that the learner has "mastered" if the output probability value is equal to or greater than a predetermined threshold (e.g., 0.5), and determines that the learner has "not mastered" if the output probability value is less than the threshold.
[0046] The server 1 also estimates the confidence level by determining whether the probability value (fixation level) output from the model falls within a predetermined numerical range. For example, the server 1 sets the numerical range to 0.25 to 0.75, and determines that the confidence level is low if the probability value falls within the numerical range, and high if the probability value does not fall within the numerical range.
[0047] Note that these thresholds are examples only, and the optimal fixation and confidence thresholds may vary depending on the trained model.
[0048] As described above, the server 1 only needs to be able to estimate the degree of fixation and the degree of certainty from the speech data, and the estimation method is not limited to the method shown in formulas (1) to (3).
[0049] When the server 1 detects a stumbling point, it determines whether or not it is necessary to intervene in the learner's speech based on the confidence level estimated above. Specifically, when the confidence level (number of samples N) is equal to or greater than a threshold and is classified as "1," the server 1 determines that it is necessary to intervene in the learner's speech.
[0050] If it is determined that intervention is necessary, the server 1 generates intervening utterance data (text) for intervening in the learner's utterance and outputs the synthesized intervening utterance data (voice) to the terminal 2. In this embodiment, templates of intervening utterance data (text) associated with learning items are stored in the learning item DB 142. The server 1 generates intervening utterance data by applying the learner's utterance data to the templates.
[0051] Here, different templates are stored in the learning item DB 142 in association with the degree of understanding (and confidence level) and the level of proficiency (see FIG. 3). The server 1 selects a different template depending on the degree of understanding estimated above and the learner's level of proficiency, and generates intervening utterance data.
[0052] Following the example above, consider a case where the learner's utterance data is "I go to the store yesterday." and the corrected utterance data is "I went to the store yesterday." In this case, "go" (went) is detected as the stumbling block, and "conjugation form: past tense" is identified as the learning item. Here, if the level of retention is "low" ("not mastered") and the learner's proficiency level is "A1," the server 1 selects the template "You should use "x" when you are talking about a past event." The server 1 generates the intervention utterance data "You should use "went" when you are talking about a past event" by replacing the variable x with the corrected word "went."
[0053] On the other hand, if the retention level is "high" ("mastered"), the server 1 selects "So you meant 'x'?" as the template. The server 1 generates "So you meant 'you went to the store yesterday'?" as intervention utterance data by replacing the subject of the correction utterance data from "I" to "you" and applying it to the variable x. In this way, the server 1 directly corrects stumbling blocks when the retention level is low, but indirectly corrects when the retention level is high, allowing the learner to become aware of the problem.
[0054] Similarly, different templates are stored in the learning item DB 142 according to the learner's level of proficiency, and the server 1 generates intervention utterance data of different difficulty levels according to the learner's level of proficiency, thereby enabling intervention in a manner appropriate for the learner's level.
[0055] In the above, the intervening utterance data is generated rule-based according to a template, but the present embodiment is not limited to this. For example, the server 1 may generate the intervening utterance data using a large-scale language model. In this case, the server 1 prepares prompts associated with learning items, etc. in the learning item DB 142, adds utterance data before and after correction to the prompts, and inputs the utterance data into the large-scale language model to generate the intervening utterance data. This makes it possible to generate a variety of intervening utterance data.
[0056] In this way, when the server 1 determines that the degree of certainty is high and that intervention in the learner's utterance is necessary, it generates intervention utterance data and outputs it to the terminal 2. On the other hand, when the server 1 detects a stumbling point but determines that the degree of certainty is low, it generates additional question data (text) related to the target learning item, and outputs the synthesized additional question data (audio) to the terminal 2.
[0057] For example, the server 1 generates additional question data by inputting a predetermined prompt into a large-scale language model. For example, if the server 1 detects the stumbling block "go" in the learner's utterance data "I go to the store yesterday." and identifies the learning item "conjugation form: past tense," but the confidence level of the learning item's retention is below a threshold, the server 1 adds the corrected utterance data "I went to the store yesterday." and the learning item "conjugation form: past tense" to the prompt and inputs them into the large-scale language model to generate additional question data such as "What did you buy there?" By outputting the additional question data to the learner, the server 1 expects the learner to use the learning item "conjugation form: past tense." If the learner correctly utters the "past tense," the retention level increases. Conversely, if the learner fails to correctly utter the "past tense," the retention level decreases.
[0058] In this way, by outputting additional question data related to the learning item, the sample number N of the target learning item can be increased, and it becomes possible to appropriately determine whether or not intervention is necessary.
[0059] As described above, according to this embodiment, stumbling points are detected from the learner's speech data, the degree of mastery of the learning item corresponding to the stumbling point and its confidence level are estimated, and it is determined whether or not intervention in the learner's speech is necessary. If it is determined that intervention in the utterance is necessary, intervening utterance data is generated and output, and if it is determined that intervention in the utterance is not necessary, additional question data is generated and output. This makes it possible to provide appropriate feedback to the learner and effectively support their language learning.
[0060] Fig. 5 is a flowchart showing an example of a processing procedure executed by the server 1. The processing executed by the server 1 will be described with reference to Fig. 5. The control unit 11 of the server 1 outputs question data to the terminal 2 according to a predetermined scenario (step S11). The control unit 11 acquires utterance data in the learning language by the learner from the terminal 2 as a response to the question data (step S12).
[0061] The control unit 11 detects stumbling points from the utterance data acquired in step S11 (step S13). Specifically, the control unit 11 inputs the acquired utterance data into a machine learning model (correction model) that has been trained to output corrected utterance data in which the stumbling point has been corrected and a tag indicating the type of stumbling point when utterance data is input, thereby outputting the corrected utterance data and the tag. The control unit 11 detects the stumbling points by comparing the output corrected utterance data with the original utterance data.
[0062] The control unit 11 identifies the learning items included in the utterance data acquired in step S12 (step S14). Specifically, the control unit 11 performs morphological analysis on the corrected utterance data output from the machine learning model and identifies the part of speech and / or conjugation form of each morpheme (word). The control unit 11 identifies the learning items included in the utterance data based on the results of identifying the part of speech and / or conjugation form.
[0063] Based on the results of detecting stumbling blocks in step S13 and the results of identifying learning items in step S14, the control unit 11 counts, for each learning item, the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly (step S15). That is, if the control unit 11 does not detect a stumbling block in the utterance data corresponding to the identified learning item, the control unit 11 increments the number of times n the learning item was used correctly. On the other hand, if the control unit 11 detects a stumbling block in the utterance data corresponding to the identified learning item, the control unit 11 increments the number of times m the learning item was used incorrectly.
[0064] The control unit 11 determines whether a stumbling block has been detected from the speech data acquired in step S12 (step S16). If it determines that a stumbling block has not been detected (S16: NO), the control unit 11 determines whether to end the dialogue (whether all question data has been output according to the scenario) (step S17). If it determines not to end the dialogue (S17: NO), the control unit 11 returns the process to step S11 and outputs the next question data according to the scenario. If it determines to end the dialogue (S17: YES), the control unit 11 updates the parameters stored in the learner DB 141 (the number of times n the learning item was used correctly, and the number of samples N, which is the sum of the number of times n the learning item was used correctly and the number of times m the learning item was used incorrectly) (step S18), and ends the series of processes.
[0065] If it is determined that a stumbling block has been detected from the speech data (S16: YES), the control unit 11 estimates the degree of mastery of the learning item corresponding to the stumbling block and a confidence factor indicating the likelihood of the degree of mastery based on the number n of times the learning item was used correctly and the number m of times the learning item was used incorrectly (step S19). Specifically, the control unit 11 calculates the degree of mastery by dividing the number n of times the learning item was used correctly by the number of samples N, which is the sum of the number n of times the learning item was used correctly and the number m of times the learning item was used incorrectly. The control unit 11 also calculates the number N of samples as the confidence factor.
[0066] The control unit 11 determines whether or not it is necessary to intervene in the learner's utterance based on the confidence factor estimated in step S19 (step S20). That is, the control unit 11 determines whether or not the confidence factor is equal to or greater than a threshold value.
[0067] If it is determined that intervention is necessary (S20: YES), the control unit 11 generates intervening utterance data for intervening in the learner's utterance and outputs it to the terminal 2 (step S21). Specifically, the control unit 11 selects a different template depending on the degree of retention estimated in step S19 and the learner's proficiency in the language being studied, and generates the intervening utterance data.
[0068] If it is determined that no intervention is necessary (S20: NO), the control unit 11 generates additional question data related to the learning item and outputs it to the terminal 2 (step S22). For example, the control unit 11 generates the additional question data by inputting the dialogue history (question data and utterance data) so far, the target learning item, etc. into a large-scale language model.
[0069] After executing the process of step S21 or S22, the control unit 11 returns the process to step S12.
[0070] As described above, according to this embodiment, it is possible to appropriately determine whether or not it is necessary to intervene in the learner's speech.
[0071] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims.
[0072] The matters described in each embodiment can be combined with each other. Furthermore, the independent claims and dependent claims described in the claims can be combined with each other in any combination, regardless of the reference format. Furthermore, the claims use a format in which a claim references two or more other claims (multiple claim format), but this is not limited to this. A multiple claim (multi-multi claim) that references at least one other multiple claim may also be used.
[0073] REFERENCE SIGNS LIST 1 Server (information processing device) 11 Control unit 12 Main memory unit 13 Communication unit 14 Auxiliary memory unit P Program 141 Learner DB 142 Learning item DB 2 Terminal
Claims
1. An information processing method in which a computer executes the following processes: acquires speech data in a language being learned by a learner; detects stumbling points from the speech data; if a stumbling point is detected, estimates from the speech data the degree of mastery of the learning item corresponding to the stumbling point and a confidence level indicating the likelihood of that mastery; and determines whether or not it is necessary to intervene in the learner's speech based on the confidence level.
2. The information processing method according to claim 1, wherein, if it is determined that intervention is necessary, intervention speech data for intervening in the learner's speech is generated and output, and, if it is determined that intervention is not necessary, additional question data related to the learning item is generated and output.
3. The information processing method according to claim 2, wherein, when it is determined that intervention is necessary, different intervention utterance data is generated depending on the degree of fixation.
4. The information processing method according to claim 3, further comprising: acquiring the learner's proficiency in the language being studied; and, if it is determined that intervention is necessary, generating different intervention utterance data according to the learner's level of mastery and proficiency.
5. An information processing method as described in claim 1, wherein the acquired speech data is input into a correction model that has been trained to output corrected speech data that corrects the stumbling block when the speech data is input, thereby outputting corrected speech data, and the stumbling block is detected by comparing the speech data with the corrected speech data.
6. The information processing method of claim 1, further comprising: identifying the learning items contained in the speech data; counting, for each learning item, the number of times the learning item was used correctly and the number of times the learning item was not used correctly based on the results of detecting the stumbling points and the results of identifying the learning items; and estimating the degree of retention and the degree of confidence of the learning item corresponding to the stumbling point based on the number of times the learning item was used correctly and the number of times the learning item was not used correctly.
7. The information processing method according to claim 6, wherein the speech data is text, the part of speech or conjugation form of each morpheme constituting the text is identified, and the learning items contained in the text are identified based on the results of the part of speech or conjugation form identification.
8. The information processing method of claim 1, wherein the acquired speech data is input into a prediction model that has been trained to output the degree of retention of the learning item contained in the speech data when the speech data is input, and the degree of retention is output by inputting the acquired speech data, and the degree of certainty is estimated by determining whether the degree of retention falls within a predetermined numerical range.
9. A program that causes a computer to execute the following process: acquire speech data in a language being learned by a learner; detect stumbling points from the speech data; if a stumbling point is detected, estimate from the speech data the degree of mastery of the learning item corresponding to the stumbling point and a confidence level indicating the likelihood of that mastery; and determine whether or not it is necessary to intervene in the learner's speech based on the confidence level.
10. An information processing device having a control unit, wherein the control unit acquires speech data in a learning language by a learner, detects stumbling points from the speech data, and when a stumbling point is detected, estimates from the speech data the degree of mastery of the learning item corresponding to the stumbling point and a confidence level indicating the likelihood of the degree of mastery, and determines whether or not it is necessary to intervene in the learner's speech based on the confidence level.
Citation Information
Patent Citations
Language information apparatus
JP1999038863A
Ubiquitous learning system, portable terminal, speech tact and information providing device, and program
JP2004287229A
Robot control device, robot, robot control method, robot control system and program
JP2017173547A
Learning support system and learning support method
JP2025031067A