Adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum
Patent Information
- Application Number
- KR1020260027854
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2026-02-11
- Filing Date
- 2026-02-11
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2046-02-11
Smart Images

Figure 112026018392117-PAT00007_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to an artificial intelligence-based language education system, and specifically, to an adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum that controls the learning curriculum and the response speed of an AI chatbot in real time by comprehensively analyzing the user's non-verbal input behavior and the hardware performance environment of the terminal. Background Technology
[0002] Due to the recent advancements in mobile devices and artificial intelligence technology, the market for educational technologies that enable language learning without the constraints of time and place is growing rapidly. Existing language learning applications provide content based on predetermined scenarios or engage in conversations with users through simple rule-based chatbots.
[0003] However, conventional technologies have the following fundamental limitations.
[0004] First, there is the limitation of 'result-oriented evaluation.' Existing systems determine progress to the next stage based solely on whether the user answered correctly (correct or incorrect). 'Processual difficulties,' such as the user repeatedly writing and erasing answers to find the correct one (hesitation) or thinking in silence for a long time, are excluded from the evaluation. As a result, learners who barely managed to get the correct answer frequently drop out of the next stage of advanced learning because they cannot withstand the overload.
[0005] This is a problem of 'hardware environment dependency.' A learner's response speed is heavily influenced not only by their learning ability but also by the performance of the device they use (display refresh rate, touch response speed, etc.) and the communication environment. However, conventional technology measures only the total time required without considering such hardware latency; consequently, an absurdity arises where skilled learners using older devices are misidentified as having 'learning difficulties' and forced into tedious, repetitive learning at a low difficulty level.
[0006] This is a problem of 'passive interaction.' Most learning apps merely require multiple-choice selection or simple text input, and fail to quantitatively evaluate metacognitive activities such as learners identifying and correcting errors on their own, or the attitude of speaking with confidence. This leads to the side effect of solidifying a passive attitude in learners who mechanically try to get only the correct answers.
[0007] Therefore, an adaptive language learning system is required that can gradually adjust the method of delivering learning content and the speed of interactive speech, etc., based on the learner's non-verbal behaviors (e.g., input correction, silence, re-recording), the physical environment (e.g., ambient noise), and the hardware status of the user terminal unit (e.g., screen refresh rate). Prior art literature
[65535] Korean Patent Publication No. 10-2025-0073854 The problem to be solved
[0008] The present invention has been devised to solve the problems of the aforementioned prior art, and the specific problems that the present invention aims to solve are as follows.
[0009] The first task aims to detect the level of 'cognitive friction' in real time by analyzing non-verbal behaviors such as deletion key input, silence, and re-recording that occur during the process of a learner deriving the correct answer, and to maintain the learner's cognitive load at a standard state by preemptively adjusting the speech speed or difficulty of the AI chatbot accordingly.
[0010] The second task aims to prevent evaluation distortion caused by device performance or simple errors and to ensure fairness in curriculum transfer by correcting response time by reflecting hardware performance information, such as the screen refresh cycle of the user terminal, and by applying differential penalties for incorrect answers according to the complexity of the task.
[0011] The third task aims to strengthen self-directed learning attitudes and provide a sophisticated reward system for them by introducing the concept of 'interaction density,' which comprehensively evaluates active correction behavior through cursor movement, clear confidence speech in contrast to noise, and the suitability of haptic feedback, going beyond simply determining the correct answer.
[0012] The problems of the present invention are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0014] According to an embodiment of the present invention for solving the above problem
[0015] A central server unit storing learning content, user log data, and artificial intelligence models; and
[0016] A user terminal unit that communicates with the central server unit above, receives voice and text input from a learner, and outputs learning content; comprising
[0017] In the above central server unit or the above user terminal unit,
[0018] A curriculum management module that provides learning content based on location, situation, and genre classification;
[0019] A chatbot execution module that performs a conversation with a user and adjusts response patterns;
[0020] A performance analysis module that analyzes the user's learning history; and
[0021] A reward management module that grants digital rewards based on learning performance; is included and operated,
[0022] The above chatbot execution module is,
[0023] A friction calculation unit that calculates a friction score quantifying the level of cognitive friction by monitoring non-verbal input behaviors occurring in real time during the process of a user performing a task; and
[0024] A speed control unit that controls the chatbot's speech speed (TTS Speed) by comparing the above-calculated friction score with a preset allowable threshold; is included,
[0025] The above speed control unit is,
[0026] Based on the cognitive load theory that the faster the current chatbot’s utterance speed, the more sensitively it must respond to even small signs of friction as the user’s cognitive processing capacity decreases, the above allowable threshold is dynamically lowered inversely proportional to the currently set utterance speed multiplier, and
[0027] It is characterized by generating a control signal that immediately reduces the speed of the next ignition by a preset unit when the above-calculated friction score exceeds the above-dynamically adjusted allowable threshold, and
[0028] The above friction calculation unit is,
[0029] It is characterized by collecting non-verbal log data including the number of edit keys (Backspace / Delete) entered during the user's text input process, the length of the final entered text, the ratio of silent intervals without valid utterances during voice input, and the number of times the user canceled the recording and retried, and utilizing this as basic data for calculating the friction score.
[0030] The above friction calculation unit is,
[0031] It is characterized by deriving an input efficiency index that corrects for the increase in the number of inputs according to text length by calculating the ratio of the number of edit key inputs to the length of the final input text, rather than the absolute value of the number of edit key inputs, and applying this as the first variable of the friction score.
[0032] The above friction calculation unit is,
[0033] It is characterized by calculating a silence ratio, which is the proportion of the interval during which the signal level is below a preset background noise level within the total voice input duration, and applying it as a second variable that increases the friction score by determining that the higher the silence ratio, the greater the psychological hesitation of the user.
[0034] The above speed control unit is,
[0035] By analyzing the location information of the above-mentioned user terminal unit and ambient noise data collected through the microphone sensor, it is determined whether the current learning environment is a noisy environment that is inconsistent with the environment recommended by the learning content, and
[0036] In the event that it is determined to be a noisy environment, the judgment criteria are relaxed by upwardly correcting the above allowable threshold by multiplying it by a weighting factor greater than 1.0 to prevent misjudging input delay caused by external factors as a lack of learning ability.
[0037] The above speed control unit is,
[0038] In determining the above allowable threshold, a calculation logic based on the value obtained by dividing the basic allowable value at the reference speed by the current firing speed is applied,
[0039] The present invention is characterized by controlling such that when the current firing speed is faster than the standard speed, the allowable threshold is lower than the basic allowable value to apply strict judgment criteria, and when the current firing speed is slower than the standard speed, the allowable threshold is higher than the basic allowable value to apply lenient judgment criteria.
[0040] The above speed control unit is,
[0041] When the friction score exceeds the allowable threshold and is determined to be in an overload state, the speed is reduced in increments of 0.1 times the current firing speed, and a lower limit is set to control the speed so that the reduced speed does not fall below the user's minimum guaranteed learning speed.
[0042] The above chatbot execution module is,
[0043] If the friction score continues to exceed the allowable threshold value in a continuous learning task even after deceleration measures by the speed control unit, it is determined that learning progress is difficult with only simple speed adjustment, and the system switches to a hint provision mode that visually highlights text hints on the screen or outputs an encouraging message via voice.
[0044] The above friction calculation unit is,
[0045] It is characterized by having an exception handling logic that prevents false positives caused by user mischief or mechanical rapid pressing by monitoring the user's typing speed per minute (APM) in real time and, when a section is detected where the typing speed exceeds a preset abnormal threshold, excluding the number of edit key inputs and text length data for that section from the friction score calculation data.
[0046] The above friction calculation unit is,
[0047] In the event that a hardware exception occurs in which the entire section of collected audio data is measured as Null due to a microphone sensor error or non-permission of authority of the user terminal unit, the second variable, the silence ratio, is forcibly assigned to 0, and the calculation method is automatically switched to calculate the friction score using only the first variable, the input efficiency index.
[0048] The above central server unit is,
[0049] It may be characterized by recording the task type, firing speed, and friction score data at the time when deceleration control is triggered by the speed control unit in a user-specific vulnerability database, and forming a feedback loop that, when creating a curriculum for the user in the future, sets the initial firing speed to the previously decelerated speed for tasks of a type similar to the recorded vulnerability.
[0050] In one embodiment, a language learning system including a central server unit and a user terminal unit performs a user-AI co-growth model, wherein
[0051] A step of collecting task performance data including text input and voice input of a learner through the above user terminal unit;
[0052] A step in which a chatbot execution module of the above system analyzes non-verbal input behaviors, including the number of edit key inputs and the ratio of silence intervals among the collected data, to calculate a cognitive friction score;
[0053] The above chatbot execution module retrieves speech rate information of the currently configured chatbot and calculates a dynamic allowable threshold having an inverse relationship in which it decreases as the speech rate is faster and increases as the speech rate is slower; and
[0054] It may include a step of determining whether the above-calculated cognitive friction score exceeds the above-calculated dynamic allowable threshold, and if it exceeds it, decelerating the chatbot's speech rate in real time.
[0055] In another embodiment, a method in which a chatbot execution module of a language learning system controls the speech rate according to the user's cognitive load,
[0056] A step of calculating a friction score by summing an input efficiency index, which is the ratio of the number of edit key inputs to the length of the user's final input text, and a silence ratio, which is the ratio of the silence period to the total recording time;
[0057] A step of measuring the noise level of the surrounding environment and, in the case of a noisy environment, upwardly correcting the allowable threshold value, which serves as the judgment criterion for the friction score, by assigning a weight greater than 1.0; and
[0058] It may include a step of applying strict overload judgment criteria by lowering the allowable threshold value as the speaking speed state increases, by correcting the standard allowable value through a calculation formula with the current chatbot speaking speed value as the denominator.
[0059] In another embodiment, a method for a language learning system to control a learning process based on non-linguistic data,
[0060] A preprocessing step that identifies hardware exceptions where the user's key input speed exceeds a preset abnormal range or a microphone input signal is not detected, and excludes data in the corresponding section from the friction score calculation;
[0061] If deceleration control is performed because the friction score calculated based on valid data exceeds a threshold dynamically set according to the current ignition speed, the step of recording the task information and speed information at that point in time in the database as user vulnerability; and
[0062] When configuring the next learning session, the method may include a step of resetting the initial utterance rate and whether to provide hints to the user by referring to the recorded vulnerability information. Effects of the invention
[0064] The adaptive language learning system based on the user-AI co-growth model and multi-axis curriculum according to the present invention provides the following effects.
[0065] It provides a significant reduction in learning dropout rates. Based on cognitive load theory, the AI performs dynamic control by automatically slowing down speech speed or providing hints, thereby minimizing the frustration experienced by learners during task performance and maintaining immersion. In particular, the logic of lowering the threshold in high-speed speech situations is highly effective in preventing burnout in advanced learners.
[0066] It provides the effect of ensuring objectivity and reliability in evaluation. By using 'pure cognitive time'—which subtracts the hardware latency of the terminal—as the evaluation standard, it is possible to accurately measure only language ability regardless of the user's economic environment (device specifications). Furthermore, the complexity-based weighting logic, which strictly judges failure on easy tasks as a lack of basic academic ability, prevents meaningless progress and guarantees substantial learning.
[0067] It provides the effects of enhancing metacognitive abilities and correcting attitudes. By highly evaluating 'cursor movement and partial correction' rather than deleting and re-entering everything, and rewarding 'confident speech' rather than mumbling, it encourages learners to monitor their own errors and form habits of speaking language confidently. This is an educational value that differentiates it from existing apps focused on rote memorization.
[0068] It provides the effect of efficient utilization of system resources. By verifying whether haptic feedback is distracting to the user and automatically tuning it, it reduces unnecessary battery consumption and automatically builds an optimal user experience (UX) environment.
[0069] The effects according to the present invention are not limited to those exemplified above, and a wider variety of effects are included within the present invention. Brief explanation of the drawing
[0071] Figure 1 illustrates the overall system operation flowchart according to the present invention. Figure 2 illustrates a dynamic threshold type speed control flowchart according to the present invention. FIG. 3 illustrates a flowchart of a hardware / complexity correction transition determination according to the present invention. Figure 4 illustrates an interaction density evaluation and compensation flowchart according to the present invention. FIG. 5 is a flowchart illustrating a logical processing procedure for calculating a matching suitability score according to an embodiment of the present invention. FIG. 6 is an exemplary diagram showing the main access landing page of a user-AI co-growth learning system according to one embodiment of the present invention. FIG. 7 is an example diagram of a login interface screen for user account authentication according to an embodiment of the present invention. FIG. 8 is an example diagram of a dashboard screen showing the selection and progress status of a learning unit based on a multi-axis curriculum according to one embodiment of the present invention. FIG. 9 is an example of a detailed statistical dashboard screen showing the trend of cognitive friction reduction and the overall learning progress rate visualized according to one embodiment of the present invention. FIG. 10 is an example of a system optimization popup screen that sets correction criteria by scanning hardware performance and ambient noise in real time before starting learning according to an embodiment of the present invention. FIG. 11 is an example of an interactive learning interface screen with AI speech speed control and cognitive friction (concern index) gauge applied according to one embodiment of the present invention. FIG. 12 is a screen example diagram showing the result of speech confidence analysis based on voice input and an example of text input according to one embodiment of the present invention. FIG. 13 is an example diagram of a detailed calculation result popup screen of a transition determination process with hardware delay correction and task complexity weighting applied according to an embodiment of the present invention. FIG. 14 is an example of a final learning completion report screen including active interaction analysis results and acquired rewards according to one embodiment of the present invention. FIG. 15 is an example of a notification screen for unlocking the next curriculum unit following the completion of the current learning unit according to one embodiment of the present invention. FIG. 16 is an example of a learning calendar screen that visualizes the learning performance status and achievement level by day according to an embodiment of the present invention using heatmap colors. FIG. 17 is an example of a detailed report popup screen including learning time, experience points, number of AI controls, etc., that is displayed when a specific date is selected in a learning calendar according to one embodiment of the present invention. FIG. 18 is an example of a screen showing a list of badges awarded and the progress of acquisition based on an evaluation of a user's learning attitude (refined correction, confident speech, etc.) according to an embodiment of the present invention. FIG. 19 is an example of a detailed information popup screen showing specific acquisition conditions and performance data of a specific achievement badge (sophisticated modifier) according to one embodiment of the present invention. FIG. 20 is an example of a setting screen that supports individual user settings for hardware delay correction, environmental noise adaptive logic, and haptic feedback intensity, etc., according to an embodiment of the present invention. Specific details for implementing the invention
[0072] Hereinafter, various embodiments are described in more detail with reference to the attached drawings. The embodiments described in this specification may be modified in various ways. Specific embodiments may be depicted in the drawings and described in detail in the detailed description. However, specific embodiments disclosed in the attached drawings are intended only to facilitate understanding of various embodiments. Accordingly, the technical concept is not limited by specific embodiments disclosed in the attached drawings, and it should be understood that it includes all equivalents or substitutions that fall within the spirit and scope of the invention.
[0073] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but these components are not limited by the aforementioned terms. The aforementioned terms are used solely for the purpose of distinguishing one component from another.
[0074] The functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0075] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using a number of training data by a learning algorithm, thereby creating predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0076] An artificial intelligence model can be composed of multiple neural network layers. Each of the multiple neural network layers has multiple nodes and weight values, and performs neural network operations through calculations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, multiple weights can be updated so that the loss value or cost value obtained by the artificial intelligence model during the learning process is reduced or minimized. Additionally, to minimize the loss value or cost value, multiple weights can be updated in a direction that minimizes the gradient associated with the loss value or cost value. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0077] A network is a network that serves as a transmission path for web pages; it may be a closed network such as a LAN (Local Area Network) or WAN (Wide Area Network), but it is desirable for it to be an open network such as the Internet. The Internet refers to a global open computer network structure that provides the TCP / IP protocol and various services existing at its upper layers, namely HTTP (HyperText Transfer Protocol), Telnet, FTP (File Transfer Protocol), DNS (Domain Name System), SMTP (Simple Mail Transfer Protocol), SNMP (Simple Network Management Protocol), NFS (Network File Service), and NIS (Network Information Service).
[0078] Terminals can be implemented in various forms. For example, the terminals described in this specification may include mobile terminals such as smartphones, tablet PCs, PDAs, portable multimedia players, and MP3 players, as well as fixed terminals such as smart TVs and desktop computers.
[0079] In this specification, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. When a component is described as being “connected” or “connected” to another component, it should be understood that it may be directly connected to or connected to that other component, or that there may be other components in between. On the other hand, when a component is described as being “directly connected” or “directly connected” to another component, it should be understood that there are no other components in between.
[0080] Meanwhile, a "module" or "part" for a component as used in this specification performs at least one function or operation. Furthermore, a "module" or "part" may perform a function or operation by hardware, software, or a combination of hardware and software. Additionally, a plurality of "modules" or a plurality of "parts," excluding a "module" or "part" that must be performed on specific hardware or on at least one processor, may be integrated into at least one module. A singular expression includes a plural expression unless the context clearly indicates otherwise.
[0081] In addition, power, power transmission, and control therefor for the following assembly configurations and embodiments, including "by control," follow conventional technology including terminals, applications, hardware control modules, etc., so they are omitted to avoid redundancy.
[0082] In addition, the operation embodiments and configurations described in a general manner without being explained in detail below follow the prior art and are omitted in order to focus on describing the purpose of the present invention and the resulting effects.
[0083] Furthermore, in describing the present invention, if it is determined that a detailed description of related known functions or configurations may unnecessarily obscure the essence of the invention, such detailed description is abbreviated or omitted.
[0084] An adaptive language learning system (100) based on a user-AI co-growth model and a multi-axis curriculum according to one embodiment of the present invention may include: a central server unit (110) that stores and operates large-capacity learning data and an artificial intelligence model as a physical hardware configuration; and a user terminal unit (120) that communicates with the central server unit, collects text, voice, and touch inputs from a learner, and outputs multimedia feedback.
[0085] Additionally, the central server unit (110) or the user terminal unit (120) may be configured to include, as software configurations for performing the core control logic of the present invention, a curriculum management module (130) that calculates the screen refresh cycle and task complexity of the user terminal to determine whether to approve axis movement of the curriculum; a chatbot execution module (140) that monitors the user's non-verbal input behavior in real time to calculate the cognitive friction level and dynamically controls the speech speed and difficulty accordingly; a performance analysis module (150) that diagnoses learning deficits by analyzing non-verbal logs of the execution process as well as learning results; and a reward management module (160) that grants differential digital rewards by verifying the suitability of the user's active correction behavior and haptic feedback.
[0086] The central server unit according to the present invention is defined as a high-performance computing device responsible for the pivotal data processing and storage of the system, equipped with a large-capacity learning content database, user-specific cumulative learning logs, and an artificial intelligence natural language processing model, and transmitting and receiving data to and from a plurality of user terminal units via a communication network. The central server unit performs the function of diagnosing the level of individual learners based on raw log data received from user terminal units, generating customized curriculums and chatbot response data optimized for them, and transmitting them to the terminal side. In particular, the server receives not only simple text data but also hardware specification information of the user terminal and real-time environmental noise data, and performs operations by partially incorporating edge computing technology to distribute the computational load on the server side and minimize response latency. As a method for acquiring data, the central server unit is connected to the user terminal unit through a communication interface, acquires the user's basic profile and device specification information at the start of learning, and acquires touch input coordinates, voice packets, and text input streams through real-time packet communication in 100-millisecond intervals during learning. In one embodiment, the central server unit can be implemented as a cloud-based distributed processing system and, even in a situation where there are more than 10,000 concurrent users, it loads learning history data including each user's correct answer rate, response speed, and preferred topics within 0.5 seconds, thereby providing a seamless continuing learning environment as soon as the user launches the application.
[0087] The user terminal unit connected to the central server unit described above is defined as a smartphone, tablet PC, or dedicated learning terminal comprising a touch screen, a microphone, a speaker, a haptic motor, and an accelerometer, serving as a contact device for a learner to use the service according to the present invention. This is not merely a simple input / output device, but rather acts as a sensing hub that primarily detects and processes the user's biological and physical responses and transmits them to the server. As a method of performing functions, the user terminal unit measures the physical screen refresh cycle of the display to calculate system latency, analyzes the waveform of a voice signal input via the microphone to extract the signal level relative to background noise, and drives the haptic motor to provide feedback according to the learning situation. Additionally, it records the user's touch input patterns in millisecond units along with timestamps to generate basic data for analyzing input efficiency. As a method of acquiring data, the user terminal unit calls an operating system-level API to acquire refresh rate information of the current display, accesses an audio buffer to acquire pulse code modulation data, and acquires coordinate and pressure value data in real time through a touch event listener. In one embodiment, the user terminal unit operates by distinguishing between a low-spec smartphone with a 30Hz refresh rate and a high-spec tablet PC with a 120Hz refresh rate, and in the low-spec device, it detects an input delay of about 33 milliseconds caused by a low touch sampling rate and performs a process of including metadata to correct this in a learning result packet and transmitting it to a server.
[0088] The curriculum management module, responsible for the logical operations of this system, is defined as a logical operation unit that generates learning content based on three classification criteria—location, situation, and genre—and determines whether to advance to the next stage based on the learner's achievement. This module is characterized by performing a fair evaluation by comprehensively considering the hardware performance of the user terminal and the complexity of the task. As a method of function execution, the curriculum management module calculates the pure cognitive time required by subtracting hardware latency, which is proportional to the terminal's screen refresh cycle, from the user's total task execution time. Additionally, in the event of a stage regression, it queries the complexity of the relevant task to calculate a transition resistance index by assigning a higher penalty weight for failures in simple tasks, and compares this with a reference value to approve or block the axis shift of the curriculum. As a method of data acquisition, the curriculum management module obtains task start and end timestamps, screen refresh cycle information, the identification number of the task performed by the user, and metadata regarding the complexity grade of the corresponding task from the user terminal unit. In one embodiment, when a user submits an incorrect answer while performing a simple word selection problem corresponding to a hospital situation complexity of 1, the module does not view this as a simple mistake but rather determines it as a lack of basic vocabulary and doubles the transfer resistance index. As a result, even if the user solves the problem quickly, the module controls the system so that the user is not able to proceed to the next topic, pharmacy, and is automatically guided to the vocabulary review course of the current stage.
[0089] The chatbot execution module responsible for interaction with the user is defined as an adaptive artificial intelligence agent that performs one-on-one conversations with the user and dynamically adjusts the speech rate and difficulty by monitoring the user's cognitive load state in real time. As a method of performing the function, the chatbot execution module calculates a cognitive friction score by analyzing the number of edit key inputs and silence time occurring during the user's text input process. At this time, a dynamic threshold algorithm is applied to set the friction tolerance threshold inversely lower as the current chatbot's speech rate increases, thereby performing the function of preemptively slowing down the speech rate or providing hints just before the user feels overloaded. As a method of acquiring data, the chatbot execution module acquires key event logs including input, deletion, and cursor movement in the text input window, the duration of silent periods during voice input, and the speed multiplier value of the currently set speech synthesis engine in real time. In one embodiment, when a pattern is detected in which a user presses the backspace key five or more times while inputting the correct answer in a situation where the chatbot asks a question at 1.2x speed, the module immediately lowers the speaking speed of the next sentence to 1.0x speed and outputs a care message such as "Shall I speak a little slower?" to perform control that reduces the psychological burden on the learner.
[0090] The performance analysis module, which analyzes learning data, is defined as a data analysis device that analyzes the user's learning history from the perspectives of short-term and long-term memory to derive the optimal review timing and learning gap points. As a method of performing this function, the performance analysis module does not merely record whether the answer is correct, but stores the friction score and transfer resistance index calculated by the chatbot execution module and curriculum management module in a time series. Based on this, the interval repetition system algorithm is refined to classify items with high friction scores as incomplete knowledge, even if the answer was correct, and set a short review cycle; conversely, items with low friction scores and confident answers are set a long review cycle. As a method of acquiring data, the performance analysis module obtains information regarding the correct answer, time taken, friction score, transfer resistance index, and date and time for each task accumulated in the database. In one embodiment, if the record confirms that a user correctly answered a word learned a week ago but there were three or more corrections during the input process at that time, the module classifies the word as requiring attention rather than being in a state of complete memorization. Accordingly, processing is performed to schedule the presentation of the word again in a quiz format at the beginning of today's learning session to reinforce memory.
[0091] The reward management module for motivation is defined as an incentive control device that provides digital rewards to motivate learners, but performs differential rewards by verifying the activeness of the process rather than simple results and the suitability of the system environment. As a method of performing the function, the reward management module detects active rewriting patterns in which the user moves a cursor to perform partial modifications and determines whether confident speech is being made by analyzing the root mean square amplitude of the microphone input signal. In addition, it performs the function of providing the maximum reward when all conditions are met by verifying whether the haptic feedback intensity at the time of learning was within the user's preferred range. As a method of acquiring data, the reward management module acquires key input sequence data, amplitude and duration data of audio signals, haptic motor operation logs, and user setting value data. In one embodiment, if a user speaks for at least 0.5 seconds with a clear voice that is at least 10 decibels higher than the surrounding noise in a task of pronouncing an English sentence, and it is verified that the simultaneously provided vibration feedback was appropriate, the module determines this as a master-grade performance. Accordingly, control is performed to boost the learner's self-esteem by providing 1.5 times the experience points for a normal correct answer along with flashy visual effects on the screen and awarding a confident speaker badge.
[0092] The adaptive language learning system of the present invention, configured as described above, does not merely provide learning content sequentially, but performs three core control processes to optimize the learning environment by analyzing the user's non-verbal input behavior, the hardware performance of the terminal, and active interaction patterns in real time.
[0093] First, the chatbot execution module performs a 'dynamic threshold-based speed control process based on cognitive load theory.' It quantifies the level of 'cognitive friction' by analyzing the frequency of edit key inputs and silence times that occur during the user's input of the correct answer, and preemptively slows down the speech rate or provides hints before the learner gives up by employing a logic that strictly lowers the overload judgment criterion (threshold) as the current AI's speech rate increases.
[0094] Next, the curriculum management module performs a 'transition determination process through hardware performance correction and task complexity weighting.' This prevents evaluation disadvantages caused by device performance by deriving the 'pure cognitive time' through subtracting the system latency based on the display refresh cycle from the user's task execution time. In addition, when a step backward (incorrect answer) occurs, a higher penalty is applied the lower the complexity of the corresponding task (the easier the problem), thereby precisely diagnosing whether there are deficiencies in basic learning and determining the axis shift of the curriculum.
[0095] Next, the reward management module performs an 'active interaction density evaluation process based on editing sequence and acoustic energy analysis.' This evaluates not only whether the answer is correct, but also 'metacognitive behaviors' such as moving a cursor to partially correct errors and 'confidence levels' such as speaking at a clear volume relative to background noise, and reinforces the learner's self-directed attitude by verifying the suitability of the haptic feedback provided at the time and providing differential rewards.
[0096] Below, the specific judgment criteria, data processing procedures, and control logic of the three core processes mentioned above are explained in detail.
[0097] The cognitive friction defined in this invention refers to the psychological hesitation and operational inefficiency experienced during the derivation process, regardless of whether the learner has derived the correct answer. The purpose of this process is not merely to slow down the speed upon an incorrect answer, but to provide the effect of preventing learning dropout and maintaining immersion by preemptively decelerating the speech speed of the AI through the detection of non-verbal cues before the learner's cognitive load reaches saturation.
[0098] To implement this, the system performs analysis by combining three logical axes to determine cognitive friction. First, it conducts input efficiency analysis to extract only the actual frequency of corrections by calculating the ratio of edit key inputs to the final text length, rather than simply counting the number of inputs. Additionally, it performs non-verbal delay weighted analysis by calculating the ratio of silent periods without valid utterances within the total recording time and combining this with the number of times the user voluntarily canceled and retried the recording. This indicates that the user is not unaware of the answer but is missing the utterance timing or lacking confidence. Furthermore, as a core technical feature, it implements current speed-based sensitivity control, which sets the allowable threshold for determining friction variably inversely proportional to the current AI's speech speed rather than fixing it. In other words, when the AI speaks quickly, the threshold is lowered to react immediately to even minor signs of friction because the learner has limited cognitive reserve; conversely, when the AI speaks slowly, the threshold is raised to prevent unnecessary speed reduction.
[0099] To make such a judgment, the control unit collects quantitative data in real time for every learning session, specifically acquiring data on the number of edit key inputs, final text length, silence interval ratio, number of re-recordings, current AI speech speed, and whether there is an environment tag mismatch. If the friction index calculated based on the collected data exceeds an allowable threshold, the control unit performs stepwise control; first, it immediately decelerates by lowering the speed in 0.1x increments starting from the next speech. If the friction index is maintained even after deceleration, it switches to a mode that visually highlights text hints and transmits the data to record the corresponding task type and speed combination as a user vulnerability in the database so that it can be reflected in the future curriculum configuration.
[0100] As a specific processing procedure in chronological order, while the user performs a task, the control unit analyzes key input events and voice waveforms. For example, if the user presses the backspace key 5 times and hesitates for 3 seconds while inputting a specific sentence, the control unit performs a step of normalizing this into a correction rate and a silence rate. Subsequently, the normalized correction rate and silence rate are summed, but if environmental noise is detected, a weight of 1.2 times is multiplied to calculate a basic friction score. At the same time, in a high-speed situation where the current AI speed is 1.2 times, the system strictly sets the allowable threshold to 0.8, whereas in a low-speed situation where the speed is 0.8 times, the allowable threshold is generously set to 1.5. Since the finally calculated friction score exceeds the current allowable threshold, the control unit determines this as cognitive overload and immediately generates a control signal to reduce the AI speech speed to 1.1 times.
[0101] As a measure for logic reliability and exception handling, data from sections with abnormally high keystrokes per minute is excluded from the friction score calculation to exclude cases where a user repeatedly presses keys merely for fun. Additionally, if the silence rate is measured as 1.0 due to reasons such as a microphone malfunction, the silence variable for that session is treated as 0, and friction is determined based solely on text input data. When comparing the operational results per scenario with prior art, the prior art treats a case where a user answers correctly but erases 5 times and hesitates for 10 seconds as a correct answer, thereby maintaining the speed or increasing the difficulty, which increases the likelihood of the user dropping out due to overload in the next problem. In contrast, the present invention detects potential risks by treating the same situation as high cognitive friction and adjusts the speed downward, thereby securing the user's cognitive margin and enabling sustainable learning.
[0102] Analyzing the 10 simulation examples, in the first and second iterations, the state is maintained because the friction score is below the threshold at the current speed of 1.0x. In the third iteration, as the speed accelerates to 1.2x, the threshold becomes stricter at 0.8, but the friction score is lowered to 0.5, maintaining a steady state. However, in the fifth iteration, under the same 1.2x speed environment, the friction score reaches 1.0 due to an increase in the correction ratio and silence ratio, exceeding the threshold of 0.8; thus, it is determined to be an overload, and deceleration control is applied at 1.1x speed. Subsequently, in the sixth iteration, as the speed decreases to 1.1x, the threshold is relaxed to 0.9, creating an effect where the system waits for the user. In the ninth iteration, the friction score records 1.3 under the 1.2x speed environment again, and deceleration control is performed.
[0103] According to the decision rule table, if the system is in a high-speed state and the friction score is high, it immediately performs a 0.1x speed reduction. Additionally, even if the system is in a high-speed state and the friction score is at a medium level, it immediately performs a 0.1x speed reduction if an inconsistency caused by a noisy environment is detected. Conversely, if the system is in a low-speed state and the friction score is high, it provides a hint and performs deceleration; if the system is in a low-speed state and the friction score is at a medium level, it decides to maintain the current state. The objectivity and necessity of these decision rules are not arbitrary but are based on the working memory capacity theory of cognitive psychology. Since the amount of information the human brain can process is limited, the faster the speed of input information, the more rapidly the cognitive resources available to correct errors or plan the next action decrease. Therefore, the threshold-variable logic, which sets a lower friction tolerance as the AI speed increases, possesses a technical necessity to ensure the system's learning efficiency.
[0104] As a technical effect of adopting the components, it detects and intervenes in the process of correction and hesitation before an incorrect answer occurs, thereby significantly reducing the rate at which learners feel frustrated and exit the app. Furthermore, rather than a fixed difficulty level, it adjusts the speed like a living partner according to the user's real-time condition, providing the effect of guaranteeing long-term learning sustainability. Finally, regarding the definition of terms and threshold criteria, valid text length is defined as the number of characters in a completed sentence excluding spaces. The dynamic threshold is a reference value that triggers deceleration if the friction score exceeds this value; in this invention, it is set to follow an inverse relationship by dividing a constant by the current speed. Environmental tag mismatch is defined as a state where GPS-based location information and decibel information collected by a microphone conflict with the recommended environment of the current learning content.
[0105] The threshold setting logic performed by the control unit of the present invention is based on the theory of 'working memory capacity (Miller's Law)' in cognitive psychology, rather than on arbitrary criteria by the author. It is a widely accepted theory in academia that since human short-term memory capacity has limits, the number of processable information chunks decreases inversely as the information input speed increases. Accordingly, the present invention does not use a fixed threshold but dynamically calculates the friction tolerance threshold through a function inversely proportional to the 'current AI speech rate'.
[0106] According to an exemplary computational model, the dynamic friction tolerance threshold can be calculated by dividing the [basic tolerance at reference speed] by the [current utterance speed multiplier] and multiplying it by the [environmental noise weighting factor]. For example, if the basic tolerance is 1.0, the current speed is 1.2 times, and the environment is not noisy, the dynamic threshold is calculated as approximately 0.83 (1.0 divided by 1.2), which is stricter than the standard. On the other hand, in the case of a noisy environment (weighting factor 1.2 applied), it is corrected to approximately 1.0 (0.83 multiplied by 1.2), thereby widening the tolerance so that input delay caused by noise is not misjudged as a lack of learning ability.
[0107] The table below shows 10 calculation data showing how the control unit changes the 'judgment baseline' according to the current state (speed, environment).
[0109]
[0110] The table below shows the results of verifying how the system responds by comparing the threshold set in [Table 1] above with the actual user's friction score (modification + silence).
[0111]
[0112] The pure cognitive time defined in this invention refers to the value obtained by subtracting system delay time caused by hardware factors, such as reduced display response speed or rendering delay, from the total physical time taken from the moment a task is presented on the user terminal until the user completes the input. The purpose of this process is to prevent algorithmic errors that unnecessarily lower the difficulty level or block curriculum transfer by misjudging the phenomenon of slow response due to device performance issues in learners using older terminals as a lack of learning ability, and to provide the effect of determining whether to approve the axis shift of the curriculum by objectively evaluating only the learner's actual language processing ability.
[0113] To implement this, the system performs calculations by combining two core logics to determine whether curriculum transition is required. First, the hardware latency subtraction logic is based on the physical law that the mechanical delay time from the user visually perceiving a task to registering a touch input increases proportionally as the screen refresh cycle of the user terminal lengthens. Therefore, the control unit performs processing to ensure evaluation fairness between devices with different performance capabilities by arithmetically subtracting a correction constant value proportional to the screen refresh cycle from the total measured time to derive the pure perception time. Additionally, the complexity inverse penalty logic performs processing that significantly increases the transition resistance value when a step-backward event occurs, classifying it as a fundamental learning deficit rather than a simple error if the linguistic complexity of the task is low. In other words, it applies pedagogical and logical weights that regard failure in basic, simple variation problems as a greater disqualifying factor for curriculum transition than failure in complex application problems.
[0114] To make such a judgment, the curriculum control module collects and analyzes data at each transition judgment point, specifically acquiring data on the total transition delay time, screen refresh cycle, number of step backwards, number of types of variation operations, and the cumulative number of axis transitions. The control unit determines an action based on whether the transition resistance index calculated from the above data exceeds a preset transition allowance threshold. If the resistance index is below the threshold, it sends a command to move the learning axis from the current school theme to the hospital theme and load a new vocabulary set. On the other hand, if the resistance index exceeds the threshold, it suspends the move to the new axis and performs control to maintain the advanced learning stage of the current theme or re-present similar variation problems. If the resistance index exceeds the threshold by more than twice, it performs control to immediately and forcibly insert the core pattern review curriculum of the previous stage.
[0115] As a specific processing procedure in chronological order, first, the control unit receives the screen refresh cycle from the user terminal and classifies the hardware performance grade; for example, it performs a step of classifying it as normal if the refresh cycle is 16ms or less, and as low performance if it is 30ms or more. Next, it measures the total time taken for the user to respond to the task; if the total time taken is 3,000ms and the terminal is low performance with a 33ms cycle, the control unit performs an operation to correct the pure perception time to 2,670ms by subtracting a preset delay factor, for example, 330ms (10 times the refresh cycle), from the total time taken. Next, it checks whether a step backward occurred because the user failed to solve the task, and if a backward occurred, it checks the complexity of the task. If a backward occurred in a very easy task with a complexity of 1, the control unit determines this as a critical defect and calculates the transition resistance value by multiplying it by a weight of 2.0, whereas if it was a very difficult task with a complexity of 5, it applies only a weight of 1.1. Finally, the transition resistance value reflecting the corrected time and weights is compared with the reference value to determine whether to approve the curriculum move to the next topic, the airport.
[0116] As a measure for logic reliability and exception handling, if server communication delays exceed 500ms, separate from the screen refresh cycle, processing is performed to prevent distortion caused by the network environment by either fully deducting the delay time from the total elapsed time or excluding the corresponding data sample itself from the evaluation. Additionally, if the elapsed time is measured to be abnormally long—for example, 10 seconds or more—after a user switches the app to the background and returns, exception handling is performed to classify this as a non-learning delay rather than a lack of learning ability, thereby excluding it from the transition judgment logic. When comparing the operation results by scenario with prior art, the prior art lowers the difficulty level if a user of an older phone responds within 3 seconds, judging the response as slow; however, the present invention considers device latency and judges the response as fast, maintaining or raising the difficulty level. Furthermore, if an incorrect answer occurs in a simple problem, the prior art treats it as a general error, whereas the present invention judges it as a lack of basic knowledge, strongly blocks the transition, and induces review, thereby providing an optimal curriculum tailored to the user's environment and skill level.
[0117] According to the judgment rule table, if the response is slow to a high-complexity task in a high-performance hardware state, it is judged as medium resistance and the current stage is maintained. However, if the response is slow to a low-complexity task in a high-performance hardware state, it is judged as very high resistance, the transition is blocked, and a review is performed. Conversely, if the response is slow to a high-complexity task in a low-performance hardware state, it is judged as corrected low resistance, and the transition to the next axis is approved; and if the response is fast to a low-complexity task in a low-performance hardware state, it is judged as very low resistance, and control is exercised to immediately transition to the advanced stage. The objectivity and necessity of these judgment rules are not based on arbitrary criteria but on the concept of input lag in digital signal processing theory. Furthermore, since it is a hardware engineering inevitability that the touch sampling rate decreases along with the display refresh rate, this is reflected in the evaluation criteria to perform accurate measurement of user ability. Additionally, the weighting based on complexity is based on pedagogical hierarchy theories such as Bloom's taxonomy, reflecting the fact that failure at a lower stage is more critical to learning progress than failure at a higher stage.
[0118] As a technical effect of adopting the components, consistent learning ability assessment results can be derived regardless of the user's economic situation or whether they change devices, thereby increasing data reliability. Furthermore, by clearly distinguishing between simple mistakes and lack of skill, meaningless repetitive learning is reduced, and learning efficiency is maximized by presenting a new curriculum at the optimal time when the user can take on challenges. Finally, regarding the definition of terms and threshold criteria, the screen refresh cycle is defined as the inverse of the number of times the user terminal's display redraws the screen per second, the number of transformation operation types is defined as the number of linguistic operations required to convert a single original sentence into a target sentence, and the transition allowance threshold is dynamically set in proportion to the total number of task attempts in the session as the upper limit of the resistance index that must not be exceeded to move to the next curriculum axis.
[0120] The hardware calibration logic of this process is based on response time guidelines in the field of Human-Computer Interaction (HCI). Typically, the minimum human response time to visual stimuli is approximately 250 milliseconds, but when the display refresh rate is low (e.g., 30Hz), physical input delay inevitably occurs due to frame latency, so this must be excluded from the evaluation.
[0121] According to an exemplary computational model, the pure cognitive time required is calculated by subtracting the value obtained by multiplying the [screen refresh cycle] by the [system delay correction factor] from the measured [total time required]. Additionally, the transition resistance index is proportional to the pure cognitive time required, but in the event of step backward, it is calculated by multiplying by a higher [inverse weight] as the [complexity] of the task decreases. That is, an inverse weighting method is applied such that if an easy problem (complexity 1) is answered incorrectly, the resistance index increases by a factor of 2, and if a difficult problem (complexity 5) is answered incorrectly, it increases by only a factor of 1.1.
[0122] The table below is correction data to evaluate the user's skill differently depending on the device performance, even if they responded in the same 3,000ms.
[0123]
[0124] The table below shows how the 'transition resistance index' changes according to task difficulty when step back (incorrect answer) occurs, and the system's decision accordingly. The reference threshold is assumed to be 5000.
[0125]
[0126] The active interaction density defined in this invention refers to the combined level of metacognitive behavior—where a learner goes beyond simply inputting the correct answer to independently recognizing errors, identifying specific sections, and correcting them—and vocal confidence—in which the learner speaks with conviction. The purpose of this process is to move away from the conventional method of evaluating solely based on the correctness of answers, to quantify how proactively a learner utilizes terminal tools and immerses themselves in learning, and to provide differential rewards appropriate to this, thereby enhancing a self-directed learning attitude.
[0127] To implement this, the system performs operations by combining three core logics to determine interaction density. First, editing sequence pattern matching distinguishes between passive actions, such as deleting an entire sentence by repeatedly pressing the delete key, and active actions, such as moving the cursor to select only the incorrect parts and modifying them in block units. Since the latter is a high-level cognitive operation that requires an understanding of the sentence structure, the control unit performs an operation that assigns a high weight when a sequence pattern of cursor movement followed by deletion and re-entry is detected in the input log. Additionally, acoustic energy persistence determination performs processing to analyze whether the root mean square amplitude of the microphone input signal—rather than a simple peak level—continues above a preset threshold for a certain period of time. This is a signal processing technique designed to verify that the learner's utterance is a confident and intentional vocalization, rather than noise or accidental impact sounds. In addition, regarding the verification of haptic feedback suitability, even if the learning performance is measured to be high, if the vibration feedback generated by the terminal at that time is outside the optimal immersion range, the performance may have been distorted by external stimuli; therefore, the control unit performs a correction logic that backtracks the haptic log and accepts the score calculated only when the feedback intensity was appropriate with 100% confidence.
[0128] To make such judgments, the reward management module acquires data in real time whenever a task is performed, specifically obtaining data on editing sequence logs, voice amplitude levels, effective speech duration, silence ratio, generated haptic intensity, and optimal haptic setting values. Based on the growth potential index calculated by synthesizing the above data, the control unit performs differential rewards; if multiple active editing sequences are detected, it performs a process to praise the learner's correction behavior by awarding digital badges, such as sophisticated modifiers. Additionally, if acoustic energy is high and sustained, it provides bonus experience points equivalent to 1.5 times the basic correct answer score, and if unsuitable judgments accumulate as a result of haptic suitability verification, it performs tuning to automatically adjust the output of the vibration motor to match the user's senses starting from the next session.
[0129] As a specific processing procedure in chronological order, assuming a situation where a user inputs a specific sentence and realizes it is partially incorrect, the control unit first performs the step of monitoring the key input stream. If it detects a series of actions where the user moves the cursor via touch input without long-pressing Backspace to delete the entire sentence, then presses Backspace four times to delete only the corresponding word and inputs the correct word, the control unit determines this as an active rewrite pattern and performs a calculation to add bonus points to the base score. Simultaneously, while the user reads the sentence aloud, the microphone sensor collects the audio signal, and the control unit performs the step of removing background noise from the collected signal and calculating the root mean square amplitude. If it is confirmed that the calculated amplitude is -15dB or higher and that state is maintained for 0.5 seconds or longer, it determines this as confident speech and grants additional bonus points. Finally, the vibration intensity of the terminal at the time of correct answer verification is checked; if the vibration at that time was Level 3 and the user's optimal range was between Level 2 and 4, this is recognized as a suitable feedback environment, the control unit confirms the final calculated growth potential score, and displays bonus points.
[0130] As a method for ensuring the reliability of the logic and handling exceptions, since sounds such as placing a terminal on a desk have high instantaneous decibels but short durations of less than 0.2 seconds, the control unit performs filtering to prevent errors in confidence evaluation by excluding high-volume signals with a duration of less than 0.5 seconds from the speech data. Additionally, even if the cursor is moved, if the modified content is merely a typo or a spacing correction unrelated to grammar, it is not considered a metacognitive correction, and thus pattern matching weights are not assigned; thus, exception handling is performed. When comparing the operation results for each scenario with prior art, the prior art treats the case where the answer is corrected by deleting the entire text and re-entering it after a typo identically to the case where the answer is corrected by moving the cursor and making partial corrections; however, the present invention determines the latter as a metacognitive operation and assigns a higher score. Furthermore, while the prior art makes no distinction between the case where the answer is corrected by mumbling in a quiet voice and the case where the answer is corrected by pronouncing clearly in a loud voice, the present invention determines the latter as high confidence and awards bonus experience points, thereby providing learning attitude correction and motivational effects through process-oriented evaluation.
[0131] According to the judgment rule table, if an active editing pattern is characterized by voice amplitude higher than the threshold and haptic conditions within the optimal range, it is judged as the highest grade, providing a badge and 1.5 times the experience points. If the pattern is active but the voice amplitude is low, it is judged as excellent, providing a badge and 1.2 times the experience points. If the pattern is passive but the voice amplitude is high, it is judged as average, providing basic experience points. If the pattern is passive, the voice amplitude is low, and the haptic condition deviates from the range, causing interference, it is judged to require correction; in this case, only basic experience points are provided, and feedback adjustment is performed. The objectivity and necessity of these judgment rules are based on the self-regulated learning theory of educational psychology and the analysis of prosodic characteristics in acoustics. Furthermore, based on the fact that a learner's act of locally identifying their own errors reflects a high level of cognitive monitoring and that the volume of speech has a positive correlation with the speaker's psychological certainty, these rules possess higher validity in predicting a learner's actual growth potential than simple accuracy rates.
[0132] As a technical effect of adopting these components, it overcomes the limitations of existing systems that focused solely on output values and induces proper learning habits by incorporating the learner's thought process of finding the correct answer as an evaluation factor. Furthermore, by including the appropriateness of haptic feedback as an evaluation variable, it provides the effect of optimizing the system environment so that physical stimuli function as a reinforcer without hindering learning immersion. Finally, regarding the definition of criteria for terms and thresholds, the active rewrite pattern is defined as a series of events in which two or more deletion inputs occur immediately after cursor movement input, followed by character input. The voice amplitude threshold is defined as a signal intensity at least +10dB higher than the average amplitude level of ambient background noise, and the optimal haptic range is defined as a range of ±1 step of the vibration intensity level designated as preferred by the user in the settings menu.
[0133] The evaluation criteria of this process are based on the 'signal-to-noise ratio (SNR)' in the field of speech signal processing and the analysis of 'metacognitive behavioral patterns' in educational technology. Environmental adaptability is ensured by setting the standard for valid utterances to a relative value higher than a certain level relative to the [background noise level] rather than an absolute decibel (dB), and accidental noise is excluded by setting the duration threshold to 500ms, which exceeds the minimum time of a human monosyllabic utterance of 200ms.
[0134] According to an exemplary computational model, the growth potential index is calculated by multiplying the sum of the [Active Editing Score] and the [Confident Speech Score] by the [Haptic Fit Coefficient]. In this case, the Active Editing Score is awarded when cursor movement and block modification are detected, and the Confident Speech Score is awarded when the voice amplitude exceeds a threshold. The Haptic Fit Coefficient corrects for performance distortion caused by environmental factors by applying damping values, such as 1.0 if the vibration feedback at the time was within the user preference range, and 0.8 if it was outside the range and became a distraction.
[0135] The table below shows data defining how scores are assigned based on user behavior patterns and hardware status.
[0136]
[0137] The table below demonstrates how the reward (growth index) varies according to the process evaluation for four cases where the same answer was correct.
[0138]
[0139] Hereinafter, the specific operational processes performed by the User-AI co-growth system according to one embodiment of the present invention during an actual learning session of Learner A are described in detail in chronological order. This embodiment assumes a situation in which Learner A learns the topic of airport immigration screening inside a noisy subway using an old tablet PC with a low refresh rate of 30 Hertz. First, as an environment recognition and initial parameter correction step, as soon as Learner A launches the application, the user terminal unit scans the device's display information to detect that the screen refresh cycle is 33 milliseconds, and simultaneously activates the microphone for 5 seconds to perform a process confirming that the average level of the surrounding background noise is 60 decibels. The central server unit receives this information and adjusts the initial setting value of this session; specifically, the curriculum management module sets 330 milliseconds, which is 10 times the screen refresh cycle, as the system delay subtraction constant according to the logic of Process 2. This is a systemic agreement that even if Learner A's response speed is measured to be slow in the future, a delay of 0.33 seconds caused by device performance will not be considered as a lack of learning ability. In addition, the chatbot execution module performs processing to prepare to make the cognitive friction judgment more lenient than usual by setting the environmental noise weighting factor of process 1 to 1.2 in consideration of subway noise.
[0141] Next, learning begins as a cognitive friction detection and AI speed adaptation phase, and the chatbot asks in English, "Could you please show me your passport?" at a fast speed of 1.2x. Learner A knows the correct answer but, due to the shaking subway environment, makes a typo while entering "passport," presses the backspace key four times in a row, and pauses input for 3 seconds. At this point, the chatbot execution module calculates the input efficiency index; although the correct answer was obtained, the ratio of edit key inputs relative to text length was high and the silence time was measured to be long. Under normal circumstances, such a delay in a high-speed speech situation at 1.2x speed should be judged as cognitive overload; however, the system applies the previously set environmental noise weighting factor of 1.2 to correct the friction score and performs control to maintain the speed, judging that the threshold has not yet been exceeded. However, when Learner A once again shows more than 5 corrections and more than 5 seconds of silence in the next sentence, the system makes a final judgment that the accumulated friction score has exceeded the dynamic threshold. Accordingly, the chatbot immediately lowers the speech speed to 1.0x, displays a text hint at the bottom of the screen, and provides voice feedback saying "Try taking your time" to prevent the learner from leaving.
[0142] Next, as part of the hardware correction-based curriculum transition judgment stage, Learner A, having completed all tasks in the immigration scenario, intends to proceed to the next axis, the baggage claim scenario. At this point, the total time taken for the final task was measured at 3.0 seconds; under the general standard of passing within 2.8 seconds, the transition would be blocked. However, the curriculum management module applies the system delay subtraction constant of 330 milliseconds set in Stage 1 to calculate the pure cognitive time as 2,670 milliseconds by subtracting 330 milliseconds from the measured 3,000 milliseconds. Since this corrected time falls within the transition allowance criterion of 2.8 seconds, the system determines that Learner A's proficiency is sufficient. Furthermore, by confirming that the incorrect answer during learning was a task involving high-complexity honorific language variations and analyzing that the step regression was not due to a lack of basic skills, the system calculates a low transition resistance index. Consequently, the system performs control to immediately approve the curriculum transition to baggage claim.
[0143] Next, as part of the active interaction evaluation and differential reward stage, Learner A composes the sentence "My bag is missing" in the new curriculum. Initially, they write "My bag is lost," then move the cursor before "is" to correct it to "has bean," and subsequently read the sentence aloud. The reward management module performs a precise analysis of this process; first, it detects logs indicating block modification via cursor movement rather than complete deletion and re-entry, classifies this as an active rewriting pattern, and assigns bonus points. Subsequently, by analyzing the voice waveform collected via the microphone, it confirms that although subway noise is present at 60 decibels, Learner A's voice persists for 0.8 seconds at a level 15 decibels higher than the 75-decibel noise, and judges this as confident speech. Finally, haptic verification is performed; at the moment of correct answer determination, the tablet vibrates at an intensity of 3, confirming that Learner A's preferred intensity setting is between 2 and 4. Accordingly, the reward module determines that all conditions, such as active correction, confident utterance, and appropriate feedback, have been met, and performs the process of awarding a Grammar Master badge with a spectacular fireworks effect on the screen and giving 1.5 times the experience points.
[0144] Finally, as part of the learning history analysis and feedback loop formation stage, the performance analysis module transmits today's learning data to the long-term memory storage after the session ends. Although Learner A answered all questions correctly, considering the high friction score in the early stages, the relevant words are classified as requiring observation rather than being fully mastered. This data is then transmitted back to the central server to generate tomorrow's curriculum; in tomorrow's learning, the passport-related vocabulary that caused friction early on will appear first as a review quiz, and the chatbot will continue to provide a stress-free learning environment by applying corrected time standards as long as Learner A continues to use the older tablet.
[0145] Although preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the invention as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention. Explanation of the symbols
[0146] Central server unit (110) User terminal unit (120) Curriculum management module (130) Chatbot execution module (140) Performance analysis module (150) Compensation management module (160)
Claims
Claim 1 A central server unit storing learning content, user log data, and an artificial intelligence model; and a user terminal unit communicating with the central server unit to receive voice and text input from a learner and output learning content; wherein the central server unit or the user terminal unit includes and operates a curriculum management module that provides learning content according to location, situation, and genre classification; a chatbot execution module that performs conversation with a user and adjusts response patterns; a performance analysis module that analyzes the user's learning history; and a reward management module that grants digital rewards based on learning performance; wherein the chatbot execution module includes a friction calculation unit that calculates a friction score quantifying the level of cognitive friction by monitoring non-verbal input behaviors occurring during the user's task performance in real time; and a speed control unit that controls the chatbot's speech speed (TTS Speed) by comparing the calculated friction score with a preset allowable threshold.The speed control unit includes, based on the cognitive load theory that the faster the current chatbot’s utterance speed, the more sensitively it must react to even small signs of friction as the user’s cognitive processing capacity decreases, dynamically lowers the allowable threshold inversely proportional to the currently set utterance speed multiplier, and if the calculated friction score exceeds the dynamically adjusted allowable threshold, generates a control signal to immediately decelerate the speed of the next utterance by a preset unit; the user terminal unit collects hardware / environment data including the display refresh cycle of the user terminal unit and the ambient noise level at the start of learning and transmits it to the central server unit; the central server unit adjusts the initial parameters of the learning session based on the hardware / environment data; the chatbot execution module controls the utterance speed or whether to provide hints based on the ambient noise level among the hardware / environment data, the friction score, and the current utterance speed; and the performance analysis module stores the friction score and speed control history as user-specific learning history and reflects it in the initial utterance speed, whether to provide hints, or the setting of content to be reviewed for the next learning session. An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum. Claim 2 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 1, the friction calculation unit collects non-verbal log data including the number of edit keys (Backspace / Delete) entered during the user's text input process, the length of the final entered text, the ratio of silent intervals without valid utterances during voice input, and the number of times the user canceled the recording and retried, and utilizes this data as basic data for calculating the friction score. Claim 3 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in paragraph 2, the friction calculation unit calculates the ratio of the number of edit key inputs to the length of the final input text, rather than the absolute value of the number of edit key inputs, to derive an input efficiency index that corrects the increase in the number of inputs according to the text length, calculates a silence ratio which is the ratio of the interval in which the signal level is below a preset background noise level within the total voice input duration, and applies the input efficiency index and the silence ratio as variables of the friction score. Claim 4 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in determining the allowable threshold, the speed control unit applies a calculation logic based on the value obtained by dividing the basic allowable value at a reference speed by the current speech speed, so that when the current speech speed is faster than the standard speed, the allowable threshold is lower than the basic allowable value to apply a strict judgment criterion, and when the current speech speed is slower than the standard speed, the allowable threshold is higher than the basic allowable value to apply a lenient judgment criterion. Claim 5 In claim 4, the speed control unit analyzes ambient noise data collected through the microphone sensor of the user terminal unit to determine whether the current learning environment is a noisy environment that is inconsistent with the environment recommended by the learning content; if it is determined to be a noisy environment, it relaxes the judgment criteria by upwardly correcting the allowable threshold by multiplying it by a weighting factor greater than 1.0 to prevent misjudging input delay caused by external factors as a lack of learning ability; if the friction score exceeds the allowable threshold and is determined to be in an overload state, it decelerates the speed in increments of 0.1 from the current speech speed, but sets a lower limit to control the decelerated speed so that it does not fall below the user's minimum guaranteed learning speed; and if the friction score continues to exceed the allowable threshold in a continuous learning task even after the deceleration measure, it determines that learning progress is difficult with only simple speed adjustment, and switches to a hint provision mode that visually highlights and displays text hints on the screen or outputs an encouraging message via voice, characterized by an adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum. Claim 6 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 5, the curriculum management module collects the total time required from the start to the completion of a task and the screen refresh cycle of the user terminal unit, and calculates the pure cognitive time required by subtracting the value obtained by multiplying the screen refresh cycle by a system delay correction coefficient from the total time required. Claim 7 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 6, the curriculum management module queries the complexity grade of the task in the event of a step backward, calculates a transition resistance index by assigning a higher penalty weight as the complexity of the task decreases, and approves or blocks movement to the next curriculum axis by comparing the transition resistance index with a reference value. Claim 8 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 7, the reward management module analyzes a key input sequence log to detect an active rewrite pattern in which two or more deletion inputs occur immediately after a cursor movement input, followed by character input. Claim 9 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 8, the compensation management module calculates the root mean square amplitude for a microphone input signal, sets a voice amplitude threshold to a signal strength that is at least +10dB higher than the average amplitude level of ambient background noise, and determines it as a confident utterance if the signal strength is maintained for at least 0.5 seconds. Claim 10 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 9, the reward management module verifies whether the intensity of haptic feedback occurring at that time is within the user preference range, calculates a growth potential index based on the suitability of the active rewrite pattern, the confidence utterance, and the haptic feedback, and differentially distributes digital rewards including digital badges or experience points. Claim 11 An adaptive language learning system based on a user-AI co-growth model and a multi-axis curriculum, wherein, in claim 10, the performance analysis module stores the friction score, the speed control history, and the transition resistance index in a time series, records the task type, speech rate, and friction score data at the time when deceleration control is triggered by the speed control unit in a user-specific vulnerability database, and resets the initial speech rate and whether to provide hints by referring to the recorded vulnerability information when configuring the next learning session. Claim 12 delete Claim 13 delete Claim 14 delete
Citation Information
Patent Citations
Systems and methods for measuring and enhancing human engagement and cognition
JP2023550846A
Methodologies and systems for foreign language learning through an intelligent virtual human tutor, including a computer program for implementation
KR1020250073854A
Method for operating lecture platform and apparatus for the same
KR102423740B1
Tablet keyboard system and control method providing firmware upgrade and wired connection function according to usage environment
KR102910372B1
System and Method for Improving Student Learning by Monitoring Student Cognitive State
US20160203726A1