Reward calculation system based on concentration analysis adaptive to learning task characteristics and method using the same
Patent Information
- Application Number
- KR1020260137758
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-10-15
Smart Images

Figure R1020260137758_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a reward calculation system based on concentration analysis adaptive to learning task characteristics and a method using the same. Background Technology
[0002] Generally, educational content was provided to learners through works produced by educators, researchers, or authors, and learners proceeded to study this content. In this context, learners were in a situation where they had no choice but to restrict their selection of content by choosing an educator, researcher, or author and consuming only the works produced by that selected entity.
[0003] Learners seek various means to easily access online educational content in order to maximize their range of choices. To address this, online platforms provide services offering content produced by mentors to learners; however, methods for determining compensation to incentivize mentors to actively produce content are currently lacking. Prior art literature
[0004] Korean Registered Patent No. 10-2488807 The problem to be solved
[0005] The objective of the present invention, which aims to solve the aforementioned problems, is to address the issues of low learner participation, limitations in developer investment, and lack of content diversity in the educational content market.
[0006] More specifically, we propose a behavioral reward-based education platform that enhances learning motivation and engagement by providing tangible rewards based on learners' behavior, induces developer participation and business performance through a platform integrating various content and functions, and offers authoring tools that enable educators or creators to easily produce educational content with applied reward features.
[0007] The technical problems to be solved in the various embodiments are not limited to those mentioned above, and other unmentioned technical problems may be considered by those skilled in the art from the various embodiments described below. means of solving the problem
[0008] To achieve the above objective, the system provides a behavior reward type platform system that calculates a reward for a scenario based on the detection of a learner's behavior according to embodiments, comprising: a content authoring tool module configured to allow a developer or administrator to configure a scenario flow and to grant reward conditions for said scenario flow; a behavior event collection module that collects a learner's behavior events when a scenario generated by said content authoring tool module is executed through a learner terminal; a reward mapping engine that determines whether said behavior events satisfy a preset reward condition; and a reward calculation module that calculates reward points according to a preset method when it is determined by said reward mapping engine that said behavior events satisfy the reward condition; wherein the content authoring tool module comprises: a scenario input unit in which a scenario, which is learning content generated by a plurality of content creators, is input to form a platform; a scenario selection unit that allows an administrator to configure a scenario flow, which is a learning process, by selecting or combining one or more of a plurality of scenarios uploaded on said platform; and a reward setting unit that individually sets reward conditions and methods for each scenario constituting said scenario flow.
[0009] The scenario selection unit may include an artificial intelligence model that performs machine learning using the scenario flow of another learner as learning information using the platform; and the artificial intelligence model may include: a preprocessing unit that preprocesses the learner's learning state information to generate a behavioral characteristic vector for the learner, wherein the preprocessing unit calculates and provides a scenario flow for the learner using the learner's learning state information; a deep learning-based prediction computation unit that predicts and outputs a scenario selection probability and a reward responsiveness using the behavioral characteristic vector as input information; and a scenario recommendation unit that configures and provides a scenario flow optimized for the learner based on the prediction result of the prediction computation unit.
[0010] The reward calculation module includes a weight calculation unit that calculates and applies weights according to a preset method when calculating the reward points; and the weight calculation unit may be configured to calculate weights through a first method of applying a relative weight to a behavioral event of a learner performing the m-th scenario based on a behavioral event of another learner performing the m-th scenario, or a second method of applying an attention index derived by analyzing the concentration on performing the m-th scenario based on the learner's behavioral event.
[0011] The above system further includes a concentration analysis module for analyzing the concentration; wherein the concentration analysis module includes a concentration data collection unit that collects at least one of the learner's gaze, reaction time, input pattern, error rate, and learning time; and a concentration index calculation unit that normalizes the data collected by the concentration data collection unit to calculate the concentration index, wherein the concentration index is calculated by considering at least one of visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb); and the concentration index calculation unit may be configured to reflect the learner's dynamic state change by combining the visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb) with a preset non-linear function.
[0012] The concentration index calculation unit further comprises: a scenario type analyzer that automatically recognizes the type of scenario being performed (e.g., video viewing, text comprehension, interactive practice); and a meta-learning-based weight optimization engine that automatically adjusts in real-time the weights of a non-linear function combining visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavior concentration (Fb) according to the analyzed scenario type; wherein the optimization engine may be characterized by dynamically applying a concentration evaluation model most suitable for task characteristics, such as increasing the weight of visual concentration (Fv) in a video viewing scenario and increasing the weight of reaction concentration (Fr) in a reaction quiz scenario.
[0013] The above system further comprises: an administrator account module for an administrator managing the learner; and a reward charging management module that manages the deduction of the charged reward points when the administrator charges reward points in advance and the learner performs the k-th scenario; wherein the reward charging management module manages the balance of the total reward points charged by the administrator, sets a reward point allocation range for the k-th scenario, and provides the reward point allocation range to the learner before the scenario is performed; and a non-linear function f(AI,E) that outputs a reward coefficient of 0 to 1 according to the concentration index and behavioral event when the performance of the k-th scenario is completed by the learner. b A reward dynamic adjustment unit that dynamically calculates the actual reward points paid to the learner based on ), wherein the reward dynamic adjustment unit, through the following mathematical formula 1, the actual reward points paid (R act ) can be computed (here, R min The minimum reward points, R max is maximum reward points, E b is a behavioral event performance indicator that includes the learner's behavior success rate, execution speed, and error rate).
[0014] [Mathematical Formula 1]
[0015]
[0016] The above reward points may be divided into common points that can be commonly paid to all learners performing the scenario, and individual points that are paid only to specific learners based on reward points charged by the manager, and the payment ratio of the common points and individual points may be dynamically changeable.
[0017] The above-mentioned learner's learning status information may include at least one of the learner's viewing time, correct answer rate, reaction time, concentration index, and past reward history.
[0018] The above system further includes an AI-based reward budget prediction module that is linked with the above administrator account module to assist in charging the administrator's reward points; the budget prediction module learns the past scenario performance history, average concentration index, and reward payment statistical data of a specific learner group to predict the total reward points to be consumed over a specific period (e.g., one month) in the future, and if the remaining points of the above administrator account module are insufficient compared to the predicted consumption amount, it may be characterized by sending a warning notification to the administrator in advance, including a recommended charging amount and an estimated time of depletion.
[0019] The above reward dynamic adjustment unit further includes a reward function personalization engine that generates a 'motivation profile' by analyzing an individual learner's long-term learning history and reward response patterns; and the personalization engine comprises: (a) a logarithmic reward function f(AI,E) in which the reward increases rapidly in the low achievement range for learners who easily give up even on small successes. b ) and (b) for learners who enjoy high challenges, an exponential reward function f(AI,E) in which rewards increase explosively in the highest level achievement range. b Applying the non-linear function f(AI,E) in accordance with the above motivation profile, such as applying ). b It can be characterized by optimizing the form of itself for each individual.
[0020] In addition, the present invention provides a method using the aforementioned system, comprising: (a1) a step of configuring a scenario, wherein learning content created by a plurality of content creators is uploaded to a scenario input section of a platform, and an administrator or developer selects or combines one or more of the uploaded scenarios to configure a scenario flow, which is a learning process, and sets reward conditions and a method of granting rewards for each scenario; (a2) a step of collecting behavioral events, wherein when the scenario flow is executed through a learner terminal, behavioral data such as the learner's viewing, clicking, input, reaction time, accuracy rate, and execution time are collected as behavioral events; (a3) a step of determining reward conditions, wherein the collected behavioral events are determined using a reward mapping engine whether they satisfy the reward conditions set for each scenario; (a4) A step of analyzing the learner's concentration, wherein data such as the learner's gaze, reaction time, input pattern, error rate, and learning time are collected to calculate at least one of visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb), and then a concentration index is calculated using a non-linear combination function; (a5) A step of calculating reward points for the learner, wherein the concentration index (AI) and the learner's behavioral event performance indicator (E b f(AI,E), a non-linear function that takes ) as input b A method comprising: (a6) a step of calculating a reward coefficient of 0 to 1 through ) and then calculating an actual reward point, wherein the function is set to be dynamically updated according to changes in concentration and performance results; and (a6) a step of paying a reward point to the learner, wherein the reward point is automatically paid to the learner by deducting the actual reward point calculated in step (a5) from the reward point charged in advance by the administrator.
[0021] Prior to the above step (a2), (a11) a step provided in advance by displaying the range of reward points and reward conditions set in the scenario on the learner's terminal before the learner starts performing the scenario; wherein the reward points are divided into common points that are commonly given to learners performing the scenario and individual points that are given only to learners who have been registered in advance by the manager based on the reward points charged by the manager, and the ratio of the common points and individual points can be set to be dynamically variable. Effects of the invention
[0022] The present invention can provide a writing tool that enhances learning motivation and immersion by providing tangible rewards based on learner behavior, induces participation and business performance from developers through a platform integrating various content and functions, and enables educators or creators to easily produce educational content with applied reward functions.
[0023] Furthermore, the present invention enables administrators or educators to freely set reward conditions for each scenario, thereby providing a customized learning environment tailored to various educational objectives and levels. Additionally, through an AI-based reward calculation method that combines behavioral event analysis and concentration index calculation, it demonstrates the effect of enabling differential rewards based on the learner's achievement and effort.
[0024] In addition, a transparent reward structure is implemented through a reward charging and management module, enabling the provision of a highly reliable reward system for both operators and learners.
[0025] The effects obtainable from various embodiments are not limited to those mentioned above, and other unmentioned effects can be clearly derived and understood by a person skilled in the art based on the following detailed description. Brief explanation of the drawing
[0026] Other aspects, features, and benefits of specific preferred embodiments of the present invention, as described above, will become more apparent from the following description in conjunction with the accompanying drawings. FIG. 1 is a block diagram showing the configuration of a behavior reward type platform system according to one embodiment of the present invention. FIG. 2 schematically shows the overall configuration of a behavior reward type platform system according to one embodiment of the present invention. FIG. 3 schematically shows the process processed in the content authoring tool module of a behavior reward type platform system according to one embodiment of the present invention. FIG. 4 schematically shows the process of calculating reward points through the reward calculation module and concentration analysis module of a behavioral reward platform system according to one embodiment of the present invention. FIG. 5 schematically shows the process processed in the reward charging management module of a behavioral reward type platform system according to one embodiment of the present invention. FIG. 6 is a flowchart of a method using a behavior reward type platform system according to one embodiment of the present invention. FIG. 7 shows an exemplary UI actually implemented in an APP on a learner terminal of a behavior reward platform system according to one embodiment of the present invention. It should be noted that in the drawings above, similar reference numbers are used to illustrate identical or similar elements, features, and structures. Specific details for implementing the invention
[0027] Embodiments of the present invention are described below with reference to the attached drawings so that those skilled in the art (hereinafter, those skilled in the art) can easily implement them. The embodiments presented in the present invention are provided to enable those skilled in the art to use or implement the contents of the present invention. Accordingly, various modifications to the embodiments of the present invention will be obvious to those skilled in the art. That is, the present invention can be embodied in various different forms and is not limited to the embodiments below.
[0028] Throughout the specification of the present invention, identical or similar reference numerals refer to identical or similar components. Additionally, to clearly explain the present invention, reference numerals in the drawings that are unrelated to the description of the present invention may be omitted.
[0029] The term "or" used in the present invention is intended to mean an implicit "or" rather than an exclusive "or." That is, unless otherwise specified in the present invention or its meaning is unclear from the context, "X uses A or B" should be understood to mean one of the natural implicit substitutions. For example, unless otherwise specified in the present invention or its meaning is unclear from the context, "X uses A or B" may be interpreted as any one of the cases where X uses A, X uses B, or X uses both A and B.
[0030] The term "at least one of A or B" used in the present invention should be interpreted as referring to A, B, and combinations of A and B.
[0031] The term "and / or" as used in the present invention should be understood to refer to and include all possible combinations of one or more of the enumerated related concepts.
[0032] The terms “comprising” and / or “comprising” as used in the present invention should be understood to mean the presence of specific features and / or components. However, the terms “comprising” and / or “comprising” should be understood not to exclude the presence or addition of one or more other features, other components and / or combinations thereof.
[0033] Where not otherwise specified in the present invention or where it is not clear from the context that the singular form indicates, the singular should generally be interpreted as including "one or more."
[0034] The term "the N (N is a natural number)" used in the present invention can be understood as an expression used to distinguish the components of the present invention from one another according to certain criteria, such as functional perspectives, structural perspectives, or convenience of explanation. For example, components performing different functional roles in the present invention may be distinguished as the first component or the second component. However, components that are substantially identical within the technical scope of the present invention but need to be distinguished for the convenience of explanation may also be distinguished as the first component or the second component.
[0035] Meanwhile, the terms "module" or "unit" used in the present invention may be understood as referring to an independent functional unit that processes computing resources, such as a computer-related entity, firmware, software or a part thereof, hardware or a part thereof, or a combination of software and hardware. In this case, "module" or "unit" may be a unit composed of a single element, or a unit expressed as a combination or set of multiple elements. For example, in a narrow sense, "module" or "unit" may refer to a hardware element of a computing device or a set thereof, an application program that performs a specific function of software, a procedure implemented through software execution, or a set of instructions for program execution. Furthermore, in a broad sense, "module" or "unit" may refer to the computing device itself that constitutes the system, or an application executed on the computing device. However, since the above-described concepts are merely examples, the concepts of "module" or "unit" may be defined in various ways within a scope understandable to those skilled in the art based on the content of the present invention.
[0036] The term "model" as used in the present invention can be understood as a system implemented using mathematical concepts and language to solve a specific problem, a set of software units to solve a specific problem, or an abstract model regarding a processing procedure to solve a specific problem. For example, a neural network "model" may refer to the entire system implemented as a neural network that possesses problem-solving capabilities through learning. In this case, the neural network can possess problem-solving capabilities by optimizing parameters connecting nodes or neurons through learning. A neural network "model" may include a single neural network or a set of neural networks composed of multiple neural networks.
[0037] The explanation of the foregoing terms is intended to aid in understanding the present invention. Therefore, it should be noted that unless the foregoing terms are explicitly stated as matters limiting the content of the present invention, they are not used to limit the technical concept of the present invention.
[0038] FIG. 1 is a block diagram showing the configuration of a behavior reward type platform system according to one embodiment of the present invention. For convenience of explanation, the 'behavior reward type platform system' will be referred to as the 'system'.
[0039] A system (100) according to one embodiment of the present invention may be a hardware device or a part of a hardware device that performs comprehensive processing and computation of data, or it may be a software-based computing environment connected to a communication network. For example, the system (100) may be a server that performs intensive data processing functions and is an entity that shares resources, or it may be a client that shares resources through interaction with the server. Additionally, the system (100) may be a cloud system that enables multiple servers and clients to interact to comprehensively process data. Since the above description is merely one example regarding the type of system (100), the type of system (100) may be configured in various ways within a range understandable to those skilled in the art based on the content of the present invention.
[0040] A system (100) according to one embodiment of the present invention may include a processor (10), a memory (20), and a network unit (30). However, since FIG. 1 is merely an example, the system (100) may include other configurations for implementing a computing environment. Additionally, only some of the disclosed configurations may be included in the system (100).
[0041] A processor (10) according to one embodiment of the present invention may be understood as a constituent unit including hardware and / or software for performing computing operations. For example, the processor (10) may read a computer program and perform data processing for machine learning. The processor (10) may process computational processes such as processing input data for machine learning, extracting features for machine learning, and calculating errors based on backpropagation. A processor (10) for performing such data processing may include a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). Since the above-described types of processors (10) are merely examples, the types of processors (10) may be configured in various ways within a range understandable to those skilled in the art based on the content of the present invention.
[0042] A memory (20) according to one embodiment of the present invention may be understood as a configuration unit comprising hardware and / or software for storing and managing data processed by a system (100). That is, the memory (20) may store data of any form generated or determined by a processor (10) and data of any form received by a network unit (30). For example, the memory (20) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory, RAM (random access memory), SRAM (static random access memory), ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), PROM (programmable read-only memory), magnetic memory, a magnetic disk, or an optical disk. Additionally, the memory (20) may include a database system that controls and manages data in a predetermined system. Since the above-described type of memory (20) is merely an example, the type of memory (20) can be configured in various ways within a range understandable to those skilled in the art based on the content of the present invention.
[0043] A network unit (30) according to one embodiment of the present invention can be understood as a configuration unit that transmits and receives data through any known form of wired or wireless communication system. For example, the network unit (30) can perform data transmission and reception using wired or wireless communication systems such as a local area network (LAN), wideband code division multiple access (WCDMA), long term evolution (LTE), wireless broadband internet (WiBro), 5th generation mobile communication (5G), ultra-wide-band wireless communication, ZigBee, radio frequency (RF) communication, wireless LAN, wireless fidelity (Wi-Fi), near field communication (NFC), or Bluetooth. Since the communication systems described above are merely examples, wired or wireless communication systems for data transmission and reception of the network unit (30) can be applied in various ways other than the examples described above.
[0044] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0045] FIG. 2 schematically shows the overall configuration of a behavior reward type platform system according to one embodiment of the present invention. With reference to FIG. 2, a system (100) according to one embodiment of the present invention will be described.
[0046] The system (100) includes a content authoring tool module (110), an action event collection module (120), a reward mapping engine (130), a reward calculation module (140), an intensity analysis module (150), and a reward charging management module (160).
[0047] The content authoring tool module (110) can upload learning content created by multiple content creators or managers into the platform and combine them to form a learning scenario. In this invention, the sequential or parallel flow structure of such learning scenarios is defined as a 'Scenario Flow,' and a single scenario flow can be composed of multiple learning contents (scenarios). That is, individual learning units (e.g., videos, quizzes, text, interactive apps, etc.) uploaded by content creators are each managed as a single scenario, and a manager or educator can combine or rearrange these scenarios to set a series of step-by-step learning paths for learners to perform. At this time, since learning conditions, evaluation criteria, and reward payment rules can be independently defined for each scenario, the platform operator can design a multi-layered reward scenario suitable for the learning purpose.
[0048] FIG. 3 schematically illustrates the process processed in the content authoring tool module of a behavioral reward platform system according to one embodiment of the present invention. With reference to FIG. 3, the content authoring tool module (110) will be described in detail.
[0049] The content authoring tool module (110) may include a scenario input section (111), a scenario selection section (112), and a reward setting section (113).
[0050] The scenario input section (111) is a core input module that constitutes the basic database of the platform and provides the function of uploading and registering various forms of learning content. The scenario input section (111) supports multiple formats including video (video lectures, practice videos, etc.), quizzes (multiple choice, short answer, true / false, etc.), text (learning materials, problem-solving explanations, etc.), and interactive content (HTML5-based learning apps, game-type learning modules, etc.), and can be easily registered through an upload screen based on a graphical user interface (GUI) in a web or app environment.
[0051] Upon registration, the system automatically assigns metadata (title, creator ID, learning difficulty, recommended learning time, default reward value, etc.) for each piece of content, and the administrator can verify or modify it. Additionally, the scenario input section (111) automatically determines the format of the uploaded file and stores it in the Content Management Storage (CMS) inside the server, and each stored piece of content is assigned a unique identifier (Scenario ID) so that it can be referenced by subsequent processing modules (scenario selection section, reward setting section).
[0052] This structure can be combined with content registration technologies based on Learning Content Management Systems (LCMS) or Learning Tools Interoperability (LTI), and can also be integrated with external learning content providers (LXP, MOOC, etc.) through the same standard interface.
[0053] The scenario selection unit (112) performs the role of selecting one or more of the uploaded scenarios, or arranging them in a sequence or parallel combination manner to configure the entire learning process, i.e., the scenario flow, to be performed by the learner.
[0054] An administrator or educator can set the sequential relationships between each scenario, progress conditions (e.g., automatic progression to the next step upon completion of the previous scenario), and evaluation criteria (e.g., branching out upon achieving a certain score) through the UI of the scenario selection section (112). These settings can be implemented using a graph structure-based node editor or a drag-and-drop interface, and the connection relationships between each scenario can be managed internally as a flow graph data structure.
[0055] Additionally, the scenario selection unit (112) can be linked with an artificial intelligence model (AI Engine) (1121) that learns the learner's level, interest, and concentration patterns based on the past scenario performance results of other learners. The artificial intelligence model (1121) can automatically calculate and suggest a combination of scenarios suitable for a specific learner through a machine learning-based recommendation algorithm, or provide supplementary recommendation information when an administrator makes a selection. As an example, high-difficulty problem scenarios can be suggested to learners with a high average concentration index (AI), and video-based scenarios can be suggested first to learners with a low concentration index. Such a scenario selection function can contribute to enhancing learner immersion and maximizing the effectiveness of customized education.
[0056] Specifically, the artificial intelligence model (1121) is configured to calculate and provide a scenario flow for the learner using the learner's learning state information, and includes a preprocessing unit (1121a), a prediction calculation unit (1121b), and a scenario recommendation unit (1121c).
[0057] The preprocessing unit (1121a) is configured to preprocess the learner's learning state information to generate a behavioral characteristic vector for the learner. The learning state information includes at least one of the learner's viewing time, correct answer rate, reaction time, concentration index, and past reward history, and the preprocessing process is configured to perform missing value correction, normalization, and feature extraction.
[0058] The prediction operation unit (1121b) is performed based on deep learning, which uses a behavioral characteristic vector as input information to predict and output the scenario selection probability and reward responsiveness.
[0059] It can be configured to predict the probability of a learner selecting a specific scenario and the reward responsiveness when performing that scenario. As an example, a deep learning-based prediction model can be configured to output the selection probability value and the expected reward value for each scenario using a sigmoid, ReLU, or softmax activation function.
[0060] The scenario recommendation unit (1121c) is configured to configure and provide a scenario flow optimized for the learner based on the prediction result of the prediction operation unit (1121b). It can determine a combination of scenarios that maximizes the predicted reward value for each learner, and arrange scenarios among multiple scenarios in a sequential or parallel manner in which the selection probability is greater than or equal to a threshold value to configure and provide a scenario flow optimized for the learner.
[0061] The reward setting unit (113) is configured to define reward conditions and payment methods for each scenario and serves as the central policy of the reward calculation logic. The administrator can set reward payment conditions for each scenario, and these conditions may include the following exemplary items.
[0062] ① Viewing Completion Rate Condition: e.g.) +5 points awarded for watching 90% or more of the video
[0063] ② Correct Answer Rate Condition: e.g.) +10 points for a quiz correct answer rate of 80% or higher
[0064] ③ Response time condition: e.g.) +3 points if average response time is 3 seconds or less
[0065] ④ Concentration Condition: e.g.) +8 points when Concentration Index (AI) is 0.75 or higher
[0066] ⑤ Overall Achievement Condition: Combination of the above items (based on AND / OR logical expressions)
[0067] The reward setting unit (113) converts the above conditions into a reward rule table and transmits it to the reward mapping engine (130). The reward setting unit (113) can define reward points of different natures. They can be divided into 'common points' assigned by a system administrator or designer and 'individual points' assigned by a learner's administrator (or educator). The reward points paid to the learner can be defined as the sum of these. Here, individual points can be directly set by the administrator in the administrator account module (102), and for this purpose, the administrator account module (102) includes a wallet unit (not shown), and can be charged to the wallet unit by the administrator using a pre-set payment method. As an example, the common points and individual points may be designated with different uses or application areas.
[0068] For example, common points can be used as a reward for learning achievements obtained throughout the system, and individual points can be used as an incentive for an administrator (or educator) to encourage specific learning behaviors or assign additional learning tasks. Additionally, the reward mapping engine (130) can check the individual point balance stored in the wallet section of the administrator account module (102) to verify whether a reward can be paid, and if necessary, perform automatic charging or reward limit adjustment through the reward charging management module (140).
[0069] The behavior event collection module (120) is configured to perform event detection, event processing, and event storage.
[0070] It can detect the learner's physical or logical input in real time by connecting to learner interface (UI) elements and sensing hardware (camera, touchscreen, microphone, gyroscope sensor, etc.) within the learner terminal. It is registered as an Event Listener within the learner terminal and can be transmitted in real time in JSON or Protobuf format along with metadata such as timestamp, event type, event value, and session ID.
[0071] The reward mapping engine (130) is configured to determine whether a behavior event satisfies a pre-set reward condition. As an example, it may operate by referring to a reward rule table. This table is composed of items defined in the reward setting section (113) of the content authoring tool module (110) and may include fields such as scenario ID, condition name, comparison operator, threshold, weight, and combination logic.
[0072] When event data is input from the behavior event collection module (120), the reward mapping engine (130) refers to a pre-set reward rule table and performs a series of processes to determine whether the event satisfies the reward conditions. The following is an exemplary processing process.
[0073] First, the reward mapping engine (130) receives event log data (Event Log) transmitted from the behavior event collection module (120) in the 'event reception stage'.
[0074] Event log data includes information such as the time of event occurrence, event type, event value, learner identifier (user_id), and scenario identifier (scenario_id).
[0075] In the 'scenario identification step,' the reward mapping engine (130) identifies which scenario the event corresponds to by referencing the scenario_id field within the event data, and can query the items of the Reward Rule Table pre-set for that scenario. This rule table records reward conditions defined for each scenario, such as viewing completion rate, correct answer rate, reaction time, and concentration index, as well as threshold values, weights, and logical operation methods for each condition.
[0076] The 'condition comparison step' can individually determine whether a condition is met by comparing the actual value of the received event data with each threshold value in the reward rule table.
[0077] The 'combination judgment step' determines whether the reward conditions for each scenario are comprehensively satisfied by combining the results of each condition according to logical operators (AND, OR, etc.) when multiple conditions exist. This process is performed according to the condition combination rules defined in the reward setting section (113) of the content authoring tool module (110), and can be designed so that partial compensation is possible even if a specific condition is not satisfied, provided that the degree of achievement of other conditions is high.
[0078] The 'Condition Fulfillment Index Calculation Step' calculates the achievement rate of each condition (e.g., actual correct answer rate / target correct answer rate) and calculates the condition fulfillment index using a weighted average method, taking into account the weights for each condition.
[0079] The reward calculation module (140) includes a weight calculation unit (141) and a reward point calculation unit (142). It is configured to calculate reward points when the reward mapping engine (130) determines that an action event satisfies the reward conditions.
[0080] The weight calculation unit (141) is configured to calculate and apply weights according to a preset method when calculating reward points, and is configured to apply the following first method and / or second method.
[0081] FIG. 4 schematically illustrates the process of calculating reward points through a reward calculation module and an attention analysis module of a behavioral reward platform system according to an embodiment of the present invention. With reference to FIG. 4, the process of calculating reward points is explained. Reward points may be configured to apply an attention index calculated in the attention analysis module (150).
[0082] The weight calculation unit (141) is configured to calculate weights by applying one or more of the first method and the second method to reflect the behavioral characteristics of individual learners and the relative performance level with other learner groups when calculating reward points.
[0083] The first method establishes a relative standard by comparing the results of a learner's behavioral events with those of other learners who performed the same scenario. For example, it calculates where Learner A's accuracy rate, reaction time, and attention index (AI) correspond to the average or median of the entire group, and assigns the result as a relative weight.
[0084] The second method is a method of applying a concentration index (AI) derived by analyzing the concentration level for performing the m-th scenario based on the learner's behavioral events. To explain the second method, a concentration analysis module (150) is described. The concentration analysis module (150) includes a concentration data collection unit (151) and a concentration index calculation unit (152).
[0085] The concentration data collection unit (151) is configured to collect at least one of the learner's eye tracking data, reaction time, input pattern (e.g., click frequency, number of scrolls), error rate, and learning time. This process can be performed by collecting data in real time from the learner's terminal camera, touch sensor, microphone, or external input device.
[0086] The concentration index calculation unit (152) is configured to normalize the data collected from the concentration data collection unit (151) and to calculate detailed concentration indicators such as visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb). Each of these indicators reflects the cognitive and behavioral immersion of the learner, and the concentration index calculation unit (152) can calculate the overall concentration index (AI) using a non-linear combination function. An exemplary non-linear function may be implemented in the form of a sigmoid, exponential, or hyperbolic tangent (Tanh). The above non-linear function is configured to reflect changes in the learner's concentration over time, so that patterns such as attention distraction or reaction delay during learning can be dynamically evaluated.
[0087] As described above, the weights calculated through the first method and / or the second method are considered when calculating reward points, thereby configuring the system to dynamically adjust the reward points, initially set by the administrator in the form of absolute values, within a predetermined variable range. That is, for a specific scenario, the administrator [sets] the minimum reward points (R min ) and maximum reward points (R max Even if ) is set in advance, the system (100) [is] the learner's behavioral performance (E b Actual payment value (R) within the corresponding range according to ) and concentration index (AI) act ) can be corrected in real time. As a result, even when performing the same scenario, it is possible to differentiate rewards based on the learner's level of immersion, accuracy rate, execution speed, error rate, etc.
[0088] Referring to FIG. 5, a reward charging management module (160) is described. The reward charging management module (160) includes a reward disclosure unit (161) and a reward dynamic adjustment unit (162).
[0089] The reward disclosure unit (161) manages the balance of total reward points charged by the administrator, sets the reward point granting range for the k-th scenario, and is configured to provide the reward point granting range to the learner before the scenario is performed. As the reward point granting range is disclosed to the learner, the learner can recognize in advance the minimum and maximum reward levels that can be expected when performing the scenario, and accordingly, can improve the motivation and concentration for learning participation. That is, the reward disclosure unit (161) functions not merely to notify the learner of the result of simply giving reward points, but to induce the learner's voluntary participation and behavior improvement by explicitly presenting the 'rewardable range' at the stage before the scenario is performed. As an example, even for the same k-th scenario, an increased reward granting range is provided to learners who have achieved an AI above a certain standard, and a relatively lower range is presented to learners who lack recent learning history, thereby promoting continuous learning inflow.
[0090] The reward dynamic adjustment unit (162) is a non-linear function f(AI,E) that outputs a reward coefficient of 0 to 1 according to the concentration index (AI) and behavioral event when the performance of the k-th scenario is completed by the learner. b Based on ), the actual reward points paid to the above learner can be dynamically calculated. In this process, by applying the following mathematical formula 1, the actual reward points paid (R act It is designed so that ) is calculated.
[0091] [Mathematical Formula 1]
[0092]
[0093] Here, here, R min is 'minimum reward points', R max is 'maximum reward points', E bmeans 'behavioral event performance indicators' including the learner's behavior success rate, performance speed, and error rate.
[0094] As mentioned above, R min is the minimum reward point guaranteed in the corresponding scenario, R max represents the maximum achievable reward point, and f(AI,Eb) is a non-linear function that takes the learner's concentration index (AI) and behavior event performance indicator (Eb: behavior success rate, performance speed, error rate, etc.) as inputs and outputs a reward coefficient α between 0 and 1.
[0095] At this time, the reward dynamic adjustment unit (162) can calculate a reward coefficient in the form of a sigmoid function, an exponential function, or a power function, rather than a simple linear proportion, by considering the mutual non-linear correlation between the concentration index (AI) and the behavior event (Eb). For example, if the concentration index (AI) rises rapidly above a certain threshold, the rate of increase of the reward coefficient is designed to saturate gradually, thereby preventing excessive reward bias. Conversely, if the concentration index (AI) is below a certain level or the behavior success rate is low, the reward coefficient is lowered rapidly, thereby minimizing the waste of rewards for the learner's inactive behavior.
[0096] Additionally, the reward dynamic adjustment unit (162) can be configured to normalize the reward coefficient or enable upward or downward adjustment based on the average value by considering the relative performance among multiple learners within the same scenario. For example, the average concentration index (AI) of the entire learner group avgA reward system based on relative evaluation can be implemented by adjusting f(AI,Eb) upward by a predetermined ratio when an individual learner's concentration index (AI) is high compared to ), and downward by adjusting it when it is low. Through this adaptive reward structure, the system can maximize the long-term learning motivation effect by automatically optimizing the reward curve in response to the learner's average achievement level, participation rate, and changes in difficulty.
[0097] The reward dynamic adjustment unit (162) is a core computational unit that calculates the final reward points to be paid at the time the learner completes the scenario execution, and performs a series of sophisticated processing steps to precisely personalize and dynamically adjust the size of the reward by comprehensively reflecting the learner's achievement results and the qualitative aspects of the process. The processing steps begin by receiving inputs of the Comprehensive Concentration Index (AI) finally calculated during the learner's scenario execution process, and the Behavioral Event Performance Index (Eb), which is expressed as a single vector by normalizing the behavioral success rate, execution speed, and error rate in the corresponding scenario. Here, Rmin represents the minimum reward points guaranteed in the corresponding scenario, and Rmax represents the maximum reward points achievable.
[0098] Next, the reward dynamic adjustment unit (162) analyzes the learner's long-term learning history and reward response pattern, refers to a previously generated 'motivation profile', and selects a non-linear function f(AI, Eb) from the reward function library that is judged to be most suitable for the profile. The function f takes the learner's concentration index (AI) and behavior event performance index (Eb) as inputs and outputs a basic reward coefficient (α) having a value between 0 and 1. At this time, the function f is implemented in the form of a sigmoid function, an exponential function, or a power function, rather than a simple linear proportional relationship, and calculates the reward coefficient by considering the complex and non-linear interaction with the concentration index (AI) and the behavior event (Eb). For example, when a learner's attention index (AI) rises rapidly beyond a certain threshold, a sigmoid function form in which the rate of increase of the reward coefficient gradually saturates can be applied to prevent excessive reward bias for abnormally high attention and ensure the stability of the reward. Conversely, in cases where the attention index (AI) is low below a certain level or the behavioral success rate is significantly low, a function in which the reward coefficient decreases rapidly to near zero is applied to minimize unnecessary reward waste for inactive behaviors that do not actually participate in learning.
[0099] Next, the reward dynamic adjustment unit (162) can perform a step of enhancing the fairness of the reward by reflecting relative evaluation elements in the calculated basic reward coefficient (α). In this step, real-time statistical data, such as the average concentration index (AIavg) of a group of peer learners who performed the same scenario, is referenced to compare where the individual learner's concentration index (AI) is located relative to the average. If the individual learner's AI is significantly higher than AIavg, a predetermined upward adjustment coefficient greater than 1 is multiplied to the basic reward coefficient (α), and conversely, if the AI is lower than AIavg, a downward adjustment coefficient less than 1 is multiplied, thereby implementing a relative evaluation-based reward system that differentially reflects relative excellence within the group in addition to the individual's absolute performance in the reward.
[0100] Ultimately, through this adaptive reward structure, the system goes beyond simply processing individual reward cases to optimizing the reward system itself from a long-term perspective. Specifically, the system continuously monitors the impact of paid rewards on learners' subsequent learning behaviors (e.g., learning persistence rate, frequency of attempting the next scenario) and uses this data as input for reinforcement learning models. By automatically fine-tuning the shapes and parameters of each non-linear reward function to maximize the long-term learning motivation effect for the entire learner group, the system is configured to actively respond to various environmental variables—such as changing average learner achievement levels, scenario difficulty, and platform participation rates—to maintain an optimal reward curve.
[0101] In the following, specific examples of how the adaptive reward structure of the aforementioned reward dynamic adjustment unit (162) operates in actual learning scenarios are described in detail through the case of a hypothetical learner with different motivation profiles.
[0102] As a first example, let us assume a case where an administrator sets the minimum reward points (Rmin) to 10 points and the maximum reward points (Rmax) to 200 points for a specific high-difficulty coding practice scenario. When Learner A, classified as having a 'Challenge Achievement' profile, participates in the above scenario, the Learner possesses strong intrinsic motivation to solve high-level tasks. During the execution of the scenario, Learner A demonstrated excellent problem-solving skills, recording a 100% behavioral success rate, a performance speed 1.5 times faster than the peer group average, and a zero error rate. Accordingly, the system normalized the Learner's behavioral event performance indicator (Eb) and calculated a high value of 0.98. At the same time, the concentration analysis module detected a state of 'continuous immersion,' in which Learner A's gaze remained fixed on the task area and input patterns were logically consistent without hesitation, and calculated the Comprehensive Concentration Index (AI) as 0.95. At the point when the above scenario is completed, the reward dynamic adjustment unit (162) first checks the 'challenge achievement type' profile of learner A and selects f(AI,Eb) from the reward function library in the form of an exponential function in which the reward value increases explosively in the highest level achievement range. Subsequently, when the AI value of 0.95 and the Eb value of 0.98 are input into the selected function, a very high base reward coefficient (α) of 0.99 is output. Furthermore, the system confirms that learner A's AI and Eb significantly exceed the average concentration index (AIavg) and average performance index (Eb_avg) of the peer learner group, and applies upward correction based on relative evaluation to determine the final reward coefficient as 1.0. Finally, the above-determined values are substituted into [Equation 1] to calculate Ract = 10 + (200 - 10) * 1.0 = 200, i.e., the maximum reward point of 200 points is calculated as the actual reward to be paid.This serves to further reinforce the learner's sense of challenge by providing maximum rewards corresponding to the highest level of effort and performance.
[0103] As a second example, we assume a case where Learner B, classified as having an 'initial encouragement-dependent' profile, participates in the same scenario. This learner has low tolerance for failure and tends to struggle in the early stages of the process. Learner B barely manages to pass the scenario after several trials and errors, recording a 70% behavioral success rate, a performance speed 1.2 times slower than average, and a high error rate; accordingly, the behavioral event performance index (Eb) was calculated as 0.65. During the performance process, Learner B's concentration index (AI) showed a 'hardship overcoming' pattern, temporarily dropping when an incorrect answer occurred and then rising again, and the average AI was measured as 0.70. In the above case, the reward dynamic adjustment unit (162) recognizes Learner B's profile and selects f(AI,Eb) in the form of a logarithmic function that guarantees a certain level of reward even in the low achievement range to prevent learning abandonment. As a result of inputting an AI value of 0.70 and an Eb value of 0.65 into the above logarithmic function, a relatively high base reward coefficient (α) of 0.58 is output by positively evaluating the process of effort, even though the absolute performance figures are not high. However, since Learner B's performance is lower than the average of the peer group, a slight downward correction based on relative evaluation is applied, and the final reward coefficient may be adjusted to 0.52. Finally, Ract = 10 + (200 - 10) * 0.52 = 10 + 98.8 = 108.8, i.e., approximately 109 points are calculated as the actual reward paid. This demonstrates the characteristics of an adaptive reward structure that encourages the learner not to lose confidence in the next learning stage by acknowledging the effort of completing the task to the end without giving up, even though the results were unsatisfactory, as a meaningful reward.
[0104] As a third example, learner C participated in the same scenario, exhibiting an 'inactive' behavioral pattern of spending time with minimal manipulation without any intention to learn. Due to random clicking and quick abandonment, the learner recorded a low behavioral success rate of 30% and a very high error rate, and the behavioral event performance indicator (Eb) was calculated as 0.15. Additionally, as a result of the concentration analysis, the gaze frequently left the screen and there were almost no input patterns, so the concentration index (AI) was also measured as a very low value of 0.10. In this case, the reward dynamic adjustment unit (162) applies f(AI,Eb) in the form of a power function, in which the reward coefficient rapidly converges to near 0 for low-level input values. As a result of inputting an AI value of 0.10 and an Eb value of 0.15 into the above function, the basic reward coefficient (α) outputs a negligible value of 0.02. Here, a downward adjustment for relative evaluation regarding performance significantly lower than the peer group average is added, allowing the final reward coefficient to be determined as 0.01. Ultimately, the calculation determines that Ract = 10 + (200 - 10) * 0.01 = 10 + 1.9 = 11.9, meaning approximately 12 points are paid, which is close to the guaranteed minimum reward point (Rmin) of 10 points. This demonstrates the process by which the system operates to maintain the fairness of the reward system and the efficiency of resources by precisely identifying inactive and abusive behaviors of learners and minimizing the payment of unnecessary rewards in such cases.
[0105] FIG. 6 is a flowchart of a method using a behavior reward platform system according to an embodiment of the present invention. Referring to FIG. 6, the present method includes steps (S10) to (S60). For convenience of explanation, descriptions that overlap with the foregoing details will be omitted.
[0106] Step (S10) is a step of configuring a scenario, wherein learning content created by multiple content creators is uploaded to the scenario input section of the platform, and an administrator or developer selects or combines one or more of the uploaded scenarios to configure a scenario flow, which is a learning process, and sets the reward conditions and granting method for each scenario. Here, prior to the start of Step (S20), additional steps may be included in advance, such as displaying the range of reward points and reward conditions set in the scenario on the learner terminal.
[0107] Step (S20) is a step for collecting behavioral events, wherein when the scenario flow is executed through a learner terminal, behavioral data such as the learner's viewing, clicking, input, response time, accuracy rate, and execution time are collected as behavioral events.
[0108] Step (S30) is a step for determining reward conditions, which is a step of determining whether the collected behavioral events satisfy the reward conditions set for each scenario using a reward mapping engine.
[0109] Step (S40) is a step for analyzing the learner's concentration, which involves collecting data such as the learner's gaze, reaction time, input pattern, error rate, and learning time, calculating at least one of visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb), and then calculating a concentration index using a non-linear combination function.
[0110] Step (S50) is a step for calculating reward points for a learner, wherein the concentration index (AI) and the learner's behavioral event performance indicator (E b f(AI,E), a non-linear function that takes ) as input b The step involves calculating a reward coefficient of 0 to 1 through ), and then calculating the actual reward points, wherein the function is configured to be dynamically updated according to changes in concentration and performance results.
[0111] Step (S60) is a step of paying reward points to a learner, wherein the reward points are automatically paid to the learner by deducting the actual reward points calculated in Step (S50) from the reward points pre-charged by the administrator.
[0112] FIG. 7 shows an exemplary UI actually implemented in an APP on a learner terminal of a behavior reward platform system according to one embodiment of the present invention.
[0113] Referring to the exemplary UI illustrated in Fig. 7, the learner can intuitively check and operate their reward accumulation history, participation events, and preferred content selection process through an application provided on a mobile terminal.
[0114] The first screen displays the learner's profile and menu interface, showing that the learner can access detailed menus such as point (cash) accumulation history, event participation, customer center, and settings.
[0115] The second screen is the reward (cash) accumulation history page, which displays the total balance of reward points earned by the learner through activities, as well as the accumulation and usage history. It also shows that detailed information, such as accumulation conditions and expiration dates for each activity type, is provided.
[0116] The third screen serves as an example of a behavioral event, demonstrating the process of accumulating reward points when a learner completes a specific learning activity or mission (e.g., watching videos, consuming content, etc.). On this screen, differential rewards may be provided based on the completion time or difficulty of each mission, and the mission progress and remaining reward points are displayed in real-time.
[0117] The fourth screen is a step for receiving the learner's preferred content (e.g., selection of music, genre, or musician), and is configured so that the system provides personalized recommended content or additional reward bonuses based on the learner's selection. At this stage, the selection UI is configured in the form of touch-based card-style buttons, demonstrating support for the learner's intuitive selection.
[0118] In other words, this invention exemplarily demonstrates that the behavioral reward platform of the present invention is linked with a reward calculation and payment system, enabling the implementation of a reward-based interface based on learner behavior data in an actual terminal environment.
[0119] In addition, one embodiment of the present invention may further include the following features.
[0120] A behavioral reward platform system according to one embodiment of the present invention may further include a social learning reward module that precisely analyzes interactions between learners occurring within a specific scenario flow, quantitatively evaluates their contribution, and grants differential rewards, as a configuration for activating a cooperative learning ecosystem formed among multiple learners beyond rewarding the learning activities of individual learners.
[0121] The above-described social learning reward module includes an interaction analysis unit that identifies cooperative behavioral events, such as question-and-answer sessions, discussions, and material sharing among learners, and collects them as structured data. The interaction analysis unit comprises a cooperative event identification and tagging unit that parses a data stream within a shared learning space in real time to identify event types, such as 'posting a question,' 'submitting an answer,' and 'adding supplementary explanations,' and tags metadata including the time of occurrence of the event, participating learner identifiers, and relevant scenario context information.
[0122] In addition, the social learning reward module includes a contribution calculation engine that receives structured data collected by the interaction analysis unit and evaluates the positive impact of a specific learner's contribution on other learners from various perspectives. The contribution calculation engine includes a semantic relevance analyzer that evaluates how well a submitted answer aligns semantically with the intent of the question using natural language processing (NLP) technology; a community response quantification unit that quantifies the number and frequency of positive feedback, such as 'likes' or 'usefulness,' obtained from other fellow learners; and an originality and influence evaluator that determines whether the contribution content possesses originality or presents a new perspective through comparison with previously submitted materials. The system is configured to calculate a single contribution score by synthesizing the evaluation results of multiple factors derived from the analyzer, unit, and evaluator according to preset weights.
[0123] The system provides a structure in which the calculated contribution score is transmitted to a reward calculation module of the system, and the reward calculation module calculates and distributes cooperative reward points differentially in proportion to the contribution score from a cooperative reward pool pre-charged by an administrator, separate from the reward points awarded based on individual learning achievement, and applies an exponentially increasing reward function to excellent interactions where the contribution score exceeds a specific threshold, thereby inducing the establishment of a learning community within the platform where high-quality knowledge is actively shared and mutually developed.
[0124] The above-mentioned scenario selection unit is configured to dynamically design and recommend an optimized learning path for individual learners by utilizing vast amounts of learner group data accumulated on the platform, and may include a machine learning-based artificial intelligence model that infers the scenario flow most suitable for a specific learner's unique learning state by using the success and failure history of scenario flows performed by multiple other learners, learning time, and patterns of change in concentration as learning data.
[0125] The artificial intelligence model described above includes a preprocessing unit that extracts multidimensional features from a learner's learning history and current interactions and converts them into input data in a normalized form. The preprocessing unit collects raw learning state information in a fragmented form, such as the learner's viewing time, correct answer rate, reaction time, concentration index, and past reward acquisition history, and performs a data normalization step that corrects missing values and unifies the scale of each data. Subsequently, it analyzes the normalized time-series data to undergo a temporal feature extraction step that reflects temporal contexts, such as the rate of change of learning speed, repetitive incorrect answer patterns for specific types of problems, and signs of decreased interest due to a reduction in recent learning activities. Finally, it performs a process of generating a high-dimensional behavioral characteristic vector that comprehensively represents the learner's current knowledge state, learning propensity, and cognitive characteristics by combining these static and dynamic features.
[0126] The behavioral characteristic vector generated by the above preprocessing unit is provided as input information to a deep learning-based prediction computing unit designed to perform various prediction tasks simultaneously. The prediction computing unit is configured to receive the behavioral characteristic vector based on a Multi-task Learning architecture and to comprehensively output multiple prediction results by including: (a) a scenario selection probability prediction model that predicts the probability of a learner selecting a particular unlearned scenario and the expected correct answer rate for each unlearned scenario existing within the platform; (b) a reward responsiveness prediction model that estimates the expected level of concentration of the learner when each scenario is performed and the expected value of the reward points that can be obtained accordingly; and (c) a knowledge gap analysis model that analyzes the learner's past incorrect answer patterns to identify scenarios related to core concepts deemed most necessary in the current state of knowledge.
[0127] Finally, the multidimensional prediction result vector output from the prediction operation unit is transmitted to a scenario recommendation unit that constructs and provides an optimal scenario sequence, or scenario flow, to be provided to the learner. The scenario recommendation unit operates based on a Multi-Objective Optimization algorithm that considers not only short-term reward maximization but also long-term learning effects, and performs a process of finding an optimal balance point by comprehensively considering the predicted scenario selection probability, reward responsiveness, and the importance of bridging the knowledge gap. Specifically, the scenario recommendation unit can provide a system characterized by prioritizing the selection of scenarios that have a high reward expectation and can most effectively bridge the learner's current knowledge gap, adjusting the difficulty level to prevent the consecutive placement of scenarios that are too difficult or too easy by considering the predicted cognitive load, and thereby finally constructing and providing a personalized scenario flow to the learner's terminal that offers a challenging learning experience without losing the learner's interest.
[0128] The reward calculation module described above is configured to include a weight calculation unit that calculates weights to be applied to the final reward value according to a preset sophisticated calculation method, as a core component for dynamically adjusting the size of the reward by evaluating the learner's performance and the quality of the process from various angles, rather than simply assigning a fixed value of reward points when receiving a condition fulfillment signal from the reward mapping engine.
[0129] The above-mentioned weight calculation unit may be configured to calculate weights using a first method, a second method, or a method combining both methods in order to evaluate the learner's achievement not only by absolute standards but also by relative standards when determining learning rewards, and furthermore, to reflect not only the result but also the level of immersion in the process.
[0130] The above-described first method is a method that applies relative weights by evaluating the level at which a current learner's performance is positioned through comparison with other learner groups that have performed the same scenario m. To this end, a peer group clustering step is first performed to dynamically identify peer learner groups exhibiting the most similar learning patterns among all learners in the system, based on a profile vector that includes the current learner's cumulative learning history, average accuracy rate, learning speed, etc. Subsequently, a step is taken to calculate the real-time statistical distribution of behavioral events (e.g., average completion time, average accuracy rate, average number of interactions) exhibited by the identified peer group while performing the corresponding scenario m, i.e., the Dynamic Performance Baseline. Finally, the results of the current learner's behavioral events are compared with the Dynamic Performance Baseline to calculate a Z-score or percentile rank indicating the relative position relative to the standard deviation. This is then converted into a relative performance weight between 0 and 1 according to a pre-set mapping function, thereby configuring the method to reflect in the reward how superior the performance was compared to others.
[0131] The second method described above is a method that directly applies an Attention Index, derived by precisely analyzing cognitive and behavioral immersion levels during the execution of the m-th scenario based on the learner's behavioral events, to weights. It operates through a temporal attention pattern analysis engine that goes beyond simply using the final Attention Index value at the time of scenario completion and analyzes changes in the Attention Index from the start to the end of the scenario as time-series data. The engine analyzes the shape of the attention graph to identify patterns such as (a) a 'continuous immersion' pattern that maintains a consistently high level of attention throughout, (b) an 'awakening' pattern in which attention rises rapidly from a specific point after overcoming low initial attention, and (c) a 'fatigue' pattern that starts with high attention but gradually declines over time. A system can be provided that assigns differential process contribution weights according to the types of the identified concentration patterns, for example, assigning high weights to 'continuous immersion' and 'awakening' patterns to reward the process of consistent effort and overcoming difficulties, and assigning relatively low weights to 'fatigue' patterns, thereby reflecting qualitative differences in the process in the reward even if the learning results are the same.
[0132] The above system may further include a concentration analysis module as a core component that comprehensively analyzes various behavioral and biological response data manifested during the learner's learning process to infer the learner's intrinsic immersion state as a quantitative indicator.
[0133] The concentration analysis module described above includes a concentration data collection unit that collects raw data necessary for concentration analysis in real time by organically linking with the hardware and software resources of the learner terminal. The concentration data collection unit is composed of a direct input layer that detects the learner's explicit input behavior and a sensor fusion layer that captures implicit bio-responses using sensors embedded in the terminal. The direct input layer collects input pattern data including keyboard input speed and patterns, mouse click coordinates and frequency, and swipe trajectories on the touchscreen, response time measured in milliseconds from the time of question presentation to the time of answer submission, and error rate data recording whether the submitted answer is correct or incorrect. At the same time, the sensor fusion layer collects gaze data by analyzing the video stream of the terminal's front camera in real time using a computer vision algorithm to extract the learner's pupil position, gaze vector, and blink rate per minute, and performs a process of precisely measuring pure learning time by distinguishing between time when active input is detected and idle time when it is not.
[0134] The heterogeneous raw data stream collected by the concentration data collection unit is transmitted to a concentration index calculation unit that converts it into a mutually comparable standardized indicator and calculates a final single concentration index. The processing of the concentration index calculation unit begins with a data normalization step in which each raw data is normalized to a value between 0 and 1, and then proceeds through a detailed index calculation step based on the normalized data to calculate multiple detailed concentration indices, such as (a) visual concentration (Fv) calculated based on the ratio and stability of the gaze retention time for the learning content area, (b) reaction concentration (Fr) calculated based on the deviation from the learner's individual average reaction time for the corresponding task type, (c) time concentration (Ft) calculated based on the ratio of pure learning time to total time spent, (d) error concentration (Fe) weighted inversely proportional to the error rate while considering the difficulty of the problem, and (e) behavioral concentration (Fb) which evaluates the quality of positive or negative interactions appearing in click frequency, scroll patterns, etc.
[0135] Finally, the multiple detailed attention indices (Fv, Fr, Ft, Fe, Fb) calculated above are integrated into a single comprehensive attention index through a non-linear combination step designed to make a comprehensive judgment most similar to changes in human dynamic states by learning the complex and non-linear interrelationships between them. The non-linear combination can be performed by an intelligent model, such as an Adaptive Neuro-Fuzzy Inference System (ANFIS) that has learned data patterns of successful learners, going beyond a pre-set simple mathematical function. The system provides a system characterized by precisely reflecting changes in the learner's dynamic cognitive and emotional states by complexly interpreting multi-dimensional data, such as inferring a specific state where 'visual attention (Fv) is high but response attention (Fr) is low' as a 'deep thinking' state rather than 'distraction'.
[0136] The concentration index calculation unit described above may be configured to overcome the limitations of applying uniform evaluation criteria to all types of learning activities and to measure immersion more precisely and fairly by simultaneously considering the unique characteristics of the task being performed and the personal tendencies of the learner, and may further include a scenario type analyzer that automatically analyzes the cognitive requirements of the scenario currently being performed and a meta-learning-based weight optimization engine that readjusts the importance of each detailed concentration index in real time based on the analysis results.
[0137] The scenario type analyzer described above sequentially or repeatedly performs a static analysis step, which analyzes the content format, metadata, and structure of the scenario at the time the learner starts the scenario, and a dynamic analysis step, which monitors the learner's actual interaction patterns to dynamically update the analysis results. In the static analysis step, it determines whether the scenario belongs to the type of video, text, interactive simulation, or mixed media, and further identifies, through tag information, whether the key cognitive ability required by the task is memorization, reasoning, problem-solving, or rapid response. Subsequently, in the dynamic analysis step, if a pattern is detected where the learner spends significantly more time on a specific type of content (e.g., video rather than text) within the mixed media scenario, it performs a process of updating the valid type of the scenario in real time based on the learner's actual behavior, such as redefining the actual type of the scenario as 'video-centered learning'.
[0138] The scenario type information identified in real-time by the scenario type analyzer is transmitted to the meta-learning-based weight optimization engine and used to dynamically reconstruct the weights of a non-linear combination function that calculates the overall concentration index. The optimization engine operates by including a global baseline model that holds a set of standard weights for each scenario type learned from the entire learner data, and a personalized meta-model that continuously learns the unique learning patterns of individual learners to fine-tune the baseline model to suit the individual. Specifically, when a learner starts a new scenario, the engine first retrieves the global baseline weights corresponding to the scenario type, and then performs internal loop optimization using the learner's personalized meta-model to fine-tune the baseline weights based on which concentration patterns the learner demonstrated in recent similar types of tasks to achieve high performance. As a result, even for a scenario such as watching a video, a personalized evaluation model is generated in real-time in which the weights for behavioral concentration (Fb) as well as visual concentration (Fv) are up-adjusted for learners for whom additional behaviors, such as note-taking, have shown a high correlation with performance.
[0139] Finally, the personalized weight set generated through the inner loop optimization is immediately applied to the non-linear combination function of the concentration index calculation unit to calculate a comprehensive concentration index optimized for both the characteristics of the task currently being performed and the learner's unique concentration pattern. Subsequently, when the execution result of the scenario (e.g., final quiz score) is confirmed, the success or failure is used as an input value for the outer loop learning that updates the individual meta-model, thereby providing a system characterized by enabling the system to progressively enhance its learning ability regarding 'how weights must be adjusted for this learner to succeed.'
[0140] The above system may further comprise an administrator account module, which serves as a core component on the administrator side for managing resources for the payment of rewards to learners, establishing reward policies, and transparently controlling the execution process, and which provides an interface for controlling the administrator's identity authentication and access rights and setting reward policies for a group of learners under their charge; and a reward charging management module that oversees the entire process of charging reward points, allocating budgets, and actually paying and deducting reward points to learners according to policies set by the administrator.
[0141] The above reward charging management module performs the function of an administrator charging reward points in advance by linking with an external payment system and managing the balance of the total charged points, but goes beyond simply managing a single balance and provides multiple virtual budget wallet functions that allow the administrator to independently allocate budgets by specific learning campaigns, learner groups, or scenario types. Subsequently, in the process of establishing a reward policy for a specific scenario, it includes a reward disclosure unit that performs the function of setting the minimum reward points (Rmin), which is the minimum reward level guaranteed to the learner when the scenario is completed, and the maximum reward points (Rmax), which is the maximum reward level obtainable when the best performance is achieved. The above-mentioned reward disclosure unit is equipped with an AI-based reward range recommendation engine that analyzes data on the average success rate and time taken by learners who have previously performed the relevant scenario to assist the administrator in setting up the system, and recommends a range of Rmin and Rmax values capable of producing the optimal motivational effect within the budget. Once the setting is complete, it clearly and visually presents the range of expected rewards to the learner's terminal just before the learner starts the k-th scenario, thereby fostering anticipation and a sense of challenge regarding the learning goal.
[0142] At the point when the execution of the k-th scenario is completed by the learner, the reward dynamic adjustment unit within the reward charging management module operates to perform a precise calculation process to calculate the actual reward points to be finally paid. The reward dynamic adjustment unit first receives as input data the final comprehensive concentration index (AI) calculated from the concentration analysis module during the execution of the scenario, and the behavioral event performance indicator (Eb) calculated by combining the learner's correct answer rate, execution speed, error rate, etc. Subsequently, the received AI and Eb values are input into a non-linear function f(AI,Eb) that outputs a reward coefficient between 0 and 1. At this time, the function f is not a fixed function applied equally to all learners, but is selected and applied from a function library whose form is optimized for each individual according to motivation profiles classified as 'challenge-seeking type', 'stability-oriented type', etc., by analyzing the learner's long-term learning patterns. Finally, the reward coefficient calculated through the above-mentioned personalized non-linear function f and the Rmin and Rmax values pre-set by the above-mentioned reward disclosure unit are substituted into [Equation 1] below to calculate the actual reward points (Ract) that accurately reflect the quality of the learner's performance and effort.
[0143]
[0144] The system provides a structure characterized by, once the above operation is completed, the reward dynamic adjustment unit deducts points equivalent to the calculated Ract value from the virtual budget wallet allocated to the corresponding scenario and generates a transaction to transfer them to the corresponding learner's account, and records all of these processes (scenario ID, applied Rmin / Rmax, AI and Eb values, final Ract) in an immutable reward ledger to ensure transparency and audit traceability of reward payments.
[0145] A system according to one embodiment of the present invention can be equipped with a sophisticated reward economic system that simultaneously achieves universal participation inducement at the platform level and goal-oriented reward design at the manager level by dualizing the resources for reward points and dynamically controlling the payout ratio. The reward points operated by the system are clearly distinguished into common points, which are generated by the platform operator for the purpose of activating the entire system and can be universally distributed to an unspecified number of learners performing scenarios, and individual points, which are distributed only to specific learners managed by an administrator as an incentive for achieving goals, based on resources privately charged into a virtual wallet within their manager account module by the administrator managing a specific group of learners. The payout ratio of the common points and individual points can be pre-set by the administrator through the reward setting section within the content authoring tool module according to the nature of each scenario and educational goals; however, the ratio is configured to be dynamically changeable, such as by increasing the ratio of individual points for mandatory assignments to ensure responsible compensation within the manager's budget, or increasing the ratio of common points for autonomous inquiry assignments to encourage participation without budget burden.
[0146] To support the stable and efficient operation of such a dual reward system, the system may further include an AI-based reward budget prediction module that is closely linked with the administrator account module to intelligently assist the administrator in charging reward points and establishing budget plans. The operation process of the budget prediction module begins with a data collection and refinement stage, in which a vast amount of learning status information generated from all learners belonging to a specific learner group managed by the administrator is collected in real-time and batch formats and accumulated in a time-series database. The collected learning status information consists of a multidimensional vector including viewing time, which indicates the learner's content consumption pattern; the correct answer rate, which directly reflects the level of knowledge acquisition; reaction time, which is related to the speed of task resolution; the concentration index, which indicates the qualitative aspect of the learning process; and past reward history, from which the level of motivation can be inferred through past reward acquisition patterns.
[0147] Next, a prediction modeling and inference step is performed to predict future reward consumption by learning the accumulated time series data. The budget prediction module utilizes a time series prediction model based on a recurrent neural network, such as LSTM (Long Short-Term Memory), which is suitable for learning complex temporal patterns, such as the periodicity of learner activities, patterns of sudden surges in activity volume according to deadlines, and changes in participation rates when new scenarios are assigned. The model learns the correlation between the past learning state information of a specific learner group and the individual points actually paid, simulates the expected learning activity volume of the group and the corresponding achievement level during a specific future period (e.g., next week, next month), and predicts the expected value of the total individual points to be paid along with a probabilistic distribution based on this.
[0148] Finally, a warning and reporting step is performed to recommend preemptive measures to the administrator by comparing the predicted future point consumption amount with the actual balance of individual points currently remaining in the administrator account module. If the predicted consumption amount exceeds the current balance and future budget depletion is anticipated, the budget prediction module automatically sends a detailed warning notification in advance through the administrator's dashboard and designated communication channels, which includes (a) a specific date and time when the current balance is predicted to be depleted, (b) a recommended top-up amount required for stable operation during a future set period (e.g., 30 days), and (c) analysis information regarding specific scenarios or learner groups expected to have the greatest impact on budget depletion, thereby supporting the administrator in recharging the budget in a timely manner and reviewing the reward policy to ensure the uninterrupted operation of the reward system.
[0149] The above reward dynamic adjustment unit may be configured to include a reward function personalization engine that, as a core component for providing a hyper-personalized reward experience in which the system actively learns and adapts to the differences in each individual learner's intrinsic tendencies and motivational factors, moving away from the method of mechanically applying the same reward system to all learners, and by deeply analyzing the learner's long-term learning history and reward response patterns to generate the learner's unique 'Motivational Profile,' and dynamically selecting and applying the optimal reward function that best fits the said profile.
[0150] The operation process of the aforementioned reward function personalization engine begins with a feature engineering step that first aggregates the learning activity logs of a specific learner over an entire period accumulated in the system—that is, longitudinal data—and generates a multidimensional feature vector to infer the learner's motivational tendencies based on this data. In the feature engineering step, multiple quantitative indicators are calculated, such as (a) a Grit Score, which is calculated by measuring the frequency of retrying a difficult task after failure; (b) a Challenge-Seeking Index, which indicates the tendency to select a more difficult task when given options of different difficulty levels; and (c) Reward Sensitivity, which analyzes whether concentration and achievement levels significantly increase in a scenario immediately after obtaining a high reward.
[0151] Subsequently, the engine performs a profiling step in which it dynamically identifies several statistically significant learner types, i.e., motivation profiles, by using the multidimensional feature vectors produced above as input and performing an unsupervised learning-based clustering algorithm on all learners within the platform, and assigns each learner to the most suitable profile. Through the clustering process, intrinsic motivation types can be automatically defined, such as, for example, a 'stable growth type' who prefers gradual growth through consistent effort, a 'challenge achievement type' who gains a great sense of accomplishment by overcoming high-difficulty tasks, and an 'early encouragement dependent type' who easily loses interest if positive feedback is not provided early in learning.
[0152] Finally, based on the motivation profile assigned to a specific learner in the profiling step described above, the system performs a function personalization and application step in which it selects the form of the non-linear function $f(AI,E_b)$ most effective for the profile from a predefined reward function library, and fine-tunes and applies the detailed parameters of the function to match the individual's feature vector. Specifically, the system provides a method for maximizing long-term learning engagement and satisfaction for learners with different tendencies by qualitatively changing and personalizing the reward payout curve itself according to the motivation profile, such as: (a) applying a logarithmic reward function to learners classified as 'initial encouragement-dependent' in which reward values increase rapidly in the low achievement range (low AI and Eb values) to prevent early dropout; (b) applying an exponential reward function to learners classified as 'challenge achievement-dependent' in which reward values increase explosively only when the highest level of achievement is reached to encourage the highest level of challenge; and (c) applying a sigmoid reward function to learners classified as 'stable growth-dependent' in which rewards are provided relatively evenly across all achievement levels while rewards for extreme performance gradually converge.
[0153] In addition, the present invention relates to a method for providing optimal rewards based on a learner's behavior by utilizing the organic components of the aforementioned system, and said method may be implemented by including the following steps.
[0154] (a1) As a scenario configuration step, the step begins with the process of uploading various forms of learning content, such as videos, text, and VR / AR modules created by multiple content creators, to the system database through the scenario input section of the platform. Subsequently, an administrator or developer configures a scenario flow having a logical flow of learning by combining the uploaded individual scenarios in a drag-and-drop manner on the visual editor of the content authoring tool module, and sets transition conditions between each scenario node. Afterward, the administrator sets the reward conditions and granting method for each scenario in detail, including the process of specifying the range of minimum and maximum reward points (Rmin, Rmax) to be paid as rewards for each scenario through the reward setting section, setting the mixing ratio of common points paid across the entire system and individual points deducted from the administrator's individual budget, and selecting a reward policy regarding which side to give more weight to between the learner's performance and effort.
[0155] (a2) As a step for collecting behavioral events, when the scenario flow configured above is executed on the learner terminal, the system activates the behavioral event collection module to collect all interaction data of the learner in real time. The step includes the process of comprehensively collecting not only explicit behavioral data such as the learner's mouse clicks, keyboard inputs, and correct answer submissions, but also implicit behavioral data including gaze vectors detected through terminal sensors, blink frequency, and actual activity time excluding idle time during total learning time, and recording them as logs along with timestamps.
[0156] (a3) As a reward condition determination step, the collected behavior event logs are transmitted in real-time to a reward mapping engine and used to determine whether pre-set reward conditions are satisfied. In this step, the reward mapping engine first checks the current proficiency profile by referring to the learner's cumulative learning history, and based on this, dynamically adjusts the basic reward conditions set by the administrator in step (a1) to suit the learner. For example, even for the same 'correct answer rate of 80% or higher' condition, the target value can be reset to 70% for beginner learners and 95% for advanced learners. Subsequently, a process is performed to finally determine whether the actual behavior events collected in step (a2) satisfy the reward conditions dynamically adjusted for each individual.
[0157] (a4) As a step for analyzing the learner's concentration, the multidimensional behavioral event data collected in step (a2) is transmitted in parallel to a concentration analysis module and converted into a single comprehensive concentration index (AI) representing the learner's intrinsic immersion state. The step includes a series of processes in which raw data such as collected gaze, reaction time, and input patterns is normalized, multiple detailed indices such as visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavioral concentration (Fb) are calculated based on this, the weights of each detailed index are optimized in real time by considering the type of scenario being performed (e.g., watching a video, interactive practice) and the learner's individual meta-learning model, and the final concentration index is calculated by inputting this into a non-linear combination function.
[0158] (a5) As a step for calculating reward points for a learner, if it is determined that the reward conditions are satisfied in step (a3), the reward dynamic adjustment unit initiates a calculation to precisely calculate the final reward points to be paid. The above step receives as input the comprehensive concentration index (AI) calculated in step (a4) and the behavioral event performance indicator (Eb) which combines the success rate and execution speed of the corresponding scenario. Subsequently, the learner's long-term motivation profile ('challenge achievement type', 'stable growth type', etc.) is queried, and a personalized non-linear function f(AI,Eb) of the form most suitable for that profile (e.g., exponential function, logarithmic function) is selected to calculate a reward coefficient between 0 and 1. Finally, the process includes calculating the actual reward points (Ract) to be paid by substituting the calculated reward coefficient and the Rmin and Rmax values set in step (a1) into [Equation 1], and creating a payment transaction for the corresponding amount.
[0159] (a6) As a step for paying reward points to the learner, the payment transaction generated in step (a5) is transmitted to the reward charging management module for final execution. In this step, the system first divides the Ract according to the common / individual point ratio set in step (a1), deducts the amount corresponding to the individual points from the reward point balance pre-charged by the administrator, and deducts the common points from the system's common funds. Subsequently, points equal to the total amount deducted are paid to the learner's account, and all transaction details (applied conditions, AI and Eb values, final payment amount, etc.) are recorded in the reward ledger to ensure transparency, and the entire method is concluded by sending a reward payment completion notification to the learner's terminal.
[0160] In addition, the method according to the present invention may further include the following step (a11) as a preliminary step to induce immersion by making the learner aware of a clear goal for the task to be performed, prior to the initiation of the (a2) behavioral event collection step.
[0161] (a11) As a reward information prior provision and awareness step, the step is initiated at the point when the learner selects an entry to start performing a specific k-th scenario within the scenario flow. Upon the learner's selection signal, the system first retrieves reward policy information in real time that the administrator has pre-set for the k-th scenario in the (a1) scenario configuration step. Subsequently, based on the retrieved information, the system performs the process of displaying a dedicated interface, such as a 'challenge briefing' pop-up window or a dedicated screen, on the screen of the learner's terminal.
[0162] The information displayed above is characterized by being visually organized to maximize the learner's intuitive understanding and motivation, rather than being merely a list of text. Specifically, the interface includes: (a) clearly displaying the minimum reward points (Rmin) guaranteed by default upon completing the scenario and the maximum reward points (Rmax) obtainable upon achieving the best performance, along with numbers and graphic elements such as bar graphs or gauges, so that the range of rewards can be visually recognized. Additionally, (b) specific reward conditions that must be met to obtain the maximum reward points, such as 'completion within a time limit of 5 minutes,' 'achievement of an accuracy rate of 90% or higher,' and 'maintenance of a concentration index of 0.8 or higher,' are listed in the form of a checklist, helping the learner to recognize the clear goals to be achieved step by step.
[0163] Furthermore, information regarding the composition of the resources for the reward points provided in the above step may be displayed. The reward points may be divided into common points derived from resources prepared by the system operator to be distributed commonly to all learners on the platform, and individual points distributed as incentives only to specific learner groups based on resources privately charged to the accounts of administrators who directly manage the learners. The interface displays details regarding the ratio of common points and individual points that make up the final reward, such as 'System Basic Reward: 70 points' and 'Administrator Special Bonus: 30 points' out of 100 Rmax points. The distribution ratio of the common points and individual points can be dynamically adjusted by the administrator according to the importance or purpose of the scenario, and by transparently disclosing this to the learners, it allows learners to directly perceive the administrator's support and encouragement, thereby inducing an effect that enhances a sense of responsibility and bonding regarding their learning. After the learner has fully recognized the contents of the briefing screen above, they begin performing the scenario by pressing the ‘Start Challenge’ or ‘Confirm’ button, thereby transitioning to the above (a2) step.
[0165] The preferred embodiments according to the present invention described above may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software.
[0166] The above embodiments in the present invention are merely examples, and the present invention is not limited thereto. Any configuration substantially identical to the technical concept described in the claims of the present invention and achieving the same functional effect is included within the technical scope of the present invention. Explanation of the symbols
[0167] 100: System 110: Content Authoring Tool Module 120: Behavior Event Collection Module 130: Reward Mapping Engine 140: Reward Calculation Module 150: Concentration Analysis Module 160: Reward Charging Management Module
Claims
Claim 1 A reward calculation system for calculating rewards for a scenario based on the detection of a learner's behavior, comprising: a content authoring tool module configured to allow a developer or administrator to configure a scenario flow and assign reward conditions for said scenario flow; a behavior event collection module that collects learner behavior events when a scenario generated by said content authoring tool module is executed through a learner terminal; a reward mapping engine that determines whether said behavior events satisfy a preset reward condition; and a reward calculation module that calculates reward points according to a preset method when it is determined by said reward mapping engine that said behavior events satisfy the reward condition; wherein the content authoring tool module comprises: a scenario input unit in which a scenario, which is learning content generated by a plurality of content creators, is input to form a platform; a scenario selection unit that allows an administrator to configure a scenario flow, which is a learning process, by selecting or combining one or more of a plurality of scenarios uploaded on said platform; and a reward setting unit that individually sets reward conditions and methods for each scenario constituting said scenario flow. The scenario selection unit includes an artificial intelligence model that performs machine learning using the scenario flow of another learner as learning information using the platform; the artificial intelligence model includes a preprocessing unit that preprocesses the learner's learning state information to generate a behavioral characteristic vector for the learner, wherein the AI model calculates and provides a scenario flow for the learner using the learner's learning state information; a deep learning-based prediction computation unit that predicts and outputs a scenario selection probability and a reward responsiveness using the behavioral characteristic vector as input information; and a scenario recommendation unit that configures and provides a scenario flow optimized for the learner based on the prediction result of the prediction computation unit; and the reward computation module includes a weighting computation unit that calculates and applies weights according to a preset method when calculating the reward points.The system comprises: a weight calculation unit configured to calculate weights through: a first method of applying a relative weight to a behavioral event of a learner performing the m-th scenario based on a behavioral event of another learner performing the m-th scenario; or a second method of applying an attention index derived by analyzing the level of concentration on performing the m-th scenario based on the learner's behavioral event; and the system further comprises an attention analysis module that analyzes the level of concentration; wherein the attention analysis module comprises: an attention data collection unit that collects at least one data among a learner's gaze, reaction time, input pattern, error rate, and learning time; and an attention index calculation unit that normalizes the data collected by the attention data collection unit to calculate the attention index, wherein the attention index is calculated by considering at least one of visual attention (Fv), reaction attention (Fr), time attention (Ft), error attention (Fe), and behavioral attention (Fb). The concentration index calculation unit comprises: a scenario type analyzer that determines the type of the scenario by analyzing the content format and metadata of the scenario being performed, and updates the determined type by monitoring the interaction pattern of the learner; and a meta-learning-based weight optimization engine that adjusts the weights of the non-linear function combining the visual concentration (Fv), reaction concentration (Fr), time concentration (Ft), error concentration (Fe), and behavior concentration (Fb) in real time according to the type of the scenario determined by the scenario type analyzer.The system further comprises: a weight optimization engine including a global baseline model having a set of standard weights for each scenario type learned from data of multiple learners, and a personal meta-model that learns the learning patterns of individual learners and adjusts the set of standard weights; a personalized weight set is generated by adjusting the set of standard weights corresponding to the type of the identified scenario using the personal meta-model and applied to the non-linear function, and the system is configured to update the personal meta-model using the result of performing the scenario as input; the system further comprises: an administrator account module for an administrator managing the learners; and a reward charging management module that manages the deduction of the charged reward points when the administrator charges reward points in advance and the learner performs the k-th scenario; wherein the reward charging management module includes a reward disclosure unit that manages the balance of the total reward points charged by the administrator, sets a reward point allocation range for the k-th scenario, and provides the reward point allocation range to the learner before performing the scenario. and when the execution of the k-th scenario is completed by the learner, a non-linear function f(AI,E) that outputs a reward coefficient of 0 to 1 according to the concentration index and behavior event; b A reward dynamic adjustment unit that dynamically calculates the actual reward points paid to the learner based on ), wherein the reward dynamic adjustment unit, through the following mathematical formula 1, calculates the actual reward points (R act Performing ) and (here, R min The minimum reward points, R max is maximum reward points, E b is a behavioral event performance indicator including the learner's behavior success rate, execution speed, and error rate), [Mathematical Formula 1] The above reward points are divided into common points that can be commonly paid to all learners performing the scenario, and individual points that are paid only to specific learners based on reward points charged by the manager, and the payment ratio of the common points and individual points can be dynamically changed, and the learner's learning status information includes at least one of the learner's viewing time, correct answer rate, reaction time, concentration index, and past reward history, and the system further includes an AI-based reward budget prediction module that is linked with the manager account module to assist the manager in charging reward points, and the budget prediction module learns the past scenario performance history, average concentration index, and reward payment statistical data of a specific learner group to predict the total reward points to be consumed over a specific period in the future, and if the remaining points of the manager account module are insufficient compared to the predicted consumption amount, it sends a warning notification to the manager in advance including a recommended charging amount and an estimated time of depletion, and the reward dynamic adjustment unit further includes a reward function personalization engine that generates a 'motivation profile' by analyzing the individual learner's long-term learning history and reward response pattern, and the The personalization engine provides, (a) a logarithmic reward function f(AI,E) in which the reward increases rapidly in the low achievement range for learners who easily give up on small successes. b ) and (b) for learners who enjoy high challenges, an exponential reward function f(AI,E) in which rewards increase explosively in the highest level achievement range. b By applying ), the non-linear function f(AI,E) is tailored to the above motivation profile. b A reward calculation system characterized by optimizing the form of itself for each individual.
Citation Information
Patent Citations
Learning value decision method to determine learning value of learners' learning activities
KR1020200001798A
Methods and programs for providing personalized learning plans using artificial intelligence
KR1020200056574A
Customized learning management system and method through artificial intelligence
KR1020240042990A
Authoring tool system to help create leveled, branched content in the form of webtoons for elementary education
KR102842433B1