A personalized learning platform user behavior data processing method and system

CN122817451APending Publication Date: 2026-09-25ZHEJIANG COLLEGE OF SECURITY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611291532.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这种跨数据源、跨时间维度信息的整合,往往因数据异构性、数据量庞大和查询延迟而导致效率低下

Benefits of technology

[0014]根据本申请实施例的技术方案,至少具有如下有益效果:本申请公开的个性化学习平台用户行为数据处理方法,通过获取期望处理的样本和物理样本,并根据预设的检测任务队列确定期望处理的样本的身份标识,将其推送至当前操作台。随后,通过扫描设备对物理样本进行解码,并获取用户对物理样本标识与期望处理的样本身份标识的核对结果。特别地,当物理样本标识磨损导致扫描设备无法完全解码时,本申请引入智能匹配机制,将扫描设备捕获的信息与期望处理的样本身份标识进行智能匹配,得到智能匹配结果。最终,根据核对结果和智能匹配结果,控制物理样本执行后续处理流程。本申请有效解决了现有技术中因样本标识磨损或手动输入错误导致的样本身份错位问题。在现有技术中,当条码标签磨损或手动输入错误时,系统往往无法识别这种“替换型”错误,导致检测数据被错误关联、覆盖或缺失,严重影响检测报告的准确性和可信度。本申请通过引入用户核对和智能匹配的双重校验机制,即使在物理样本标识受损的情况下,也能通过智能算法对模糊信息进行匹配,大大降低了人工干预的错误率。当核对结果或智能匹配结果显示不一致时,系统能够及时预警并阻止错误样本进入后续流程,从源头上避免了数据污染。这不仅提高了检测数据的准确性和可追溯性,还显著提升了质量控制的效率和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817451A_ABST
    Figure CN122817451A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of personalized learning platform user behavior data processing method and system, it is related to personalized learning platform technical field, method includes obtaining sample and physical sample of expected processing;According to the identity of the sample of expected processing determined by the preset detection task queue, and the identity of the sample of expected processing is pushed to current operating platform;Physical sample is decoded by scanning equipment;Obtain the collation result that user collates the identity of physical sample with the identity of the sample of expected processing through current operating platform;When the identity of physical sample is abraded and causes scanning equipment to be unable to decode completely, information captured by scanning equipment is intelligently matched with the identity of the sample of expected processing, and intelligent matching result is obtained;According to collation result and intelligent matching result, control physical sample to execute subsequent processing procedure.The application can improve the efficiency of online learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of personalized learning platform technology, and in particular to a method and system for processing user behavior data in a personalized learning platform. Background Technology

[0002] In related technologies, online learning platforms collect vast amounts of user behavior data to provide customized learning experiences. However, with the introduction of deep learning functions such as open-ended question answering and essay writing, users are generating massive amounts of unstructured text data with highly diverse content and structure. Existing systems struggle to perform efficient deep semantic analysis when processing this complex text, and they also find it difficult to quickly integrate the analysis results with the user's dynamic learning background information. This results in personalized recommendations failing to respond promptly and accurately to the user's immediate needs, making the recommendations appear lagging and irrelevant. When users input open-ended text on personalized learning platforms, such as writing essays or submitting detailed answers, existing systems face multiple challenges in processing these continuously generated and dynamically changing unstructured text streams. First, extracting deep semantic information from this text in real-time and efficiently, such as core concepts, argument structures, potential knowledge errors, and logical loopholes, requires significant computational resources and time. Second, quickly and accurately associating and integrating these complex semantic analysis results with dynamic background information such as the user's current learning task status, past learning trajectory, and real-time conversation status is a bottleneck that existing systems struggle to overcome. The integration of information across data sources and time dimensions often results in inefficiency due to data heterogeneity, massive data volumes, and query latency. Therefore, existing technologies struggle to identify potential gaps in a user's knowledge or deviations in their thinking before or during input, and thus fail to provide adaptive learning resource suggestions that closely align with the user's thought process. This lag significantly diminishes the value of personalized recommendations, preventing the achievement of truly adaptive learning support that "synchronizes with the user's thinking." Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method and system for processing user behavior data in a personalized learning platform, aiming to improve the efficiency of online learning.

[0004] In a first aspect, embodiments of this application provide a method for processing user behavior data on a personalized learning platform, including: When a user inputs text on the platform, the system acquires the user's text input operation events and updates the text input content based on the operation events. The operation events include character input, deletion, cursor movement, text selection, and pasting. Perform semantic analysis on the updated text input to obtain the semantic analysis results; The system acquires the user's current learning task information, past learning performance information, and current session status information, and integrates these information into a real-time learning status snapshot. Based on the semantic analysis results and the real-time learning state snapshot, the user's knowledge comprehension bias or insufficient thinking is identified, and the identification result is obtained; Get the current text position of the user's cursor; Based on the recognition results and the current text position of the cursor, the intervention method and timing are determined to select learning resources from a preset learning resource library and push the learning resources to the user.

[0005] According to some embodiments of this application, the step of performing semantic analysis on the updated text input to obtain semantic analysis results includes: When the text input is a paragraph, semantic analysis is performed on the updated text input based on a pre-trained language model to obtain core concepts, argumentation logic, and specific knowledge points. Based on the core concepts, the argumentation logic, and the specific knowledge points, the semantic analysis results are obtained.

[0006] According to some embodiments of this application, semantic analysis results are obtained based on the core concepts, the argumentation logic, and the specific knowledge points, including: Obtain the semantic vector of the updated text input content; The semantic vector and the preset initial semantic vector are calculated to obtain the dot product between the semantic vector and the preset initial semantic vector, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic deviation threshold is obtained based on the dot product, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic analysis results are obtained by considering the core concepts, the argumentation logic, the specific knowledge points, and the semantic deviation threshold.

[0007] According to some embodiments of this application, the step of obtaining the user's current learning task information, past learning performance information, and current session state information, and integrating the current learning task information, past learning performance information, and current session state information into a real-time learning state snapshot, includes: Obtain the user's current learning task information, wherein the current learning task information includes the topic, knowledge points, and core concepts of the current learning task; Obtain past learning performance information, which includes the accuracy rate in exercises, the viewing time in video courses, and the assessment of the content of notes and the degree of understanding of concepts; Obtain current session state information, wherein the current session state information includes the number of characters of text entered by the user, the duration of the current session, and the area where the user's cursor is located in the text; The current learning task's topic, knowledge points, core concepts, accuracy rate in exercises, viewing time in video courses, comprehension assessment of notes and concepts, number of words in user-inputted text, duration of the current session, and the area where the user's cursor is positioned in the text are integrated into a real-time learning status snapshot.

[0008] According to some embodiments of this application, the step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: When the semantic analysis results show that the user's understanding of the core concepts of the current learning task information differs from the standard definition in the real-time learning state snapshot, the text similarity of the difference is calculated to obtain the current difference value; If the current difference value is less than the preset difference value, the identification result is a concept comprehension deviation.

[0009] According to some embodiments of this application, the step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: When the semantic analysis results identify inconsistencies or contradictions between the argument logic in the text input and the standard argument logic in the real-time learning state snapshot, the identification result is that the argument logic is missing.

[0010] According to some embodiments of this application, the step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: If a user's current learning task is required to cover a specific knowledge point in the real-time learning state snapshot, and the semantic analysis results and the text entered by the user do not contain the specific knowledge point, the identification result is that the knowledge point is missing.

[0011] According to some embodiments of this application, the step of determining the intervention form and timing based on the recognition result and the current text position of the cursor to filter learning resources from a preset learning resource library and push the learning resources to the user includes: Obtain the semantic completeness of the user's input and the user's past learning preferences; Based on the semantic completeness of the user's input, the user's past learning preferences, the recognition results, and the current text position of the cursor, the intervention method and timing are determined to select learning resources from a preset learning resource library and push the learning resources to the user.

[0012] According to some embodiments of this application, determining the intervention form and timing to filter learning resources from a preset learning resource library and push the learning resources to the user includes: The intervention method and timing are determined to select learning resources that highly match the user's current intent and context based on the preset learning resource library, determine the type, difficulty and relevance score of the learning resources, and push the learning resources to the user.

[0013] Secondly, embodiments of this application provide a personalized learning platform user behavior data processing system, comprising: The acquisition and update module is used to acquire the user's text input operation events when the user inputs text on the platform and update the content of the text input according to the operation events. The operation events include character input, deletion, cursor movement, text selection and pasting. The semantic analysis module is used to perform semantic analysis on the updated text input and obtain the semantic analysis results; The integration module is used to obtain the user's current learning task information, past learning performance information, and current session status information, and integrate the current learning task information, past learning performance information, and current session status information into a real-time learning status snapshot; The identification module is used to identify user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtain the identification results. The text position acquisition module is used to obtain the current text position of the user's cursor; The filtering and push module is used to determine the form and timing of intervention based on the recognition results and the current text position of the cursor, so as to filter learning resources from a preset learning resource library and push the learning resources to the user.

[0014] The technical solution according to the embodiments of this application has at least the following beneficial effects: The personalized learning platform user behavior data processing method disclosed in this application obtains the sample to be processed and the physical sample, and determines the identity identifier of the sample to be processed according to the preset detection task queue, and pushes it to the current operation table. Subsequently, the physical sample is decoded by a scanning device, and the user's verification result of the physical sample identifier and the identity identifier of the sample to be processed is obtained. In particular, when the physical sample identifier is worn and the scanning device cannot fully decode it, this application introduces an intelligent matching mechanism to intelligently match the information captured by the scanning device with the identity identifier of the sample to be processed, and obtains an intelligent matching result. Finally, based on the verification result and the intelligent matching result, the physical sample is controlled to execute the subsequent processing flow. This application effectively solves the problem of sample identity misalignment caused by sample identifier wear or manual input errors in the prior art. In the prior art, when the barcode label is worn or there is a manual input error, the system often cannot identify this "substitution" error, resulting in the detection data being incorrectly associated, covered or missing, which seriously affects the accuracy and credibility of the detection report. This application introduces a dual verification mechanism of user verification and intelligent matching. Even when physical sample identifiers are damaged, intelligent algorithms can still match ambiguous information, significantly reducing the error rate of manual intervention. When the verification results or intelligent matching results show inconsistencies, the system can promptly issue warnings and prevent erroneous samples from entering subsequent processes, avoiding data contamination at the source. This not only improves the accuracy and traceability of detection data but also significantly enhances the efficiency and reliability of quality control.

[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0016] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0017] Figure 1 A flowchart illustrating a user behavior data processing method for a personalized learning platform provided in one embodiment of this application; Figure 2 This is a schematic diagram of a user behavior data processing system for a personalized learning platform provided in one embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical methods, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that the meaning of "multiple" (or "more than") in the description of the embodiments of this application refers to two or more, and "greater than," "less than," "exceeding," etc. are understood to exclude the number itself, while "above," "below," "within," etc. are understood to include the number itself. If "first," "second," etc. are used in the description, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.

[0020] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: the existence of a alone, the existence of b alone, the existence of c alone, the simultaneous existence of a and b, the simultaneous existence of a and c, the simultaneous existence of b and c, or the simultaneous existence of a, b, and c, where a, b, and c can be single or multiple.

[0021] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0022] Based on the above, this application proposes a method and system for processing user behavior data on a personalized learning platform, aiming to improve the efficiency of online learning.

[0023] The personalized learning platform user behavior data processing method provided in this application embodiment can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms; the software can be an application that implements the personalized learning platform user behavior data processing method, but is not limited to the above forms.

[0024] This application can be applied to numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via communication networks. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices. It should be noted that in various specific embodiments of this invention, when processing data related to the characteristics of an object (e.g., a user's attribute information or a set of attribute information) is required, permission or consent from the corresponding object is obtained first, and the collection, use, and processing of this data comply with relevant laws and standards. Furthermore, when the embodiments of the present invention need to obtain the attribute information of an object, they will obtain the separate permission or separate consent of the corresponding object through pop-up windows or redirection to a confirmation page. After obtaining the separate permission or separate consent of the corresponding object, they will then obtain the relevant data of the object necessary for the embodiments of the present invention to operate normally.

[0025] See Figure 1 , Figure 1 This is a flowchart illustrating a method for processing user behavior data in a personalized learning platform according to an embodiment of this application. The method includes, but is not limited to, steps S110 to S160, which will be described in detail below.

[0026] Step S110: When a user inputs text on the platform, obtain the user's text input operation events and update the text input content according to the operation events. Operation events include character input, deletion, cursor movement, text selection, and pasting. Step S120: Perform semantic analysis on the updated text input to obtain the semantic analysis results; Step S130: Obtain the user's current learning task information, past learning performance information, and current session status information, and integrate the current learning task information, past learning performance information, and current session status information into a real-time learning status snapshot; Step S140: Based on the semantic analysis results and real-time learning status snapshots, identify user knowledge comprehension biases or insufficient thinking, and obtain the identification results; Step S150: Obtain the current text position of the user's cursor; Step S160: Based on the recognition results and the current text position of the cursor, determine the form and timing of intervention to select learning resources from the preset learning resource library and push the learning resources to the user.

[0027] It's important to note that operational events refer to various interactive behaviors generated by the user during text input, such as typing characters, deleting text, moving the cursor within the text, selecting text content, and pasting text. These operational events are key data points for capturing the user's real-time input intentions and thought processes. Semantic analysis results are information extracted from the user's input text content through in-depth processing, including the text's meaning, core concepts, argument structure, and potential knowledge points. A real-time learning state snapshot is a comprehensive data set formed by integrating the user's learning task information, past learning performance information, and current session state information at a specific point in time. This snapshot comprehensively reflects the user's current learning background and cognitive state. The identification result, based on the semantic analysis results and the real-time learning state snapshot, determines whether the user has any biases in knowledge understanding or shortcomings in thinking. Intervention forms refer to the methods of pushing learning resources to the user, such as text prompts, video links, exercise recommendations, and concept explanations. Intervention timing refers to when pushing learning resources during the user's learning process will achieve the best learning effect; this could be after the user inputs specific keywords, after the cursor hovers over a certain area for a period of time, or when an obvious error is detected. The pre-set learning resource library is a collection of learning materials of various types, difficulties, and topics. These materials are pre-categorized and labeled to facilitate the system's filtering and recommendation based on user needs. The core of the user behavior data processing method for the personalized learning platform in this application lies in the real-time capture and in-depth analysis of user text input behavior, combined with intelligent intervention based on the user's dynamic learning status.

[0028] Specifically, when a user inputs text on the platform, the system acquires the user's text input operation events in real time. These operation events can include character typing, deletion, cursor movement, text selection, and pasting. For example, these operations can be captured by listening to keyboard events, mouse events, and clipboard events. When a user types a character, the system records the character and its position; when a user deletes a character, the system records the deletion operation and the affected text range; when a user moves the cursor, the system records the new cursor position; when a user selects a piece of text, the system records the selected text content and range; when a user pastes text, the system records the pasted content. Based on these operation events, the system dynamically updates the text input content. Each typing or deletion operation results in a real-time change in the text content. Subsequently, semantic analysis is performed on the updated text input content to obtain semantic analysis results. A rule-based approach can be used to identify core concepts and argument structures in the text through predefined keywords and grammatical patterns. When a user inputs "The products of photosynthesis are oxygen and glucose," the system can identify core concepts such as "photosynthesis," "oxygen," and "glucose." Alternatively, statistical methods can be used, such as Term Frequency-Inverse Document Frequency (TF-IDF), to extract important words and phrases from the text. Next, the system acquires the user's current learning task information, past learning performance information, and current session status information, integrating this information into a real-time learning status snapshot. Current learning task information may include the topic of the assignment the user is currently completing and the scope of knowledge points involved. Past learning performance information may include the user's accuracy rate on historical exercises, viewing time in video courses, and assessments of their understanding of concepts in their notes. Current session status information may include the number of words the user has entered, the duration of the current session, and the area where the cursor is positioned in the text. This information can be obtained by querying the user database, learning logs, and front-end interaction data.

[0029] In one embodiment, based on semantic analysis results and real-time learning state snapshots, the system identifies user knowledge comprehension biases or insufficient thinking, yielding identification results. If the semantic analysis results show that the user mentions a concept multiple times in the text, but their expression differs significantly from the standard definition of that concept in the real-time learning state snapshot, it is identified as a concept comprehension bias. If the semantic analysis results reveal jumps or contradictions in the user's argumentation logic, and the real-time learning state snapshot indicates a weakness in the user's logical reasoning ability on the relevant topic, it is identified as an argumentation logic omission. After identifying user knowledge comprehension biases or insufficient thinking, the system obtains the current text position of the user's cursor. The line number and column number of the cursor in the text editor can be obtained in real time through a front-end JavaScript event listener. Finally, based on the identification results and the current text position of the cursor, the system determines the form and timing of intervention to select learning resources from a preset learning resource library and push the learning resources to the user. If the identification result is "concept comprehension bias" and the cursor is hovering near the concept, the system may immediately push a short explanation of the concept or a related video clip. If the identification result is "argument logic omission" and the user's cursor is at the end of the argument paragraph, the system will recommend a logical reasoning exercise or an example of an argument structure. The intervention can take the form of a pop-up prompt, a sidebar recommendation, or a suggestion inserted directly below the text. The timing of the intervention can be an immediate push notification, a delayed push notification, or a push notification after the user completes the current input paragraph.

[0030] In this regard, this application further proposes the steps of performing semantic analysis on the updated text input to obtain the semantic analysis results, including: When the text input is a paragraph, semantic analysis is performed on the updated text input based on a pre-trained language model to obtain core concepts, argumentation logic, and specific knowledge points. Based on core concepts, argumentation logic, and specific knowledge points, semantic analysis results are obtained.

[0031] Specifically, a pre-trained language model can be understood as a deep learning model trained on a large-scale text corpus, possessing powerful text understanding and generation capabilities. Models such as BERT, GPT series, RoBERTa, or T5 can be used. These models, by learning from massive amounts of text data, can capture the contextual relationships, syntactic structures, and semantic information between words, thus exhibiting excellent performance in handling specific tasks. Core concepts are the main ideas, themes, or key terms expressed in a paragraph. In a paragraph about "photosynthesis," "photosynthesis," "chloroplasts," and "energy conversion" can all be identified as core concepts. Argumentation logic refers to the reasoning relationships and structures between different viewpoints, facts, or arguments in a paragraph, such as causal relationships, parallel relationships, and adversative relationships. Its purpose is to assess the coherence and rigor of the user's thinking. Specific knowledge points are specific knowledge units that are highly relevant to the current learning task and mentioned in the paragraph. For example, when learning "Newton's First Law," "inertia" and "the relationship between force and motion" mentioned in the paragraph are specific knowledge points. By identifying these elements, we can gain a more comprehensive and in-depth understanding of the text content entered by the user.

[0032] This application's solution overcomes the limitations of traditional generalized semantic analysis in handling complex text by introducing a pre-trained language model and performing specialized semantic analysis on paragraph-type text input. When user input is identified as a paragraph, the pre-trained language model is used for deep parsing of that paragraph. Leveraging its training experience on large-scale corpora, the model can effectively identify the core concepts within the paragraph, analyze its internal argumentation logic, and extract specific knowledge points relevant to the learning task.

[0033] Specifically, the steps to obtain semantic analysis results based on core concepts, argumentation logic, and specific knowledge points include: Obtain the semantic vector of the updated text input; The semantic vector and the preset initial semantic vector are calculated to obtain the dot product between the semantic vector and the preset initial semantic vector, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic deviation threshold is obtained based on the dot product, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic analysis results are obtained by considering core concepts, argumentation logic, specific knowledge points, and semantic deviation thresholds.

[0034] The semantic vector is a numerical representation of text content mapped to a high-dimensional vector space using word embedding or sentence embedding techniques. It captures the semantic information and contextual relationships of the text. The preset initial semantic vector is a semantic vector generated from reference text related to the current learning task or standard answer, serving as a benchmark to measure the degree of semantic deviation of the user's text. The dot product is a measure of similarity between two vectors; a larger value generally indicates that the two vectors are closer in direction and have higher semantic similarity. The Euclidean norm, also known as the L2 norm, represents the length or size of a vector and can be used to normalize vectors or calculate the distance between vectors. The semantic deviation threshold is a value calculated based on the dot product and the Euclidean norm. It quantifies the degree of difference between the user's text and the preset initial semantic vector. When the difference exceeds this threshold, it may indicate that the user has a bias in knowledge comprehension or insufficient thinking.

[0035] This application's solution, by introducing semantic vectors and a preset initial semantic vector, can transform abstract text content into a quantifiable numerical representation. By calculating the dot product and Euclidean norm between the user's text's semantic vector and the preset initial semantic vector, the similarity and deviation between the user's text and the standard or expected semantics can be accurately assessed. Specifically, the dot product reflects the similarity of the two vectors in direction, while the Euclidean norm provides information about the vector's magnitude. Combining these calculation results, a semantic deviation threshold can be dynamically generated, which can more precisely define the differences between the user's text and the standard in terms of core concepts, argumentation logic, and specific knowledge points. Thus, the generation of semantic analysis results no longer relies solely on the identification of concepts and logic but incorporates quantitative semantic similarity assessment, making the judgment of user knowledge misunderstanding deviations or insufficient thinking more objective and accurate.

[0036] Specifically, the steps described above—obtaining the user's current learning task information, past learning performance information, and current session state information, and integrating these information into a real-time learning state snapshot—can be further refined as follows: First, the system obtains the user's current learning task information. This information can include the topic, knowledge points, and core concepts of the current learning task. The topic refers to the subject or area where the user is currently engaged in learning, such as "basics of calculus" or "matrix operations in linear algebra." Knowledge points refer to specific knowledge units related to the topic of the current learning task, such as "the definition of the derivative" or "the inverse of a matrix." Core concepts refer to key concepts or theories that the user needs to focus on understanding and mastering in the current learning task, such as "limits" or "eigenvalues." Second, the system obtains past learning performance information. This information can include the user's accuracy rate in exercises, viewing time in video courses, and assessment of the user's understanding of notes and concepts. The accuracy rate in exercises reflects the user's mastery and application ability of specific knowledge points. Viewing time in video courses can indicate the user's level of engagement with learning materials and potential depth of understanding. Assessment of the user's understanding of notes and concepts can be determined by analyzing the user's notes during the learning process or by the system's assessment of the user's concept comprehension, thus judging the user's internalization of knowledge and cognitive level. Third, the system obtains the current session state information. The current session status information can include the number of words the user has entered, the duration of the current session, and the area where the user's cursor is hovering over the text. The number of words entered reflects the user's level of engagement and thought in the current task. The duration of the current session indicates the amount of time the user has spent on the current learning task. The area where the user's cursor is hovering over the text can infer the specific text paragraph or concept the user is currently focusing on or thinking about, thus revealing the user's immediate points of interest or confusion. Finally, the topic, knowledge points, core concepts, accuracy rate in exercises, viewing time in video courses, assessment of note content and concept comprehension, the number of words entered by the user, the duration of the current session, and the area where the user's cursor is hovering over the text are integrated into a real-time learning status snapshot. By integrating this multi-dimensional, granular information, a comprehensive and dynamic view of the user's learning status can be formed.

[0037] This application's solution, by acquiring and integrating detailed information on the user's current learning task, past learning performance, and current session state, can construct a more comprehensive and detailed snapshot of the user's real-time learning status. Specifically, the current learning task information provides the macro-context and micro-focus of the user's learning; past learning performance information reflects the user's knowledge mastery level and learning habits from a historical perspective; and the current session state information captures the user's behavioral details and attention distribution during the real-time learning process. This comprehensive consideration of all this information allows the system to gain a deep understanding of the user's learning status from multiple perspectives, providing a solid data foundation for subsequently identifying user knowledge comprehension biases or insufficient thinking.

[0038] In this regard, this application further proposes the following steps for identifying user knowledge comprehension biases or insufficient thinking based on semantic analysis results and real-time learning state snapshots, to obtain the identification results: When semantic analysis results show that the user's understanding of the core concepts of the current learning task differs from the standard definitions in the real-time learning state snapshot, the text similarity of the difference is calculated to obtain the current difference value. If the current difference value is less than the preset difference value, the identification result is a concept comprehension deviation.

[0039] Specifically, when semantic analysis results show a difference between the user's understanding of the core concepts of the current learning task and the standard definitions in the real-time learning snapshot, it means there is a semantic inconsistency between the core concepts expressed by the user in the text input and the standard core concepts defined in the system's preset or course. This difference can be detected by comparing the core concepts extracted from the user's text with the standard definitions obtained from the current learning task information. Calculating the text similarity of the difference to obtain the current difference value means that various text similarity calculation methods can be used to quantify this semantic difference. For example, cosine similarity technology can be used to calculate the degree of similarity between the user's expressed core concepts and the standard definitions. The calculated similarity value is the current difference value, which reflects the closeness between the user's understanding and the standard understanding. In practical applications, determining the identification result as a concept comprehension bias based on the current difference value being less than a preset difference value means comparing the calculated current difference value with a preset difference threshold. The preset difference value is an empirical value or a threshold determined through training a model, used to define what degree of difference is considered a "concept comprehension bias". When the current difference value is lower than the preset difference value, it indicates that the user's understanding of the core concept deviates significantly from the standard definition. At this time, the system can identify that the user has a conceptual misunderstanding.

[0040] In this regard, this application further proposes the following steps for identifying user knowledge comprehension biases or insufficient thinking based on semantic analysis results and real-time learning state snapshots, to obtain the identification results: When the semantic analysis results identify inconsistencies or contradictions between the argument logic in the text input and the standard argument logic in the real-time learning state snapshot, the identification result is that the argument logic is missing.

[0041] Specifically, the statement "inconsistencies or contradictions exist between the user's argument logic and the standard argument logic in the real-time learning state snapshot" refers to the user's argument logic extracted through in-depth analysis of the user's text input by the semantic analysis module. When compared with the preset standard argument logic for the current learning task, logical breaks or jumps are found, or conflicting viewpoints exist in the reasoning chain. For example, when discussing a concept, the user's arguments or derivation processes may be inconsistent, or the conclusion may lack necessary logical support from its premises. "Standard argument logic" can be understood as a generally accepted, rigorous, and consistent reasoning process and argument structure for a specific learning task or knowledge point, which can be pre-defined and stored in the learning resource library. "Omissions in argument logic" refers to the system identifying that the user failed to construct their argument process completely or correctly in the text input, resulting in insufficient logical support or obvious logical flaws in the expressed viewpoint. Such omissions may manifest as missing key arguments, incomplete reasoning steps, or deviations in the direction of argumentation. This application's solution compares the argumentation logic of the user's text input, identified in the semantic analysis results, with the standard argumentation logic contained in the real-time learning state snapshot. This allows for the precise identification of logical inconsistencies or contradictions in the user's argumentation process. This meticulous comparison enables the system to recognize potential omissions in the user's logical reasoning, thus compensating for the limitations of merely identifying core concepts or knowledge point deviations. In this way, the system can more comprehensively assess the depth of the user's thinking and the rigor of their logic, providing a more accurate basis for subsequent personalized intervention.

[0042] In this regard, this application further proposes the following steps for identifying user knowledge comprehension biases or insufficient thinking based on semantic analysis results and real-time learning state snapshots, to obtain the identification results: If a user's current learning task requires that the real-time learning state snapshot cover specific knowledge points, and the semantic analysis results show that the user's input text does not contain specific knowledge points, the recognition result is that the knowledge point is missing.

[0043] Specifically, the requirement for a user's current learning task to cover specific knowledge points in the real-time learning state snapshot means that when constructing the real-time learning state snapshot, the system identifies the key knowledge points that must be mastered or mentioned for the task, based on its nature and objectives. These specific knowledge points can be predefined in the task template or extracted through natural language processing of the task description. The real-time learning state snapshot integrates the user's current learning task information, thus reflecting the task's specific requirements for knowledge points. Furthermore, the semantic analysis results and the display of specific knowledge points not included in the user's input text are obtained after semantic analysis of the user's text input. The system extracts the core concepts, argumentation logic, and specific knowledge points actually contained in the user's text. Subsequently, the system compares these extracted knowledge points with the specific knowledge points required to be covered in the real-time learning state snapshot. If it finds that one or more specific knowledge points required in the real-time learning state snapshot are not reflected in the semantic analysis results of the user's input text, it is considered that the user has omitted knowledge points. Therefore, when the above conditions are met, the system will identify the knowledge point as missing, which indicates that the user has not fully covered or mentioned the key knowledge content required by the task in the text input of the current learning task.

[0044] In this regard, this application further proposes that the steps for determining the form and timing of intervention include: Obtain the semantic completeness of the user's input and the user's past learning preferences; Based on the semantic completeness of the user's input, the user's past learning preferences, the recognition results, and the current position of the cursor in the text, the intervention method and timing are determined to select learning resources from a pre-set learning resource library and push the learning resources to the user.

[0045] Specifically, the semantic completeness of user-inputted content refers to assessing the semantic integrity and coherence of the text currently being entered by the user. This can be understood as whether the user's current input has formed a relatively complete and logically clear expression, or is still in the conceptual or fragmented input stage. Semantic completeness can be quantified by analyzing indicators such as text length, syntactic structure, keyword density, and relevance to the current learning task topic. Its purpose is to determine the user's current thought process and avoid unnecessary interruptions before the user's thinking is fully developed or nearing completion. The user's past learning preferences can be understood as the user's inclination towards specific learning resource types, learning methods, difficulty levels, and interaction modes during their historical learning process. Users may prefer watching video tutorials to reading long texts, or prefer interactive exercises to purely theoretical explanations. This preference information can be collected and modeled by analyzing data such as the user's historical learning records, resource click-through rates, completion rates, and feedback. Its purpose is to make the pushed learning resources more aligned with the user's personal habits and acceptance levels, thereby improving the effectiveness of intervention and the user's learning motivation.

[0046] This application's solution, by incorporating the semantic completeness of the user's input and their past learning preferences, can more comprehensively assess the user's immediate learning status and personalized needs. When it identifies a user's knowledge comprehension bias or insufficient thinking, the system no longer intervenes solely based on the recognition result and cursor position, but further considers semantic completeness to determine the timing of intervention. For example, if the semantic completeness is low, it indicates that the user may be in the early stages of thinking or encountering significant obstacles, making timely and direct intervention more effective; conversely, if the semantic completeness is high, a lighter or delayed intervention may be necessary to avoid interrupting the user's established train of thought. Simultaneously, by considering the user's past learning preferences, the system can select learning resources from a pre-set learning resource library that are more aligned with the user's habits in terms of type, difficulty, and presentation, thereby improving the attractiveness of the resources and the learning effect. It is precisely because of these additional factors that the intervention strategy can be more intelligent and personalized, effectively compensating for the limitations of relying solely on recognition results and cursor position for intervention.

[0047] In some embodiments, assuming a user is writing a short essay about "photosynthesis" on a personalized learning platform, the system identifies logical inconsistencies in the user's discussion of the "light reaction" section through semantic analysis, and the cursor is currently positioned at the end of that paragraph. At this point, the system first assesses the semantic completeness of the user's input. If the evaluation shows that the currently inputted "light reaction" paragraph has low semantic completeness, incomplete sentence structure, and missing key concepts, this may indicate that the user is encountering significant comprehension difficulties or a thinking bottleneck in this section. Simultaneously, the system queries the user's past learning preferences and finds that the user prefers to understand complex concepts by watching short animated videos. Based on this information, the system determines an immediate and intuitive form and timing of intervention. For example, the system might immediately pop up a small window on the right side of the user interface, pushing a link to a 3-minute animated video about the "light reaction process of photosynthesis," along with a brief tip: "You may have some conceptual confusion when describing the light reaction; this video might help you clarify your thoughts." In this way, the system not only identifies the problem but also provides timely, accurate, and user-friendly learning resources based on the user's current input state and personal preferences. This effectively avoids blindly pushing irrelevant or disliked resources, thereby improving the effectiveness of learning intervention and user satisfaction.

[0048] The above-mentioned determination of intervention methods and timing, in order to select learning resources from a pre-set learning resource library and push these resources to users, specifically includes: The intervention method and timing are determined to select learning resources that are highly matched with the user's current intention and context based on the pre-set learning resource library, determine the type, difficulty and relevance score of the learning resources, and push the learning resources to the user.

[0049] Specifically, "learning resources that highly match the user's current intent and context" are those selected that accurately respond to the specific needs or questions the user exhibits in the current learning task and have a strong semantic relevance to the text the user is inputting. "User's current intent" can be understood as the goal the user hopes to achieve or the question they are pondering in the current learning activity; for example, the user may be trying to understand a concept, argue a point, or find supporting materials for a specific knowledge point. This intent can be inferred through a comprehensive analysis of the user's real-time learning state snapshot, semantic analysis results, and the current position of the cursor in the text. If the identification results show that the user has a conceptual misunderstanding, the user's intent may be to seek clarification or supplementary explanation of that concept. "Context" refers to information such as the content of the user's current text input, completed learning tasks, the duration of the current session, and the area where the cursor is located in the text; these information collectively constitute the context of the user's current learning situation. Furthermore, after selecting highly matching learning resources, it is also necessary to "determine the type, difficulty, and relevance score of the learning resources." The "type of learning resources" can include, but is not limited to, various forms such as text explanations, video tutorials, interactive exercises, case studies, and mind maps. The selection can be determined based on the user's learning preferences, the type of recognition result (e.g., text explanations or videos may be more suitable for conceptual misunderstandings, while case studies may be more suitable for logical inconsistencies), and the intervention method. "Difficulty" refers to the complexity and depth of knowledge contained in the learning resources, and should be appropriate to the user's past learning performance and current knowledge level, avoiding the recommendation of overly difficult or easy resources. "Relevance score" is a quantitative indicator used to assess the degree of matching between the learning resources and the user's current intent and context; a higher score indicates a higher degree of matching. This score is derived by comprehensively calculating factors such as the semantic similarity between the learning resources and the user's current text input, the coverage of knowledge points in the current learning task, and the relevance to the recognition result. Therefore, after determining the intervention method and timing, the system will select the highest-scoring learning resource that best meets the user's needs from the pre-set learning resource library based on the above screening and evaluation results, and push it to the user.

[0050] This application's solution addresses the potential generalization problem in resource selection inherent in basic solutions by introducing in-depth analysis of the user's current intent and context, and by quantitatively evaluating the type, difficulty, and relevance of learning resources. Specifically, when a user is identified as having a knowledge misunderstanding bias or insufficient thinking, the system no longer simply selects resources relevant to the identification result from the resource library. Instead, it first combines a snapshot of the user's real-time learning state, semantic analysis results, and cursor position to accurately infer the user's current learning intent, such as needing conceptual clarification, supporting arguments, or supplementing knowledge points. Simultaneously, the system analyzes the context of the user's input text to ensure that the recommended resources closely align with the user's current thought process and context. Based on this, the system calculates the degree of matching between each candidate resource in the pre-set learning resource library and the user's current intent and context, and assesses whether its type and difficulty are suitable for the current intervention.

[0051] See Figure 2 , Figure 2 This is a schematic diagram of a personalized learning platform user behavior data processing system provided in one embodiment of this application. The personalized learning platform user behavior data processing system 200 includes: The acquisition and update module 210 is used to acquire the user's text input operation events when the user inputs text on the platform and update the text input content according to the operation events. The operation events include character input, deletion, cursor movement, text selection and pasting. Semantic analysis module 220 is used to perform semantic analysis on the updated text input and obtain semantic analysis results; The integration module 230 is used to obtain the user's current learning task information, past learning performance information and current session status information, and integrate the current learning task information, past learning performance information and current session status information into a real-time learning status snapshot; The recognition module 240 is used to identify user knowledge comprehension biases or insufficient thinking based on semantic analysis results and real-time learning state snapshots, and obtain recognition results. The text position acquisition module 250 is used to acquire the current text position of the user's cursor. The filtering and push module 260 is used to determine the form and timing of intervention based on the recognition results and the current text position of the cursor, so as to filter learning resources from the preset learning resource library and push the learning resources to the user.

[0052] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0053] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0054] The foregoing has provided a detailed description of the preferred embodiments of this application. However, this application is not limited to the above-described embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined in this application.

Claims

1. A method for processing user behavior data on a personalized learning platform, characterized in that, include: When a user inputs text on the platform, the system acquires the user's text input operation events and updates the text input content based on the operation events. The operation events include character input, deletion, cursor movement, text selection, and pasting. Perform semantic analysis on the updated text input to obtain the semantic analysis results; The system acquires the user's current learning task information, past learning performance information, and current session status information, and integrates these information into a real-time learning status snapshot. Based on the semantic analysis results and the real-time learning state snapshot, the user's knowledge comprehension bias or insufficient thinking is identified, and the identification result is obtained; Get the current text position of the user's cursor; Based on the recognition results and the current text position of the cursor, the intervention method and timing are determined to select learning resources from a preset learning resource library and push the learning resources to the user.

2. The method according to claim 1, characterized in that, The step of performing semantic analysis on the updated text input to obtain semantic analysis results includes: When the text input is a paragraph, semantic analysis is performed on the updated text input based on a pre-trained language model to obtain core concepts, argumentation logic, and specific knowledge points. Based on the core concepts, the argumentation logic, and the specific knowledge points, the semantic analysis results are obtained.

3. The method according to claim 2, characterized in that, Based on the core concepts, the argumentation logic, and the specific knowledge points, the semantic analysis results are obtained, including: Obtain the semantic vector of the updated text input content; The semantic vector and the preset initial semantic vector are calculated to obtain the dot product between the semantic vector and the preset initial semantic vector, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic deviation threshold is obtained based on the dot product, the Euclidean norm of the semantic vector, and the Euclidean norm of the preset initial semantic vector. The semantic analysis results are obtained by considering the core concepts, the argumentation logic, the specific knowledge points, and the semantic deviation threshold.

4. The method according to claim 1, characterized in that, The step of acquiring the user's current learning task information, past learning performance information, and current session state information, and integrating the current learning task information, past learning performance information, and current session state information into a real-time learning state snapshot, includes: Obtain the user's current learning task information, wherein the current learning task information includes the topic, knowledge points, and core concepts of the current learning task; Obtain past learning performance information, which includes the accuracy rate in exercises, the viewing time in video courses, and the assessment of the content of notes and the degree of understanding of concepts; Obtain current session state information, wherein the current session state information includes the number of characters of text entered by the user, the duration of the current session, and the area where the user's cursor is located in the text; The current learning task's topic, knowledge points, core concepts, accuracy rate in exercises, viewing time in video courses, comprehension assessment of notes and concepts, number of words in user-inputted text, duration of the current session, and the area where the user's cursor is positioned in the text are integrated into a real-time learning status snapshot.

5. The method according to claim 2, characterized in that, The step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: When the semantic analysis results show that the user's understanding of the core concepts of the current learning task information differs from the standard definition in the real-time learning state snapshot, the text similarity of the difference is calculated to obtain the current difference value; If the current difference value is less than the preset difference value, the identification result is a concept comprehension deviation.

6. The method according to claim 2, characterized in that, The step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: When the semantic analysis results identify inconsistencies or contradictions between the argument logic in the text input and the standard argument logic in the real-time learning state snapshot, the identification result is that the argument logic is missing.

7. The method according to claim 2, characterized in that, The step of identifying user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtaining the identification results, includes: If a user's current learning task is required to cover a specific knowledge point in the real-time learning state snapshot, and the semantic analysis results and the text entered by the user do not contain the specific knowledge point, the identification result is that the knowledge point is missing.

8. The method according to claim 1, characterized in that, The step of determining the intervention form and timing based on the recognition result and the current text position of the cursor to select learning resources from a preset learning resource library and push the learning resources to the user includes: Obtain the semantic completeness of the user's input and the user's past learning preferences; Based on the semantic completeness of the user's input, the user's past learning preferences, the recognition results, and the current text position of the cursor, the intervention method and timing are determined to select learning resources from a preset learning resource library and push the learning resources to the user.

9. The method according to claim 1, characterized in that, The process of determining the intervention method and timing to select learning resources from a pre-set learning resource library and push those resources to the user includes: The intervention method and timing are determined to select learning resources that highly match the user's current intent and context based on the preset learning resource library, determine the type, difficulty and relevance score of the learning resources, and push the learning resources to the user.

10. A user behavior data processing system for a personalized learning platform, characterized in that, include: The acquisition and update module is used to acquire the user's text input operation events when the user inputs text on the platform and update the content of the text input according to the operation events. The operation events include character input, deletion, cursor movement, text selection and pasting. The semantic analysis module is used to perform semantic analysis on the updated text input and obtain the semantic analysis results; The integration module is used to obtain the user's current learning task information, past learning performance information, and current session status information, and integrate the current learning task information, past learning performance information, and current session status information into a real-time learning status snapshot; The identification module is used to identify user knowledge comprehension biases or insufficient thinking based on the semantic analysis results and the real-time learning state snapshot, and obtain the identification results. The text position acquisition module is used to obtain the current text position of the user's cursor; The filtering and push module is used to determine the form and timing of intervention based on the recognition results and the current text position of the cursor, so as to filter learning resources from a preset learning resource library and push the learning resources to the user.