Software operation data intelligent processing and analyzing system based on big data
By assessing multi-skill mastery and identifying inter-stage deviations in online English learning software, and combining knowledge point tags and skill dependencies, the teaching schedule is adjusted. This solves the problems of overly coarse learning outcome assessment and insufficient knowledge point dependency analysis in existing technologies, achieving refined and interpretable personalized teaching support, improving learning efficiency and reducing cognitive load.
Patent Information
- Application Number
- CN202511833914.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-01-02
AI Technical Summary
Existing online English learning software fails to decouple and analyze multiple skill dimensions such as listening comprehension, oral expression, reading comprehension, and written expression. This results in an overly coarse-grained learning effectiveness evaluation mechanism that lacks analysis of the dependencies and relationships between language knowledge points. Consequently, learners are forced to move on to advanced content without a solid foundation, increasing cognitive load and frustration.
By collecting users' historical learning records, analyzing their mastery of each skill level, identifying deviation units, and adjusting the teaching schedule through knowledge point tag matching and skill dependency analysis to prevent cumulative progress risks, we can achieve refined and interpretable personalized teaching support.
It enables refined assessment and proactive adjustment of language skills, reduces the phenomenon of students advancing to higher levels with weak foundations, reduces comprehension barriers and cognitive load, and improves learning efficiency.
Smart Images

Figure CN121258751A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of English software running processing, particularly relates to English online teaching processing technology, and specifically discloses a software running data intelligent processing and analysis system based on big data. BACKGROUND
[0002] With the continuous growth of the demand for English language ability, online English learning software has been widely used and rapidly popularized. Such platforms provide learners with flexible and on-demand accessible language learning paths through digital teaching environments. During the user learning process, the system can automatically collect and record multi-modal software running data to form a fine-grained learning process log. Based on these running data, the system can analyze the user learning effectiveness in real time and dynamically adjust the teaching content, aiming to achieve personalized teaching adaptation.
[0003] There are related technical solutions in the prior art. For example, a Chinese invention patent with publication number CN119474687A discloses an English learning system based on a mobile device, which collects learning data sets including score data, unit data, and time length data for data preprocessing and analysis, obtains learning reference values, and provides real-time feedback and effectiveness evaluation based on clustering classification results.
[0004] For another example, a Chinese invention patent with publication number CN120316351A proposes an English-oriented learning content recommendation method, which collects basic information when the user registers, constructs a personal learning profile including course learning records, learning duration, interruption times, practice completion conditions, etc., and dynamically adjusts the recommendation strategy based on learning content recommendation evaluation indicators.
[0005] The above two solutions can reflect the overall learning state of the user from a macro perspective, but both ignore the fact that English is a comprehensive language discipline, and its ability composition covers multiple independent and collaborative development skill dimensions such as listening comprehension, oral expression, reading comprehension, and written expression. Each skill has different cognitive load. However, the existing technology generally analyzes the learning effectiveness as a whole, and fails to decouple and analyze the ability mastery state of each skill link, resulting in a coarse granularity of the learning effectiveness evaluation mechanism.
[0006] In addition, although the second scheme can adjust content pushing according to learning completion degree, the adjustment logic mainly depends on the current learning progress and target achievement rate of the user, and does not deeply model the dependent relationship between language knowledge points in historical learning and subsequent teaching content. Since language knowledge has significant accumulative and preposition dependency, insufficient mastery of early knowledge points may cause a chain of understanding obstacles in subsequent learning. However, the prior art lacks analysis of such knowledge transmission effects, making it difficult to achieve forward-looking and preventive adjustment of teaching progress, resulting in learners entering high-level content without a solid foundation, exacerbating cognitive load imbalance and learning frustration. SUMMARY
[0007] In view of this, the present application aims to propose an online English learning software based on big data software running data intelligent processing and analysis system that can deeply integrate multi-skill mastery evaluation, link deviation identification and forward risk prediction, to realize truly fine, interpretable and forward-looking personalized teaching support.
[0008] The purpose of the present application can be achieved by the following technical solutions: a big data based software running data intelligent processing and analysis system, comprising the following modules: a learning data acquisition module: acquiring user historical learning language unit learning record data in four links of listening comprehension, oral expression, reading comprehension and written expression from an online English learning software, extracting system scores and practice time of each link.
[0009] A language mastery analysis module: based on the practice time and interactive behavior of the learning record, effective learning records are screened, and the four-link scores of each language unit in the effective learning records are analyzed in the order of teaching progress, and the mastery ability of each language unit in each link is calculated.
[0010] A mastery deviation identification module: compares the mastery ability difference between the four links for each language unit, and identifies target language units with mastery deviation.
[0011] A unit correlation analysis module: analyzes the correlation strength of the identified target language unit with subsequent language units through knowledge point tag matching and skill dependency relationship.
[0012] A teaching progress adjustment module: based on the correlation strength analysis result, cumulative progress risk is judged, and then the teaching progress arrangement with cumulative progress risk is adjusted.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: 1. The present application collects the learning record data of the user's historical learned language units in the four aspects of listening comprehension, oral expression, reading comprehension and written expression, combines the system automatic scoring mechanism to quantitatively analyze the ability mastery of each aspect, on this basis, by comparing the differences in the ability mastery of the four aspects under the same language unit, the units with mastery deviation are identified, the decoupled modeling and differentiated evaluation of language mastery ability are realized, and the granularity of learning effectiveness analysis is significantly improved.
[0014] 2. On the basis of completing the learning effectiveness analysis, the present application identifies the correlation strength of the mastery deviation units and the subsequent content through knowledge point label matching and skill dependency relationship mining, and then adjusts the teaching progress which has cumulative progress risk based on the correlation strength analysis result, realizes the prospective and preventive adjustment of the teaching progress, can significantly reduce the phenomenon of taking sick progress due to weak foundation, avoids the learners to enter high-order content learning in the case of lacking necessary skill support, thereby reducing understanding obstacles, cognitive load imbalance and learning frustration. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 The system composition schematic diagram in the present application.
[0017] Figure 2 The effective learning record filtering implementation flowchart of the practice duration and interactive behavior based on learning record in the present application.
[0018] Figure 3 The skill dependency graph atlas in the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0020] The present application provides a software running data intelligent processing and analysis system based on big data, which comprises a learning data acquisition module, a language mastery analysis module, a mastery deviation identification module, a unit correlation analysis module and a teaching progress adjustment module.
[0021] Referring to Figure 1 As shown, the modules are sequentially connected at the ends to form an intelligent processing closed loop integrating perception, analysis, judgment and intervention.
[0022] The learning data collection module is configured to collect, from an English online learning software, learning record data of a user's historical learned language unit in four aspects of listening comprehension, oral expression, reading comprehension and written expression, and extract system scores and practice time length of each aspect.
[0023] Since English online teaching usually adopts a task-driven and human-computer interaction training mode, the system automatically generates structured learning records during the process of completing the four types of skill tasks of listening, speaking, reading and writing. Such records, as a core component of software operation data, completely record the user's behavior trajectory and result feedback, which contains the practice time length of each aspect task and the ability score automatically rated by the system based on preset rules or models, providing a quantifiable and traceable data basis for subsequent multi-dimensional learning effectiveness analysis.
[0024] In the specific implementation of the above scheme, the learning record data includes the following collection process: extracting the completed language unit sequence in the time sequence according to the teaching schedule.
[0025] In English online teaching, the teaching organization is usually designed based on language units as basic teaching modules. Each language unit integrates listening, speaking, reading and writing skill training tasks around a specific theme or functional scenario to form a complete learning closed loop.
[0026] For each completed language unit, the automatically rated listening comprehension score is collected from the learning record of the listening training aspect.
[0027] The automatically rated oral expression score is extracted from the learning record of the oral practice aspect.
[0028] The automatically rated reading comprehension score is extracted from the learning record of the reading training aspect.
[0029] The automatically rated written expression score is extracted from the learning record of the writing task aspect.
[0030] The actual practice time length of each aspect in the learning record is recorded.
[0031] The system scores and practice time length of the four aspects are associated and stored according to the language unit.
[0032] The system scores of each skill link under each language unit of the user above reflect the knowledge mastery level and skill execution accuracy of the user in a specific language ability dimension, which is a quantitative representation of learning effectiveness; while the practice duration reflects the learning input of the learner in this link, embodying the continuity and concentration of his learning behavior. Both of them constitute the basis for multi-dimensional evaluation of the learning process.
[0033] The language mastery analysis module is used to filter effective learning records based on the practice duration and interactive behavior of the learning records, and analyze the four-link scores under each language unit in the effective learning records in the order of teaching progress, and calculate the mastery ability of each language unit in each link.
[0034] Referring to Figure 2 In the preferred implementation of the above module, the filtering of effective learning records based on the practice duration and interactive behavior of the learning records is as follows: comparing the practice duration of each language unit in each link with the complete practice duration set by the learning task.
[0035] In particular, the complete practice duration set by the learning task is set by the teaching design specification, representing the minimum reasonable time required to complete the task. If the actual practice duration is lower than this duration, it is considered as not completing the basic training process.
[0036] The user operation event sequence is extracted from the learning record, including the timestamp of each interaction, and the time interval between adjacent operation events is calculated based on the timestamp sequence, and the arithmetic mean of the average user interaction interval duration is obtained, which is compared with the normal interaction interval reference value set by the learning task.
[0037] Understandably, in the English online teaching environment, there are frequent human-computer interaction operations between the learner and the system, including clicking, sliding, voice input, text submission, answer selection, etc. Such interactive behavior is a core component in the process of digital learning, and real cognitive participation is usually manifested as a continuous interaction sequence with orderly rhythm and task-related.
[0038] Based on the above understanding, by analyzing the time interval sequence between user operation events, the attention concentration level of the user in the learning task can be indirectly evaluated. The average interaction interval duration reflects the response rhythm of the user within a unit task, which is an important behavioral indicator for measuring learning concentration: shorter interval usually corresponds to high attention state, while longer interval may indicate distraction or mental stagnation.
[0039] The normal interaction interval reference value of the learning task setting represents the typical response mode of the user's operation under ideal cognitive participation conditions. The reference value can be obtained by collecting operation time interval data of a large number of valid learning samples under the current task type, and calculating the median or mode as the reference value to represent the normal interaction rhythm of the task, which provides a reference standard for subsequent invalid learning record screening.
[0040] When any learning record meets any of the following conditions, it is determined to be an invalid learning record: (1) the practice duration does not reach the complete practice duration of the learning task setting.
[0041] (2) the average user interaction interval duration exceeds the normal interaction interval reference value of the learning task setting.
[0042] The condition (1) follows the minimum task completion principle, that is, the learner needs to invest sufficient time to complete the core learning activity to ensure the integrity of the learning process.
[0043] Condition (2) follows the cognitive participation consistency principle, that is, effective learning behavior should show a consistent and rhythmic interaction pattern that matches the task requirements.
[0044] All learning records that are not determined to be invalid are classified as valid learning records.
[0045] The above-mentioned dual verification mechanism based on time integrity and interaction continuity is implemented on the collected user learning records to identify and filter invalid or low-quality learning behavior as the input data source for subsequent language proficiency and teaching progress optimization analysis.
[0046] In further preferred implementation of the above module, the proficiency of each language unit at each link is calculated as follows: S1, for each link of the completed language unit, extract the system score of all valid learning records at the link.
[0047] The above operation only extracts the system score of the valid learning record, excluding low-quality data caused by interruption, lack of attention or non-substantial interaction.
[0048] S2, calculate the average of the system scores of all valid records at each link as the proficiency indicator of the link.
[0049] The above takes the average of the performance of multiple rounds of practice in the same link of the same language unit to reduce the influence of accidental errors and improve the stability and reliability of the evaluation.
[0050] S3, according to the qualified proficiency standard score of each link of the language unit teaching goal, the ratio of the proficiency indicator of each link to the qualified proficiency standard score is taken as the proficiency value of the language unit at each link.
[0051] The ability mastery value obtained by the ratio operation of the actual mastery degree index and the qualified mastery standard can effectively represent the target achievement of the learner in a specific skill link. The proportional form reflects that even if the absolute score is high, if the preset ability threshold is not reached, the mastery value is still at a low level, thereby strengthening the rigid constraints on the basic requirements of teaching.
[0052] The mastery deviation identification module is used to compare the ability mastery differences among the four links for each language unit, and identify the target language unit with mastery deviation.
[0053] The specific implementation of the above module is as follows: for each learned language unit, the ability mastery values in the four links are standardized and mapped into a four-column ability distribution graph arranged in a preset link order, and each column height corresponds to the standardized mastery score.
[0054] In the above, the ability mastery values of each link are uniformly mapped to a normalized interval, aiming to eliminate the differences in the original score scales among different skill dimensions, realize the dimensional unification and comparability alignment of multi-skill ability values, and provide a consistent measurement benchmark for cross-skill analysis. The preset skill order uses listening comprehension → speaking expression → reading comprehension → written expression as the dimension arrangement axis, which conforms to the cognitive development law of input priority and output follow-up in language acquisition.
[0055] The mapped four-column ability distribution graph presents the mastery level distribution of the four skills under the same language unit in a visual columnar structure, which directly reflects the relative strength and weakness of each skill.
[0056] In the ideal state based on the same ability scores of all skill links, the ideal ability distribution median is determined.
[0057] Considering that the ideal learning path should promote the balanced development of each skill link, that is, to achieve relatively consistent ability improvement in the four skill dimensions of listening comprehension, speaking expression, reading comprehension and written expression, which is ideal. In this ideal state, the ability scores of all skill links are the same, ensuring that the learner is balanced in all aspects of training and development.
[0058] In the four-column ability distribution graph, this ideal state of balanced development is manifested as equal column heights, and the weighted center position of the distribution is located in the middle position. Specifically, under the assumption of sequence weight system of listen = 1, speak = 2, read = 3, and write = 4, the ideal ability distribution median should be located at 2.5. This median represents the virtual balance point formed when the four skill scores are completely consistent, which is used to compare the deviation of the actual ability distribution in the subsequent, so as to judge whether there is an unbalanced problem in skill development.
[0059] The weighted center position of the actual ability distribution is calculated according to the mastery scores of the four skill links for a given language unit.
[0060] The weighted center line position of the actual ability distribution is expressed as follows: , wherein represents the ability mastery value of the i th skill link, represents the sequence number weight of the i th skill link, represents the sequence number weight of the i th skill link, represents the sequence number weight of the i th skill link, represents the weighted center line position of the actual ability distribution of the current language unit.
[0061] Understandably, the weighted average method is used to calculate the actual center position in the four-column ability distribution graph, which can effectively represent the comprehensive ability distribution center of the user in the listening, speaking, reading, and writing skill links. This method takes the sequence number of each skill link as the weight and the corresponding ability mastery value as the observation variable, and quantifies the spatial concentration trend of the ability distribution through the weighted mean model.
[0062] The weighted center line position of the actual ability distribution calculated for each language unit is compared with the ideal ability distribution center line. If the weighted center line position of the actual ability distribution of a language unit deviates from the ideal ability distribution center line, where the deviation represents that the center lines do not overlap, it is determined that the ability distribution of the language unit has a center line deviation, and it is marked that the language unit has a link mastery deviation.
[0063] The ability mastery values of the four links under each language unit are compared with the group baseline of the corresponding skill link. If the ability mastery value of a link is lower than the group baseline, the link is identified as a weak link, indicating that the development of the skill link lags behind the average level of learners at the same stage.
[0064] The group baseline in the above is an external reference standard, usually using the statistical distribution such as the mean of the ability mastery values of a user group of the same age, similar learning stage, or similar learning track in the corresponding skill link as the baseline. This baseline reflects the typical development level of the skill item under a specific teaching background, and has group representativeness and stage characteristics.
[0065] Specifically, the setting can be realized by the following way: continuously collecting historical learning data of a homogeneous learner group during system operation, grouping and statistically analyzing the ability mastery values of each skill link under each language unit, and determining the baseline value by using the mean after filtering out outliers.
[0066] The language unit with a link mastery deviation or a weak link is determined as a target language unit.
[0067] The present application aims to comprehensively capture the structural imbalance problem existing in the learning process by means of the dual judgment mechanism of mastering deviation recognition and weak link detection, and the core purpose is to realize the fine positioning of learning difficulties. Compared with the traditional method which only relies on the comprehensive score or single threshold judgment, the present application effectively avoids the misjudgment and omission caused by the heterogeneity of ability distribution, and improves the sensitivity and specificity of evaluation. At the same time, the accurate anchoring of the target language unit provides clear and operable intervention targets for the subsequent teaching progress adjustment module.
[0068] The unit relevance analysis module is used to analyze the association strength of the identified target language unit with subsequent language units through knowledge point label matching and skill dependency analysis.
[0069] In the specific implementation of the present application, the knowledge point label matching content is as follows: a knowledge point label set is established for each language unit, and the knowledge point label of the target language unit is extracted.
[0070] Since each language unit is assigned a set of structured knowledge point labels at the course design stage, the purpose of the label is to realize the explicit modeling of the knowledge structure of the teaching content. The label set covers dimensions such as grammatical structure, vocabulary theme, functional expression, etc., which describes the core language knowledge and communicative ability elements carried by the language unit from multiple granularity levels. Through systematic labeling, the knowledge point label set becomes the semantic fingerprint of the language unit, which not only supports the fine organization and retrieval of teaching content, but also provides a structured data basis for subsequent knowledge association analysis.
[0071] According to the teaching progress, the language units that have not been learned by the user after positioning the target language unit from the teaching plan constitute a forward to-be-learned unit set, and the knowledge point label set of each forward to-be-learned unit is extracted as a query reference for subsequent forward association analysis.
[0072] Each language unit in the forward to-be-learned unit set is traversed to count the occurrence of the knowledge point label contained in the target language unit in its knowledge point label set, and the occurrence frequency of the label is accumulated as the recurrence frequency of the knowledge point of the target language unit in subsequent teaching.
[0073] By forward scanning the knowledge point label set of the target language unit, the present application can quantify the forward penetration strength of the unit knowledge content in the course system by counting the label recurrence frequency of the target language unit in the subsequent untaught language unit. High recurrence frequency indicates that the knowledge points covered by it will be repeated in future teaching, which has high teaching continuity and cognitive basis, and therefore has strong association with subsequent learning content. When the language unit has mastering deviation, it indicates that the knowledge defect may be repeatedly exposed in the subsequent high-frequency recurrence scene, thereby causing persistent understanding difficulties and learning efficiency decline.
[0074] Therefore, on the basis of the identified target language unit, the recurrence frequency is taken as a key indicator for measuring the breadth of its knowledge influence, which is used to evaluate the potential interference degree of the preposed weak knowledge on the future learning path, and to provide data support and decision basis for the generation of preposed intervention strategies for the teaching progress adjustment module.
[0075] In another embodiment of the present application, the skill dependency relationship is determined as follows: based on the course knowledge system and teaching logic, a directed acyclic graph form skill dependency graph is established, each node in the graph represents a language unit, and the edge represents the preposed and postposed dependency relationship of skill development. The graph reflects the cognitive order in the process of language ability advancement.
[0076] As an embodiment of the above-mentioned scheme, it is assumed that a certain English course contains the following language units, as shown in Table 1.
[0077] Table 1: Phonological units and their knowledge points
[0078] The process of constructing the skill dependency graph is as follows: mastering the verb be and subject-verb agreement is the prerequisite for understanding the general present tense of the substantive verb, and the edge LU1→LU2 is established.
[0079] The general present tense and the present progressive tense form a contrast in form and semantics, and mastering the former helps to understand the structural difference of be+verb-ing in the latter, and the edge LU2→LU3 is established.
[0080] Mastering the verb-ing form rule provides an analogy basis for learning the past form-ed, and the edge LU3→LU4 is established.
[0081] Mastering the past form of regular verbs provides a cognitive anchor for learning irregular verbs, and the edge LU4→LU5 is established.
[0082] The skilled use of the past tense is the basis for understanding the past experience meaning in the present perfect tense, and the edge LU5→LU6 is established.
[0083] After mastering the basic structure of the present perfect tense, further learning of its continuous use and the difference from the general past tense can be carried out, and the edge LU6→LU7 is established.
[0084] The structure be+doing of the present progressive tense is the direct preposed knowledge for understanding the past progressive tense was / were+doing, and the edge LU3→LU8 is established.
[0085] The skill dependency graph constructed by applying the above example is shown in Figure 3 .
[0086] In the constructed skill dependency graph, the target language unit is taken as the starting node, and all the subsequent language unit sets reachable from the starting node are extracted to form a forward propagation link. The propagation link reflects the range of subsequent learning paths supported by the target language unit.
[0087] The number of nodes on the longest directed path from the target language unit to all its successor nodes is counted from the forward propagation link, and is defined as the forward skill influence path length of the target language unit.
[0088] The forward skill influence path length in the above is used as a representation of the teaching range and cognitive extension depth covered by the knowledge and skills required to be mastered in future learning. A longer forward path length means that the language unit constitutes the basis for the ability of subsequent multiple language units, and has a higher influence range in the skill development chain. When the unit is not mastered, its cognitive defects will be transmitted through the directed edges in the skill dependency graph, causing understanding obstacles or application difficulties in subsequent dependent units, leading to decreased learning efficiency and error pattern solidification, forming a risk accumulation effect.
[0089] The present application identifies the target language unit and constructs a multi-dimensional correlation strength evaluation model by fusing its knowledge point recurrence frequency in subsequent teaching and forward skill influence path length. The recurrence frequency reflects the occurrence density of the unit involved knowledge points in future courses, and embodies the horizontal penetration strength of knowledge content; the forward skill influence path length reflects the farthest teaching level that can be influenced in the skill dependency graph, and embodies the vertical depth of ability development. The two constitute a knowledge density-structure depth two-dimensional correlation strength index, which is used to quantify the overall dependency between the target language unit and subsequent learning content.
[0090] The teaching progress adjustment module is used to make cumulative progress risk judgment based on the correlation strength analysis result, and then adjust the teaching progress arrangement with cumulative progress risk.
[0091] Optionally, the cumulative progress risk judgment based on the correlation strength analysis result refers to the following process: taking the recurrence frequency of the target language unit knowledge point in subsequent teaching and the forward skill influence path length as the correlation strength representation.
[0092] The correlation strength representation is compared with the risk judgment boundary. If any of them exceeds the risk judgment boundary, it is determined that there is a cumulative progress risk.
[0093] The risk judgment boundary in the above can be set by language teaching experts according to the course structure and cognitive rules to set an empirical threshold. For example: if a language unit is repeated ≥5 times in subsequent teaching, or its forward skill influence path length ≥3 nodes, it is considered to have a significant influence and is included in the risk attention range.
[0094] If both correlation strength representations exceed the risk determination boundary, the cumulative progress is evaluated as high risk, indicating that the unit is a critical bottleneck node, and its lack of mastery will lead to systemic learning blockage, and if only one exceeds the risk determination boundary, the cumulative progress is evaluated as low risk, indicating that the unit knowledge is isolated and has limited impact, and intervention can be delayed.
[0095] Further optionally, the teaching progress arrangement with cumulative progress risk is adjusted as follows: for the subsequent language unit evaluated as high cumulative progress risk, a review link is inserted before the start of its teaching, forming a closed-loop mechanism of first filling in the foundation and then advancing.
[0096] For the subsequent language unit evaluated as low cumulative progress risk, lightweight knowledge prompts are embedded in the learning process to achieve implicit filling without affecting the overall pace.
[0097] In a specific example, the lightweight knowledge prompts are grammar prompts, speaking rule prompts, vocabulary prompts, etc.
[0098] The present application realizes adaptive optimization of personalized learning paths, and different intervention strategies are implemented for subsequent teaching progress arrangements according to the cumulative progress risk level of the target language unit, so as to ensure that the learner has the necessary pre-ability before entering new knowledge learning.
[0099] The above embodiments can be realized by software, hardware, firmware or any combination thereof, in whole or in part. When realized by software, the above embodiments can be realized in the form of a computer program product in whole or in part.
[0100] Those skilled in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized by hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0101] In addition, the functional modules in each embodiment of the present application can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0102] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0103] Finally, the above merely describes preferred embodiments of the present application and is not used to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A software operation data intelligent processing and analysis system based on big data, characterized in that, Includes the following modules: Learning data collection module: Collects learning record data of users' historical language units in the four aspects of listening comprehension, oral expression, reading comprehension and written expression from the online English learning software, and extracts the system score and practice time for each aspect; Language mastery analysis module: Based on the practice time and interaction behavior of the learning records, effective learning records are selected, and the scores of the four stages under each language unit in the effective learning records are analyzed in the order of teaching progress, and the mastery ability of each language unit under each stage is calculated. Mastery deviation identification module: For each language unit, compare the differences in mastery ability among the four stages to identify the target language unit where mastery deviation exists; Unit Relationship Analysis Module: Analyzes the strength of the relationship between the identified target language units and subsequent language units through knowledge point tag matching and skill dependency analysis; Teaching schedule adjustment module: Based on the results of correlation strength analysis, the module determines the cumulative schedule risk and then adjusts the teaching schedule arrangements that have cumulative schedule risks.
2. The software operation data intelligent processing and analysis system based on big data as described in claim 1, characterized in that: The learning record data includes the following collection process: Extract the sequence of completed language units according to the time arrangement of the teaching schedule; For each completed language unit, the listening comprehension score is automatically scored from the learning record of the corresponding listening training session. Extract the automatically scored oral expression score from the learning records of the oral practice session; Extract the automatically scored reading comprehension scores from the learning records of the reading training session; Extract the automatically scored written expression score from the learning records of the writing task phase; Simultaneously, the practice time for each step in the actual learning process is recorded; The system scores and practice time for the four stages are stored in association according to the corresponding language units.
3. The software operation data intelligent processing and analysis system based on big data as described in claim 1, characterized in that: The effective learning record filtering based on practice duration and interaction behavior is described below: Compare the practice time in the corresponding learning records of each completed language unit with the total practice time set for the learning task. Extract user operation event sequences from learning records, including timestamps for each interaction, calculate the time interval between adjacent operation events based on the timestamp sequence, take the arithmetic mean to obtain the average user interaction interval duration, and compare it with the normal interaction interval benchmark value set in the learning task. A learning record is considered invalid if it meets any of the following conditions: (1) The practice time did not reach the full practice time set for the learning task; (2) The average user interaction interval exceeds the normal interaction interval baseline value set in the learning task; All learning records that were not deemed invalid were classified as valid learning records.
4. The software operation data intelligent processing and analysis system based on big data as described in claim 1, characterized in that: The process for calculating the mastery ability of each language unit at each stage is as follows: For each completed language unit, extract the system score of all valid learning records for that stage; Calculate the average system score of all valid records under each stage, and use it as an indicator of the mastery level of that stage; Based on the passing score for each stage of the language unit's teaching objectives, the ratio of the mastery level indicator for each stage to the passing score is used as the ability mastery value for that language unit at each stage.
5. The software operation data intelligent processing and analysis system based on big data as described in claim 4, characterized in that: The specific details regarding the comparison of ability mastery differences among the four stages for each language unit are as follows: For each completed language unit, its ability mastery value in the four stages is standardized and mapped into a four-bar ability distribution chart arranged in the order of the preset stages, with the height of each bar corresponding to the standardized mastery score; Based on the ideal state where all skill levels have the same ability score, determine the ideal ability distribution midline; For a given language unit, calculate the weighted center position of the actual ability distribution based on the mastery scores of the four stages; The weighted median position of the actual ability distribution calculated for each completed language unit is compared with the median of the ideal ability distribution. If the weighted median position of the actual ability distribution corresponding to a certain language unit deviates from the median of the ideal ability distribution, the language unit is marked as having inter-stage mastery deviation.
6. The software operation data intelligent processing and analysis system based on big data as described in claim 5, characterized in that: The process of identifying target language units with mastery bias includes the following steps: Under each language unit, the competency mastery values of the four aspects are compared with the group baseline of the corresponding skill aspect. If the competency mastery value of a certain aspect is lower than the group baseline, then that aspect is identified as a weak aspect. The language units where there are gaps or weaknesses in the understanding between different stages are identified as target language units.
7. The software operation data intelligent processing and analysis system based on big data as described in claim 1, characterized in that: The content matched by the knowledge point tags is as follows: Establish a set of knowledge point tags for each language unit, and extract the knowledge point tags for the target language unit; Based on the teaching progress, locate the language units that have not yet been learned by the user after the target language unit in the teaching plan, form a set of units to be learned in the future, and extract the set of knowledge point tags for each unit to be learned in the future. Iterate through each language unit in the forward learning unit set, count the occurrence of the knowledge point tags contained in the target language unit in its knowledge point tag set, and accumulate the frequency of tag occurrence as the recurrence frequency of the target language unit knowledge points in subsequent teaching.
8. The software operation data intelligent processing and analysis system based on big data as described in claim 7, characterized in that: The skill dependency relationship is determined as follows: Based on the curriculum knowledge system and teaching logic, a skill dependency graph in the form of a directed acyclic graph is established. Each node in the graph represents a language unit, and the edges represent the pre- and post-dependent relationships of skill development. In the constructed skill dependency graph, starting with the target language unit as the starting node, extract all its reachable subsequent language unit sets to form a forward propagation link starting from that node; The number of nodes on the longest directed path from the target language unit to all its successor nodes in the forward propagation chain is defined as the forward skill influence path length of the target language unit.
9. The software operation data intelligent processing and analysis system based on big data as described in claim 8, characterized in that: The process for determining cumulative schedule risk based on correlation strength analysis results is as follows: The frequency of recurrence of target language unit knowledge points in subsequent teaching and the length of the forward skill influence path are used as the representation of the association strength. By comparing the correlation strength indicator with the risk assessment boundary, if either indicator exceeds the risk assessment boundary, it is determined that there is a cumulative schedule risk.
10. The software operation data intelligent processing and analysis system based on big data as described in claim 1, characterized in that: The adjustments to the teaching schedule that pose a risk of cumulative progress are as follows: If both correlation strength indicators exceed the risk assessment boundary, the cumulative progress is assessed as high risk; if only one exceeds the risk assessment boundary, the cumulative progress is assessed as low risk. For subsequent language units that are assessed as having a high risk of cumulative progress, a review session should be inserted before the start of their instruction. For subsequent language units that are assessed as having low cumulative progress risk, lightweight knowledge cues are embedded during the learning process.
Citation Information
Patent Citations
English learning system and method based on mobile device
CN119474687A
Test and evaluation method for English learning process
CN119692871A
English-oriented learning content recommendation method and system
CN120316351A
Online course learning management method and system based on knowledge graph
CN120563068A
Teaching information management method based on big data
CN120746039A