Intra-disciplinary equilibrium combination interval constitutive analysis method and system

Through standardized processing of subject learning data and combining interval division, a discipline core literacy relationship matrix is constructed, and the problem of in-depth analysis of learning quality differences in the traditional evaluation method is solved, and accurate subject quality assessment and personalized teaching path optimization are achieved.

CN120373945AActive Publication Date: 2025-07-25SOUTH CHINA NORMAL UNIV

Patent Information

Application Number
CN202510444678.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The existing technology cannot deeply analyze the differences in learning quality among various components within the discipline, especially the relationship between different levels of achievement, making it difficult for education managers to provide targeted and effective improvement measures.

Method used

By obtaining subject learning data, standardizing data, dividing combination intervals, building a combination interval model, comparing subject-level grade distribution, generating hierarchical feature data sets, establishing a discipline core literacy relationship matrix, and conducting subject learning quality evaluation and learning path optimization.

Benefits of technology

It has achieved accurate assessment of learning quality at different levels in the subject, identified weak links, provided personalized teaching path optimization, and improved the pertinence and accuracy of teaching quality monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373945A_ABST
    Figure CN120373945A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of education monitoring, in particular to an intra-subject equilibrium combination interval constitutive analysis method and system. The method comprises the following steps: obtaining subject learning data, and carrying out data standardization processing to obtain standardized subject learning data; performing interval structured modeling based on the standardized subject learning data to obtain a combined interval model; performing subject learning quality difference evaluation according to the combined interval model to obtain a hierarchical feature data set; performing subject core accomplishment relation modeling based on the standardized subject learning data to obtain a subject core accomplishment relation matrix; and performing subject learning quality evaluation according to the subject core literacy relation matrix and the hierarchical feature data set, and performing learning path optimization to obtain a subject optimization strategy. The pertinence and accuracy of teaching quality monitoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of educational monitoring, and in particular, to a method and system for analyzing the composition of balanced combination intervals within a subject. Background Art

[0002] Currently, the assessment of subject learning quality in the field of education generally relies on one-dimensional statistical means, such as calculating the mean and ranking analysis of students' exam scores. This traditional technical method mainly focuses on the overall scores of different regions, schools or student groups. Although it can provide a general score distribution, it lacks in-depth analysis of the internal composition of the subject and the relationship between different learning levels. The existing technology mainly calculates the mean through score data, and cannot reveal the differences in learning quality among the various components within the subject, especially the relationship between different score levels.

[0003] Traditional assessment methods usually rely on a single exam score indicator and ignore the structural analysis within the subject. The learning quality of a subject not only includes the total score, but also should involve the mastery of knowledge points, the achievement of core literacy, and the performance of students on questions at different difficulty levels. The existing methods cannot reveal the internal connections between different parts of the subject, nor can they understand the balance of the content examined in the subject, thus making it difficult for educational administrators to comprehensively understand the subject quality problems, and further affecting the pertinence and effectiveness of decision-making.

[0004] In addition, the traditional methods have not effectively achieved a comprehensive analysis of the various parts within the subject. Especially in multi-level and multi-dimensional quality diagnosis, how to identify the relationships between the various parts within the subject and provide data support for improvement remains an urgent problem to be solved. Therefore, the existing technology cannot provide accurate subject analysis for educational decision-makers, and cannot propose effective improvement measures according to the characteristics of different learning levels and student groups. Summary of the Invention

[0005] Based on this, it is necessary for the present invention to provide a method and system for analyzing the composition of balanced combination intervals within a subject to solve at least one of the above technical problems.

[0006] To achieve the above object, a method for analyzing the composition of balanced combination intervals within a subject includes the following steps:

[0007] Step S1: Obtain subject learning data, and perform data standardization processing on the subject learning data to obtain standardized subject learning data;

[0008] Step S2: Based on the standardized subject learning data, perform combination interval division to obtain subject combination interval data; perform interval structure modeling on the subject combination interval data to obtain a combination interval model;

[0009] Step S3: Compare the subject-level score distributions according to the combined interval model to obtain subject-level score deviation data; conduct an assessment of the differences in subject learning quality based on the subject-level score deviation data to obtain a hierarchical feature dataset;

[0010] Step S4: Build a model for the relationship between subject core competences based on the standardized subject learning data to obtain a subject core competence relationship matrix;

[0011] Step S5: Conduct an assessment of subject learning quality according to the subject core competence relationship matrix and the hierarchical feature dataset to obtain a subject learning quality assessment report, and optimize the learning path for the standardized subject learning data based on the subject learning quality assessment report to obtain a subject optimization strategy.

[0012] Through the standardized processing of subject learning data, the present invention can ensure that various types of data are maintained within a unified scale range, thereby avoiding inaccurate evaluations caused by differences in data units, making subsequent analyses more comparable and scientific. The standardized data can serve as an effective basis in subsequent steps, providing high-quality inputs for subsequent analyses of subject-level grade distributions and modeling of the relationships between subject core competencies. When dividing subject combination intervals, the use of combination intervals can effectively stratify students' grades, and generate a combination interval model through interval structured modeling, thereby providing a refined partition perspective for subsequent analyses of subject-level grade distributions. By analyzing the grade performance of students within each interval, the differences in learning quality at different levels within the subject can be deeply revealed, helping education administrators identify the unevenness of grade distributions and providing a basis for targeted teaching and resource allocation. In interval structured modeling, the selection of combination interval parameters can help accurately divide different-level intervals, so that differences in subject performance within each grade level can be precisely captured during analysis, avoiding omission of details. By generating subject-level grade deviation data based on the combination interval model, and further through a hierarchical feature dataset for evaluating differences in subject learning quality, the differences in learning quality in different parts of the subject can be effectively reflected. Through hierarchical analysis, the differences between different dimensions within the subject can be revealed, helping to identify weak links that need to be optimized, rather than simply making evaluations based on total scores. In addition, the generation of the hierarchical feature dataset helps to precisely lock in the groups or subject content that require key support for educational resources, so as to conduct refined control during the teaching improvement process. The construction of the subject core competency relationship matrix utilizes the interaction characteristics between core competencies within the subject. By analyzing the relationships between various dimensions within the subject, it can reveal the interactions between subject knowledge points, core competencies, and student performance. This step can help identify the structural associations between different parts of the subject, especially the interaction relationships between different competency dimensions, by constructing a core competency relationship diagram in matrix form, enabling evaluations not to be limited to single grades, but to comprehensively evaluate students' learning effects from broader perspectives such as knowledge mastery and ability development. During the generation of the subject learning quality assessment report, the comprehensive use of the core competency relationship matrix and the hierarchical feature dataset can provide accurate subject analyses for education decision-makers, helping to identify aspects within the subject that need to be optimized and improved. Through this step, education administrators can optimize personalized and hierarchical teaching paths based on specific subject quality assessment reports, thereby targetedly improving teaching effects.

[0013] Optionally, step S1 is specifically as follows:

[0014] Step S11: Collect multi-source original subject learning data through the regional teaching management platform, and perform unified encoding on the data frame structure of the multi-source original subject learning data to obtain subject learning data;

[0015] Step S12: Detect missing values in the subject learning data, and fill the missing values detected with the category mode to obtain cleaned data;

[0016] Step S13: Normalize the fields of the cleaned data, and standardize the encoding of the categorical fields for the normalization result to generate transformed data;

[0017] Step S14: Perform scale standardization on the transformed data to obtain standardized subject learning data.

[0018] By unifying the encoding of the structures of multi-source original subject learning data, the present invention solves the problem of inconsistent data formats from different sources, providing a consistent data framework for subsequent analysis. Filling missing values with the category mode can maintain the representativeness of the data distribution without introducing abnormal deviations, and is applicable to most processing scenarios of educational data. Field normalization and standardization of categorical field encoding ensure that numerical fields and categorical fields have a unified expression scale, which is conducive to the effective learning of features by the model. Scale standardization further maps the data to the same range, reducing the dimensional influence between dimensions and improving the accuracy of clustering and model analysis. The overall process improves the data quality and expression consistency, laying a solid foundation for subsequent modeling and evaluation.

[0019] Optionally, the combination interval division described in step S2 is specifically as follows:

[0020] Extract the student achievement characteristics from the standardized subject learning data to obtain the student subject achievement data, and perform statistical analysis on the student subject achievement data to obtain the standardized total score data and the student achievement proportion data;

[0021] Conduct low-achieving student proportion division based on the student achievement proportion data to obtain the low-achieving student proportion data;

[0022] Set combination variables based on the standardized total score data and the low-achieving student proportion data, where the horizontal axis variable is set as the standardized total score data and the vertical axis variable is set as the low-achieving student proportion data, and construct a two-dimensional interaction coordinate system;

[0023] Extract the standardized distribution characteristics of the spatial combination analysis basis, and set the mean ± 0.67σ in the standardized distribution characteristics as the division threshold to divide the two-dimensional interaction coordinate system, and divide the horizontal axis and the vertical axis of the two-dimensional interaction coordinate system into three sections respectively to form a two-dimensional combination space label matrix;

[0024] Map the standardized subject learning data to the corresponding combination labels in the two-dimensional combination space label matrix according to the two-dimensional interaction coordinate system, mark the interval categories, and count the number of samples and the proportion in each interval to generate subject combination interval data.

[0025] The present invention extracts the characteristics of student scores and constructs a two-dimensional interactive coordinate system based on the proportion of low-scoring students, effectively realizing the horizontal and vertical joint evaluation of subject quality, and making up for the one-sidedness of the traditional analysis method based only on total scores. Setting the standardized total score as the horizontal axis and the proportion of low-scoring students as the vertical axis can reveal the synergistic relationship between the overall level and the distribution of weak students. The division threshold uses the mean ±0.67σ, which conforms to the 68% interval characteristics of the normal distribution, so that the interval division has statistical robustness and discrimination, which is conducive to identifying marginal and extreme groups. Mapping the sample data in this coordinate system and forming a combined space label matrix can intuitively identify the risk areas and advantage areas of learning quality, provide structured input for subsequent models, and enhance the accuracy and diagnostic depth of data stratification analysis.

[0026] Optionally, the subject-level score distribution comparison in step S3 is specifically as follows:

[0027] Perform spatial label aggregation processing on the interval combination model, classify the sample data under the same interval label hierarchically, and combine the subject combination interval data to generate the combined label student distribution data;

[0028] According to the distribution of standardized total score data, the interval boundary thresholds of low stratification, middle stratification and high stratification are set, and the combined label student distribution data is interval-stratified to generate a hierarchical distribution label matrix;

[0029] The standardized subject learning data is double-labeled mapped according to the combined label student distribution data and the level distribution label matrix to generate a multi-label level score data set, and the number of samples at each level under different combined labels in the multi-label level score data set is normalized and counted to obtain the level score frequency matrix;

[0030] Based on the hierarchical score frequency matrix, the difference measurement calculation is performed to evaluate the hierarchical score distribution deviation between different combination labels and obtain the subject-level score deviation data.

[0031] The present invention introduces a dual-label mapping mechanism to jointly classify students according to combined labels and grade levels, thereby enhancing the hierarchy and interpretability of the data structure. When setting the grade level, the boundary thresholds of the low, middle and high layers are delineated based on the standardized total score distribution to ensure that the hierarchical division has a statistical basis and balance, which helps to reveal the distribution patterns of students at different levels under each combined label. By normalizing and counting the multi-label hierarchical grade data set and constructing a hierarchical score frequency matrix, the bias caused by sample size differences can be effectively removed, thereby improving the fairness of the comparison. The final distribution deviation calculation can reveal the specific differences in hierarchical grades among the various groups, making up for the inability of traditional methods to depict the distribution structure of groups at different learning levels, and improving the refinement of subject quality diagnosis.

[0032] Optionally, the assessment of the subject learning quality difference described in step S3 is specifically as follows:

[0033] Perform distribution statistics and central tendency analysis on the subject-level score deviation data, calculate the mean, range, and standard deviation of the corresponding levels of each combination label, and generate a deviation distribution statistical matrix;

[0034] Based on the deviation distribution statistical matrix, conduct horizontal regional balance analysis to evaluate the performance consistency of different regions at each level of scores, and generate regional balance characteristic data;

[0035] Perform vertical level gradient analysis on the deviation distribution statistical matrix, calculate the continuity of score improvement and the probability of level transition, and generate a level gradient change model;

[0036] Fuse the regional balance characteristic data and the level gradient change model for weighted comprehensive evaluation to obtain subject learning quality difference assessment data;

[0037] Based on the subject learning quality difference assessment data and the multi-label level score data set, integrate the student level characteristics to obtain a hierarchical characteristic data set.

[0038] The present invention uses statistical parameters such as the mean, range, and standard deviation in the deviation distribution statistical matrix to describe the central tendency and dispersion degree of scores at each level of different combination labels, which helps to identify the performance stability and extreme value situations, thus making up for the defect that traditional mean analysis ignores the discrete characteristics. The regional balance analysis improves the fairness of subject assessment and the visibility of regional differences by horizontally comparing the level performance consistency of different regions; the level gradient analysis sets a transition probability index to measure the potential trend of students' development from a lower level to a higher level, providing a basis for dynamically evaluating the learning growth path. In the fusion analysis, a weighted method is used to comprehensively consider the regional balance and the gradient model, making the quality difference assessment results have both spatial coverage and longitudinal development. Finally, the integrated hierarchical characteristic data set provides high-granularity and multi-dimensional data support for subsequent core literacy modeling and path optimization.

[0039] Optionally, step S4 is specifically as follows:

[0040] Step S41: Identify subject element labels for the standardized subject learning data, extract student scores, question difficulties, practice scores, and core literacy labels, construct an element label vector set, and thus generate a literacy structured data set;

[0041] Step S42: Perform dimensionality reduction and feature redundancy removal processing on the literacy structured data set, extract highly correlated literacy association features, and generate a refined literacy feature set;

[0042] Step S43: Construct a collaborative performance tensor model of literacy indicators based on the refined literacy feature set, and decompose the performance intensity and mutual relationship of core literacy dimensions by using the collaborative performance tensor model of literacy indicators to generate a literacy interaction feature matrix;

[0043] Step S44: Construct a multi-dimensional relationship map based on the literacy interaction feature matrix, identify the hidden correlation structure between core literacy dimensions, and generate a core subject literacy relationship map;

[0044] Step S45: Perform a structural matrix representation on the core subject literacy relationship map to obtain a core subject literacy relationship matrix.

[0045] The present invention constructs literacy structured data by identifying students' scores, question difficulties, practical scores, and core literacy labels in standard subject learning data, effectively breaking through the single limitation of the traditional score dimension and realizing the multi-dimensional expression of subject elements. By removing redundant features and reducing dimensions, highly correlated literacy indicators are extracted, effectively reducing the model complexity and improving the subsequent analysis efficiency. The tensor model is used to depict the collaborative performance between core literacies, and the parameter decomposition process can extract the strong and weak relationships and complementary features between literacies, enhancing the depth of interactive understanding. Constructing a multi-dimensional relationship map can reveal the hidden structure between literacies, making up for the neglect of the internal logic of literacies in traditional analysis. The final matrix representation of the relationship map not only improves the structural readability of the data but also provides a standard input format for subsequent modeling, achieving the unity of high structural stability and high semantic integrity.

[0046] Optionally, step S44 is specifically:

[0047] Step S441: Perform a high-dimensional space node mapping on the literacy interaction feature matrix. Set the number of core literacy dimensions to 8, define each core literacy dimension as a graph node, and calculate the collaborative coefficient of the performance intensity of student labels in each core literacy dimension as the initial edge weight to generate an initial literacy node graph;

[0048] Step S442: Set the Gaussian kernel scale factor to 0.5, construct an adjacency matrix for the initial literacy node graph, and perform a non-linear transformation on the edge weights between nodes to generate a high-order literacy adjacency matrix;

[0049] Step S443: Set the number of encoding layers to 2 layers, and the number of neurons in each layer to [64, 32] respectively. Based on the high-order literacy adjacency matrix, perform graph structure encoding to optimize the semantic discrimination of node representations and generate a literacy semantic embedding graph;

[0050] Step S444: Perform a structural community detection on the literacy semantic embedding graph, identify hidden correlation subgroups, and generate a literacy subgroup graph;

[0051] Step S445: Merge the literacy subgroup graph and perform matrix structure reorganization to generate a core subject literacy relationship matrix.

[0052] The present invention maps the core literacy dimensions to graph nodes, and constructs a literacy node graph with the collaborative performance of student tags as edge weights, effectively depicting the actual interaction relationships between literacies. The core literacy dimensions are set to 8, which can cover mainstream subject assessment frameworks and ensure the integrity of the graph structure; by setting the Gaussian kernel scale factor to 0.5 for non-linear edge weight conversion, the adjacency matrix smooths the weak connections, enhancing the robustness of the graph structure. A two-layer coding structure is adopted and the number of neurons is set to [64, 32], which can enhance the semantic discrimination of nodes while ensuring computational efficiency, generating a more representative embedded graph. By graph community detection, the potential aggregation structure between literacies is mined, making up for the deficiencies of traditional linear relationship analysis. The final structure reorganization and matrix merging process retains the subgroup relationship and semantic embedding features in the graph, providing multi-dimensional structural support for subsequent learning quality diagnosis.

[0053] Optionally, step S5 is specifically as follows:

[0054] Step S51: Conduct a joint structural analysis on the subject core literacy relationship matrix and the hierarchical feature dataset to generate a quality deviation feature vector set;

[0055] Step S52: Conduct a clustering diagnosis analysis on the quality deviation feature vector set, identify the imbalance points and deviation regions, and generate a learning quality assessment report;

[0056] Step S53: Extract the teaching resource features and subject target requirement features from the standardized subject learning data;

[0057] Step S54: Construct a path weight recommendation model based on the learning quality assessment report, teaching resource features, and subject target requirement features;

[0058] Step S55: Conduct a student group portrait analysis based on the standardized subject learning data to obtain regional student group portrait data, and use the path weight recommendation model to perform path matching and optimization on the regional student group portrait data to obtain subject optimization strategies.

[0059] By jointly analyzing the core literacy matrix and hierarchical feature data, the present invention comprehensively depicts the structural features of subject quality deviation, making the diagnosis results more accurate and interpretable. Identifying the deviation areas through clustering helps to accurately locate the imbalance problems in the subject; setting up a multi-center clustering method can adapt to diverse student feature distributions. Extracting teaching resource and subject objective information as auxiliary factors enhances the content matching degree of path recommendation. The path weight recommendation model is constructed based on multi-source feature fusion, improving the personalization and rationality of path generation; the introduction of the attention mechanism can automatically allocate the weights of key factors, improving the recommendation accuracy. Finally, combining with the regional student group portrait for path optimization can provide differential strategy support for the learning foundation differences in different regions, significantly improving the efficiency of educational resource allocation and the accuracy of intervention.

[0060] Optionally, step S52 is specifically as follows:

[0061] Step S521: Perform principal component analysis and feature orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum dominant dimension number to 6 to extract the dominant deviation dimension factors, generating the deviation main cause tensor data;

[0062] Step S522: Perform density peak estimation on the deviation main cause tensor data, set the minimum sample number to 10, and the clustering distance threshold to 0.75 to identify potential deviation clustering centers, generating the subject imbalance clustering map;

[0063] Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the subject imbalance clustering map, and set the neighborhood number to 20 to measure the anomaly intensity of each clustering unit in the subject imbalance clustering map, identifying the deviation areas and generating the local imbalance anomaly distribution map;

[0064] Step S524: Combine the local imbalance anomaly distribution map and the deviation main cause tensor data, set the deviation factor weight threshold to 0.6 to calculate the deviation factors of each literacy dimension, generating the imbalance feature weight data set;

[0065] Step S525: Perform semantic reconstruction and structural induction on the imbalance feature weight data set, and output the learning quality assessment report.

[0066] The present invention compresses high-dimensional quality deviation features through principal component analysis. Under the conditions of setting an 85% variance contribution rate and at most six dominant dimensions, it effectively retains the main information while eliminating redundant dimensions, thus improving the analysis efficiency. Density peak estimation is used for deviation clustering, with the minimum number of samples set to 10 and the clustering distance threshold set to 0.75, ensuring that the clustering is representative and can accurately identify the core area of deviation. By setting the boundary spacing to 0.05 and the number of neighborhoods to 20, the local imbalance distribution is carefully characterized, which helps to discover edge anomalies and structural holes. The deviation factor weight threshold is set to 0.6, which can effectively screen out the significantly influential literacy deviation dimensions, making the diagnosis more targeted. Finally, through semantic reconstruction, the accuracy and interpretability of learning quality diagnosis are comprehensively improved.

[0067] Optionally, this specification also provides an in-discipline balanced combination interval constitutive analysis system for performing the in-discipline balanced combination interval constitutive analysis method as described above. The in-discipline balanced combination interval constitutive analysis system includes:

[0068] A data acquisition module for obtaining in-discipline learning data and performing data standardization processing on the in-discipline learning data to obtain standardized in-discipline learning data;

[0069] A combination interval generation module for performing combination interval division based on the standardized in-discipline learning data to obtain in-discipline combination interval data; performing interval structure modeling on the in-discipline combination interval data to obtain a combination interval model;

[0070] A combination interval analysis module for comparing the in-discipline hierarchical performance distribution according to the combination interval model to obtain in-discipline hierarchical performance deviation data; performing in-discipline learning quality difference evaluation based on the in-discipline hierarchical performance deviation data to obtain a hierarchical feature data set;

[0071] A core literacy relationship modeling module for performing in-discipline core literacy relationship modeling based on the standardized in-discipline learning data to obtain an in-discipline core literacy relationship matrix;

[0072] A quality diagnosis module for performing in-discipline learning quality evaluation according to the in-discipline core literacy relationship matrix and the hierarchical feature data set to obtain an in-discipline learning quality evaluation report, and optimizing the learning path of the standardized in-discipline learning data based on the in-discipline learning quality evaluation report to obtain an in-discipline optimization strategy.

[0073] The in-discipline balanced combination interval constitutive analysis system of the present invention can implement any in-discipline balanced combination interval constitutive analysis method of the present invention. It is a medium for coordinating the operations and signal transmissions between various modules to complete the in-discipline balanced combination interval constitutive analysis method. The internal modules of the system cooperate with each other, thereby improving the pertinence and accuracy of teaching quality monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non - limiting embodiments read in conjunction with the accompanying drawings:

[0075] Figure 1 It is a schematic flowchart of the steps of the method for analyzing the composition of the balanced combination interval within the discipline of the present invention;

[0076] Figure 2 It is a detailed flowchart of step S1 in the present invention;

[0077] The realization of the object of the present invention, functional features, and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed Embodiments

[0078] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0079] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0080] It should be understood that although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed related items.

[0081] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a method for analyzing the composition of the balanced combination interval within a discipline, and the method includes the following steps:

[0082] Step S1: Obtain the subject learning data, and perform data standardization processing on the subject learning data to obtain the standardized subject learning data;

[0083] In this embodiment, subject learning data from third-year junior high school students in a certain area are first collected, including student personal student numbers, grades in various subjects (Chinese, mathematics, English, physics, chemistry, etc.), full marks for questions, scores, knowledge point codes to which questions belong, question difficulty coefficients (ranging from 0 to 1), answering time and whether they are completed, subject goal requirements, and core literacy labels to which questions belong. The collected original subject learning data is cleaned, fields with a missing rate greater than 30% are removed, and the continuous score fields are standardized using the Z-score processing method. The standardized threshold is set to ±3σ, and values exceeding this range are considered to be abnormal values for correction or elimination. Finally, a standardized subject learning data set containing 12,823 students, 6 courses, and an average of 100 questions per course is obtained. After standardization, all score variables are converted to a normal distribution with a mean of 0 and a standard deviation of 1 to ensure comparability and interpretability in subsequent analysis.

[0084] Step S2: performing combination interval division based on the standardized subject learning data to obtain subject combination interval data; performing interval structured modeling on the subject combination interval data to obtain a combination interval model;

[0085] In this embodiment, based on the obtained standardized subject learning data, the students' standardized total scores (weighted average by weight) and the proportion of questions with scores below 50 points are first extracted as two variables for dividing the two-dimensional interactive coordinate system. The horizontal axis is set as the standardized total score (μ=0,σ=1), and the vertical axis is the proportion of low-score questions (μ=0.35,σ=0.12). According to the mean ±0.67σ, it is divided into three sections of "low, medium, and high", corresponding to the intervals: [-∞,-0.67σ), [-0.67σ,+0.67σ], (+0.67σ,∞]. Based on this, a 3×3 two-dimensional combination space label matrix is constructed, with a total of 9 intervals (such as LL, MH, etc.), where "LH" indicates a low total score and a high proportion of low scores. Subsequently, the standardized subject learning data is mapped to the combination interval matrix, and the number and proportion of samples in each interval are counted, and finally a subject combination interval data table containing each interval label, number of samples, average score and standard deviation is generated. The KMeans algorithm is further used to perform structural modeling on the combination interval data, setting the cluster number K=4, the initialization method is k-means++, and the maximum number of iterations is 300, and the combination interval model is obtained for subsequent stratified comparison.

[0086] Step S3: Compare the subject-level score distribution according to the combined interval model to obtain subject-level score deviation data; evaluate the subject learning quality differences based on the subject-level score deviation data to obtain a hierarchical feature data set;

[0087] In this embodiment, first, based on the combined interval model, the samples are divided into four types of student groups, namely, the "high-performance area", "medium area", "area to be improved", and "key attention area". Then, hierarchical performance comparison is carried out according to the score distributions of each group in different subjects. Taking the average scores of students in the "high-performance area" in mathematics and physics as the benchmark, the mean deviation and variance change of the "area to be improved" in the same subjects are compared. The results show that there is an average difference of -1.2σ in mathematics and the variance increases by 22.3%. The above deviation features are extracted to construct the subject-level performance deviation data. Further, based on this data, the quality differences between different subjects are analyzed. The quality balance index G (the absolute value of the standardized mean difference) is introduced. When G > 1.0, it is judged as unbalanced. The least squares regression is used to calculate the complementary intensity between subjects, and combined with the hierarchical distribution uniformity E (defined as the standard deviation of the sample numbers in each layer), finally, a hierarchical feature dataset is generated as the feature expression of the learning quality structure.

[0088] Step S4: Based on the standardized subject learning data, establish a relationship model of subject core literacy to obtain a subject core literacy relationship matrix;

[0089] In this embodiment, based on the standardized subject learning data, first, the "knowledge point - performance - literacy label" triple between the questions and students is identified to construct a preliminary literacy structured dataset. The score performances of each student in the core literacy labels of the questions (such as logical thinking, language expression, experimental operation, comprehensive application, etc.) are extracted. The information gain method is used to remove redundant literacy label features, and the features with the top 10 mutual information values are retained to construct a refined literacy feature set. The performances of students in different literacy dimensions are tensorized to form a three-dimensional collaborative performance tensor of "student × literacy × question type". The CP decomposition method (Rank is set to 5) is used to extract potential interaction factors and generate a literacy interaction feature matrix. Then, based on this matrix, a multi-dimensional relationship graph is constructed. Each node in the graph represents a literacy dimension, and the edge weight represents the collaborative performance intensity between the dimensions. The Louvain algorithm is used for community detection to visualize and strengthen the relationship of the implicit association subgroups. Finally, through the structural matrix expression, a subject core literacy relationship matrix is obtained, with a dimension of 9×9, and each cell identifies the potential association intensity between the literacy dimensions.

[0090] Step S5: According to the subject core literacy relationship matrix and the hierarchical feature dataset, conduct a subject learning quality assessment to obtain a subject learning quality assessment report, and based on the subject learning quality assessment report, optimize the learning path of the standardized subject learning data to obtain a subject optimization strategy.

[0091] In this embodiment, first, the obtained matrix of subject core literacy relationships and the obtained hierarchical feature dataset are subjected to principal component fusion, and the combination of literacy dimensions with the largest performance fluctuations is extracted as the quality deviation feature vector. DBSCAN clustering analysis (eps = 0.4, min_samples = 10) is used to identify imbalances in the abnormal clustering points in the deviation space, and a learning quality assessment report is output. Subsequently, the standard subject learning data is subjected to feature annotation, and the question source, the level of examined ability (memory, understanding, application, etc.), the resource type (text, video, experiment, etc.) are extracted. At the same time, the Bloom level required by the subject target and the key abilities of the curriculum standard are extracted to construct a target matching vector. Based on the matching path between the diagnostic result and the teaching target, a regulation path network is constructed, and the weight of the path node is set as the deviation degree × matching degree. The conversion probability uses the historical path recommendation confidence (range [0.3 - 0.9]) to generate a path weight recommendation model. Finally, according to the model, path recommendation matching is performed on the regional group portrait data, and a subject optimization strategy table containing "strategy sequence, strategy implementation difficulty, expected improvement direction" is output to guide personalized teaching and path intervention.

[0092] Optionally, step S1 is specifically as follows:

[0093] Step S11: Collect multi-source original subject learning data through the regional teaching management platform, and perform unified coding on the data frame structure of the multi-source original subject learning data to obtain subject learning data;

[0094] In this embodiment, through the data interface of the regional teaching management platform, the original subject learning data in multiple data sources is collected, including data from online homework platforms (such as intelligent homework systems), classroom teaching platforms (such as learning situation analysis systems), and paper exam score entry systems, covering 5 subjects including Chinese, mathematics, English, science, and history, and a total of 132 schools, 8,547 teaching classes, approximately 246,000 student learning records and class performance records are included. There are problems with inconsistent field naming and data formats in the data of different platforms. A unified data access module is used for field structure identification and mapping. By defining a unified data structure template (fields include student ID, subject code, question ID, answer result, score, knowledge point label, answer time, answer duration, class performance score, subject core literacy, etc., a total of 18 items), all original data is standardized into a JSON format object with a unified structure, and finally a multi-dimensional fusion dataset containing subject learning behavior, grades, and process performance is generated and uniformly output in CSV format as the input data for subsequent cleaning.

[0095] Step S12: Detect missing values in the subject learning data, and fill the missing value detection results with the category mode to obtain cleaned data;

[0096] In this embodiment, missing value detection is performed on the subject learning data generated in step S11, and the missing rates of each field are respectively counted. For numerical fields such as "score" and "answering duration", the record counting method is used to detect null values. For categorical fields such as "knowledge point label" and "answering status", the regular expression + default marker recognition method is used to mark invalid items. The detection results show that the missing rate of the field "answering duration" is 12.4%, and the missing rate of the field "knowledge point label" is 9.1%. For the missing values of categorical fields, the category mode is used for filling. For example, the mode of the "knowledge point label" field is "recognition of geometric figures", and all missing items are replaced with this value. For the "answering duration" field, the missing threshold is set to 15%. Since it is lower than the threshold range, this field is retained and filled with the median of the overall distribution (45 seconds in this data). The missing rate of the finally obtained cleaned data is less than 1%, meeting the requirements of subsequent normalization and standardization modeling.

[0097] Step S13: Perform field normalization on the cleaned data, and perform categorical field encoding standardization on the normalization result to generate transformed data;

[0098] In this embodiment, field normalization processing is performed on the obtained cleaned data. For numerical fields (such as "score", "answering duration", "question difficulty", etc., a total of 9 items), the min-max normalization method is used to linearly scale all numerical values to the [0,1] interval. The normalization formula is: (X - X_min) / (X_max - X_min). For example, for the "score" field, the original score range is [0,100], and the highest value after normalization is 1.0 and the lowest value is 0.0. Subsequently, unified encoding processing is performed on the categorical fields (such as "knowledge point label", "answering status", "question type") contained in the data. The label encoding method (LabelEncoding) is used to map each category value to a unique integer value. For example, in the "question type" field, "multiple-choice question", "fill-in-the-blank question", and "answer question" are encoded as 0, 1, 2, ensuring that all input data can be used for subsequent machine learning model or graph construction processing. The format of the transformed data is a matrix structure, including normalized continuous variables and uniformly encoded discrete variables, forming a complete transformed data set for subsequent standardization operations.

[0099] Step S14: Perform scale standardization processing on the transformed data to obtain standardized subject learning data.

[0100] In this embodiment, scale normalization processing is performed on the generated conversion data. The normalization uses the Z-score normalization method to convert all numerical fields into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula is: (X - μ) / σ, where μ is the sample mean and σ is the sample standard deviation. For example, the original mean of the "response duration" field is 45 seconds, and the standard deviation is 13.2 seconds. After normalization, the mean becomes 0 and the standard deviation becomes 1, and outliers (such as entries with a response duration exceeding 90 seconds) are compressed to the tail of the distribution. A normality test is performed on the normalization result, and the Shapiro-Wilk method is used for distribution testing. The significance level is set to α = 0.05. The results show that most fields follow an approximate normal distribution, which is suitable for subsequent principal component extraction and structure modeling. The normalized subject learning data is stored in matrix form, containing a total of 243 feature fields, and the data size is 247382×243, serving as the basic input data for subsequent learning path analysis and feature modeling.

[0101] Optionally, the combination interval division described in step S2 is specifically:

[0102] Student subject performance features are extracted from the normalized subject learning data to obtain student subject performance data, and statistical analysis is performed on the student subject performance data to obtain normalized total score data and student performance proportion data;

[0103] In this embodiment, when extracting student subject performance features from the normalized subject learning data, field filtering is used to extract fields directly related to performance such as "final exam score", "usual score", "process evaluation", "homework score", and "test score" in each subject, excluding text-based remarks and non-quantitative fields of behavior types. The above fields are weighted and fused according to the weight coefficients (such as the weight of the final exam score is 0.5, the usual score is 0.2, and the rest is 0.3 in total) to obtain normalized student subject performance data. On this basis, the describe() function in the Pandas library is used to obtain basic statistics such as the mean, standard deviation, quantiles, maximum and minimum values, and the proportion of students in each grade interval (such as below 60 points, 60 - 80 points, above 80 points) is statistically calculated based on the score segments to form normalized total score data and student performance proportion data.

[0104] The proportion of low-performance students is divided according to the student performance proportion data to obtain the proportion data of low-performance students;

[0105] In this embodiment, when dividing the proportion of low-achieving students according to the proportion data of students' grades, first, the low-achievement standard is set as students with a total score lower than 60 points or in the bottom 20% percentile of the grade, and grouped statistics are carried out in the sample sets of each region, school or class. Using the threshold T_low = 60 or P_low = 20% as the dividing line, the proportion of the number of students below this threshold in each analysis unit to the total number of students is statistically obtained to get the proportion data of low-achieving students. To ensure stability, for analysis units with a sample size less than 30, a sliding window is used for mean smoothing processing.

[0106] Based on the standardized total score data and the proportion data of low-achieving students, a combined variable is set. Among them, the horizontal axis variable is set as the standardized total score data, and the vertical axis variable is set as the proportion data of low-achieving students, to construct a two-dimensional interaction coordinate system;

[0107] In this embodiment, when setting the combined variable based on the standardized total score data and the proportion data of low-achieving students, first, the standardized total score is determined as the horizontal axis variable, and the proportion of low-achieving students is determined as the vertical axis variable, to construct a two-dimensional interaction coordinate system. The data points of all analysis units are mapped into this two-dimensional coordinate system, and each unit data point represents the combined performance of a school, class or region. To ensure the consistency of variable dimensions, both variables are processed by Z-score standardization, where Z = (X - μ) / σ, to ensure that the coordinate coefficient value space distribution is symmetric and suitable for subsequent distribution analysis.

[0108] Extract the standardized distribution characteristics of the spatial combination analysis base, and set the mean ± 0.67σ in the standardized distribution characteristics as the division threshold to divide the two-dimensional interaction coordinate system. The horizontal axis and the vertical axis of the two-dimensional interaction coordinate system are respectively divided into three sections to form a two-dimensional combined space label matrix;

[0109] In this embodiment, when extracting the standardized distribution characteristics of the spatial combination analysis base, based on the two-dimensional coordinate point set, the distribution mean μ and standard deviation σ of the horizontal axis and vertical axis variables are respectively extracted, and ± 0.67σ is used as the interval division threshold. According to the empirical rule of normal distribution, each axis is divided into three sections. Taking the horizontal axis as an example: less than μ - 0.67σ is defined as the low section, the middle between μ ± 0.67σ is defined as the middle section, and greater than μ + 0.67σ is defined as the high section; the vertical axis is divided in the same way, and finally a 3×3 nine-grid combined space is formed. Each combination interval is set with an independent identification code, such as "L-M" (low achievement - medium waiting for improvement), for subsequent mapping and label classification.

[0110] According to the two-dimensional interaction coordinate system, map the standardized subject learning data to the corresponding combined label in the two-dimensional combined space label matrix, mark the interval category, and count the sample quantity and proportion in each interval to generate subject combination interval data.

[0111] In this embodiment, when mapping the standard subject learning data to the corresponding combined labels in the two-dimensional combined space label matrix according to the two-dimensional interaction coordinate system, first, the coordinate classification of the (total score, low score ratio) value pairs of each analysis unit (a subset obtained by splitting the standard subject learning data according to specific organizational dimensions such as schools, classes, school districts, student groups, etc.) is performed. By determining the section numbers on the X-axis and Y-axis, it is mapped to a specific interval in the 3×3 label matrix, and the combined label category to which the unit belongs is marked. Subsequently, all analysis units are classified and summarized according to the labels, and the number of samples and their proportion in the total samples in each label interval are counted to form structured subject combination interval data, providing basic coordinate information support for subsequent difference analysis and literacy modeling.

[0112] Of particular importance is that forming the two-dimensional combined space label matrix specifically involves:

[0113] Perform normal distribution parameter estimation on the standardized total score data, calculate the mean and standard deviation σ1 of the horizontal axis, and divide the score level interval with the mean ± 0.67σ1 as the threshold. Define respectively: the value of the horizontal axis less than -0.67σ1 is the low area of the horizontal axis, the value of the horizontal axis in the range [-0.67σ1, +0.67σ1] is the middle area of the horizontal axis, and the value of the horizontal axis greater than -0.67σ1 is the high area of the horizontal axis, thus generating the horizontal axis label set;

[0114] In this embodiment, normal distribution parameter estimation processing is performed on the standardized total score data. The maximum likelihood estimation method is used to fit the score sample distribution curve. The calculated standardized total score mean μ1 and standard deviation σ1 are 75.6 and 6.2 respectively. On this basis, the threshold parameter range is set as μ1 ± 0.67σ1, that is, the interval range is [71.44, 79.76]. The standardized total score data of all analysis units are divided into three sections according to this interval: low area (<71.44), middle area ([71.44, 79.76]), and high area (>79.76), and are respectively marked as "L" (Low), "M" (Medium), "H" (High) as the horizontal axis label set. During the division process, a continuous interval splitting strategy is adopted to avoid misjudgment of critical points, and at the same time, the division results are visually verified to ensure the balance of sample coverage.

[0115] Perform distribution parameter estimation on the low score ratio data, calculate the mean and standard deviation σ2 of the vertical axis, and divide the score level interval with the mean ± 0.67σ2 as the threshold. Define respectively: the value of the vertical axis less than -0.67σ2 is the low area of the vertical axis, the value of the vertical axis in the range [-0.67σ2, +0.67σ2] is the middle area of the vertical axis, and the value of the vertical axis greater than -0.67σ2 is the high area of the vertical axis, thus generating the vertical axis label set;

[0116] In this embodiment, the kernel density estimation is used to fit the distribution trend of the proportion data of low-achieving students. After excluding asymmetric abnormal data by combining the skewness coefficient detection, the mean μ2 of the proportion of low-achieving students on the vertical axis is calculated to be 23.4% and the standard deviation σ2 is 5.8%. Similarly, taking μ2 ± 0.67σ2 as the division threshold, the segmented interval is calculated to be [19.51%, 27.29%]. Accordingly, the vertical axis is divided into three segments: low zone (<19.51%), middle zone ([19.51%, 27.29%]), and high zone (>27.29%), and they are respectively marked as "L" (low proportion), "M" (medium proportion), and "H" (high proportion) to construct the vertical axis label set. During the division process, outlier compression processing is performed on the low-proportion fluctuation samples, and the z-score normalization method is used to make the distribution closer to the normal model, improving the label stability.

[0117] Based on the horizontal axis label set and the vertical axis label set, a two-dimensional combination space is constructed, and 9 interval labels are generated by cross-combination. Then the 9 interval labels are matrix-converted and output in matrix form to obtain the two-dimensional combination space label matrix.

[0118] In this embodiment, the generated horizontal axis label set {L, M, H} is cross-combined with the vertical axis label set {L, M, H} generated in step S222. The Cartesian product method is used to construct the two-dimensional combination space label set, and a total of 9 combination labels are obtained, including "L-L", "L-M", "L-H", "M-L", "M-M", "M-H", "H-L", "H-M", and "H-H". Subsequently, the label set is mapped to the spatial position matrix to construct a 3×3 two-dimensional combination space label matrix, where the row direction is the vertical axis section and the column direction is the horizontal axis section. Finally, the following matrix form is output: The label matrix serves as the spatial basis for the subsequent combination interval mapping, supports mapping the actual data of each analysis unit to the above 9 label regions, and is used for statistical analysis and interval proportion modeling.

[0119] Optionally, the comparison of the subject-level achievement distribution described in step S3 is specifically as follows:

[0120] Perform spatial label aggregation processing on the interval combination model, classify the sample data under the same interval label hierarchically, and combine the subject combination interval data to generate the combined label student distribution data;

[0121] In this embodiment, clustering and grouping are performed on the 9 constructed combined spatial labels (such as L-M, M-H, etc.). For the sample set within each label region, grouping is carried out according to the student ID, and key indicators such as the scores in each subject dimension, the frequency of course participation, and the completion of homework are uniformly encoded to construct a structured sample index table. "Sample proportion" is introduced as an aggregation weight factor into the sample data under the label to hierarchically aggregate the sample group characteristics, and the aggregation granularity is set to no less than 50 student samples under each label. After the aggregation is completed, combined with the original subject combination interval data, a combined label student distribution data set is generated through a label-sample mapping function to represent the number of students, the average score level, and their degree of difference contained in each combination interval, providing a basis for subsequent stratified analysis.

[0122] Set the interval boundary thresholds for the low-stratified, middle-stratified, and high-stratified levels according to the standardized total score data distribution, and perform interval stratification on the combined label student distribution data to generate a hierarchical distribution label matrix;

[0123] In this embodiment, based on the standardized total score data distribution, a three-segment interval division is performed on the student group using the quantile method. The low-stratified level is set as the students with a total score below the 30th percentile (threshold μ - σ), the middle-stratified level is set as the students between 30% and 70% (μ ± σ), and the high-stratified level is set as the students with a percentile above 70% (threshold μ + σ). The set parameters are: low score threshold ≤ 68.4, middle score interval [68.4, 81.2], high score threshold ≥ 81.2. Traverse the combined label student distribution data according to the combined label dimension, classify the student group data within each combined label according to the above score segments, complete the "low-middle-high" stratified label annotation corresponding to each type of label, and embed it into a unified data structure to form a hierarchical distribution label matrix with a dimension of 9×3, where each combined label corresponds to a three-layer structure representing the number and proportion distribution of different stratifications.

[0124] Map the standardized subject learning data according to the combined label student distribution data and the hierarchical distribution label matrix to generate a multi-label hierarchical score data set, and perform normalized statistics on the number of samples at each level under different combined labels in the multi-label hierarchical score data set to obtain a hierarchical score frequency matrix;

[0125] In this embodiment, the standard subject learning data is bidirectionally matched according to the combined label (such as L-M) and the hierarchical label (such as high level), and each student is uniquely located in the combined interval and the score level, realizing the double-label mapping. On this basis, a multi-label hierarchical performance data set is generated. Each record in the data set contains the combined interval label, the hierarchical label, the standardized performance value, and the corresponding student code. The multi-label hierarchical performance data set is statistically normalized, and the Min-Max normalization is used to uniformly scale the number of samples in each level to the interval [0,1] to ensure the comparability of the sample distribution comparison between different combined labels. Finally, a hierarchical score frequency matrix with a structure of 9×3 is formed. This matrix has the combined label as the row and the score level as the column, and each cell represents the normalized sample proportion value at the corresponding level.

[0126] Based on the hierarchical score frequency matrix, the difference metric calculation is carried out to evaluate the deviation of the hierarchical score distribution between different combined labels, and the subject hierarchical performance deviation data is obtained.

[0127] In this embodiment, the distribution difference analysis is carried out according to the hierarchical score frequency matrix. The Kullback-Leibler Divergence (KL divergence) is used to calculate the pairwise difference of the score frequency vectors between different combined labels to measure the hierarchical distribution deviation. The difference sensitivity threshold is set to 0.15. Any pair of combined labels with a KL divergence value higher than this threshold is considered to have a significant hierarchical deviation. At the same time, the Jensen-Shannon distance is combined for auxiliary verification. The output result is a structural subject hierarchical performance deviation data set, which records the deviation score and its relative standard deviation of each combined label at each level score, and at the same time identifies the label combinations with prominent deviations, providing a deviation positioning basis for subsequent path recommendation and quality diagnosis.

[0128] Optionally, the evaluation of the subject learning quality difference in step S3 is specifically as follows:

[0129] The distribution statistics and central tendency analysis are carried out on the subject hierarchical performance deviation data, and the mean, range, and standard deviation of the corresponding levels of each combined label are calculated to generate a deviation distribution statistical matrix;

[0130] In this embodiment, based on the subject-level performance deviation data, statistical analysis methods are used to measure the central tendency and dispersion degree of the hierarchical performance distribution under each combined label. First, the low, medium, and high-level scores corresponding to each combined label (such as L-M, M-H, etc.) are aggregated, and the mean, range (the difference between the maximum and minimum values), and standard deviation of the scores within each level are calculated respectively to measure the concentration and fluctuation degree of the performance distribution. The analysis accuracy is set to two decimal places, and extreme outliers outside 3 standard deviations are automatically filtered during the range calculation. The output deviation distribution statistical matrix is a 9×3-dimensional structure, where the rows represent 9 combined labels, and the columns represent the mean, range, and standard deviation statistical indicators of each level, which are used for subsequent horizontal and vertical analyses.

[0131] Based on the deviation distribution statistical matrix, horizontal regional balance analysis is carried out to evaluate the performance consistency of different regions in the scores of each level, and regional balance characteristic data are generated;

[0132] In this embodiment, according to the generated deviation distribution statistical matrix, horizontal division is carried out according to the regional dimension (such as streets, villages and towns), and analysis of variance (ANOVA) is performed on the mean values of the combined labels of each region under the same level dimension to measure the consistency degree of different regions in the level performance. Calculate the mean variance of different combined labels within the region. If the variance within the region is lower than the set threshold of 0.03, it is considered that the region has high consistency in the performance of this level. Map the consistency degrees of different regions at the three levels into a scoring matrix, and finally output the regional balance characteristic data, which includes the "region-level" two-dimensional index, the score balance index (value range [0,1]), and the mean shift rate index, as the horizontal index basis for subsequent gradient analysis and difference evaluation.

[0133] Perform vertical level gradient analysis on the deviation distribution statistical matrix, calculate the continuity of performance improvement and the probability of level transition, and generate a level gradient change model;

[0134] In this embodiment, using the level score means in the deviation distribution statistical matrix, level gradient vectors (low→medium→high) are constructed for each combined label, and the performance improvement amplitude and transition probability from the low level to the medium level and from the medium level to the high level are calculated respectively. The transition probability is based on the change in the proportion of the number of people between levels, and the transition probability Pij = Nij / (Nij + Nii) is defined, where Nij represents the number of students transitioning to the higher level, and Nii represents the number of students remaining at the current level. A threshold of more than 0.5 is set to determine a positive transition trend. Finally, a level gradient change model is generated, which includes the level improvement path, transition probability matrix (3×3), and score continuity index (such as improvement slope, etc.) for each combined label, and is used to evaluate the naturalness and controllability of students' performance improvement under this combination.

[0135] Integrate the regional balance characteristic data and the hierarchical gradient change model for weighted comprehensive evaluation to obtain the evaluation data of the subject learning quality difference;

[0136] In this embodiment, the regional balance characteristic data and the hierarchical gradient change model are integrated to conduct a weighted comprehensive evaluation of the performance of each combined label. The regional balance weight is set to 0.4, and the hierarchical gradient continuity weight is set to 0.6. A linear weighted model is used to sum and calculate the two types of indicators, and the comprehensive score evaluation index Gij = 0.4×Rij + 0.6×Lij is generated, where Rij is the regional consistency score and Lij is the normalized value of the hierarchical gradient slope. The output evaluation data of the subject learning quality difference includes the combined label, the comprehensive quality score, the main imbalance reason type (such as regional imbalance, low hierarchical jump, etc.) and the improvement direction prompt information, which serves as an important basis for student portrait integration and path recommendation.

[0137] Integrate the student hierarchical characteristics based on the evaluation data of the subject learning quality difference and the multi-label hierarchical achievement dataset to obtain the hierarchical characteristic dataset.

[0138] In this embodiment, based on the aforementioned generated evaluation data of the subject learning quality difference and the multi-label hierarchical achievement dataset, the information such as the combined label, the hierarchical position, and the score fluctuation trend of each student is fused to construct a five-tuple data structure including the student ID, the subject label, the hierarchical label, the comprehensive quality score, and the difference type. During the fusion process, a score fluctuation smoothing algorithm (the moving average window width is set to 3) is introduced to process the score continuity trend of the students and reduce the influence of extreme fluctuations. The finally generated hierarchical characteristic dataset includes the positions of all students in the combined space, the levels they are in, the score change patterns, the learning quality deviation types, etc., providing refined input features for subsequent path recommendation and resource regulation.

[0139] Optionally, step S4 is specifically as follows:

[0140] Step S41: Identify the subject element labels for the standardized subject learning data, extract the student scores, question difficulties, practice scores, and core literacy labels, and construct an element label vector set to generate a literacy structured dataset;

[0141] In this embodiment, based on standard subject learning data, the question fields, student answer records, scoring items, and evaluation criteria in the data are first parsed, and subject element tags are identified through natural language keyword matching and knowledge point label extraction algorithms. The specific extraction fields include: student score (score), question difficulty level (difficulty_level), experimental and practical score (practice_score), and the corresponding core literacy dimension tags of the questions (such as "information awareness", "practical ability", "innovative thinking", etc.). A word vector model is used to uniformly encode the core literacy tags, and the vector dimension is set to 128 dimensions. A structured element vector is constructed with students-questions as the unit, forming a triple set containing student IDs, question IDs, and element label vectors. Finally, a literacy structured dataset is generated, and the data structure is in the form of a two-dimensional nested dictionary. The first-dimensional index is the student ID, and the second-dimensional index is the question ID. The corresponding data is the complete element label vector.

[0142] Step S42: Perform dimensionality reduction and feature redundancy removal on the literacy structured dataset, extract highly correlated literacy association features, and generate a refined literacy feature set;

[0143] In this embodiment, in the generated literacy structured dataset, first, a dimension compression algorithm based on PCA (Principal Component Analysis) is used for preliminary dimensionality reduction, and the retained variance explanation rate is set to 90%. The original 128-dimensional literacy element vector is reduced to between about 30 and 40 dimensions. Then, the correlation threshold method is used for feature redundancy removal. The Pearson correlation coefficient between each feature is calculated, and the redundancy removal threshold is set to 0.92. If the correlation between two features is higher than this value, the one with the higher information entropy is retained. For categorical features, the mutual information coefficient is used for redundancy determination, and those with an information value less than 0.02 are directly removed. Finally, a refined literacy feature set is generated, containing the key performance feature vectors that are most discriminative and representative of each student in the literacy dimension, as the input for subsequent tensor modeling.

[0144] Step S43: Construct a literacy index collaborative performance tensor model based on the refined literacy feature set, and use the literacy index collaborative performance tensor model to decompose the performance intensity and mutual relationship of the core literacy dimensions to generate a literacy interaction feature matrix;

[0145] In this embodiment, a refined literacy feature set is used as the input to construct a three-dimensional literacy performance tensor model. The tensor dimensions are set as: the number of students × the number of literacy dimensions × the performance intensity level (such as low, medium, and high levels). The CP (CANDECOMP / PARAFAC) tensor decomposition algorithm is used to decompose the tensor, and the decomposition order is set to 5. The dominant performance factors of each literacy dimension in all student groups and their collaborative patterns with other dimensions are extracted. The interaction intensity between literacy dimensions is calculated through the factor loading values in the tensor decomposition matrix, and a literacy interaction feature matrix is constructed. Each element of the matrix represents the joint performance correlation degree between two literacy dimensions in the student group, and the value range is [0,1], where those greater than 0.7 are regarded as strong correlation relationships.

[0146] Step S44: Based on the literacy interaction feature matrix, construct a multi-dimensional relationship graph, identify the hidden correlation structure between core literacy dimensions, and generate a core subject literacy relationship graph;

[0147] In this embodiment, based on the generated literacy interaction feature matrix, a multi-dimensional relationship graph is constructed. The nodes of the graph represent each core literacy dimension, and the edge weights represent the collaborative correlation degree between two literacy dimensions. The graph modeling framework NetworkX is used for modeling to realize the visual management of the interaction relationship. The retention condition of the edge is that the edge weight is greater than 0.3, and the edges with edge weights greater than 0.7 are marked as strongly dependent edges. Further, the Louvain community detection algorithm is used to identify the hidden grouping structure between literacy dimensions. For example, if it is found that "mathematical modeling ability" and "logical reasoning ability" frequently cluster in the same subgraph, it indicates that there is a deep interaction between them in the current student group. The finally output core subject literacy relationship graph is presented in the form of a directed weighted graph, which is used to reveal the essential relationship and development path between the dimensions of literacy.

[0148] Step S45: Perform a structural matrix representation of the core subject literacy relationship graph to obtain a core subject literacy relationship matrix.

[0149] In this embodiment, the core subject literacy relationship graph is transformed into a structural matrix representation form. Assuming the total number of literacy dimensions is N, the relationship matrix is a symmetric or asymmetric matrix of N×N dimensions. The element value of the matrix is the edge weight value between literacy dimensions i and j, that is, the collaborative correlation intensity. If there is no direct connection relationship between literacy dimensions, the corresponding element value is set to 0. For the convenience of subsequent modeling processing, the matrix is normalized, and the Min-Max method is used to uniformly compress all correlation intensities into the [0,1] interval. The finally generated core subject literacy relationship matrix is stored in the format of a two-dimensional array, and a dimension label mapping dictionary for the rows and columns of the matrix is output as a supporting file for the path recommendation algorithm and the graph neural network processing module to call.

[0150] Optionally, step S44 is specifically:

[0151] Step S441: Perform high-dimensional space node mapping on the literacy interaction feature matrix. Set the number of core literacy dimensions to 8, define each core literacy dimension as a graph node, calculate the performance intensity correlation coefficient of the student label in each core literacy dimension as the initial edge weight, and generate an initial literacy node graph;

[0152] In this embodiment, based on the literacy interaction feature matrix, the total number of core literacy dimensions is set to 8, namely: information awareness, logical reasoning, innovation ability, problem-solving, mathematical modeling, practical operation, collaborative communication, and learning transfer. Each dimension is mapped to a node in a graph. Taking students as units, calculate the score correlation coefficient of their scores on any two literacy dimensions, measure it using the Pearson correlation coefficient, and calculate the consistency degree of the collaborative performance between each pair of dimensions in the sample population. The correlation coefficient is used as the edge weight, and the edge weight range is set to [-1, 1]. Set the edge weight of the negative correlation relationship to 0, and retain the positive correlation part to construct an initial literacy node graph. The node graph is represented in an undirected graph structure, with a total of 8 nodes and a maximum number of 28 edges. The edge weight values are used as the basis for graph construction to generate an initial graph structure with the ability to express strong and weak connections.

[0153] Step S442: Set the Gaussian kernel scale factor to 0.5, construct an adjacency matrix for the initial literacy node graph, and perform non-linear transformation on the edge weights between nodes to generate a high-order literacy adjacency matrix;

[0154] In this embodiment, in the already constructed initial literacy node graph, set the Gaussian kernel scale factor σ to 0.5, and perform non-linear mapping transformation on the edge weights using the Gaussian kernel function. The Gaussian function expression is: where w ij is the original edge weight, and the transformed edge weight w i ′ j tends to a non-linear expression with stronger discrimination. This transformation is used to strengthen the discrimination of strong collaborative relationships and at the same time compress the influence of weak collaborations. The generated high-order literacy adjacency matrix is an 8×8 dimensional symmetric matrix, and the main diagonal is set to 0. The sparsity of the adjacency matrix is controlled within 30% to retain sufficient structural information for the subsequent graph learning stage and at the same time enhance the semantic expression ability of the graph structure.

[0155] Step S443: Set the number of coding layers to 2 layers, with the number of neurons in each layer being [64, 32] respectively. Based on the high-order literacy adjacency matrix, perform graph structure encoding to optimize the semantic discrimination of node representations and generate a literacy semantic embedding graph;

[0156] In this embodiment, the constructed high-order adjacency matrix of literacy is input into a preset graph neural network encoding module. Specifically, it is set as follows: the encoding layer has 2 layers, the number of neuron nodes in the first layer is 64, and the second layer is 32. The ReLU activation function is used. The model adopts the GCN (Graph Convolutional Network) architecture, and the training objective is to maximize the semantic distinguishability between nodes, and different functional blocks are separated through the node embedding vectors in the Euclidean space. The training data input is the adjacency matrix and the initial node feature vector (set as the unit vector or the initial principal component vector). The number of training iterations is set to 300 rounds, the Adam optimizer is used, and the learning rate is set to 0.001. Finally, a 32-dimensional semantic embedding vector of each node is output, and a literacy semantic embedding graph is constructed to represent the semantic features possessed by each literacy dimension in the graph structure.

[0157] Step S444: Perform structural community detection on the literacy semantic embedding graph, identify latent associated subgroups, and generate a literacy subgroup atlas;

[0158] In this embodiment, the generated literacy semantic embedding graph is used to perform community structure detection on the node embedding vectors therein, and the Louvain clustering algorithm is used to automatically identify the potential sub-community relationships in the graph structure. The maximum number of communities is set to 4, and the minimum community size is set to 2. The detection strategy is based on maximizing the community modularity based on node semantic similarity. If it is found during the detection that "mathematical modeling", "problem-solving", and "logical reasoning" frequently form a type of sub-community, it is regarded as having an internal association mechanism. Finally, a literacy subgroup atlas is output, which identifies the internal structure of the subgroup in units of communities and retains the node-community mapping relationship, providing a structural basis for the construction of the final relationship matrix. The subgroup atlas can be used to identify which literacy dimensions are more integrative and which have a tendency to differentiate.

[0159] Step S445: Merge the literacy subgroup atlas and reorganize the matrix structure to generate a disciplinary core literacy relationship matrix.

[0160] In this embodiment, the literacy subgroup atlas is subjected to structural merging processing, and a structural matrix is reconstructed by aggregating nodes within the group and strengthening the connection between nodes between groups. The matrix dimension is maintained at 8×8, the edge weights of the internal connections of the subgroup are normalized and enhanced, and the weight increase is set to 1.2 times; the cross-subgroup connections are diluted, and the edge weight reduction is set to 0.8 times. After performing 0-1 normalization processing on all elements in the matrix, the final disciplinary core literacy relationship matrix is output. This matrix retains the community structure characteristics, node semantic embedding characteristics, and the original collaborative relationship, forming a comprehensive expression structure with high semantics, high correlation, and high structural stability. The final output structure will be used as the input layer matrix feature of the learning quality diagnosis model. The disciplinary core literacy relationship matrix can be expressed as: where the diagonal element r ii (such as r 11, r 22 , …, r 88 ) represents the relationship strength between each core literacy dimension and itself, that is, the self - association of this dimension. The non - diagonal element r ij (such as r 12 , r 23 , …, r 78 ) represents the interaction relationships between different core literacy dimensions, and these relationships are reflected through clustering and edge - weight adjustment. During the construction of the matrix, the node relationships within the subgroup are enhanced, resulting in an increase in the relationship strength of the corresponding elements (such as r 12 or r 21 ), while the connections across subgroups are diluted, reducing the relationship strength of the corresponding elements (such as: r 13 or r 31 ). After normalization, all elements in the matrix are adjusted to the range of 0 to 1 to ensure a standardized expression of the results.

[0161] Optionally, step S5 is specifically as follows:

[0162] Step S51: Conduct a joint structure analysis on the subject core literacy relationship matrix and the hierarchical feature dataset to generate a quality deviation feature vector set;

[0163] In this embodiment, the input subject core literacy relationship matrix is an 8×8 - dimensional semantic structure matrix, representing the association strength between core literacies; the hierarchical feature dataset is derived from the structured feature set generated by the two - dimensional mapping of combined labels and performance levels in standard subject learning data, and contains the vector representation of students' hierarchical performance in each literacy dimension. These two sets of data are subjected to a tensor concatenation operation, merged along the feature axis in a tensor splicing manner, unified into the spatial dimension, and dimensionality reduction is performed through principal component analysis (PCA), retaining the first K principal components (such as setting K = 12) with a cumulative contribution rate exceeding 90%, and the output is the quality deviation feature vector set. This vector set takes into account the interaction features between the literacy structure and the hierarchical structure and is used for subsequent deviation clustering and anomaly detection.

[0164] Step S52: Conduct a clustering diagnosis analysis on the quality deviation feature vector set to identify the imbalance points and deviation regions, and generate a learning quality assessment report;

[0165] In this embodiment, for the output quality deviation feature vector set, the density-based DBSCAN clustering algorithm is used to perform abnormal clustering recognition and analysis. Set the clustering radius ε = 0.3 and the minimum number of samples MinPts = 5 to identify the boundary samples and isolated samples of the cluster as potential imbalance points. Calculate the Euclidean distance between each cluster center and other center points, and set the deviation threshold to 1.5 times the average distance to mark the deviation area. At the same time, evaluate the internal consistency of the cluster based on the within-class variance index of each type of sample to determine whether there is a group with significantly inconsistent learning quality. The final output result is presented in the form of a structured evaluation report, including the imbalance dimension, typical deviation point examples, feature deviation direction, and spatial distribution map, to construct a visual learning quality evaluation report for the path optimization model to call.

[0166] Step S53: Extract the teaching resource features and subject objective requirement features from the standardized subject learning data;

[0167] In this embodiment, hierarchical semantic features are extracted from the question fields in the standardized subject learning data. Teaching resource features are extracted based on attributes such as question type, knowledge point label, answering time, and correct rate, including the question coverage range, skill directivity, and resource difficulty level, to construct a teaching resource feature vector set. Secondly, target ability labels and hierarchical structures are extracted according to curriculum standard requirements or subject objective libraries. For example, the targets are divided into three levels: knowledge mastery, ability transfer, and comprehensive application. Each level contains several dimensional target items, and a target requirement feature vector set is constructed in one-hot form. The final output feature data is in the form of a two-dimensional matrix, with rows representing each resource / target sample and columns representing different feature dimensions (not less than 20 dimensions), providing a data basis for the structural input part of the path recommendation model.

[0168] Step S54: Construct a path weight recommendation model based on the learning quality evaluation report, teaching resource features, and subject objective requirement features;

[0169] In this embodiment, the learning quality evaluation report is used as the constraint input, and the teaching resource features and subject objective requirement features are used as semantic inputs. A weighted multi-objective path modeling strategy is adopted to construct a path weight recommendation model. The total score of each path is set as the linear weighted sum of three parts: resource fitness R, target matching degree G, and learning deviation correction potential P. The model form is: Scorepath = α × R + β × G + γ × P; where α, β, and γ are set to [0.3, 0.4, 0.3] and are optimized through 10-fold cross-validation. The path search uses the shortest path method, with the knowledge point mastery graph as the path graph and the path score as the evaluation function to output the optimal path set and the alternative path set (top-k paths, k = 5). This recommendation model provides a matching score for each student group portrait for subsequent path optimization and intervention suggestion generation.

[0170] Step S55: Analyze the student group portrait based on the standardized subject learning data to obtain the regional student group portrait data, and use the path weight recommendation model to perform path matching and optimization on the regional student group portrait data to obtain the subject optimization strategy.

[0171] In this embodiment, the standardized subject learning data is aggregated to the student level, and feature dimensions such as age, gender, region, score, behavior frequency, and task completion time are extracted. K-means++ is used for clustering analysis, the number of clustering centers is set to 8, and the regional student groups are divided according to the Euclidean distance minimization strategy. Each group generates corresponding group portrait labels, including the main weak dimension, the literacy tendency structure, and the learning stability score. These group portraits are input into the path weight recommendation model constructed in step S54, and path matching is performed for each group. During the matching process, the resource distribution density and the target level complexity are considered, and the matching path and improvement suggestions for each group are output. Finally, a subject optimization strategy is generated, including: key adjustment of the literacy dimension, priority recommendation of resource types, and the target achievement path map, providing a structured strategy reference for regional teaching decision-making.

[0172] Optionally, step S52 is specifically:

[0173] Step S521: Perform principal component analysis and feature orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum number of dominant dimensions to 6 to extract the dominant deviation dimension factor, generating the deviation main cause tensor data;

[0174] In this embodiment, the quality deviation feature vector set (with a shape of [N×M], where N is the number of student samples and M is the number of deviation features) is standardized so that the mean of each dimension is 0 and the standard deviation is 1. The principal component analysis (PCA) method is used to reduce the dimension of the standardized data, and the variance contribution rate threshold is set to 85%, that is, the first K principal components are selected so that the cumulative explained variance is not less than 85%. To control the model complexity, the maximum number of dominant dimensions K≤6 is limited, so as to extract the dominant deviation dimension factor in the generated K-dimensional orthogonal feature space. These dominant dimensions are used as tensor axes, and the deviation main cause tensor data is generated by combining the sample index dimension and the dimension factor dimension. The tensor shape is [N×K], which is used for subsequent density modeling and clustering analysis.

[0175] Step S522: Perform density peak estimation on the deviation main cause tensor data, set the minimum number of samples to 10, and the clustering distance threshold to 0.75 to identify potential deviation clustering centers, generating the subject imbalance clustering map;

[0176] In this embodiment, the obtained deviation main cause tensor data is used as input samples, and the method based on Density Peak Clustering (DPC) is adopted to identify potential clustering centers. The minimum number of samples is set to 10 as the minimum unit for local density evaluation, and the clustering distance threshold is set to 0.75 to judge the minimum distance condition between clustering candidate centers. The specific operations include: calculating the Euclidean distance matrix between all samples, estimating the local density ρ value of each sample point based on the kernel density estimation method, and at the same time combining the minimum distance δ value to judge high-density isolated points. Finally, the points that simultaneously satisfy high ρ and high δ are selected as the deviation clustering centers, and the output is a two-dimensional space coordinate map, called the subject imbalance clustering map, which reflects the abnormal clustering trend of the main cause dimension in the sample space.

[0177] Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the subject imbalance clustering map, and set the number of neighborhoods to 20 to measure the anomaly intensity of each clustering unit in the subject imbalance clustering map, identify the deviation area and generate the local imbalance anomaly distribution map;

[0178] In this embodiment, after obtaining the subject imbalance clustering map, local structure refinement is performed on its boundary region. Set the reorganization boundary spacing to 0.05, and segment and slice the continuously distributed points in the map according to this spacing to form a distribution grid. Use the sample density within the grid to construct the boundary reorganization curve, and smooth the boundary of each clustering unit through probability density function fitting. Subsequently, set the number of neighborhoods to 20, and use the K-nearest neighbor algorithm to evaluate the local anomaly intensity of each point in the map, which is defined as the density difference ratio between the point and the sample with the minimum density in its neighborhood. Samples higher than the set threshold are identified as local deviation points. The final output result is the local imbalance anomaly distribution map, marking all the spatial regions determined to be density anomalies and the corresponding main cause dimension indexes.

[0179] Step S524: Combine the local imbalance anomaly distribution map with the deviation main cause tensor data, set the deviation factor weight threshold to 0.6 to calculate the deviation factors of each literacy dimension, and generate the imbalance characteristic weight data set;

[0180] In this embodiment, based on the local imbalance anomaly distribution map, the feature traceability is performed for each identified abnormal clustering unit, and traced back to the dimension factor influence weight in the deviation main cause tensor data. To quantify the deviation degree of each literacy dimension in the anomaly, the weight normalization index method is used to construct a deviation factor calculation model. The deviation factor weight threshold is set to 0.6. If the ratio of the average deviation value of a certain literacy dimension in the abnormal area to the total deviation value exceeds this threshold, it is regarded as the main imbalance factor. Finally, the imbalance feature weight data set is output, and each record contains 5 fields: sample index, imbalance dimension code, deviation direction (positive / negative), deviation intensity (real value), and standardized deviation level (high / medium / low), which are used for the structured modeling of learning quality diagnosis.

[0181] Step S525: Perform semantic reconstruction and structural induction on the imbalance feature weight data set, and output a learning quality assessment report.

[0182] In this embodiment, semantic reconstruction is performed on the imbalance feature weight data set, that is, the numerical deviation information is converted into descriptive labels easy to understand in the education field through label mapping, such as "leapfrog ability deviation" and "application layer literacy disadvantage". Secondly, a structured pattern is aggregated according to the imbalance dimension, for example, "deviation from the literacy combination centered on collaborative ability" and "the main imbalance group is concentrated in the interdisciplinary transfer dimension". Stratified induction is carried out in combination with context features such as region and group portrait, and finally output in the form of a report. The report includes three parts: 1) Summary table of the main causes of deviation; 2) Heat map of the imbalance clustering distribution; 3) Optimization suggestion cards based on the literacy dimension (no less than 2 suggestions for each dimension). This report constitutes the complete learning quality diagnosis result and provides a decision-making basis for subsequent path matching and resource intervention.

[0183] Especially importantly, step S54 is specifically:

[0184] Step S541: Perform structured analysis on the learning quality assessment report, extract the deviation direction, deviation amplitude and clustering imbalance label of each literacy dimension, combine the content difficulty, presentation method and adaptation type in the teaching resource characteristics, construct a literacy deviation regulation demand matrix, and generate a regulation demand feature set;

[0185] In this embodiment, field extraction is performed on the structured results of the learning quality assessment report, and keywords such as "literacy dimension name", "deviation direction (positive / negative)", "deviation amplitude (normalized Z-value)", "cluster label encoding (such as C01, C02)" are identified and extracted using a regularization template. And this information is used as the structural input. Combining the "resource content difficulty score (0-1)", "resource presentation mode label (such as text, diagram, interactive)", and "adaptation type identifier (individual / group / region)" in the teaching resource feature data, a three-dimensional matrix is constructed through the dimension alignment method. The matrix axes are the literacy dimension, the resource matching type, and the deviation amplitude level respectively, forming a literacy deviation regulation demand matrix. After construction, the non-zero value items in the matrix are extracted and weighted and sorted according to the deviation amplitude to form a regulation demand feature set. Example: When the deviation of the "interdisciplinary transfer" literacy dimension is negative and the amplitude is 0.82, and a resource of the diagram type with a content difficulty of 0.6 is matched, the feature item ["interdisciplinary transfer", "negative", 0.82, "diagram", 0.6, "group"] is constructed.

[0186] Step S542: Perform semantic vector encoding on the subject target requirement features, extract the ability level, knowledge dimension, and task-oriented label of each literacy target, and perform semantic correlation analysis with the regulation demand feature set to generate a literacy target matching relationship map;

[0187] In this embodiment, for the subject target requirement document, the BERT model is used to perform embedding encoding on the text description, and three types of semantic labels are extracted: ability level label (such as "application", "comprehensive", "creation"), knowledge dimension (such as "conceptual", "procedural", "implicit transfer"), and task-oriented label (such as "problem solving", "cooperative inquiry"). After encoding, a target semantic vector space is constructed. The regulation demand feature set formed in step S541 is also encoded as a regulation vector, and based on the cosine similarity method, the semantic matching threshold is set to 0.75, and the elements in the two vector spaces are paired and analyzed. A semantic matching record table is established by constructing a triple <literacy dimension, matching target label, similarity>. Based on the triple set, a literacy target matching relationship map is further generated, where the nodes are the literacy dimension and the task label, the edges are the semantic similarity relationships, and the edge weights are the similarity scores, which are used to support the subsequent path map construction.

[0188] Step S543: According to the regulation demand feature set and the literacy target matching relationship map, construct a literacy regulation path map, use each literacy dimension node in the path as a graph node, and use the regulation order between the paths as a directed edge relationship to form a literacy regulation path network;

[0189] In this embodiment, based on the literacy objective matching relationship graph generated in the previous step, a regulation path graph is constructed. The literacy dimensions in the regulation requirement feature set are set as the nodes in the graph, and the pair of targets with the highest semantic matching degree is extracted as the end point of the regulation path target. The edges of the directed graph are constructed, and the edge direction is defined as "from the literacy node to be regulated to the target literacy task node". The regulation order is determined by the deviation amplitude, and they are connected in sequence from high deviation to low deviation. The "path association density" is used as the condition for constructing the relay edge. If two literacy dimensions have a common target orientation and their deviation directions are the same, a relay edge is automatically generated and connected to form a "relay regulation path". Finally, a literacy regulation path network is formed, and the average degree of nodes is controlled within 2.3 to avoid dense stacking of paths.

[0190] Step S544: Evaluate the importance of paths and weight the structure of the literacy regulation path network. Calculate the weights of literacy nodes based on the deviation amplitude, node centrality, and resource matching degree of the literacy regulation path network. At the same time, calculate the literacy conversion probability to generate a set of path weight vectors.

[0191] In this embodiment, based on the generated regulation path network, the path weights and conversion probabilities need to be calculated. The deviation amplitude weight factor is set as α = 0.5, the node centrality weight factor is β = 0.3, and the resource matching degree factor is γ = 0.2. Combining the above three indicators, calculate the weight score Wi of each literacy node as Wi = αZi + βCi + γRi, where Zi is the standardized deviation amplitude, Ci is the PageRank centrality of the node in the regulation network, and Ri is the average matching degree of the resources of this dimension. To model the literacy conversion potential in the regulation path, use the path Markov jump probability modeling method to construct a conversion probability matrix through the historical learning case path, and define the conversion probability value between each pair of nodes. The final output result is a set of path weight vectors, and each record contains [starting node, ending node, node weight, edge conversion probability].

[0192] Step S545: Embed the set of path weight vectors into the regulation path network structure, perform probabilistic regulation sorting and diversity recombination, construct a literacy regulation path library, and use the regulation effect score and path confidence of the literacy regulation path library as the path recommendation evaluation criteria to generate a path weight recommendation model.

[0193] In this embodiment, the path weight vector set is re-embedded into the original regulation path network, and a literacy regulation path library is generated through reordering and structural optimization. The Top-K probabilistic regulation sorting algorithm is adopted, the number of candidate paths is set to K=5, and the path score of each path is calculated as ∑(Wi×Pij), where Wi is the node weight and Pij is the conversion probability of the edge in the path. The path confidence threshold is set to 0.7, low-confidence paths are eliminated, and path diversity is reorganized based on the maximum structural diversity criterion (path node overlap <50%). Each path is accompanied by evaluation indicators including regulation coverage, resource adaptability and path expected improvement value, and finally a literacy regulation path library is constructed. Based on the dual indicators of path score and confidence, the recommendation evaluation function is defined: R(path)=αScore+βConfidence, and the paths with the top three evaluation scores are recommended as the final output to form a path weight recommendation model for the teaching intervention system to call. The path weight recommendation model is built on the graph attention network (GAT), with core literacy dimensions as graph nodes and regulation paths as directed edges, integrating features such as student deviation amplitude, resource adaptability and conversion probability for graph structure encoding and semantic embedding. The model dynamically learns the importance of each node in the path through the attention mechanism, and combines the path conversion probability and structural centrality for comprehensive score ranking, and finally generates a literacy regulation path recommendation result that takes into account both diversity and adaptability.

[0194] Optionally, the present specification also provides a system for analyzing the composition of balanced combinations within a subject, which is used to execute the above-mentioned method for analyzing the composition of balanced combinations within a subject. The system for analyzing the composition of balanced combinations within a subject includes:

[0195] The data collection module is used to obtain subject learning data and perform data standardization on the subject learning data to obtain standardized subject learning data;

[0196] A combination interval generation module is used to divide the combination interval based on the standardized subject learning data to obtain subject combination interval data; perform interval structured modeling on the subject combination interval data to obtain a combination interval model;

[0197] The combined interval analysis module is used to compare the distribution of subject-level scores according to the combined interval model to obtain subject-level score deviation data; based on the subject-level score deviation data, the subject learning quality difference is evaluated to obtain a hierarchical feature data set;

[0198] The core literacy relationship modeling module is used to model the subject core literacy relationship based on standardized subject learning data to obtain the subject core literacy relationship matrix;

[0199] A quality diagnosis module is used to evaluate the quality of subject learning according to the subject core literacy relationship matrix and the hierarchical feature dataset, obtain a subject learning quality evaluation report, and optimize the learning path of standardized subject learning data based on the subject learning quality evaluation report to obtain a subject optimization strategy.

[0200] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application document are intended to be encompassed within the present invention.

[0201] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A method for analyzing the constitutiveness of an equilibrium combination interval within a discipline, characterized in that, Including the following steps: Step S1: Obtain subject learning data, and perform data standardization processing on the subject learning data to obtain standardized subject learning data; Step S2: Based on the standardized subject learning data, perform combined interval division to obtain subject combined interval data; perform interval structured modeling on the subject combined interval data to obtain a combined interval model; Step S3: According to the combined interval model, conduct a comparison of the subject hierarchical score distribution to obtain subject hierarchical score deviation data; based on the subject hierarchical score deviation data, conduct an assessment of the subject learning quality difference to obtain a hierarchical feature data set; Step S4: Based on the standardized subject learning data, perform modeling of the relationship between subject core qualities to obtain a subject core quality relationship matrix; Step S5: According to the subject core quality relationship matrix and the hierarchical feature data set, conduct an assessment of the subject learning quality to obtain a subject learning quality assessment report, and based on the subject learning quality assessment report, optimize the learning path of the standardized subject learning data to obtain a subject optimization strategy.

2. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 1, wherein Specifically, step S1 is as follows: Step S11: Collect multi-source original subject learning data through the regional teaching management platform, and perform unified coding on the data framework structure of the multi-source original subject learning data to obtain subject learning data; Step S12: Detect missing values in the subject learning data, and fill in the missing values with the category mode based on the detection results to obtain cleaned data; Step S13: Normalize the fields of the cleaned data, and standardize the encoding of the categorical fields in the normalization results to generate transformed data; Step S14: Perform scale standardization processing on the transformed data to obtain standardized subject learning data.

3. The method for analyzing the composition of the balanced combination interval within a discipline according to claim 1, wherein Specifically, the combined interval division described in step S2 is as follows: Extract the student score characteristics from the standardized subject learning data to obtain student subject score data, and conduct statistical analysis on the student subject score data to obtain standardized total score data and student score proportion data; Based on the student score proportion data, conduct a division of the proportion of low-scoring students to obtain low-scoring student proportion data; Based on the standardized total score data and the low-scoring student proportion data, set combined variables, where the horizontal axis variable is set as the standardized total score data, and the vertical axis variable is set as the low-scoring student proportion data, to construct a two-dimensional interactive coordinate system; Extract the standardized distribution characteristics of the spatial combination analysis base, and set the mean ± 0.67σ in the standardized distribution characteristics as the division threshold to divide the two-dimensional interactive coordinate system. Divide the horizontal axis and the vertical axis of the two-dimensional interactive coordinate system into three sections respectively to form a two-dimensional combined space label matrix; Map the standardized subject learning data to the corresponding combination labels in the two-dimensional combined space label matrix according to the two-dimensional interactive coordinate system, mark the interval categories, and count the number of samples and the proportion in each interval to generate subject combined interval data.

4. The method for analyzing the constitutive analysis of the intra-disciplinary balanced combination interval according to claim 1, wherein Specifically, the comparison of the subject hierarchical score distribution described in step S3 is as follows: Perform spatial label aggregation processing on the interval combination model, classify the sample data under the same interval label hierarchically, and combine with the subject combined interval data to generate combined label student distribution data; Set the interval boundary thresholds for the low, medium, and high stratification levels according to the distribution of the standardized total score data, and perform interval stratification on the combined label student distribution data to generate a hierarchical distribution label matrix; Perform double-label mapping on the standardized subject learning data according to the combined label student distribution data and the hierarchical distribution label matrix to generate a multi-label hierarchical score dataset, and perform normalized statistics on the number of samples at each level under different combined labels in the multi-label hierarchical score dataset to obtain a hierarchical score frequency matrix; Perform differential metric calculation based on the hierarchical score frequency matrix to evaluate the deviation of the hierarchical score distribution between different combined labels, and obtain the subject hierarchical score deviation data.

5. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 1, wherein The evaluation of the subject learning quality difference described in step S3 is specifically: Perform distribution statistics and central tendency analysis on the subject hierarchical score deviation data, calculate the mean, range, and standard deviation of each level corresponding to each combined label, and generate a deviation distribution statistical matrix; Perform horizontal regional balance analysis based on the deviation distribution statistical matrix to evaluate the performance consistency of different regions in the scores of each level, and generate regional balance characteristic data; Perform vertical hierarchical gradient analysis on the deviation distribution statistical matrix, calculate the continuity of score improvement and the probability of hierarchical transition, and generate a hierarchical gradient change model; Fuse the regional balance characteristic data and the hierarchical gradient change model for weighted comprehensive evaluation to obtain the subject learning quality difference evaluation data; Integrate the student hierarchical characteristics based on the subject learning quality difference evaluation data and the multi-label hierarchical score dataset to obtain a hierarchical characteristic dataset.

6. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 1, characterized in that Step S4 is specifically: Step S41: Identify the subject element labels of the standardized subject learning data, extract the student scores, question difficulties, practice scores, and core literacy labels, construct an element label vector set, and thus generate a literacy structured dataset; Step S42: Perform dimensionality reduction and feature redundancy removal on the literacy structured dataset, extract highly correlated literacy association features, and generate a refined literacy feature set; Step S43: Construct a literacy index co-performance tensor model according to the refined literacy feature set, and use the literacy index co-performance tensor model to decompose the performance intensity and mutual relationship of the core literacy dimensions to generate a literacy interaction feature matrix; Step S44: Construct a multi-dimensional relationship graph based on the literacy interaction feature matrix, identify the hidden association structure between each core literacy dimension, and generate a subject core literacy relationship graph; Step S45: Perform structural matrix expression on the subject core literacy relationship graph to obtain a subject core literacy relationship matrix.

7. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 6, characterized in that Step S44 is specifically: Step S441: Perform high-dimensional space node mapping on the literacy interaction feature matrix, set the number of core literacy dimensions to 8, define each core literacy dimension as a graph node, and calculate the performance intensity co-efficient of the student label in each core literacy dimension as the initial edge weight to generate an initial literacy node graph; Step S442: Set the Gaussian kernel scale factor to 0.5, construct an adjacency matrix for the initial literacy node graph, and perform non-linear transformation on the edge weights between nodes to generate a literacy high-order adjacency matrix; Step S443: Set the number of encoding layers to 2, with the number of neurons in each layer being [64, 32] respectively, perform graph structure encoding based on the literacy high-order adjacency matrix, optimize the semantic distinction of node representation, and generate a literacy semantic embedding graph; Step S444: Perform structural community detection on the literacy semantic embedding graph, identify invisible related subgroups, and generate a literacy subgroup graph; Step S445: Merge the literacy subgroup maps and reorganize the matrix structure to generate a subject core literacy relationship matrix.

8. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 1, characterized in that Step S5 is specifically as follows: Step S51: Perform joint structural analysis on the subject core literacy relationship matrix and the hierarchical feature data set to generate a quality deviation feature vector set; Step S52: performing cluster diagnosis analysis on the quality deviation feature vector set, identifying imbalance points and deviation areas, and generating a learning quality assessment report; Step S53: extracting teaching resource features and subject target requirement features from the standardized subject learning data; Step S54: constructing a path weight recommendation model based on the learning quality assessment report, teaching resource characteristics and subject goal requirement characteristics; Step S55: Perform student group portrait analysis based on standardized subject learning data to obtain regional student group portrait data, and use the path weight recommendation model to perform path matching and optimization on the regional student group portrait data to obtain a subject optimization strategy.

9. The method for analyzing the constitutive analysis of the balanced combination interval within a discipline according to claim 8, wherein Step S52 is specifically as follows: Step S521: Perform principal component analysis and feature orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum number of dominant dimensions to 6 to extract the dominant deviation dimension factors, and generate the deviation main cause tensor data; Step S522: Estimating the density peak of the main cause tensor data of the deviation, setting the minimum sample number to 10 and the cluster distance threshold to 0.75 to identify the potential deviation cluster center, and generating a discipline imbalance cluster map; Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the discipline imbalance cluster map, and set the number of neighborhoods to 20 to measure the abnormal intensity of each cluster unit in the discipline imbalance cluster map, identify the deviation area and generate a local imbalance abnormal distribution map; Step S524: combining the local imbalance anomaly distribution map with the deviation main cause tensor data, setting the deviation factor weight threshold to 0.6 to calculate the deviation factor of each literacy dimension, and generating an imbalance feature weight data set; Step S525: Perform semantic reconstruction and structural induction on the imbalanced feature weight data set, and output a learning quality assessment report.

10. A system for analyzing the composition of an equilibrium combination interval within a discipline, characterized in that, For executing the method for analyzing the composition of balanced combinations within a subject as claimed in claim 1, the composition analysis system for analyzing the composition of balanced combinations within a subject comprises: The data collection module is used to obtain subject learning data and perform data standardization on the subject learning data to obtain standardized subject learning data; A combination interval generation module is used to divide the combination interval based on the standardized subject learning data to obtain subject combination interval data; perform interval structured modeling on the subject combination interval data to obtain a combination interval model; The combined interval analysis module is used to compare the grade distributions at the subject level according to the combined interval model, and obtain the subject-level grade deviation data; based on the subject-level grade deviation data, it conducts an assessment of the differences in subject learning quality and obtains a hierarchical feature data set; The core literacy relationship modeling module is used to model the relationship between subject core literacies based on the standardized subject learning data, and obtain the subject core literacy relationship matrix; The quality diagnosis module is used to evaluate the subject learning quality according to the subject core literacy relationship matrix and the hierarchical feature data set, obtain the subject learning quality assessment report, and optimize the learning path of the standardized subject learning data based on the subject learning quality assessment report to obtain the subject optimization strategy.

Citation Information

Patent Citations

  • College entrance examination subject selection model and establishment method thereof

    CN111861106A

  • Student quality and quality prediction system based on reinforcement learning

    CN113962444A

  • System and method for measuring subject core accomplishment based on production technology

    CN115660469A

  • Multi-task discipline core accomplishment level mining method and system

    CN118410079A

  • Education quality evaluation method based on machine learning model

    CN119228599A

Cited By

  • Application type college talent cultivation quality tracking service system based on GIS

    CN120655173A

  • A GIS-based application-oriented university talent training quality tracking service system

    CN120655173B

  • Teaching quality evaluation method and system based on big data

    CN121981618A

  • Big data-based teaching quality evaluation method and system

    CN121981618B