A method and system for constructing a balanced combination interval within a subject
By standardizing and combining interval analysis of subject learning data, a subject core competency relationship matrix is constructed, which solves the problem that traditional assessment methods cannot deeply analyze the differences in learning quality within a subject, and achieves accurate assessment and optimization of learning quality within a subject.
Patent Information
- Application Number
- CN202510444678.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Existing technologies in the field of education cannot deeply analyze the differences in learning quality among the various components of a subject, especially the relationship between different achievement levels, making it difficult for education administrators to provide targeted and effective improvement measures.
By acquiring subject learning data, standardizing the data, dividing it into combination intervals, constructing a combination interval model, comparing the distribution of subject-level scores, generating a hierarchical feature dataset, establishing a subject core competency relationship matrix, generating a subject learning quality assessment report, and optimizing the learning path.
It enables an in-depth understanding of the differences in learning quality at different levels within a subject, helps identify weaknesses, provides precise suggestions for teaching and resource allocation, and improves the relevance and efficiency of teaching effectiveness.
Smart Images

Figure CN120373945B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational monitoring technology, and in particular to a method and system for analyzing the composition of balanced combination intervals within a subject. Background Technology
[0002] Currently, the assessment of subject learning quality in the education field generally relies on single-dimensional statistical methods, such as calculating the mean and ranking of student test scores. This traditional approach primarily focuses on statistically analyzing the overall performance of different regions, schools, or student groups. While it can provide a general distribution of scores, it lacks in-depth analysis of the internal structure of subjects and the relationships between different learning levels. Existing technologies mainly calculate the mean using score data, failing to reveal the differences in learning quality among the various components of a subject, especially the relationships between different performance levels.
[0003] Traditional assessment methods often rely on single test score indicators, neglecting structural analysis within a subject. Subject learning quality includes not only overall grades but also mastery of knowledge points, achievement of core competencies, and students' performance on questions of varying difficulty. Existing methods fail to reveal the intrinsic connections between different parts of a subject or understand the balance of content assessed, making it difficult for education administrators to comprehensively understand subject quality issues and consequently affecting the targetedness and effectiveness of decision-making.
[0004] Furthermore, traditional methods have failed to effectively achieve comprehensive analysis of all parts within a subject, especially in multi-level and multi-dimensional quality diagnosis. Identifying the relationships between different parts of a subject and providing data support for improvement remains a pressing issue. Therefore, existing technologies cannot provide education policymakers with accurate subject analysis, nor can they propose effective improvement measures based on different learning levels and student group characteristics. Summary of the Invention
[0005] Therefore, it is necessary for the present invention to provide a method and system for structural analysis of balanced combination intervals within a discipline, in order to solve at least one of the above-mentioned technical problems.
[0006] To achieve the above objectives, a method for constitutive analysis of equilibrium combination intervals within a discipline is proposed, comprising the following steps:
[0007] Step S1: Obtain subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data;
[0008] Step S2: Divide the subject combination intervals based on standardized subject learning data to obtain subject combination interval data; perform interval structure modeling on the subject combination interval data to obtain the combination interval model;
[0009] Step S3: Compare the subject-level performance distribution based on the combined interval model to obtain subject-level performance deviation data; evaluate the differences in subject learning quality based on the subject-level performance deviation data to obtain a hierarchical feature dataset.
[0010] Step S4: Model the relationship between core competencies in a subject based on standardized subject learning data to obtain a matrix of core competencies in a subject.
[0011] Step S5: Conduct subject learning quality assessment based on the subject core competency relationship matrix and hierarchical feature dataset to obtain a subject learning quality assessment report, and optimize the learning path of standardized subject learning data based on the subject learning quality assessment report to obtain subject optimization strategies.
[0012] This invention, through the standardization of subject learning data, ensures that all types of data remain within a uniform scale, thereby avoiding inaccurate assessments caused by differences in data units and making subsequent analyses more comparable and scientific. The standardized data serves as a valid foundation in subsequent steps, providing high-quality input for subsequent subject-level performance distribution analysis and modeling of the relationship between core subject competencies. When dividing subjects into combined intervals, using combined intervals effectively stratifies student performance, and generating combined interval models through interval structured modeling provides a detailed zoning perspective for subsequent subject-level performance distribution analysis. Analyzing student performance within each interval reveals in-depth differences in learning quality at different levels within a subject, helping education administrators identify uneven performance distribution and providing a basis for targeted teaching and resource allocation. In interval structured modeling, the selection of combined interval parameters helps to accurately divide different level intervals, thus precisely capturing subject performance differences within each performance level during analysis and avoiding omissions of details. By generating subject-level performance deviation data based on a combined interval model, and further using a hierarchical feature dataset for assessing subject learning quality differences, the differences in learning quality across different parts of a subject can be effectively reflected. Hierarchical analysis reveals the differences between different dimensions within a subject, helping to identify weak areas requiring optimization, rather than simply assessing based on total scores. Furthermore, the generation of the hierarchical feature dataset helps to precisely pinpoint groups or subject content requiring focused support from educational resources, enabling refined adjustments during teaching improvement. The construction of a subject-specific core competency relationship matrix utilizes the interaction characteristics between core competencies within a subject. By analyzing the relationships between various dimensions within a subject, it reveals the interaction between subject knowledge points, core competencies, and student performance. This step, by constructing a matrix-style core competency relationship diagram, helps identify the structural connections between different parts of a subject, particularly the interaction relationships between different competency dimensions. This allows assessment to go beyond a single score, comprehensively evaluating student learning outcomes from a broader perspective of knowledge acquisition and ability development. In the process of generating subject learning quality assessment reports, the comprehensive use of core competency relationship matrices and hierarchical feature datasets can provide education decision-makers with precise subject analysis, helping to identify which aspects within a subject need optimization and improvement. Through this step, education administrators can optimize personalized and hierarchical teaching pathways based on specific subject quality assessment reports, thereby improving teaching effectiveness in a targeted manner.
[0013] Optionally, step S1 specifically includes:
[0014] Step S11: Collect multi-source subject learning raw data through the regional teaching management platform, and uniformly encode the multi-source subject learning raw data into a data framework structure to obtain subject learning data;
[0015] Step S12: Perform missing value detection on the subject learning data, and impute the missing value detection results using the class mode to obtain cleaned data;
[0016] Step S13: Perform field normalization on the cleaned data, and standardize the normalization results by encoding categorical fields to generate transformed data;
[0017] Step S14: Perform scale standardization on the transformed data to obtain standardized subject learning data.
[0018] This invention solves the problem of inconsistent data formats from different sources by unifying the structure of raw learning data from multiple disciplines, providing a consistent data framework for subsequent analysis. Imputing missing values using the mode of categories maintains the representativeness of the data distribution without introducing outliers, making it suitable for most educational data processing scenarios. Field normalization and categorical field encoding standardization ensure a unified scale of expression for numerical and categorical fields, facilitating effective feature learning by the model. Scale standardization further maps data to the same range, reducing the influence of dimensions and improving the accuracy of clustering and model analysis. The overall process improves data quality and consistency of expression, laying a solid foundation for subsequent modeling and evaluation.
[0019] Optionally, the combined interval division described in step S2 is specifically as follows:
[0020] Student performance characteristics are extracted from standardized subject learning data to obtain student subject performance data. Statistical analysis is then performed on the student subject performance data to obtain standardized total performance data and student performance percentage data.
[0021] Based on the percentage of students with low grades, the percentage of students with low grades is divided to obtain the percentage of students with low grades.
[0022] Based on standardized total score data and low score student percentage data, a combination variable is set, where the horizontal axis variable is set to standardized total score data and the vertical axis variable is set to low score student percentage data, and a two-dimensional interactive coordinate system is constructed.
[0023] The standardized distribution features of the spatial combination analysis basis are extracted, and the mean ±0.67σ in the standardized distribution features is set as the dividing threshold to divide the two-dimensional interactive coordinate system. The horizontal and vertical axes of the two-dimensional interactive coordinate system are divided into three segments to form a two-dimensional combination spatial label matrix.
[0024] Based on the two-dimensional interactive coordinate system, standardized subject learning data is mapped to the corresponding combination labels in the two-dimensional combination space label matrix, the interval categories are marked, and the number and proportion of samples in each interval are counted to generate subject combination interval data.
[0025] This invention constructs a two-dimensional interactive coordinate system by extracting student performance characteristics and combining them with the proportion of students with low scores. This effectively achieves a combined horizontal and vertical assessment of subject quality, overcoming the limitations of traditional analysis methods that rely solely on total scores. Setting the standardized total score as the horizontal axis and the proportion of students with low scores as the vertical axis reveals the synergistic relationship between overall performance and the distribution of weaker students. The threshold value used is the mean ± 0.67σ, which conforms to the 68% interval characteristic of a normal distribution, ensuring statistical robustness and discriminative power in interval division, facilitating the identification of marginal and extreme groups. Mapping sample data under this coordinate system forms a combined spatial label matrix, which can intuitively identify learning quality risk areas and strength areas, providing structured input for subsequent models and enhancing the accuracy and diagnostic depth of data stratification analysis.
[0026] Optionally, the subject-level score distribution comparison mentioned in step S3 is specifically as follows:
[0027] Spatial label aggregation is performed on the interval combination model, and the sample data under the same interval label are hierarchically classified. Combined with the subject combination interval data, student distribution data with combination labels is generated.
[0028] Based on the distribution of standardized total score data, set interval boundary thresholds for low, medium, and high strata, and perform interval stratification on the combined label student distribution data to generate a hierarchical distribution label matrix.
[0029] Standardized subject learning data is mapped to student distribution data with combined labels and hierarchical distribution label matrix using dual labels to generate multi-label hierarchical score dataset. The number of samples at each level under different combined labels in the multi-label hierarchical score dataset is normalized and statistically analyzed to obtain hierarchical score frequency matrix.
[0030] The difference measurement is calculated based on the hierarchical score frequency matrix to evaluate the hierarchical score distribution deviation between different combination labels and obtain subject hierarchical score deviation data.
[0031] This invention introduces a dual-label mapping mechanism to jointly categorize students based on combined labels and performance levels, enhancing the hierarchy and interpretability of the data structure. When setting performance levels, boundary thresholds for low, medium, and high strata are defined based on the standardized total score distribution, ensuring that the stratification is statistically sound and balanced, which helps reveal the distribution patterns of students at different levels under each combined label. By normalizing the multi-label stratified performance dataset and constructing a stratified score frequency matrix, biases caused by sample size differences can be effectively removed, improving the fairness of comparisons. Finally, the calculation of distribution bias reveals the specific differences in performance across different groups at each level, compensating for the shortcomings of traditional methods in characterizing the distribution structure of different learning levels and improving the precision of subject quality diagnosis.
[0032] Optionally, the subject learning quality difference assessment described in step S3 specifically includes:
[0033] The distribution statistics and central tendency analysis of the subject-level performance deviation data are performed, and the mean, range and standard deviation of each combination label corresponding to the level are calculated to generate a deviation distribution statistical matrix.
[0034] Based on the deviation distribution statistical matrix, a horizontal regional balance analysis is conducted to assess the consistency of performance across different regions at each level and generate regional balance characteristic data.
[0035] A vertical hierarchical gradient analysis is performed on the deviation distribution statistical matrix to calculate the continuity of performance improvement and the probability of hierarchical transition, thereby generating a hierarchical gradient change model.
[0036] By integrating regional equilibrium characteristic data and hierarchical gradient change models to conduct a weighted comprehensive evaluation, data on the differences in subject learning quality are obtained.
[0037] Based on the subject learning quality difference assessment data and the multi-label hierarchical score dataset, student hierarchical features are integrated to obtain a hierarchical feature dataset.
[0038] This invention characterizes the central tendency and dispersion of different label combinations at various levels through statistical parameters such as mean, range, and standard deviation in the deviation distribution statistical matrix. This helps identify performance stability and extreme values, thus overcoming the deficiency of traditional mean analysis in neglecting dispersion characteristics. Regional balance analysis improves the fairness of subject assessment and the visibility of regional differences by comparing the consistency of performance across different regions at different levels. The hierarchical gradient analysis sets a transition probability index to measure the potential trend of students developing from lower to higher levels, providing a basis for dynamically assessing learning growth paths. The fusion analysis uses a weighted approach to integrate regional balance and gradient models, ensuring that the quality difference assessment results have both spatial coverage and longitudinal developmental characteristics. The resulting hierarchical feature dataset provides high-granularity, multi-dimensional data support for subsequent core competency modeling and path optimization.
[0039] Optionally, step S4 specifically includes:
[0040] Step S41: Perform subject element label recognition on standardized subject learning data, extract student scores, question difficulty, practice scores and core competency labels, construct element label vector set, and thus generate a competency structured dataset;
[0041] Step S42: Perform dimensionality reduction and feature redundancy removal on the structured dataset of literacy, extract highly relevant literacy association features, and generate a concise literacy feature set;
[0042] Step S43: Construct a tensor model of collaborative performance of literacy indicators based on the simplified literacy feature set, and use the tensor model of collaborative performance of literacy indicators to decompose the performance intensity and interrelationship of core literacy dimensions to generate a literacy interaction feature matrix.
[0043] Step S44: Construct a multi-dimensional relationship graph based on the competency interaction feature matrix, identify the implicit association structure between the dimensions of each core competency, and generate a subject core competency relationship graph;
[0044] Step S45: Represent the subject core competency relationship diagram in a structured matrix to obtain the subject core competency relationship matrix.
[0045] This invention constructs structured competency data by identifying student scores, question difficulty, practical scores, and core competency tags in standardized subject learning data. This effectively breaks through the single-dimensional limitations of traditional performance metrics, achieving a multi-dimensional expression of subject elements. By extracting highly relevant competency indicators through feature redundancy removal and dimensionality reduction, the complexity of the model is effectively reduced, improving the efficiency of subsequent analysis. Tensor models are used to characterize the synergistic performance among core competencies; the parameter decomposition process extracts the strength and weakness relationships and complementary features between competencies, enhancing the depth of interactive understanding. Constructing a multi-dimensional relationship graph reveals the implicit structure between competencies, compensating for the neglect of the inherent logic of competencies in traditional analysis. Finally, the matrix representation of the relationship graph not only improves the structural readability of the data but also provides a standard input format for subsequent modeling, achieving a balance between high structural stability and high semantic integrity.
[0046] Optionally, step S44 specifically includes:
[0047] Step S441: Map high-dimensional spatial nodes to the competency interaction feature matrix, set the number of core competency dimensions to 8, define each core competency dimension as a graph node, calculate the performance intensity synergy coefficient of student tags in each core competency dimension as the initial edge weight, and generate the initial competency node graph.
[0048] Step S442: Set the Gaussian kernel scaling factor to 0.5, construct the adjacency matrix for the initial literacy node graph, perform nonlinear transformation on the edge weights between nodes, and generate a higher-order literacy adjacency matrix.
[0049] Step S443: Set the number of encoding layers to 2, with the number of neurons in each layer being [64, 32]. Perform graph structure encoding based on the literacy high-order adjacency matrix, optimize the semantic discriminability of node representations, and generate a literacy semantic embedding graph.
[0050] Step S444: Perform structural community detection on the literacy semantic embedding graph, identify hidden association subgroups, and generate a literacy subgroup graph;
[0051] Step S445: Merge the competency subgroup graphs and reorganize the matrix structure to generate a core competency relationship matrix.
[0052] This invention maps core competency dimensions to graph nodes and constructs a competency node graph using the collaborative performance of student labels as edge weights, effectively depicting the actual interaction relationships between competencies. Setting the core competency dimensions to 8 ensures coverage of mainstream subject assessment frameworks and guarantees the integrity of the graph structure. By setting a Gaussian kernel scaling factor of 0.5 for nonlinear edge weight transformation, the adjacency matrix smooths weak connections, improving the robustness of the graph structure. Employing a two-layer encoding structure and setting the number of neurons to [64, 32] enhances the semantic distinguishability of nodes while ensuring computational efficiency, generating a more representative embedding graph. Graph community detection uncovers potential clustering structures between competencies, compensating for the shortcomings of traditional linear relationship analysis. The final structural reorganization and matrix merging process preserves the subgroup relationships and semantic embedding features in the graph, providing multidimensional structural support for subsequent learning quality diagnosis.
[0053] Optionally, step S5 specifically includes:
[0054] Step S51: Perform joint structural analysis on the subject core competency relationship matrix and hierarchical feature dataset to generate a quality deviation feature vector set;
[0055] Step S52: Perform cluster diagnostic analysis on the quality deviation feature vector set to identify imbalance points and deviation areas, and generate a learning quality assessment report;
[0056] Step S53: Extract teaching resource features and subject objective requirements features from standardized subject learning data;
[0057] Step S54: Construct a path weight recommendation model based on the learning quality assessment report, teaching resource characteristics, and subject objective requirements.
[0058] Step S55: Based on standardized subject learning data, conduct student group profile analysis to obtain regional student group profile data, and use the path weight recommendation model to perform path matching and optimization on the regional student group profile data to obtain subject optimization strategies.
[0059] This invention comprehensively characterizes the structural features of subject quality deviations by jointly analyzing the core competency matrix and hierarchical feature data, making the diagnostic results more accurate and interpretable. Identifying deviation areas through clustering helps to accurately locate imbalances within subjects; the multi-center clustering method can adapt to diverse student characteristic distributions. Extracting teaching resources and subject goal information as auxiliary factors enhances the content matching accuracy of path recommendations. The path weight recommendation model is constructed based on multi-source feature fusion, improving the personalization and rationality of path generation; the introduction of an attention mechanism can automatically allocate key factor weights, improving recommendation accuracy. Finally, combining regional student group profiles for path optimization can provide differentiated strategy support for differences in learning foundations across different regions, significantly improving the efficiency of educational resource allocation and the precision of intervention.
[0060] Optionally, step S52 specifically includes:
[0061] Step S521: Perform principal component analysis and orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum number of dominant dimensions to 6 to extract the dominant deviation dimension factors and generate deviation principal cause tensor data.
[0062] Step S522: Perform density peak estimation on the principal cause tensor data of the bias, set the minimum number of samples to 10 and the cluster distance threshold to 0.75 to identify potential bias cluster centers and generate a subject imbalance cluster map;
[0063] Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the subject imbalance cluster map, and set the number of neighborhoods to 20 to measure the anomaly intensity of each cluster unit in the subject imbalance cluster map, identify the deviation area and generate a local imbalance anomaly distribution map.
[0064] Step S524: Combining the local imbalance anomaly distribution map and the deviation principal cause tensor data, set the deviation factor weight threshold to 0.6 to calculate the deviation factor of each literacy dimension and generate an imbalance feature weight dataset.
[0065] Step S525: Perform semantic reconstruction and structural induction on the imbalanced feature weight dataset, and output a learning quality evaluation report.
[0066] This invention compresses high-dimensional quality deviation features through principal component analysis. Under the conditions of an 85% variance contribution rate and a maximum of six dominant dimensions, it effectively retains key information while eliminating redundant dimensions, thus improving analytical efficiency. Deviation clustering is performed using density peak estimation, with a minimum sample size of 10 and a clustering distance threshold of 0.75, ensuring representative clustering and accurate identification of core deviation regions. By setting a boundary spacing of 0.05 and a neighborhood size of 20, local imbalance distributions are meticulously characterized, aiding in the discovery of marginal anomalies and structural voids. A deviation factor weight threshold of 0.6 effectively filters significantly influential literacy deviation dimensions, making the diagnosis more targeted. Finally, semantic reconstruction comprehensively improves the accuracy and interpretability of learning quality diagnosis.
[0067] Optionally, this specification also provides a system for analyzing the constitutive composition of intra-disciplinary balanced combination intervals, used to perform the method for analyzing the constitutive composition of intra-disciplinary balanced combination intervals as described above. This system includes:
[0068] The data acquisition module is used to acquire subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data.
[0069] The combined interval generation module is used to divide combined intervals based on standardized subject learning data to obtain subject combined interval data; and to perform interval structure modeling on the subject combined interval data to obtain the combined interval model.
[0070] The combined interval analysis module is used to compare the distribution of subject-level scores based on the combined interval model to obtain subject-level score deviation data; and to evaluate the differences in subject learning quality based on the subject-level score deviation data to obtain a hierarchical feature dataset.
[0071] The core competency relationship modeling module is used to model the relationship between core competencies in a subject based on standardized subject learning data, and obtain the core competency relationship matrix.
[0072] The quality diagnosis module is used to assess the quality of subject learning based on the subject core competency relationship matrix and hierarchical feature dataset, generate a subject learning quality assessment report, and optimize the learning path of standardized subject learning data based on the subject learning quality assessment report to obtain subject optimization strategies.
[0073] The present invention provides a subject-specific balanced combination interval compositional analysis system. This system can implement any of the subject-specific balanced combination interval compositional analysis methods of the present invention. It is used to combine the operation and signal transmission media between various modules to complete the subject-specific balanced combination interval compositional analysis method. The modules within the system cooperate with each other, thereby improving the pertinence and accuracy of teaching quality monitoring. Attached Figure Description
[0074] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0075] Figure 1 This is a schematic diagram of the steps in the method for analyzing the constitutive composition of balanced combination intervals within the discipline of this invention.
[0076] Figure 2 This is a detailed flowchart of step S1 in the present invention;
[0077] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0078] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0079] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0080] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0081] To achieve the above objectives, please refer to Figures 1 to 2 This invention provides a method for constitutive analysis of balanced combination intervals within a discipline, the method comprising the following steps:
[0082] Step S1: Obtain subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data;
[0083] In this embodiment, subject learning data was first collected from junior high school students in a certain region. This included student ID numbers, scores in each subject (Chinese, Mathematics, English, Physics, Chemistry, etc.), full marks for each question, individual scores, the knowledge point code of each question, the difficulty coefficient of each question (range 0-1), answering time and completion status, subject objective requirements, and the core competency tags associated with each question. The collected raw subject learning data was cleaned, removing fields with a missing rate greater than 30%. Continuous score fields were standardized using Z-scores, with a standardization threshold of ±3σ. Values exceeding this range were considered outliers and corrected or removed. The final standardized subject learning dataset contained 12,823 students, 6 courses, and an average of 100 questions per course. After standardization, all score variables were converted to a normal distribution with a mean of 0 and a standard deviation of 1 to ensure comparability and interpretability in subsequent analyses.
[0084] Step S2: Divide the subject combination intervals based on standardized subject learning data to obtain subject combination interval data; perform interval structure modeling on the subject combination interval data to obtain the combination interval model;
[0085] In this embodiment, based on the obtained standardized subject learning data, the students' total standardized score (weighted average) and the percentage of questions with scores below 50 are first extracted as two variables to divide the two-dimensional interactive coordinate system. The horizontal axis represents the standardized total score (μ = 0, σ = 1), and the vertical axis represents the percentage of low-scoring questions (μ = 0.35, σ = 0.12). The scores are divided into three segments: "low," "medium," and "high," based on a mean ± 0.67σ, corresponding to intervals: [-∞, -0.67σ), [-0.67σ, +0.67σ], and (+0.67σ, ∞), respectively. A 3×3 two-dimensional combined spatial label matrix is constructed, containing nine intervals (e.g., LL, MH), where "LH" indicates a low total score and a high percentage of low scores. The standardized subject learning data is then mapped to the combined interval matrix, and the sample size and percentage for each interval are calculated. This results in a subject combined interval data table containing the label, sample size, mean score, and standard deviation for each interval. The KMeans algorithm is then used to model the structure of this combined interval data, with K = 4 clusters, k-means++ initialization, and a maximum iteration count of 300. This combined interval model is then used for subsequent stratified comparisons.
[0086] Step S3: Compare the subject-level performance distribution based on the combined interval model to obtain subject-level performance deviation data; evaluate the differences in subject learning quality based on the subject-level performance deviation data to obtain a hierarchical feature dataset.
[0087] In this embodiment, the sample is first divided into four student groups—"High-Performance Zone," "Medium-Performance Zone," "Needs Improvement Zone," and "Key Focus Zone"—based on a combined interval model. Then, hierarchical performance is compared based on the score distribution of each group across different subjects. Using the average scores of students in the "High-Performance Zone" in mathematics and physics as a benchmark, the mean deviation and variance changes of students in the "Needs Improvement Zone" in the same subjects are compared. The results show that mathematics has a mean deviation of -1.2σ and a variance increase of 22.3%. These deviation characteristics are extracted to construct subject-level performance deviation data. Further analysis of the quality differences between different subjects is conducted based on this data. A quality balance index G (the absolute value of the standardized mean difference) is introduced. When G > 1.0, an imbalance is identified. Least squares regression is used to calculate the complementarity strength between subjects, combined with the hierarchical distribution uniformity E (defined as the standard deviation of the sample size in each layer), ultimately generating a hierarchical feature dataset as a feature representation of the learning quality structure.
[0088] Step S4: Model the relationship between core competencies in a subject based on standardized subject learning data to obtain a matrix of core competencies in a subject.
[0089] In this embodiment, based on standardized subject learning data, the "knowledge point-performance-competency label" triplet between questions and students is first identified to construct a preliminary structured dataset of competency data. Each student's score performance on the core competency label (such as logical thinking, language expression, experimental operation, and comprehensive application) of the question is extracted. Redundant competency label features are removed using the information gain method, retaining the top-10 features with mutual information values to construct a concise competency feature set. Students' performance on different competency dimensions is tensorized to form a three-dimensional collaborative performance tensor of "student × competency × question type." The CP decomposition method (rank set to 5) is used to extract latent interaction factors, generating a competency interaction feature matrix. A multidimensional relationship graph is then constructed based on this matrix, where each node represents a competency dimension, and the edge weight represents the collaborative performance strength between dimensions. The Louvain algorithm is used for community detection, visualizing and strengthening the relationships of implicitly related subgroups. Finally, a subject core competency relationship matrix is obtained through structured matrix representation, with a dimension of 9×9, where each cell indicates the potential correlation strength between competency dimensions.
[0090] Step S5: Conduct subject learning quality assessment based on the subject core competency relationship matrix and hierarchical feature dataset to obtain a subject learning quality assessment report, and optimize the learning path of standardized subject learning data based on the subject learning quality assessment report to obtain subject optimization strategies.
[0091] In this embodiment, the obtained core competency relationship matrix of the subject is first fused with the obtained hierarchical feature dataset using principal component analysis. The combination of competency dimensions with the largest performance fluctuation is extracted as the quality deviation feature vector. DBSCAN clustering analysis (eps=0.4, min_samples=10) is used to identify imbalances in abnormal clusters in the deviation space, and a learning quality assessment report is output. Subsequently, the standardized subject learning data is feature-annotated to extract the source of questions, the level of assessment ability (memory, understanding, application, etc.), and the type of resources (text, video, experiment, etc.). At the same time, the Bloom level of subject goal requirements and the key competencies of the curriculum standard are extracted to construct a goal matching vector. Based on the diagnostic results and the matching path of teaching goals, a regulation path network is constructed. The weight of the path node is set as the deviation degree × matching degree, and the conversion probability is adopted using the historical path recommendation confidence level (range [0.3-0.9]), generating a path weight recommendation model. Finally, the model is used to perform path recommendation matching on the regional group profile data, and a subject optimization strategy table containing "strategy sequence, strategy implementation difficulty, and expected improvement direction" is output to guide personalized teaching and path intervention.
[0092] Optionally, step S1 specifically includes:
[0093] Step S11: Collect multi-source subject learning raw data through the regional teaching management platform, and uniformly encode the multi-source subject learning raw data into a data framework structure to obtain subject learning data;
[0094] In this embodiment, raw subject learning data is collected from multiple data sources through the data interface of the regional teaching management platform. This includes data from online homework platforms (such as the smart homework system), classroom teaching platforms (such as the learning analysis system), and paper-based exam score entry systems. The data covers five subjects: Chinese, mathematics, English, science, and history, encompassing 132 schools, 8547 classes, and approximately 246,000 student learning records and classroom performance records. The data from different platforms exhibits inconsistencies in field naming and data format. A unified data access module is used for field structure identification and mapping. By defining a unified data structure template (including 18 fields such as student ID, subject code, question ID, answer result, score, knowledge point tag, answer time, answer duration, classroom performance score, and subject core competencies), all raw data is standardized into a uniformly structured JSON format object. This ultimately generates a multi-dimensional fusion dataset containing subject learning behavior, grades, and process performance, which is output in CSV format as input data for subsequent cleaning.
[0095] Step S12: Perform missing value detection on the subject learning data, and impute the missing value detection results using the class mode to obtain cleaned data;
[0096] In this embodiment, missing value detection is performed on the subject learning data generated in step S11, and the missing rate of each field is calculated. For numerical fields such as "score" and "response time", a record counting method is used to detect null values. For categorical fields such as "knowledge point tags" and "response status", a regular expression + default label recognition method is used to mark invalid items. The detection results show that the missing rate of the "response time" field is 12.4%, and the missing rate of "knowledge point tags" is 9.1%. For missing values in categorical fields, the mode of the category is used for imputation. For example, the mode of the "knowledge point tags" field is "geometric shape recognition", and all missing items are replaced with this value. For the "response time" field, a missing threshold of 15% is set. Since it is below the threshold range, the field is retained and imputed with the median of the overall distribution (45 seconds in this data). The missing rate of the final cleaned data is less than 1%, which meets the requirements of subsequent normalization and standardized modeling.
[0097] Step S13: Perform field normalization on the cleaned data, and standardize the normalization results by encoding categorical fields to generate transformed data;
[0098] In this embodiment, the obtained cleaned data undergoes field normalization. For numerical fields (such as "score", "answer time", "question difficulty", etc., a total of 9 items), a minimum-maximum normalization method is used to linearly scale all values to the [0,1] interval. The normalization formula is: (X-X_min) / (X_max-X_min). For example, for the "score" field, the original score range is [0,100], and after normalization, the highest value is 1.0 and the lowest value is 0.0. Subsequently, the categorical fields (such as "knowledge point label", "answer status", "question type") contained in the data are uniformly encoded using label encoding, mapping each category value to a unique integer value. For example, in the "question type" field, "multiple choice", "fill in the blank", and "problem-solving" are encoded as 0, 1, and 2, respectively, ensuring that all input data can be used for subsequent machine learning model or graph construction processing. The transformed data format is a matrix structure, containing normalized continuous variables and uniformly encoded discrete variables, forming a complete transformed dataset for subsequent standardization operations.
[0099] Step S14: Perform scale standardization on the transformed data to obtain standardized subject learning data.
[0100] In this embodiment, the generated transformed data undergoes scaling. Standardization employs the Z-score standardization method, converting all numerical fields into a distribution with a mean of 0 and a standard deviation of 1. The calculation is: (X-μ) / σ, where μ is the sample mean and σ is the sample standard deviation. For example, the original mean of the "response duration" field was 45 seconds, with a standard deviation of 13.2 seconds. After standardization, the mean becomes 0, and the standard deviation becomes 1, compressing outliers (such as entries with response durations exceeding 90 seconds) to the tails of the distribution. The standardization results are then tested for normality using the Shapiro-Wilk method, with a significance level set at α = 0.05. The results show that most fields follow an approximately normal distribution, suitable for subsequent principal component extraction and structural modeling. The standardized subject learning data is stored in matrix form, containing 243 feature fields with a data size of 247382 × 243, serving as the foundational input data for subsequent learning path analysis and feature modeling.
[0101] Optionally, the combined interval division described in step S2 is specifically as follows:
[0102] Student performance characteristics are extracted from standardized subject learning data to obtain student subject performance data. Statistical analysis is then performed on the student subject performance data to obtain standardized total performance data and student performance percentage data.
[0103] In this embodiment, when extracting student performance features from standardized subject learning data, a field filtering method is used to extract fields directly related to grades, such as "final exam score," "class participation score," "process evaluation," "homework score," and "test score," excluding text-based notes and behavioral non-quantitative fields. These fields are then weighted and merged according to weight coefficients (e.g., final exam score has a weight of 0.5, class participation score has a weight of 0.2, and the rest have a total weight of 0.3) to obtain standardized student subject performance data. Based on this, the `describe()` function from the Pandas library is used to obtain basic statistics such as mean, standard deviation, quantiles, maximum, and minimum values. Furthermore, based on score ranges, the proportion of students in each grade interval (e.g., below 60, 60-80, and above 80) is calculated to form standardized total score data and student score percentage data.
[0104] Based on the percentage of students with low grades, the percentage of students with low grades is divided to obtain the percentage of students with low grades.
[0105] In this embodiment, when dividing the proportion of students with low scores based on student performance data, the standard for low scores is first set as students with a total score below 60 or those in the bottom 20% percentile of their grade. The data is then grouped and statistically analyzed within the sample set of each region, school, or class. Using a threshold T_low = 60 or P_low = 20% as the dividing line, the proportion of students with scores below this threshold within each analysis unit is calculated to obtain the proportion of students with low scores. To ensure stability, a sliding window is used for mean smoothing for analysis units with a sample size of less than 30.
[0106] Based on standardized total score data and low score student percentage data, a combination variable is set, where the horizontal axis variable is set to standardized total score data and the vertical axis variable is set to low score student percentage data, and a two-dimensional interactive coordinate system is constructed.
[0107] In this embodiment, when setting combined variables based on standardized total score data and the percentage of students with low scores, the standardized total score is first determined as the horizontal axis variable, and the percentage of students with low scores is determined as the vertical axis variable, constructing a two-dimensional interactive coordinate system. All data points of the analysis units are mapped to this two-dimensional coordinate system, with each unit data point representing the combined performance of a school, class, or region. To ensure consistency in variable dimensions, both variables are standardized using Z-score, where Z = (X - μ) / σ, to ensure a symmetrical spatial distribution of coordinate coefficient values and to adapt to subsequent distribution analysis.
[0108] The standardized distribution features of the spatial combination analysis basis are extracted, and the mean ±0.67σ in the standardized distribution features is set as the dividing threshold to divide the two-dimensional interactive coordinate system. The horizontal and vertical axes of the two-dimensional interactive coordinate system are divided into three segments to form a two-dimensional combination spatial label matrix.
[0109] In this embodiment, when extracting standardized distribution features from the spatial combination analysis basis, the mean μ and standard deviation σ of the horizontal and vertical axis variables are extracted based on the two-dimensional coordinate point set. ±0.67σ is used as the interval division threshold, and each axis is divided into three segments according to the empirical rule of normal distribution. Taking the horizontal axis as an example: values less than μ-0.67σ are defined as the low segment, values between μ±0.67σ are defined as the middle segment, and values greater than μ+0.67σ are defined as the high segment. The vertical axis is divided in the same way, ultimately forming a 3×3 nine-square grid combination space. Each combination interval is assigned an independent identifier, such as "LM" (low performance - medium, waiting for improvement), for subsequent mapping and label classification.
[0110] Based on the two-dimensional interactive coordinate system, standardized subject learning data is mapped to the corresponding combination labels in the two-dimensional combination space label matrix, the interval categories are marked, and the number and proportion of samples in each interval are counted to generate subject combination interval data.
[0111] In this embodiment, when mapping standardized subject learning data to corresponding combination labels in a two-dimensional combination space label matrix based on a two-dimensional interactive coordinate system, the coordinates of the (total score, percentage of low scores) values of each analysis unit (a subset of standardized subject learning data split according to specific organizational dimensions (such as school, class, school district, student group, etc.)) are first classified. By determining the segment number of the unit on the X-axis and Y-axis, it is mapped to a specific interval in the 3×3 label matrix, and the combination label category of the unit is labeled. Subsequently, all analysis units are classified and summarized by label, and the number of samples in each label interval and its proportion of the total sample are counted to form structured subject combination interval data, providing basic coordinate information support for subsequent difference analysis and literacy modeling.
[0112] Most importantly, the formation of the two-dimensional combined spatial label matrix is specifically as follows:
[0113] Normal distribution parameters are estimated for the standardized total score data. The mean and standard deviation σ1 of the horizontal axis are calculated. The score level intervals are divided with the mean ± 0.67σ1 as the threshold. The following definitions are made: the horizontal axis value is less than -0.67σ1 as the low zone, the horizontal axis value is [-0.67σ1, +0.67σ1] as the middle zone, and the horizontal axis value is greater than -0.67σ1 as the high zone, thus generating a set of horizontal axis labels.
[0114] In this embodiment, the standardized total score data is processed by normal distribution parameter estimation. The maximum likelihood estimation method is used to fit the score sample distribution curve, and the calculated mean μ1 and standard deviation σ1 of the standardized total score are 75.6 and 6.2, respectively. Based on this, the threshold parameter range is set to μ1 ± 0.67σ1, i.e., the interval range is [71.44, 79.76]. The standardized total score data of all analysis units are divided into three segments according to this interval: low (<71.44), medium ([71.44, 79.76]), and high (>79.76), and labeled as "L" (Low), "M" (Medium), and "H" (High) as the horizontal axis label set, respectively. A continuous interval segmentation strategy is used during the segmentation process to avoid misjudgment of critical points. The segmentation results are also visually verified to ensure balanced sample coverage.
[0115] The distribution parameters of the low-score data are estimated, and the mean and standard deviation σ2 of the vertical axis are calculated. The score level intervals are divided with the mean ± 0.67σ2 as the threshold. The following are defined: the vertical axis value is less than -0.67σ2 as the low zone, the vertical axis value is [-0.67σ2, +0.67σ2] as the middle zone, and the vertical axis value is greater than -0.67σ2 as the high zone, thus generating a set of vertical axis labels.
[0116] In this embodiment, kernel density estimation is used to fit the distribution trend of the low-achieving student percentage data. After excluding asymmetric outliers by combining skewness coefficient detection, the mean μ2 of the low-achieving percentage on the vertical axis is calculated to be 23.4%, and the standard deviation σ2 is 5.8%. Similarly, using μ2±0.67σ2 as the dividing threshold, the segmented intervals are calculated to be [19.51%, 27.29%]. Based on this, the vertical axis is divided into three segments: low (<19.51%), medium ([19.51%, 27.29%]), and high (>27.29%), and labeled as "L" (low percentage), "M" (medium percentage), and "H" (high percentage) respectively, constructing a vertical axis label set. During the segmentation process, outlier compression is performed on the low-percentage fluctuating samples, and z-score normalization is used to make the distribution closer to the normal model, improving label stability.
[0117] A two-dimensional combination space is constructed based on the set of labels on the horizontal axis and the set of labels on the vertical axis. Nine interval labels are generated by cross-combination, and the nine interval labels are transformed into matrices and output in matrix form to obtain the two-dimensional combination space label matrix.
[0118] In this embodiment, the generated horizontal axis label set {L,M,H} is cross-combined with the vertical axis label set {L,M,H} generated in step S222. A two-dimensional combined spatial label set is constructed using the Cartesian product method, resulting in 9 combined labels, including "LL", "LM", "LH", "ML", "MM", "MH", "HL", "HM", and "HH". Subsequently, the label set is mapped using a spatial position matrix to construct a 3×3 two-dimensional combined spatial label matrix, where the rows represent the vertical axis segments and the columns represent the horizontal axis segments. The final output is in the following matrix form: The label matrix serves as the spatial basis for subsequent combined interval mapping, supporting the mapping of actual data from each analysis unit to the aforementioned nine label regions, and is used for statistical analysis and interval proportion modeling.
[0119] Optionally, the subject-level score distribution comparison mentioned in step S3 is specifically as follows:
[0120] Spatial label aggregation is performed on the interval combination model, and the sample data under the same interval label are hierarchically classified. Combined with the subject combination interval data, student distribution data with combination labels is generated.
[0121] In this embodiment, the nine pre-constructed combined spatial labels (such as LM, MH, etc.) are clustered and grouped. For the sample set within each label region, students are grouped by their student IDs, and key indicators such as their grades in each subject dimension, course participation frequency, and homework completion are uniformly encoded to construct a structured sample index table. A "sample percentage" is introduced as an aggregation weight factor into the sample data under each label to hierarchically aggregate the characteristics of the sample group. The aggregation granularity is set to at least 50 student samples under each label. After aggregation, combined with the original subject combination interval data, a combined label student distribution dataset is generated using a label-sample mapping function. This dataset represents the number of students, their mean grade level, and the degree of difference within each combination interval, providing a foundation for subsequent stratified analysis.
[0122] Based on the distribution of standardized total score data, set interval boundary thresholds for low, medium, and high strata, and perform interval stratification on the combined label student distribution data to generate a hierarchical distribution label matrix.
[0123] In this embodiment, based on the standardized total score data distribution, the student group is divided into three segments using the quantile method. The low stratum is defined as students whose total scores are below the top 30% quantile (threshold μ-σ), the middle stratum is defined as students between 30% and 70% (μ±σ), and the high stratum is defined as students above the 70% quantile (threshold μ+σ). The parameters are set as follows: low score threshold ≤ 68.4, middle stratum [68.4, 81.2], and high score threshold ≥ 81.2. The student distribution data with combined labels is traversed according to the combined label dimension. The student group data within each combined label is classified according to the above score range, completing the "low-middle-high" stratified label labeling for each category of label. This is then embedded into a unified data structure, forming a 9×3 stratified distribution label matrix, where each combined label corresponds to a three-layer structure representing the number and proportion of students in different strata.
[0124] Standardized subject learning data is mapped to student distribution data with combined labels and hierarchical distribution label matrix using dual labels to generate multi-label hierarchical score dataset. The number of samples at each level under different combined labels in the multi-label hierarchical score dataset is normalized and statistically analyzed to obtain hierarchical score frequency matrix.
[0125] In this embodiment, standardized subject learning data is bidirectionally matched based on combination labels (e.g., LM) and hierarchical labels (e.g., higher level), uniquely locating each student within both the combination interval and the score level, thus achieving dual-label mapping. Based on this, a multi-label hierarchical score dataset is generated. Each record in the dataset simultaneously contains a combination interval label, a hierarchical label, a standardized score value, and the corresponding student code. The multi-label hierarchical score dataset undergoes statistical normalization. Min-Max normalization is used to uniformly scale the sample size of each level to the [0,1] interval, ensuring the comparability of sample distribution comparisons between different combination labels. The final result is a 9×3 hierarchical score frequency matrix. This matrix uses combination labels as rows and score levels as columns, with each cell representing the normalized sample proportion value under the corresponding level.
[0126] The difference measurement is calculated based on the hierarchical score frequency matrix to evaluate the hierarchical score distribution deviation between different combination labels and obtain subject hierarchical score deviation data.
[0127] In this embodiment, a distribution difference analysis is performed based on the hierarchical score frequency matrix. KL divergence (KL divergence) is used to calculate pairwise differences between the score frequency vectors of different label combinations, measuring hierarchical distribution bias. A difference sensitivity threshold of 0.15 is set; label combinations with KL divergence values higher than this threshold are considered to have significant hierarchical bias. Jensen-Shannon distance is also used for auxiliary verification. The output is a structured subject hierarchical score bias dataset, recording the bias score and relative standard deviation of each label combination at each level. It also identifies label combinations with prominent biases, providing a basis for bias localization in subsequent path recommendation and quality diagnosis.
[0128] Optionally, the subject learning quality difference assessment described in step S3 specifically includes:
[0129] The distribution statistics and central tendency analysis of the subject-level performance deviation data are performed, and the mean, range and standard deviation of each combination label corresponding to the level are calculated to generate a deviation distribution statistical matrix.
[0130] In this embodiment, based on subject-level score deviation data, statistical analysis methods are used to measure the central tendency and dispersion of stratified score distributions under each combined label. First, the low, medium, and high-level scores corresponding to each combined label (such as LM, MH, etc.) are aggregated, and the mean, range (difference between the maximum and minimum values), and standard deviation of the scores within each level are calculated to measure the central tendency and volatility of the score distribution. The analysis precision is set to two decimal places, and extreme outliers beyond three standard deviations are automatically filtered in the range calculation. The output deviation distribution statistical matrix is a 9×3-dimensional structure, with rows representing the nine combined labels and columns representing the mean, range, and standard deviation statistical indicators of each level, used for subsequent horizontal and vertical analysis.
[0131] Based on the deviation distribution statistical matrix, a horizontal regional balance analysis is conducted to assess the consistency of performance across different regions at each level and generate regional balance characteristic data.
[0132] In this embodiment, based on the generated deviation distribution statistical matrix, the system is horizontally divided according to regional dimensions (such as streets and towns). Analysis of variance (ANOVA) is performed on the mean of the combined labels for each region at the same hierarchical level to measure the consistency of different regions in their hierarchical performance. The mean variance of different combined labels within a region is calculated. If the variance within a region is lower than a set threshold of 0.03, it is considered that the region has high consistency in its performance at that level. The consistency of different regions at the three levels is mapped to a scoring matrix, ultimately outputting regional balance characteristic data, including a two-dimensional index of "region-hierarchy," a score balance index (range [0,1]), and a mean deviation rate index, serving as the basis for subsequent gradient analysis and difference evaluation.
[0133] A vertical hierarchical gradient analysis is performed on the deviation distribution statistical matrix to calculate the continuity of performance improvement and the probability of hierarchical transition, thereby generating a hierarchical gradient change model.
[0134] In this embodiment, the mean scores of each level in the deviation distribution statistical matrix are used to construct the level gradient vector (low → medium → high) under each combination label. The improvement rate and leap probability from low to medium level and from medium to high level are calculated respectively. The leap probability is based on the change in the proportion of students between levels, and is defined as Pij = Nij / (Nij + Nii), where Nij represents the number of students who leap to the next higher level, and Nii represents the number of students who remain at the current level. A threshold of 0.5 or higher is considered a positive leap trend. Finally, a level gradient change model is generated, which includes the level improvement path under each combination label, the leap probability matrix (3×3), and score continuity indicators (such as the slope of improvement), used to evaluate the naturalness and controllability of student performance improvement under this combination.
[0135] By integrating regional equilibrium characteristic data and hierarchical gradient change models to conduct a weighted comprehensive evaluation, data on the differences in subject learning quality are obtained.
[0136] In this embodiment, regional balance characteristic data and hierarchical gradient change models are integrated to perform a weighted comprehensive evaluation of the performance of each combined label. The weight of regional balance is set to 0.4, and the weight of hierarchical gradient continuity is set to 0.6. A linear weighted model is used to sum the two types of indicators to generate a comprehensive score evaluation index Gij = 0.4 × Rij + 0.6 × Lij, where Rij is the regional consistency score and Lij is the normalized value of the hierarchical gradient slope. The output subject learning quality difference assessment data includes combined labels, comprehensive quality scores, main imbalance cause types (such as regional imbalance, low hierarchical jump, etc.), and improvement direction suggestions, serving as an important basis for student profile integration and path recommendation.
[0137] Based on the subject learning quality difference assessment data and the multi-label hierarchical score dataset, student hierarchical features are integrated to obtain a hierarchical feature dataset.
[0138] In this embodiment, based on the aforementioned subject learning quality difference assessment data and multi-label hierarchical performance dataset, information such as each student's combined labels, hierarchical position, and performance fluctuation trends are fused to construct a five-tuple data structure containing student ID, subject label, hierarchical label, comprehensive quality score, and difference type. During the fusion process, a performance fluctuation smoothing algorithm (with a moving average window width set to 3) is introduced to process the continuous trend of students' performance and reduce the impact of extreme fluctuations. The final generated hierarchical feature dataset contains all students' positions in the combination space, their hierarchical levels, performance change patterns, and learning quality deviation types, providing refined input features for subsequent path recommendation and resource allocation.
[0139] Optionally, step S4 specifically includes:
[0140] Step S41: Perform subject element label recognition on standardized subject learning data, extract student scores, question difficulty, practice scores and core competency labels, construct element label vector set, and thus generate a competency structured dataset;
[0141] This embodiment, based on standardized subject learning data, first analyzes the test question fields, student answer records, scores, and evaluation criteria in the data. It then identifies subject element tags using natural language keyword matching and knowledge point tag extraction algorithms. Specific extracted fields include: student score, question difficulty level, experiment and practice score, and the corresponding core competency dimension tags (such as "information awareness," "practical ability," and "innovative thinking"). A word vector model is used to uniformly encode the core competency tags, with a vector dimension of 128. A structured element vector is constructed based on student-test question units, forming a set of triplets containing student ID, test question ID, and element tag vectors. Finally, a structured competency dataset is generated, with a two-dimensional nested dictionary structure. The first dimension index is the student ID, the second dimension index is the test question ID, and the corresponding data is the complete element tag vector.
[0142] Step S42: Perform dimensionality reduction and feature redundancy removal on the structured dataset of literacy, extract highly relevant literacy association features, and generate a concise literacy feature set;
[0143] In this embodiment, the generated literacy structured dataset is first processed using a PCA (Principal Component Analysis)-based dimensionality reduction algorithm, with a 90% variance explanation rate retained, reducing the original 128-dimensional literacy element vector to approximately 30-40 dimensions. Then, a correlation threshold method is used for feature redundancy removal. The Pearson correlation coefficient between each feature is calculated, and a redundancy removal threshold of 0.92 is set. If the correlation between two features is higher than this value, the feature with higher information entropy is retained. For categorical features, mutual information coefficients are used for redundancy determination; features with information values less than 0.02 are directly discarded. Finally, a simplified literacy feature set is generated, containing the most distinctive and representative key performance feature vectors for each student across literacy dimensions, which serve as input for subsequent tensor modeling.
[0144] Step S43: Construct a tensor model of collaborative performance of literacy indicators based on the simplified literacy feature set, and use the tensor model of collaborative performance of literacy indicators to decompose the performance intensity and interrelationship of core literacy dimensions to generate a literacy interaction feature matrix.
[0145] In this embodiment, a simplified literacy feature set is used as input to construct a three-dimensional literacy performance tensor model. The tensor dimensions are set as: number of students × number of literacy dimensions × performance intensity level (e.g., low, medium, and high). The tensor is decomposed using the CP (CANDECOMP / PARAFAC) tensor decomposition algorithm, with a decomposition order of 5. The dominant performance factors of each literacy dimension in all student groups and their collaborative patterns with other dimensions are extracted. The interaction strength between literacy dimensions is calculated using the factor loading value in the tensor decomposition matrix, and a literacy interaction feature matrix is constructed. Each element of the matrix represents the joint performance correlation of two literacy dimensions in the student group, with a value range of [0,1], where values greater than 0.7 are considered strong correlations.
[0146] Step S44: Construct a multi-dimensional relationship graph based on the competency interaction feature matrix, identify the implicit association structure between the dimensions of each core competency, and generate a subject core competency relationship graph;
[0147] In this embodiment, a multi-dimensional relationship graph is constructed based on the generated competency interaction feature matrix. Nodes in the graph represent various core competency dimensions, and edge weights represent the synergistic correlation between two competency dimensions. The NetworkX graph modeling framework is used for modeling to achieve visualized management of interaction relationships. Edges are retained if their weight is greater than 0.3; edges with a weight greater than 0.7 are marked as strongly dependent edges. Furthermore, the Louvain community partitioning algorithm is used to identify implicit grouping structures among competency dimensions. For example, it was found that "mathematical modeling ability" and "logical reasoning ability" frequently cluster in the same subgraph, indicating a deep interaction between the two in the current student population. The final output core competency relationship graph is presented as a directed weighted graph to reveal the essential relationships and development paths between each competency dimension.
[0148] Step S45: Represent the subject core competency relationship diagram in a structured matrix to obtain the subject core competency relationship matrix.
[0149] In this embodiment, the core competency relationship graph is transformed into a structural matrix representation. The total number of competency dimensions is set to N, resulting in an N×N symmetric or asymmetric matrix. The element values of the matrix represent the edge weights between competency dimensions i and j, i.e., the synergistic association strength. If there is no direct connection between competency dimensions, the corresponding element value is set to 0. To facilitate subsequent modeling, the matrix is normalized by using the Min-Max method to uniformly compress all association strengths to the [0,1] interval. The final generated core competency relationship matrix is stored in a two-dimensional array format, and a corresponding dimension label mapping dictionary is output for use by the path recommendation algorithm and graph neural network processing module.
[0150] Optionally, step S44 specifically includes:
[0151] Step S441: Map high-dimensional spatial nodes to the competency interaction feature matrix, set the number of core competency dimensions to 8, define each core competency dimension as a graph node, calculate the performance intensity synergy coefficient of student tags in each core competency dimension as the initial edge weight, and generate the initial competency node graph.
[0152] In this embodiment, based on the competency interaction feature matrix, the total number of core competency dimensions is set to 8: information awareness, logical reasoning, innovation ability, problem solving, mathematical modeling, practical operation, collaborative communication, and learning transfer. Each dimension is mapped to a node in a graph. Taking students as units, the synergy coefficient of their scores on any two competency dimensions is calculated, using the Pearson correlation coefficient as a measure to calculate the consistency of synergistic performance between each pair of dimensions in the overall sample. The synergy coefficient is used as the edge weight, with the weight range set to [-1, 1]. Negative correlations are assigned a weight of 0, while positive correlations are retained, constructing an initial competency node graph. The node graph is represented as an undirected graph structure, containing 8 nodes and a maximum of 28 edges. The edge weight values serve as the basis for graph construction, generating an initial graph structure with strong and weak relationship expression capabilities.
[0153] Step S442: Set the Gaussian kernel scaling factor to 0.5, construct the adjacency matrix for the initial literacy node graph, perform nonlinear transformation on the edge weights between nodes, and generate a higher-order literacy adjacency matrix.
[0154] In this embodiment, the Gaussian kernel scaling factor σ is set to 0.5 in the constructed initial literacy node graph, and the Gaussian kernel function is used to perform a nonlinear mapping transformation on the edge weights. The expression for the Gaussian function is: Where w ij The original edge weights, and the transformed edge weights w i ′ j This transformation tends to produce a more discriminative nonlinear representation. It enhances the discriminative power of strong cooperative relationships while compressing the influence of weak cooperative relationships. The generated literacy higher-order adjacency matrix is an 8×8 dimensional symmetric matrix with its main diagonal set to 0. The sparsity of the adjacency matrix is controlled to within 30% to retain sufficient structural information for subsequent graph learning stages, while also enhancing the semantic expressive power of the graph structure.
[0155] Step S443: Set the number of encoding layers to 2, with the number of neurons in each layer being [64, 32]. Perform graph structure encoding based on the literacy high-order adjacency matrix, optimize the semantic discriminability of node representations, and generate a literacy semantic embedding graph.
[0156] In this embodiment, the constructed higher-order adjacency matrix of literacy is input into a preset graph neural network encoding module. Specifically, the encoding layer has two layers: the first layer has 64 neurons and the second layer has 32, with ReLU activation. The model adopts a GCN (Graph Convolutional Network) architecture, and the training objective is to maximize the semantic discriminability between nodes by separating different functional blocks in Euclidean space through node embedding vectors. The training data input consists of the adjacency matrix and initial feature vectors of nodes (set as unit vectors or initial principal component vectors). The training iterations are set to 300 rounds, using the Adam optimizer with a learning rate of 0.001. The final output is a 32-dimensional semantic embedding vector for each node, constructing a literacy semantic embedding graph to represent the semantic features of each literacy dimension within the graph structure.
[0157] Step S444: Perform structural community detection on the literacy semantic embedding graph, identify hidden association subgroups, and generate a literacy subgroup graph;
[0158] In this embodiment, the generated literacy semantic embedding graph is used to detect community structure in the node embedding vectors. The Louvain clustering algorithm is employed to automatically identify potential sub-community relationships within the graph structure. The maximum number of communities is set to 4, and the minimum community size is 2. The detection strategy maximizes community modularity based on node semantic similarity. If "mathematical modeling," "problem solving," and "logical reasoning" frequently form a sub-community, it is considered to have an inherent correlation mechanism. The final output is a literacy subgroup graph, identifying the internal structure of subgroups by community and preserving the node-community mapping relationship, providing a structural foundation for the final relationship matrix construction. The subgroup graph can be used to identify which literacy dimensions are more integrative and which have a tendency to differentiate.
[0159] Step S445: Merge the competency subgroup graphs and reorganize the matrix structure to generate a core competency relationship matrix.
[0160] In this embodiment, the competency subgroup graph is structurally merged, and the structural matrix is reconstructed by aggregating nodes within subgroups and strengthening connections between subgroups. The matrix dimension is maintained at 8×8. The weights of connections within subgroups are normalized and enhanced, with a weight increase of 1.2 times; cross-subgroup connections are diluted, with a weight decrease of 0.8 times. After normalizing all elements in the matrix to 0-1, the final core competency relation matrix is output. This matrix retains community structural features, node semantic embedding characteristics, and original collaborative relationships, forming a comprehensive expression structure with high semantics, high relevance, and high structural stability. The final output structure will be used to learn the input layer matrix features of the quality diagnostic model. The core competency relation matrix can be represented as: The diagonal element r ii (e.g. r) 11,r 22 ,…,r 88 The ) represents the strength of the relationship between each core competency dimension and itself, i.e., the self-association of that dimension. Off-diagonal element r ij (e.g. r) 12 ,r 23 ,…,r 78 The matrix represents the interaction relationships between different core competency dimensions, which are reflected through clustering and edge weight adjustment. During matrix construction, node relationships within subgroups are strengthened, leading to an increase in the relationship strength of corresponding elements (e.g., r). 12 or r 21 Connections across subgroups are diluted, reducing the strength of the relationship between corresponding elements (e.g., r). 13 or r 31 After normalization, all elements in the matrix will be adjusted to the range of 0 to 1, ensuring a standardized representation of the result.
[0161] Optionally, step S5 specifically includes:
[0162] Step S51: Perform joint structural analysis on the subject core competency relationship matrix and hierarchical feature dataset to generate a quality deviation feature vector set;
[0163] In this embodiment, the input subject core competency relationship matrix is an 8×8 dimensional semantic structure matrix, representing the correlation strength between each core competency. The hierarchical feature dataset is a structured feature set generated from standardized subject learning data through a two-dimensional mapping of combined labels and grade levels, containing vector representations of students' hierarchical performance across each competency dimension. These two sets of data are concatenated using tensors, merged along the feature axes, and uniformly embedded in spatial dimensions. Dimensionality reduction is performed using Principal Component Analysis (PCA), retaining the top K principal components (e.g., K=12) with a cumulative contribution rate exceeding 90%. The output is a quality deviation feature vector set. This vector set takes into account the interaction features between the competency structure and the hierarchical structure, and is used for subsequent deviation clustering and anomaly detection.
[0164] Step S52: Perform cluster diagnostic analysis on the quality deviation feature vector set to identify imbalance points and deviation areas, and generate a learning quality assessment report;
[0165] In this embodiment, the density-based DBSCAN clustering algorithm is used to identify and analyze abnormal clusters in the output quality deviation feature vector set. The cluster radius ε = 0.3 and the minimum number of samples MinPts = 5 are set to identify cluster boundary samples and isolated samples as potential imbalance points. For each cluster center, the Euclidean distance between it and other center points is calculated, and a deviation threshold of 1.5 times the average distance is set to mark the deviation region. Simultaneously, the internal consistency of the clusters is evaluated based on the variance index within each class of samples to determine whether there are groups with significantly inconsistent learning quality. The final output is presented in the form of a structured evaluation report, including the imbalance dimension, typical deviation point examples, feature offset direction, and spatial distribution map, constructing a visualized learning quality evaluation report for use by the path optimization model.
[0166] Step S53: Extract teaching resource features and subject objective requirements features from standardized subject learning data;
[0167] In this embodiment, hierarchical semantic feature extraction is performed on the question field in standardized subject learning data. Teaching resource features are extracted based on attributes such as question type, knowledge point tags, answering time, and accuracy, including question coverage, skill orientation, and resource difficulty level, constructing a teaching resource feature vector set. Secondly, target ability tags and level structures are extracted based on curriculum standards or subject target databases. For example, targets are divided into three levels: knowledge mastery, ability transfer, and comprehensive application. Each level contains several dimensional target items, encoded in one-hot encoding to construct a target requirement feature vector set. The final output feature data is in the form of a two-dimensional matrix, with rows representing each resource / target sample and columns representing different feature dimensions (no less than 20 dimensions), providing a data foundation for the structural input of the path recommendation model.
[0168] Step S54: Construct a path weight recommendation model based on the learning quality assessment report, teaching resource characteristics, and subject objective requirements.
[0169] In this embodiment, a learning quality assessment report is used as the constraint input, and teaching resource characteristics and subject objective requirements are used as semantic input. A weighted multi-objective path modeling strategy is employed to construct a path weight recommendation model. The total score of each path is defined as a linear weighted sum of three parts: resource fit R, objective alignment G, and learning deviation correction potential P. The model form is: Scorepath = α × R + β × G + γ × P; where α, β, and γ are set to [0.3, 0.4, 0.3], and optimized through 10-fold cross-validation. The path search adopts the shortest path method, using the knowledge point mastery graph as the path graph and the path score as the evaluation function, outputting the optimal path set and the candidate path set (top-k paths, k = 5). This recommendation model provides a matching score for each student group profile, which is used for subsequent path optimization and intervention suggestion generation.
[0170] Step S55: Based on standardized subject learning data, conduct student group profile analysis to obtain regional student group profile data, and use the path weight recommendation model to perform path matching and optimization on the regional student group profile data to obtain subject optimization strategies.
[0171] In this embodiment, standardized subject learning data is aggregated to the student level, and features such as age, gender, region, score, behavior frequency, and task completion time are extracted. K-means++ clustering analysis is used, with a cluster center number of 8, and regional student groups are divided according to the Euclidean distance minimization strategy. A corresponding group profile label is generated for each group, including key weaknesses, competency structure, and learning stability score. These group profiles are input into the path weight recommendation model constructed in step S54, and path matching is performed for each group. During the matching process, resource distribution density and target level complexity are considered, and the matching path and improvement suggestions for each group are output. Finally, a subject optimization strategy is generated, including: focusing on adjusting competency dimensions, prioritizing resource types, and a target achievement path map, providing a structured strategy reference for regional teaching decisions.
[0172] Optionally, step S52 specifically includes:
[0173] Step S521: Perform principal component analysis and orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum number of dominant dimensions to 6 to extract the dominant deviation dimension factors and generate deviation principal cause tensor data.
[0174] In this embodiment, the quality deviation feature vector set (of shape [N×M], where N is the number of student samples and M is the number of deviation features) is standardized, with the mean of each dimension set to 0 and the standard deviation to 1. Principal component analysis (PCA) is used to reduce the dimensionality of the standardized data, setting a variance contribution rate threshold of 85%, i.e., selecting the top K principal components whose cumulative explained variance is no less than 85%. To control model complexity, the maximum number of dominant dimensions K is limited to ≤ 6, thereby extracting the dominant deviation dimension factors in the generated K-dimensional orthogonal feature space. These dominant dimensions are used as tensor axes, and deviation principal cause tensor data are generated by combining the sample index dimension and the dimension factor dimension. The tensor shape is [N×K], which is used for subsequent density modeling and cluster analysis.
[0175] Step S522: Perform density peak estimation on the principal cause tensor data of the bias, set the minimum number of samples to 10 and the cluster distance threshold to 0.75 to identify potential bias cluster centers and generate a subject imbalance cluster map;
[0176] In this embodiment, the obtained bias principal causal tensor data is used as input samples, and a method based on Density Peak Clustering (DPC) is employed to identify potential cluster centers. A minimum sample size of 10 is set as the minimum unit for local density assessment, and a clustering distance threshold of 0.75 is set to determine the minimum distance condition between candidate cluster centers. Specific operations include: calculating the Euclidean distance matrix between all samples; estimating the local density ρ value of each sample point based on the kernel density estimation method; and simultaneously identifying high-density isolated points by combining the minimum distance δ value. Finally, points that simultaneously satisfy high ρ and high δ are selected as bias cluster centers, and the output is a two-dimensional spatial coordinate map, called the subject imbalance clustering map, reflecting the anomalous clustering trend of the principal causal dimension in the sample space.
[0177] Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the subject imbalance cluster map, and set the number of neighborhoods to 20 to measure the anomaly intensity of each cluster unit in the subject imbalance cluster map, identify the deviation area and generate a local imbalance anomaly distribution map.
[0178] In this embodiment, after obtaining the subject-specific imbalance clustering map, local structural refinement is performed on its boundary regions. The renormalization boundary spacing is set to 0.05, and the continuous distribution points in the map are segmented and sliced according to this spacing to form a distribution grid. Boundary renormalization curves are constructed using the sample density within the grid, and the boundaries of each cluster unit are smoothed by fitting a probability density function. Subsequently, the number of neighborhoods is set to 20, and the K-nearest neighbor algorithm is used to evaluate the local anomaly intensity of each point in the map, defined as the ratio of the density difference between the point and the lowest-density sample in its neighborhood. Samples exceeding a set threshold are identified as local deviation points. The final output is a local imbalance anomaly distribution map, marking all spatial regions identified as density anomalies and their corresponding principal causal dimension indices.
[0179] Step S524: Combining the local imbalance anomaly distribution map and the deviation principal cause tensor data, set the deviation factor weight threshold to 0.6 to calculate the deviation factor of each literacy dimension and generate an imbalance feature weight dataset.
[0180] This embodiment uses a local imbalance anomaly distribution map as a basis to trace the features of each identified anomaly cluster unit back to the influence weights of dimensional factors in the principal cause tensor data of the deviation. To quantify the degree of deviation of each competency dimension in the anomaly, a deviation factor calculation model is constructed using the weight normalization exponent method. A deviation factor weight threshold of 0.6 is set. If the ratio of the average deviation value of a competency dimension in the anomaly region to the total deviation value exceeds this threshold, it is considered a major imbalance factor. The final output is an imbalance feature weight dataset, where each record contains five fields: sample index, imbalance dimension code, deviation direction (positive / negative), deviation intensity (real value), and standardized deviation level (high / medium / low), used for structured modeling of learning quality diagnosis.
[0181] Step S525: Perform semantic reconstruction and structural induction on the imbalanced feature weight dataset, and output a learning quality evaluation report.
[0182] In this embodiment, the imbalanced feature weight dataset undergoes semantic reconstruction. This involves transforming the numerical offset information into easily understandable descriptive labels in the education field, such as "leapfrog ability offset" or "application-level competency disadvantage." Secondly, structured patterns are aggregated based on the imbalance dimensions, such as "deviation from a competency combination centered on collaborative ability" or "the main imbalanced groups are concentrated in the interdisciplinary transfer dimension." Layered summarization is then performed using contextual features such as region and group profiles, ultimately outputting a report. The report includes three parts: 1) a summary table of the main causes of the deviation; 2) a heatmap of the imbalance cluster distribution; and 3) optimization suggestion cards based on competency dimensions (at least two suggestions per dimension). This report constitutes a complete learning quality diagnostic result, providing a basis for subsequent path matching and resource intervention decisions.
[0183] Of particular importance is that step S54 specifically includes:
[0184] Step S541: Perform structured analysis on the learning quality assessment report, extract the deviation direction, deviation magnitude and cluster imbalance labels of each competency dimension, and combine the content difficulty, presentation mode and adaptation type in the teaching resource characteristics to construct a competency deviation regulation demand matrix and generate a regulation demand feature set;
[0185] In this embodiment, the structured results of the learning quality assessment report are processed by extracting fields. A regularized template is used to identify and extract key fields such as "competency dimension name," "deviation direction (positive / negative)," "deviation magnitude (standardized Z-score)," and "clustering label code (e.g., C01, C02)." This information is then used as structural input, combined with the "resource content difficulty score (0-1)," "resource presentation method label (e.g., text, illustration, interactive)," and "adaptation type identifier (individual / group / region)" from the teaching resource feature data. A three-dimensional matrix is constructed using dimension alignment, with the axes representing the competency dimension, resource matching type, and deviation magnitude level, forming a competency deviation control demand matrix. After construction, non-zero values are extracted from the matrix and sorted by deviation magnitude to form a control demand feature set. Example: When the "interdisciplinary transfer" competency dimension deviation is negative and the magnitude is 0.82, and a resource with illustration and content difficulty of 0.6 is matched, the feature item ["interdisciplinary transfer," "negative," 0.82, "illustrated," 0.6, "group"] is constructed.
[0186] Step S542: Semantic vector encoding is performed on the subject target requirements features, the ability level, knowledge dimension and task orientation label of each literacy target are extracted, and semantic association analysis is performed with the regulatory demand feature set to generate a literacy target matching relationship map;
[0187] In this embodiment, for documents requiring subject objectives, the BERT model is used to embed and encode the text descriptions, extracting three types of semantic tags: ability-level tags (such as "application," "integration," and "creation"), knowledge dimensions (such as "conceptual," "procedural," and "implicit transfer"), and task-oriented tags (such as "problem-solving" and "collaborative inquiry"). After encoding, a target semantic vector space is constructed. The regulatory requirement feature set formed in step S541 is also encoded into a regulatory vector, and based on the cosine similarity method, a semantic matching threshold of 0.75 is set to perform pairing analysis on elements in the two vector spaces. A semantic matching record table is established by constructing triples <competency dimension, matching target tag, similarity>. Based on the triple set, a competency target matching relationship graph is further generated, where nodes represent competency dimensions and task tags, edges represent semantic similarity relationships, and edge weights are similarity scores, used to support subsequent path graph construction.
[0188] Step S543: Based on the matching relationship graph between the feature set of regulation needs and the literacy goals, construct a literacy regulation path graph, take each literacy dimension node in the path as a graph node, and take the regulation order between paths as a directed edge relationship to form a literacy regulation path network.
[0189] In this embodiment, a regulation path graph is constructed based on the competency target matching relationship graph generated in the previous step. The competency dimensions in the set of regulation demand features are set as nodes in the graph, and the target with the highest semantic matching degree for each pair is extracted as the endpoint of the regulation path. An edge of the directed graph is constructed, defining the edge direction as "from the competency node to be regulated to the target competency task node." The regulation order is set to be determined by the deviation magnitude, connecting from high deviation to low deviation sequentially, and "path association density" is used as the condition for constructing relay edges. If two competency dimensions have a common target orientation and their deviation directions are consistent, a relay edge is automatically generated to connect them, constructing a "relay regulation path." Finally, a competency regulation path network is formed, with the average node degree controlled within 2.3 to avoid dense path stacking.
[0190] Step S544: Evaluate the importance of the literacy regulation path network and perform structural weighting. Calculate the literacy node weights based on the magnitude of the deviation, node centrality, and resource matching degree of the literacy regulation path network. Simultaneously, calculate the literacy conversion probability and generate a path weight vector set.
[0191] In this embodiment, based on the generated control path network, path weights and conversion probabilities need to be calculated. The deviation magnitude weight factor is set to α = 0.5, the node centrality weight factor to β = 0.3, and the resource matching degree factor to γ = 0.2. Combining these three indicators, the weight score Wi = αZi + βCi + γRi for each competency node is calculated, where Zi is the standardized deviation magnitude, Ci is the PageRank centrality of the node in the control network, and Ri is the average resource matching degree for that dimension. To model the competency conversion potential in the control path, the path Markov jump probability modeling method is used. A conversion probability matrix is constructed through historical learning case paths, defining the conversion probability value between each pair of nodes. The final output is a path weight vector set, where each record contains [starting node, ending node, node weight, edge conversion probability].
[0192] Step S545: Embed the path weight vector set into the regulation path network structure, perform probability regulation ranking and diversity reorganization, construct a literacy regulation path library, and use the regulation effect score and path confidence of the literacy regulation path library as the path recommendation evaluation criteria to generate a path weight recommendation model.
[0193] In this embodiment, the path weight vector set is re-embedded into the original regulation path network, and a literacy regulation path library is generated through re-ranking and structural optimization. A Top-K probabilistic regulation ranking algorithm is used, with K = 5 candidate paths. The path score for each path is calculated as ∑(Wi × Pij), where Wi is the node weight and Pij is the conversion probability of the edges in the path. A path confidence threshold of 0.7 is set, and low-confidence paths are eliminated. Simultaneously, path diversity is reorganized based on the maximum structural diversity criterion (path node overlap < 50%). Each path is accompanied by evaluation indicators including regulation coverage, resource adaptability, and expected path improvement value, ultimately constructing a literacy regulation path library. Based on the dual indicators of path score and confidence, a recommendation evaluation function is defined: R(path) = αScore + βConfidence. The top three recommended evaluation scores are used as the final output, forming a path weight recommendation model for use by the teaching intervention system. The path weight recommendation model is built on a Graph Attention Network (GAT). It uses core competency dimensions as graph nodes and control paths as directed edges, integrating features such as student deviation magnitude, resource suitability, and conversion probability for graph structure encoding and semantic embedding. The model dynamically learns the importance of each node in the path through an attention mechanism and ranks them based on a comprehensive score combining path conversion probability and structural centrality, ultimately generating competency-based control path recommendation results that balance diversity and suitability.
[0194] Optionally, this specification also provides a system for analyzing the constitutive composition of intra-disciplinary balanced combination intervals, used to perform the method for analyzing the constitutive composition of intra-disciplinary balanced combination intervals as described above. This system includes:
[0195] The data acquisition module is used to acquire subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data.
[0196] The combined interval generation module is used to divide combined intervals based on standardized subject learning data to obtain subject combined interval data; and to perform interval structure modeling on the subject combined interval data to obtain the combined interval model.
[0197] The combined interval analysis module is used to compare the distribution of subject-level scores based on the combined interval model to obtain subject-level score deviation data; and to evaluate the differences in subject learning quality based on the subject-level score deviation data to obtain a hierarchical feature dataset.
[0198] The core competency relationship modeling module is used to model the relationship between core competencies in a subject based on standardized subject learning data, and obtain the core competency relationship matrix.
[0199] The quality diagnosis module is used to assess the quality of subject learning based on the subject core competency relationship matrix and hierarchical feature dataset, generate a subject learning quality assessment report, and optimize the learning path of standardized subject learning data based on the subject learning quality assessment report to obtain subject optimization strategies.
[0200] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0201] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for analyzing the constitutive composition of balanced combination intervals within a discipline, characterized in that, Includes the following steps: Step S1: Obtain subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data; Step S2: Divide the subject combination intervals based on standardized subject learning data to obtain subject combination interval data; perform interval structure modeling on the subject combination interval data to obtain the combination interval model; Step S3: Compare the subject-level performance distribution based on the combined interval model to obtain subject-level performance deviation data; evaluate the differences in subject learning quality based on the subject-level performance deviation data to obtain a hierarchical feature dataset. Step S4: Based on standardized subject learning data, perform subject core competency relationship modeling to obtain the subject core competency relationship matrix; Step S4 specifically includes: Step S41: Perform subject element label recognition on standardized subject learning data, extract student scores, question difficulty, practice scores and core competency labels, construct element label vector set, and thus generate a competency structured dataset; Step S42: Perform dimensionality reduction and feature redundancy removal on the structured dataset of literacy, extract highly relevant literacy association features, and generate a concise literacy feature set; Step S43: Construct a tensor model of collaborative performance of literacy indicators based on the simplified literacy feature set, and use the tensor model of collaborative performance of literacy indicators to decompose the performance intensity and interrelationship of core literacy dimensions to generate a literacy interaction feature matrix. Step S44: Construct a multi-dimensional relationship graph based on the competency interaction feature matrix, identify the implicit association structure between the dimensions of each core competency, and generate a subject core competency relationship graph; Step S45: Represent the subject core competency relationship diagram in a structured matrix to obtain the subject core competency relationship matrix; Step S5: Based on the subject core competency relationship matrix and hierarchical feature dataset, conduct a subject learning quality assessment to obtain a subject learning quality assessment report. Then, based on the subject learning quality assessment report, optimize the learning path of the standardized subject learning data to obtain a subject optimization strategy. Step S5 specifically involves: Step S51: Perform joint structural analysis on the subject core competency relationship matrix and hierarchical feature dataset to generate a quality deviation feature vector set; Step S52: Perform cluster diagnostic analysis on the quality deviation feature vector set to identify imbalance points and deviation areas, and generate a learning quality assessment report; Step S53: Extract teaching resource features and subject objective requirements features from standardized subject learning data; Step S54: Construct a path weight recommendation model based on the learning quality assessment report, teaching resource characteristics, and subject objective requirements. Step S55: Based on standardized subject learning data, conduct student group profile analysis to obtain regional student group profile data, and use the path weight recommendation model to perform path matching and optimization on the regional student group profile data to obtain subject optimization strategies.
2. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, Step S1 is as follows: Step S11: Collect multi-source subject learning raw data through the regional teaching management platform, and uniformly encode the multi-source subject learning raw data into a data framework structure to obtain subject learning data; Step S12: Perform missing value detection on the subject learning data, and impute the missing value detection results using the class mode to obtain cleaned data; Step S13: Perform field normalization on the cleaned data, and standardize the normalization results by encoding categorical fields to generate transformed data; Step S14: Perform scale standardization on the transformed data to obtain standardized subject learning data.
3. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, The specific division of the combined intervals mentioned in step S2 is as follows: Student performance characteristics are extracted from standardized subject learning data to obtain student subject performance data. Statistical analysis is then performed on the student subject performance data to obtain standardized total performance data and student performance percentage data. Based on the percentage of students with low grades, the percentage of students with low grades is divided to obtain the percentage of students with low grades. Based on standardized total score data and low score student percentage data, a combination variable is set, where the horizontal axis variable is set to standardized total score data and the vertical axis variable is set to low score student percentage data, and a two-dimensional interactive coordinate system is constructed. The standardized distribution features of the spatial combination analysis basis are extracted, and the mean ±0.67σ in the standardized distribution features is set as the dividing threshold to divide the two-dimensional interactive coordinate system. The horizontal and vertical axes of the two-dimensional interactive coordinate system are divided into three segments to form a two-dimensional combination spatial label matrix. Based on the two-dimensional interactive coordinate system, standardized subject learning data is mapped to the corresponding combination labels in the two-dimensional combination space label matrix, the interval categories are marked, and the number and proportion of samples in each interval are counted to generate subject combination interval data.
4. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, The comparison of subject-level score distribution mentioned in step S3 is as follows: Spatial label aggregation is performed on the interval combination model, and the sample data under the same interval label are hierarchically classified. Combined with the subject combination interval data, student distribution data with combination labels is generated. Based on the distribution of standardized total score data, set interval boundary thresholds for low, medium, and high strata, and perform interval stratification on the combined label student distribution data to generate a hierarchical distribution label matrix. Standardized subject learning data is mapped to student distribution data with combined labels and hierarchical distribution label matrix using dual labels to generate multi-label hierarchical score dataset. The number of samples at each level under different combined labels in the multi-label hierarchical score dataset is normalized and statistically analyzed to obtain hierarchical score frequency matrix. The difference measurement is calculated based on the hierarchical score frequency matrix to evaluate the hierarchical score distribution deviation between different combination labels and obtain subject hierarchical score deviation data.
5. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, The subject learning quality difference assessment mentioned in step S3 specifically includes: The distribution statistics and central tendency analysis of the subject-level performance deviation data are performed, and the mean, range and standard deviation of each combination label corresponding to the level are calculated to generate a deviation distribution statistical matrix. Based on the deviation distribution statistical matrix, a horizontal regional balance analysis is conducted to assess the consistency of performance across different regions at each level and generate regional balance characteristic data. A vertical hierarchical gradient analysis is performed on the deviation distribution statistical matrix to calculate the continuity of performance improvement and the probability of hierarchical transition, thereby generating a hierarchical gradient change model. By integrating regional equilibrium characteristic data and hierarchical gradient change models to conduct a weighted comprehensive evaluation, data on the differences in subject learning quality are obtained. Based on the subject learning quality difference assessment data and the multi-label hierarchical score dataset, student hierarchical features are integrated to obtain a hierarchical feature dataset.
6. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, Step S44 is as follows: Step S441: Map high-dimensional spatial nodes to the competency interaction feature matrix, set the number of core competency dimensions to 8, define each core competency dimension as a graph node, calculate the performance intensity synergy coefficient of student tags in each core competency dimension as the initial edge weight, and generate the initial competency node graph. Step S442: Set the Gaussian kernel scaling factor to 0.5, construct the adjacency matrix for the initial literacy node graph, perform nonlinear transformation on the edge weights between nodes, and generate a higher-order literacy adjacency matrix. Step S443: Set the number of encoding layers to 2, with the number of neurons in each layer being [64, 32]. Perform graph structure encoding based on the literacy high-order adjacency matrix, optimize the semantic discriminability of node representations, and generate a literacy semantic embedding graph. Step S444: Perform structural community detection on the literacy semantic embedding graph, identify hidden association subgroups, and generate a literacy subgroup graph; Step S445: Merge the competency subgroup graphs and reorganize the matrix structure to generate a core competency relationship matrix.
7. The method for analyzing the constitutive composition of balanced combination intervals within a discipline according to claim 1, characterized in that, Step S52 is as follows: Step S521: Perform principal component analysis and orthogonal transformation on the quality deviation feature vector set, set the principal component variance contribution rate threshold to 85%, and set the maximum number of dominant dimensions to 6 to extract the dominant deviation dimension factors and generate deviation principal cause tensor data. Step S522: Perform density peak estimation on the principal cause tensor data of the bias, set the minimum number of samples to 10 and the cluster distance threshold to 0.75 to identify potential bias cluster centers and generate a subject imbalance cluster map; Step S523: Set the reorganization boundary spacing to 0.05, perform boundary reorganization and distribution estimation on the subject imbalance cluster map, and set the number of neighborhoods to 20 to measure the anomaly intensity of each cluster unit in the subject imbalance cluster map, identify the deviation area and generate a local imbalance anomaly distribution map. Step S524: Combining the local imbalance anomaly distribution map and the deviation principal cause tensor data, set the deviation factor weight threshold to 0.6 to calculate the deviation factor of each literacy dimension and generate an imbalance feature weight dataset. Step S525: Perform semantic reconstruction and structural induction on the imbalanced feature weight dataset, and output a learning quality evaluation report.
8. A system for analyzing the constitutive composition of balanced combination intervals within a discipline, characterized in that, For performing the intradisciplinary balanced combination interval constitutive analysis method as described in claim 1, the intradisciplinary balanced combination interval constitutive analysis system comprises: The data acquisition module is used to acquire subject learning data and perform data standardization processing on the subject learning data to obtain standardized subject learning data. The combined interval generation module is used to divide combined intervals based on standardized subject learning data to obtain subject combined interval data; and to perform interval structure modeling on the subject combined interval data to obtain the combined interval model. The combined interval analysis module is used to compare the distribution of subject-level scores based on the combined interval model to obtain subject-level score deviation data; and to evaluate the differences in subject learning quality based on the subject-level score deviation data to obtain a hierarchical feature dataset. The core competency relationship modeling module is used to model the relationship between core competencies in a subject based on standardized subject learning data, and obtain the core competency relationship matrix. The quality diagnosis module is used to assess the quality of subject learning based on the subject core competency relationship matrix and hierarchical feature dataset, generate a subject learning quality assessment report, and optimize the learning path of standardized subject learning data based on the subject learning quality assessment report to obtain subject optimization strategies.
Citation Information
Patent Citations
System and method for measuring subject core accomplishment based on production technology
CN115660469A
Multi-task discipline core accomplishment level mining method and system
CN118410079A