A sparse concept prerequisite relationship prediction method and device based on hypergraph learning
Patent Information
- Application Number
- CN202611198294.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-11
AI Technical Summary
缺陷:仅依赖显式成对先决关系构建图结构,无法表达多概念间的高阶关联;在稀疏场景下节点连接不足,信息传播范围有限,模型表示学习不充分
1、本发明从根源缓解概念关系稀疏问题,结构补全能力显著提升现有技术仅依赖人工标注的显式成对先决关系开展学习,在关系占比不足3%的极稀疏场景下会出现节点孤立、信息传播中断、表示学习失效等问题。本发明通过K-means聚类挖掘概念语义相似性+多归属超图构建,不依赖稀缺标注数据,直接从文档-概念共现信息中自动生成高阶关联结构,利用超边同时连接多个相似概念,为孤立节点补充连通路径,使模型在缺乏显式关系时仍可完成稳定的信息聚合与表示学习,从数据层面彻底改善结构稀疏带来的学习瓶颈。
Smart Images

Figure CN122734573A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of education and artificial intelligence technology, specifically relating to a method and apparatus for predicting sparse concept prerequisites based on hypergraph learning. Background Technology
[0002] With the deep integration of artificial intelligence and education, smart education, online learning platforms, and educational knowledge graphs have become core directions in the construction of educational informatization. Concept prerequisite relation prediction, as a key technology in educational data mining and knowledge graph construction, is mainly used to automatically identify directed dependencies between concepts, providing structural support for applications such as learning path planning, personalized resource recommendation, learning status diagnosis, and intelligent question bank construction. In recent years, the rapid development of technologies such as deep neural networks, graph representation learning, and natural language processing has significantly improved the ability to model and identify concept relationships, and related research has been widely applied in real-world educational scenarios such as MOOCs, university courses, online lecture notes, and subject knowledge bases.
[0003] Currently, public datasets such as MOOCs, University Courses, and LectureBank have become standard benchmarks for predicting concept prerequisites, with related technical approaches mainly falling into two categories: traditional machine learning methods and deep learning methods. However, in real-world educational knowledge networks, the number of concepts is enormous, and manual annotation is extremely costly. Annotated prerequisites typically account for less than 3% of all concept pairs, exhibiting extreme sparsity, uneven distribution, and a lack of connections. This sparsity directly leads to problems in traditional graph neural networks during representation learning, such as limited information propagation, convergent node representations, and decreased generalization ability. This results in existing models exhibiting low prediction accuracy and poor robustness in sparse scenarios, making it difficult to meet the needs of large-scale educational knowledge graph construction and deployment in practical teaching systems.
[0004] Meanwhile, technologies such as hypergraph learning, multi-view fusion, and multi-task learning have shown significant advantages in complex relationship modeling, heterogeneous information fusion, and few-shot learning, but have not yet formed a mature, unified, and feasible technical solution for sparse concept-predetermined relation prediction. The industry urgently needs prediction methods that can effectively alleviate relation sparsity, make full use of higher-order associations, and adaptively fuse multi-source information.
[0005] Existing technologies mainly focus on learning concept prerequisites based on graph neural networks. These technologies can be categorized into four types, all of which have significant drawbacks: 1. Conceptual Relationship Prediction Based on Traditional Graph Neural Networks These methods model explicit prerequisite relationships between concepts using ordinary graphs (binary edges) and learn node representations through networks such as GCN and GAT. Representative works include ConLearn (Sun H et al., 2022, SDM conference) and HGAPNet (Jia C et al., 2021, NAACL conference). Limitations: Relying solely on explicit pairwise prerequisite relationships to construct the graph structure fails to express higher-order associations between multiple concepts; in sparse scenarios, node connections are insufficient, limiting information propagation and resulting in inadequate model representation learning.
[0006] 2. Learning the relationship between heterogeneous graphs and global optimization These methods construct heterogeneous networks containing nodes of various types, such as concepts, documents, and courses, and employ global optimization strategies to improve representation performance. A representative work is GKROM (Zhang M et al., 2025, AAAI conference), which is currently the best-performing baseline method in this field. Limitations: It still relies on ordinary graphs and does not introduce hypergraph structures to mine implicit semantic relationships; the fusion of textual semantic and structural information is relatively simple and lacks an adaptive weighting mechanism; and it does not utilize multi-task supervision signals to alleviate sparsity issues.
[0007] 3. Representation learning based on variational graph autoencoders These methods employ variational inference and multi-head attention to enhance the robustness of graph representations. A representative work is MHAVGAE (Zhang J et al., 2022, WSDM conference). Limitations: They rely on sparse explicit relations for graph modeling, failing to address the sparsity problem of the underlying structure; they do not perform semantic clustering and hypergraph reconstruction of concepts, thus failing to supplement higher-order relational information.
[0008] 4. Sparse optimization based on lightweight graph and learning path These methods improve efficiency by pruning redundant edges and simplifying the graph structure. A representative work is LCPRE (Sun J et al., 2024, CIKM conference).
[0009] Limitations: It only optimizes existing edges and does not actively explore potential semantic similarities between concepts; it does not integrate multi-view features, resulting in a significant performance drop in extremely sparse scenarios.
[0010] In summary, all existing closest techniques suffer from common problems such as reliance on sparse explicit relations, inability to model higher-order associations, coarse fusion of multi-source information, and single supervision signals, making it impossible to stably and accurately predict concept prerequisite relations under conditions of scarce annotations. Summary of the Invention
[0011] To address the problems raised in the background art, a first aspect of the present invention provides a sparse concept prerequisite relation prediction method based on hypergraph learning, comprising: acquiring co-occurrence data of concept nodes and associated resource carriers in a target object set; performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph to represent high-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis; extracting structural features of the concept nodes based on the concept hypergraph, and acquiring textual semantic features of the concept nodes; fusing the structural features and the textual semantic features through an adaptive gating mechanism to generate a concept fusion representation; performing multi-task joint learning based on the concept fusion representation with concept-level prerequisite relation prediction as the main task and resource carrier-level prerequisite relation prediction as the auxiliary task, and outputting the prerequisite relation prediction results between the concept nodes.
[0012] In some embodiments of the present invention, the step of performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph to represent higher-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis includes: converting the co-occurrence data into an association matrix and inputting it into an unsupervised feature mapping module to extract dense concept features; using a clustering algorithm to divide the dense concept features to generate multiple semantic clusters; and using a multi-attribution strategy to assign each concept node to at least one semantic cluster and map each semantic cluster as a hyperedge to construct the concept hypergraph.
[0013] Furthermore, the step of converting the co-occurrence data into an association matrix and inputting it into an unsupervised feature mapping module to extract dense concept features includes: constructing an unsupervised representation network containing an encoder and a decoder; inputting the association matrix as a sparse co-occurrence representation into the encoder, performing feature dimensionality reduction through multi-layer nonlinear transformations, and outputting low-dimensional dense concept features; inputting the dense concept features into the decoder to reconstruct the association matrix, and calculating the reconstruction error between the reconstructed matrix and the original association matrix; and obtaining the dense concept features finally output by the encoder in response to the convergence of backpropagation of the network aimed at minimizing the reconstruction error.
[0014] In some embodiments of the present invention, the step of extracting the structural features of the concept nodes based on the concept hypergraph includes: performing bidirectional information propagation along the nodes and hyperedges in the concept hypergraph to capture local high-order structural information; and introducing a global attention mechanism to establish a global dependency relationship across hyperedges based on the local high-order structural information to output the structural features.
[0015] In some embodiments of the present invention, the step of fusing the structural features and the text semantic features through an adaptive gating mechanism to generate a concept fusion representation includes: projecting the structural features and the text semantic features to a unified dimensional space; calculating a dimensional gating weight matrix based on the aligned features; and performing adaptive weighted fusion of the structural features and the text semantic features according to the gating weight matrix to generate the concept fusion representation.
[0016] In some embodiments of the present invention, the step of performing multi-task joint learning based on the concept fusion representation and outputting the prediction results of the prerequisite relationships between the concept nodes includes: for the concept node pair to be predicted, constructing a composite feature to characterize the directed asymmetry based on the concept fusion representation, the composite feature including its own features, difference features, and element-product features; calculating the concept-level main task loss and the resource carrier-level auxiliary task loss respectively, and calculating the weighted overall loss according to a preset balance coefficient; and outputting the directed prerequisite probability of the concept pair to be predicted in response to the convergence of the joint optimization based on the weighted overall loss.
[0017] A second aspect of the present invention provides a sparse concept prerequisite relation prediction device based on hypergraph learning, comprising: an acquisition module for acquiring co-occurrence data of concept nodes and associated resource carriers in a target object set; a construction module for performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph for representing high-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis; an extraction module for extracting structural features of the concept nodes based on the concept hypergraph, and acquiring textual semantic features of the concept nodes; a generation module for fusing the structural features and the textual semantic features through an adaptive gating mechanism to generate a concept fusion representation; and an output module for performing multi-task joint learning based on the concept fusion representation, with concept-level prerequisite relation prediction as the main task and resource carrier-level prerequisite relation prediction as the auxiliary task, and outputting the prerequisite relation prediction results between the concept nodes.
[0018] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the sparse concept prerequisite relation prediction method based on hypergraph learning provided in the first aspect of the present invention.
[0019] In a fourth aspect, the present invention provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the sparse concept prerequisite relation prediction method based on hypergraph learning provided in the first aspect of the present invention.
[0020] The beneficial effects of this invention are: 1. This invention addresses the root cause of the sparsity problem in concept relationships, significantly improving structural completion capabilities. Existing technologies rely solely on manually labeled explicit pairwise prerequisites for learning, leading to issues such as isolated nodes, interrupted information propagation, and representation learning failure in extremely sparse scenarios where relationships account for less than 3%. This invention utilizes K-means clustering to mine semantic similarity of concepts and constructs a multi-attribute hypergraph. It does not rely on scarce labeled data, but automatically generates high-order association structures directly from document-concept co-occurrence information. Hyperedges simultaneously connect multiple similar concepts, supplementing connected paths for isolated nodes. This allows the model to achieve stable information aggregation and representation learning even in the absence of explicit relationships, fundamentally improving the learning bottleneck caused by structural sparsity at the data level.
[0021] 2. Supports high-order multi-concept association modeling, with structural expressive power far superior to traditional graph neural networks. Traditional graph neural networks can only represent binary dependencies, failing to depict the group-based associations of "same topic, same module, multiple concept dependencies" in educational knowledge. This invention uses hypergraph learning as the core structural modeling tool, explicitly modeling high-order dependencies between multiple concepts through the "node-hyperedge-node" information propagation method. Combined with multi-head self-attention, it further captures global topological relationships across hyperedges, enabling structural embedding to simultaneously include local community features and long-distance dependency information, better reflecting the organizational form of real-world subject knowledge, and significantly improving the richness and rationality of structural representation.
[0022] 3. Adaptive Weighted Fusion of Multi-View Information: This approach offers stronger discriminative power and lower redundancy. Existing methods often employ fixed methods like simple concatenation and summation to fuse structural and textual features, failing to dynamically allocate information weights based on conceptual characteristics. This invention designs a multi-view gating fusion mechanism that uses a learnable gating matrix to finely weight structural and textual embeddings. When structural information is abundant, topological features are emphasized; when textual information is complete, semantic features are emphasized. This avoids biases caused by insufficient information from a single view and eliminates dimensionality explosion and information redundancy resulting from simple concatenation, significantly improving the accuracy and discriminative power of the fused representation.
[0023] 4. Maintaining High Accuracy in Small Sample / Weakly Supervised Scenarios Through Multi-Task Supervision and Transfer Learning: Existing technologies rely solely on concept-level prerequisite relationships for single-task learning, which is prone to overfitting and poor generalization due to scarce annotations. This invention introduces a multi-task learning framework, using document-level prerequisite relationship prediction, which is easier to obtain and has richer samples, as an auxiliary task. Leveraging the inherent transitivity of "document relationships reflecting concept dependencies," it provides additional regularization constraints and supervision signals for the main concept task, achieving cross-task knowledge transfer. This enables the model to converge stably and significantly improve prediction accuracy even under extremely sparse annotation conditions.
[0024] 5. The model exhibits strong robustness and lower dependence on external knowledge bases and high-quality text. Similar deep learning methods heavily rely on complete concept descriptions from external knowledge bases such as Wikipedia; performance drops sharply once textual information is missing. This invention, supported by hypergraph structure enhancement, maintains prediction performance even with incomplete textual information, relying on high-quality structural representations. It maintains stable output even in real-world educational data with low concept coverage and high textual noise, making it more applicable and practical.
[0025] 6. The overall solution is modular, scalable, and stable in training, making it suitable for engineering deployment. This invention adopts a standardized process of "data structuring - feature learning - hypergraph construction - representation fusion - multi-task prediction". The modules are clearly decoupled and can be directly adapted to various educational datasets such as MOOCs, university courses, and lecture notes, without the need to redesign the structure for specific scenarios. At the same time, it improves training stability through strategies such as residual connections, layer normalization, fixed random seeds, and weighted loss. It can converge quickly on ordinary GPU devices and has outstanding advantages such as low cost, easy reproducibility, and suitability for industrial deployment.
[0026] 7. The prediction accuracy and generalization performance are significantly better than the existing state-of-the-art methods. On three publicly available standard datasets, MOOC, University Course, and LectureBank, the present invention achieves an F1 score improvement of up to 7.1% compared to the current state-of-the-art baseline model, and also achieves significant improvements in metrics such as ACC and AUC. Under an extremely sparse setting with a positive-to-negative sample ratio of 1:6, the performance degradation is much smaller than that of existing methods, demonstrating strong sparsity adaptability. It can be directly used for advanced applications such as large-scale educational knowledge graph construction, learning path planning, and intelligent recommendation. Attached Figure Description
[0027] Figure 1 This is a schematic diagram of the basic process of a sparse concept prerequisite relation prediction method based on hypergraph learning in some embodiments of the present invention. Figure 2 This is a schematic diagram of the K-means-based hypergraph construction process in some embodiments of the present invention; Figure 3 This is a schematic diagram of the architecture of a sparse concept prerequisite relation prediction method based on hypergraph learning in some embodiments of the present invention. Figure 4 This is a schematic diagram of a sparsity analysis experiment of a sparse concept prerequisite relation prediction method based on hypergraph learning in some embodiments of the present invention. Figure 5 This is a schematic diagram of ablation experiment results in some embodiments of the present invention; Figure 6 This is one of the schematic diagrams of parameter sensitivity analysis in some embodiments of the present invention; Figure 7 This is a second schematic diagram of parameter sensitivity analysis in some embodiments of the present invention; Figure 8 This is the third schematic diagram of parameter sensitivity analysis in some embodiments of the present invention; Figure 9 This is a schematic diagram of the structure of a sparse concept prerequisite relation prediction device based on hypergraph learning in some embodiments of the present invention. Figure 10 This is a schematic diagram of the structure of an electronic device in some embodiments of the present invention. Detailed Implementation
[0028] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0029] Example 1 refer to Figures 1 to 3 In a first aspect of the present invention, a sparse concept prerequisite relation prediction method based on hypergraph learning is provided, comprising: S100. obtaining co-occurrence data of concept nodes and associated resource carriers in a target object set; S200. performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph to represent high-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis; S300. extracting structural features of the concept nodes based on the concept hypergraph, and obtaining textual semantic features of the concept nodes; S400. fusing the structural features and the textual semantic features through an adaptive gating mechanism to generate a concept fusion representation; S500. performing multi-task joint learning based on the concept fusion representation with concept-level prerequisite relation prediction as the main task and resource carrier-level prerequisite relation prediction as the auxiliary task, and outputting the prerequisite relation prediction results between the concept nodes.
[0030] This invention addresses four core shortcomings in educational knowledge networks: extremely sparse annotation of concept prerequisite relationships, the inability of traditional graphs to model higher-order associations, coarse multi-source information fusion, and a single supervision signal. It proposes an integrated technical solution of "structural completion—multi-view fusion—multi-task collaborative optimization." The overall technical approach is as follows: First, this invention utilizes K-means clustering to mine semantic similarity of concepts from document-concept co-occurrence data, actively constructing a hypergraph to supplement higher-order associations between concepts, thus alleviating the problem of sparse annotation of concept prerequisite relationships at its root. Then, it learns concept structure embeddings through hypergraph convolution combined with a multi-head self-attention mechanism, while simultaneously using a pre-trained language model to obtain semantic embeddings of concept texts. A multi-view gating fusion mechanism is employed to adaptively weight structural and semantic information, generating a more discriminative concept fusion representation. Based on this, a multi-task learning framework is introduced, using document-level prerequisite relationship prediction as an auxiliary task to provide additional constraints and supervision for concept-level prerequisite relationship prediction with sparse supervision signals. Finally, by jointly optimizing the overall loss function, high-precision and robust concept prerequisite relationship prediction in sparse scenarios is achieved.
[0031] It should be noted that this method is implemented in the form of a computer program, and the required hardware and software operating environment is as follows: For hardware, an Intel or AMD 64-bit multi-core CPU is used, paired with an NVIDIA discrete graphics card supporting CUDA acceleration (preferably RTX 3060 or higher), with at least 16GB of RAM and at least 50GB of available storage space for dataset loading, model training, and intermediate result storage. For software, Ubuntu 18.04 and Windows 10 or later operating systems are supported. The system is built on the PyTorch and PyTorchGeometric frameworks, relying on libraries such as scikit-learn, numpy, pandas, transformers, and nltk, and uses bge-large-zh-v1.5 as the text encoder to complete semantic feature extraction.
[0032] In step S100 of some embodiments of the present invention, co-occurrence data of concept nodes and associated resource carriers in the target object set are obtained; Specifically, the input teaching document set D = {d1, d2, ..., d...} m} and the normalized core concept set C={c1,c2,…,c n This paper employs precise word matching and entity linking to achieve accurate alignment between concepts and documents, constructing a binary document-concept association matrix K∈Rm×n, where element k ji =1 indicates the concept c i Appears in document d j In the middle, k ji=0 indicates that the node did not appear. Row and column normalization is performed on the matrix to filter out low-frequency concepts, empty documents, and isolated nodes, transforming unstructured educational resources into standardized structured data. This matrix simultaneously supports the parallel construction of concept hypergraphs and document hypergraphs, preserving the co-occurrence patterns of concepts in teaching documents and providing a unified foundation for subsequent unsupervised feature learning and dual-hypergraph structure modeling. The above steps complete the structuring of heterogeneous educational resources; provide a unified data basis for dual-hypergraph construction; and eliminate noise and invalid nodes.
[0033] Assuming the educational knowledge system includes One teaching document and One core concept. Each concept... With each document The relationship is formalized as a binary relation. Based on this, a document-concept association matrix is defined. for:
[0034] in, Representing concepts With Documents The association relationship, with a value of 1 representing a concept. In the document 0 indicates that it appears in the middle, and 0 indicates that it does not appear.
[0035] The row vectors of this matrix represent the concept inclusion of a single document. For a document... Its corresponding row vector It is n A binary vector is a set of documents in which the positions of the non-zero elements indicate all the concepts covered by the document. Similarly, the column vectors of a matrix represent the distribution of individual concepts within the document set. Its corresponding column vector It is m A two-dimensional vector, where the positions of non-zero elements indicate all documents containing the concept. This matrix representation transforms the original discrete relationships into a structured numerical form, laying the data foundation for subsequent statistical analysis, feature extraction, and cluster modeling.
[0036] In step S200 of some embodiments of the present invention, the step of performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph to represent higher-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis includes: S201. The co-occurrence data is converted into an association matrix and input into the unsupervised feature mapping module to extract dense concept features; Furthermore, the step of converting the co-occurrence data into an association matrix and inputting it into an unsupervised feature mapping module to extract dense concept features includes: constructing an unsupervised representation network containing an encoder and a decoder; inputting the association matrix as a sparse co-occurrence representation into the encoder, performing feature dimensionality reduction through multi-layer nonlinear transformations, and outputting low-dimensional dense concept features; inputting the dense concept features into the decoder to reconstruct the association matrix, and calculating the reconstruction error between the reconstructed matrix and the original association matrix; and obtaining the dense concept features finally output by the encoder in response to the convergence of backpropagation of the network aimed at minimizing the reconstruction error.
[0037] Specifically, a three-layer symmetric autoencoder is constructed. The input layer corresponds to the document-concept matrix dimension, the hidden layer dimensions are 512, 256, and 128 respectively, and the latent space dimension is fixed at 64. Mean squared error is used as the reconstruction loss function, and ReLU activation and Adam optimizer are employed for unsupervised pre-training. This pre-training does not rely on any pre-defined relationships from manual annotations, but learns the latent semantic distribution of concepts solely through co-occurrence information. After training, the encoder parameters are fixed, and the decoder is discarded, mapping the high-dimensional sparse concept-document distribution vector to a 64-dimensional low-dimensional dense feature vector zi. This dense feature is used for both concept and document clustering, achieving dimensionality reduction, denoising, and semantic compaction of the high-dimensional sparse data, providing a high-quality representation for subsequent hypergraph construction.
[0038] Without loss of generality, this application first pre-trains a three-layer symmetric autoencoder. It consists of two parts: an encoder and a decoder. The encoder part processes the input... The sparse binary vector is progressively compressed into a low-dimensional latent space, and the decoder attempts to reconstruct the original input from this low-dimensional representation. Let the encoder function be... The decoder function is ,in For the potential representation dimension, and Let be the learnable parameters of the encoder and decoder, respectively. If the matrix consists of the document distribution vectors of all concepts, then the overall reconstruction process can be described as follows: ,
[0039] The training objective of an autoencoder is to minimize the reconstruction error, that is, to obtain the reconstructed vector. As close as possible to the original input This application uses mean squared error as the loss function: ,
[0040] Parameters are optimized using the backpropagation algorithm. An autoencoder can learn to map the original high-dimensional sparse column vector representation to a low-dimensional dense representation, thereby reducing the feature dimensionality while preserving the main structural information of concepts in the document distribution. After pre-training, the encoder... It can be used as a feature extraction module to extract the original data. 2D sparse vector Convert to 3D dense feature vector .
[0041] In summary, this application employs a two-stage process of pre-training and feature extraction. First, it uses the document distribution vectors of all concepts... The autoencoder is fully pre-trained, and the encoder parameters are fixed. This is used as a feature extractor, enabling the encoder to learn more accurate concept representations from high-dimensional, sparse document-concept inclusion relationships. For each concept... , and its document distribution vector Input encoder, output This is a low-dimensional dense feature representation of the concept: ,
[0042] By introducing an autoencoder-based pre-trained feature extraction method, the problems caused by the original data representation can be alleviated to some extent. Firstly, during the unsupervised training phase, the model can automatically learn latent structural information from large-scale unlabeled document-concept co-occurrence data without relying on limited prior relation annotations. Simultaneously, this process transforms the original high-dimensional sparse binary vectors... Mapped to a low-dimensional dense continuous representation This approach reduces data dimensionality and alleviates sparsity issues. Furthermore, the entire feature extraction process is primarily based on the distribution patterns of concepts within the document, without introducing prior relation annotations as a priori information, making the method more applicable to different data scales and application scenarios.
[0043] The low-dimensional representation obtained after nonlinear mapping by the autoencoder mitigates the impact of the high-dimensional space to some extent, while also reflecting the underlying semantic structure of concepts in the document distribution. Based on this, combining K-means clustering analysis yields more stable and discriminative concept partitioning results, thus providing more reliable support for subsequent hypergraph construction.
[0044] The above steps address the problems of high dimensionality, sparsity, and high noise in the original data by learning latent semantics in an unsupervised manner, and provide unified semantic features for the hypergraph.
[0045] S202. A clustering algorithm is used to divide the dense concept features into multiple semantic clusters; Specifically, using the dense features output by the autoencoder as input, K-means clustering is performed to obtain Q semantic clusters, where Q is adaptively selected between 16 and 32 based on the dataset. A multi-attribution strategy is adopted, assigning each concept to the nearest s = Q / 4 clusters, allowing a concept to belong to multiple hyperedges, which aligns with the multi-topic nature of knowledge. Each cluster is mapped to a hyperedge, constructing a concept hypergraph association matrix H∈{0,1}n×Q, and filtering out small-scale hyperedges with fewer than a minimum threshold τmin=5 to optimize the structure.
[0046] S203. Using a multi-attribution strategy, each concept node is assigned to at least one semantic cluster, and each semantic cluster is mapped as a hyperedge to construct the concept hypergraph.
[0047] Specifically, the document-side hypergraph is constructed simultaneously using the exact same process, forming a parallel structure of concept hypergraph and document hypergraph. Clustering prioritizes the use of cosine distance to measure high-dimensional semantic similarity. The entire hypergraph construction process does not rely on manual annotation of prerequisite relationships, but establishes high-order associations purely based on semantic similarity.
[0048] After completing standard K-means clustering and determining the location of each cluster center, instead of using a hard decision of "nearest is best," a new clustering algorithm is used for each concept. Calculate it to all Distance between cluster centers After sorting these distances in ascending order, select the first... The cluster corresponding to the minimum distance is taken as the belonging cluster of this concept. In this application, we set... This means that each concept is allowed to belong to an average of one-quarter of the semantic clusters. Even if a concept does not directly share a hyperedge with another concept, it may still be indirectly connected through multiple hyperedge paths, thus enriching the structural associations between concepts. This multi-attribution clustering strategy enables concepts to participate in hypergraph modeling through multiple hyperedges, further alleviating the connectivity problem caused by the sparsity of the original prerequisite relationships.
[0049] Through the improved clustering process described above, the originally scattered and abstract concepts are organized into several semantically related sets of concepts with relatively clear boundaries but allowing for some overlap. Each cluster can, to some extent, correspond to a potential knowledge topic, teaching unit, or cognitive module. The concepts within these clusters often exhibit common features in teaching materials or have certain inherent connections in the knowledge structure. These clustering results are further mapped into a hypergraph structure in subsequent stages: each concept cluster... A hyperedge in the corresponding hypergraph The superedge connects the cluster All concept nodes within the hypergraph. Through this process, the similarity of concepts in the document distribution feature space is transformed into topological connections in the hypergraph, thereby establishing higher-order relationships between concepts in the absence of explicit prerequisite relationship annotations, providing key support for alleviating the problem of structural sparsity.
[0050] Based on the improved K-means knowledge entity clustering, a concept group division was obtained. Building upon this foundation, the task at this stage is to transform these semantically related sets of concepts into a formalized hypergraph structure, in order to construct a mathematical representation capable of characterizing higher-order relationships between concepts.
[0051] First, define the concept set as... , which contains all A conceptual entity to be modeled. Based on the clustering obtained in the previous step. Cluster of concepts A corresponding set of hyperedges can be constructed. Each of the super edges With a cluster of concepts Correspondingly, it includes all concept vertices within the cluster: ,
[0052] This mapping process transforms semantic clustering results into a hypergraph structure: a set of concepts formed based on similarity in the feature space corresponds to a hyperedge connecting multiple concept nodes in the hypergraph. In this way, the clustering results and the hyperedge structure are effectively linked, enabling the constructed hypergraph to not only express the relationships between nodes but also reflect the semantic similarity features between concepts. The hypergraph constructed using this modeling approach... It can explicitly express higher-order relationships between concepts, even between two concepts. and There is no direct prerequisite relationship between them, as long as they belong to a common concept cluster. Alternatively, structural connections can be formed in the hypergraph through concept clusters.
[0053] For ease of calculation and analysis, the hypergraph structure is represented as a hypergraph matrix. , where matrix elements The definition of is: ,
[0054] In the correlation matrix In the diagram, rows correspond to concept nodes, and columns correspond to hyperedges. The positions of the non-zero elements in a row indicate the concept. All concept clusters to which it belongs, the first The positions of the non-zero elements in the column indicate the superedge. All the concept nodes that are connected.
[0055] However, the initially constructed hypergraph structure may suffer from significant differences in the quality of different hyperedges. Some clusters, due to the randomness of initialization or the influence of data distribution characteristics, may contain only a small number of concept nodes. Such hyperedges have limited effectiveness in information propagation and may even introduce noise interference. To improve the stability and information efficiency of the hypergraph structure, this application further introduces a hypergraph optimization strategy: setting a minimum hyperedge size threshold. Remove superedges containing fewer than a certain threshold of vertices. Specifically, define the optimized set of superedges as follows: ,
[0056] in Indicates the superedge The number of connected vertices. This is set in the experimental setup. This means that only hyperedges connecting at least 5 concept vertices are retained. This filtering operation not only removes small-scale hyperedges with insufficient information, but also ensures that the final generated hypergraph... More compact and stable.
[0057] Through the above steps, the original sparse document-concept relationships are transformed into a more structurally complete concept hypergraph with relatively clear semantic relationships. In this hypergraph structure, higher-order associations between concepts are represented through hyperedges. Even when annotations are scarce, concept nodes can still obtain richer structural context information by sharing hyperedges, thereby mitigating the impact of data sparsity to some extent.
[0058] In step S300 of some embodiments of the present invention, extracting the structural features of the concept nodes based on the concept hypergraph includes: S301. Perform bidirectional information propagation along the nodes and hyperedges in the conceptual hypergraph to capture local higher-order structural information; S302. Introduce a global attention mechanism to establish a global dependency relationship across hyperedges based on the local high-order structural information, so as to output the structural features.
[0059] Specifically, based on the concept hypergraph, L=6 layers of hypergraph convolution are performed, employing a bidirectional information propagation mode of "node-hyperedge-node," and introducing residual connections with residual coefficients η∈[0,1] to avoid over-smoothing of multiple convolution layers. The outputs of multiple convolution layers are averaged and aggregated to obtain the basic structure embedding, and then an h=8-head multi-head self-attention mechanism is superimposed to model global long-distance dependencies across hyperedges, finally obtaining the concept structure embedding Cgraph containing local high-order structure and global association information. Simultaneously, a structure embedding Dgraph is generated for the document, with the process being completely consistent with that of the concept. Among them, hypergraph convolution is responsible for capturing local high-order structure information, and multi-head attention is responsible for capturing global dependencies; the two complement each other to achieve a comprehensive characterization of structural features.
[0060] The above steps, by aggregating higher-order structural information, overcome the limitation of traditional graphs in only modeling binary relationships; and make up for the defect that local aggregation of hypergraphs cannot cover global dependencies.
[0061] In step S400 of some embodiments of the present invention, the step of fusing the structural features and the text semantic features through an adaptive gating mechanism to generate a concept fusion representation includes: S401. Project the structural features and the text semantic features into a unified dimensional space; Specifically, in the knowledge relationship clustering module, this application has constructed a concept hypergraph. and its correlation matrix This hypergraph organizes semantically similar concepts together in the form of hyperedges, with each hyperedge connecting multiple concept nodes, explicitly characterizing the higher-order semantic relationships between concepts. However, the hypergraph structure itself only records the relationship between nodes and hyperedges. How to transform this structural information into node embeddings that can be used by downstream tasks is the core problem of structural modeling. Building on this, this application further introduces a hypergraph convolutional network, which learns the higher-order structural representation of concepts in the hypergraph structure by aggregating the feature information of concepts within the same hyperedge.
[0062] The core idea of hypergraph convolution is to achieve joint modeling of high-order relationships among multiple nodes through a bidirectional information propagation mechanism of "node-hyperedge-node". Specifically, it defines the vertex degree matrix. and hypermarginality matrix ,in 、 All are column vectors consisting of only 1s. Let the hyperedge weight matrix be... This is the identity matrix. The... The hypergraph convolution operation is defined as follows: ,
[0063] in For the first The conceptual feature matrix of a layer represents the residual connectivity coefficients, used to preserve original feature information and prevent oversmoothing. In the above equation, Normalized information aggregation from nodes to hyperedges has been achieved. The information after hyperedge aggregation is then fed back to the nodes, forming a normalized hypergraph Laplacian operator. Residual term The introduction of residual connections is to alleviate the oversmoothing phenomenon after multiple convolutions. That is, as the number of network layers increases, the node representations tend to become consistent. Residual connections preserve the original features of the current layer, enabling the model to maintain the individual characteristics of nodes while utilizing higher-order neighborhood information.
[0064] By stacking In a multi-layer hypergraph convolution, the structural embedding matrix of all concept nodes can be calculated as the average of the outputs of each layer. ,
[0065] This multi-layer output averaging strategy can, on the one hand, comprehensively utilize the structural information of different granularities captured by shallow and deep convolutions, and on the other hand, effectively reduce the variance in the training process and enhance the stability of the embedding representation compared to using only the output of the last layer.
[0066] However, the information propagation range in hypergraph convolutional networks is limited by the hyperedge connection method, mainly focusing on the aggregation of structural information between nodes within the same hyperedge. Specifically, the aggregation process of hypergraph convolution is usually limited to the set of nodes sharing the same hyperedge. While this local neighborhood-based aggregation mechanism can effectively characterize the group characteristics of nodes within a semantic cluster, it has certain shortcomings in modeling global dependencies between different hyperedges. For example, two concepts, even if they do not belong to the same semantic cluster, may still form indirect connections through multi-hop paths or semantic relevance, and this type of cross-cluster global dependency information is also valuable in predicting prerequisite relationships. To overcome this limitation of local neighborhood-based aggregation, this application further introduces a multi-head self-attention mechanism to model global dependencies between different concepts at the concept level.
[0067] Let the number of attention heads be , No. The calculation of each attention point is as follows: ,
[0068] in and For a learnable projection matrix, The core advantage of multi-head self-attention mechanisms lies in the fact that each attention head independently calculates the attention weights between nodes in a different semantic subspace. The combined effect of these elements enables the model to capture global dependencies between nodes from multiple perspectives. Compared to the local aggregation of hypergraph convolution, self-attention can directly model the interaction between any two concept nodes, without being restricted by hyperedge boundaries, thus effectively supplementing the global semantic information in structural embedding.
[0069] To preserve the original structural information while utilizing multi-head self-attention to model global dependencies, and to mitigate gradient decay issues that may occur during deep network training, this application introduces residual connections into the multi-head attention module. Specifically, after concatenation and linear mapping, the output of multi-head attention is further processed through residual connections and layer normalization to obtain the final concept structure embedding representation. ,
[0070] in This is the output projection matrix. Residual connections ensure that the introduction of multi-head attention does not disrupt the original structural representation, enabling the model to achieve a balance between local structure and global dependencies; layer normalization helps stabilize the training process and accelerate model convergence.
[0071] At the document level, this application employs the same modeling process as conceptual structure embedding generation, within the document hypergraph. and its correlation matrix By applying hypergraph convolution and attention enhancement, document structure embeddings can also be obtained. Through the above design, the structural embedding of concepts and documents not only includes the group characteristics of nodes in the local structure of the hypergraph, but also incorporates global dependency information between nodes, providing a high-quality structured representation foundation for subsequent multi-view fusion.
[0072] S402. Calculate the dimension-wise gated weight matrix based on the aligned features; S403. Based on the gating weight matrix, the structural features and the text semantic features are adaptively weighted and fused to generate the concept fusion representation.
[0073] Specifically, by using linear transformations and nonlinear activation functions, the two views are embedded and projected into the same feature space, achieving feature dimension alignment and representation unification. ,
[0074] ,
[0075] in, , , For learnable parameters, To merge the dimensions of the space. Through this projection step, the embeddings of the two views are mapped to a unified plane. In 3D space, the issue of inconsistency between the dimensions of structural embedding and text embedding is resolved, and the foundation for subsequent element-by-element gating operations is laid. (Adopting...) The activation function can normalize the embedding values of the two views to The interval ensures that the two values are consistent in terms of numerical range, thus avoiding the influence of differences in magnitude on the judgment of the gating network.
[0076] After feature dimension alignment, a sigmoid-based gating mechanism is used to adaptively adjust the weight allocation of each view in the final embedding. The gating vector is calculated as follows: ,
[0077] in, and For gating network parameters, Indicates feature splicing, The sigmoid activation function is used. The gate matrix is... Each element represents the proportion of structural information retained in the corresponding feature dimension. The sigmoid function compresses the gate value to... The range allows for a clear probabilistic interpretation: the closer a value is to 1, the more structural information that feature dimension tends to be preserved; the closer it is to 0, the more semantic information is preserved. This fine-grained gating mechanism enables the model to make independent decisions on different feature dimensions, rather than assigning a single fusion weight to the entire concept.
[0078] The final concept fusion embedding is obtained through gated weighted computation: ,
[0079] in, This represents element-wise multiplication, where This is a matrix of all 1s. Through the weighted calculation in the above formula, the fusion embedding is constrained within the space spanned by the structural view and the text view: when the gate value approaches 1, the fusion embedding mainly comes from the structural view; when the gate value approaches 0, it mainly comes from the text view. This gating mechanism allows the model to adaptively weigh the contribution of structural and semantic information according to the specific contextual characteristics of each concept, avoiding the dimensionality expansion problem caused by simple concatenation.
[0080] Similarly, by applying the same multi-view fusion process to document embedding, the final fused document embedding can be obtained. Document-level fusion and concept-level fusion are completely identical in computational process, but the gating network parameters are learned independently to adapt to the differences in textual and structural features between documents and concepts. Through the aforementioned multi-view gating fusion mechanism, the model can dynamically integrate structural and semantic information to generate more comprehensive and discriminative node representations, laying the foundation for subsequent prerequisite relationship prediction tasks.
[0081] More specifically, the first 400 words of descriptive text for each concept are extracted from Wikipedia. When a term is missing, the concept name is used as a fallback. A 1024-dimensional concept text embedding (Ctext) is generated using a pre-trained model (bge-large-zh-v1.5). For documents, the full text content is used to generate the document text embedding (Dtext). Through linear transformation and tanh activation, the structural and text embeddings are projected onto a unified 128-dimensional fusion space for dimensional alignment. A dimensional gating matrix is calculated based on the Sigmoid function, and adaptive fusion is achieved according to weighting rules. The document embedding undergoes the same gating fusion to obtain Dfused. The gating independently assigns weights to each feature dimension, dynamically determining whether it relies more on structural or semantic information.
[0082] The above steps achieve complementary enhancement of structural and semantic information; avoid information redundancy and dimensional explosion caused by simple splicing; and improve the discriminative power and robustness of concept representation.
[0083] In some embodiments of the present invention, step S500 involves predicting the prerequisite relationships at the concept level as the main task and predicting the prerequisite relationships at the resource carrier level as an auxiliary task. Multi-task joint learning is performed based on the concept fusion representation to output the prediction results of the prerequisite relationships between the concept nodes.
[0084] The step of performing multi-task joint learning based on the concept fusion representation and outputting the prediction results of the prerequisite relationships between the concept nodes includes: S501. For the concept node pair to be predicted, construct a composite feature to characterize the directed asymmetry based on the concept fusion representation, wherein the composite feature includes self-feature, difference feature and element-product feature; Specifically, a nonlinear transformation is performed on the fusion embedding of each concept: ,
[0085] in Representing concepts Fusion embedding, and For learnable parameters, For the hidden layer dimension. The introduction of activation functions brings non-linear expressive power to the model, enabling it to capture complex feature interaction patterns in concept embeddings.
[0086] S502. Calculate the loss of the conceptual-level main task and the loss of the resource carrier-level auxiliary task respectively, and calculate the weighted overall loss according to the preset balance coefficient; Specifically, for concept pairs Construct four different types of interactive features: concepts Features ,concept Features Differences and characteristics Element-wise product characteristics These four feature designs each have their own emphasis: single-concept features preserve the semantic information of the concept itself; difference features characterize the directional offset between two concepts in the embedding space, helping to capture the asymmetry of preconditions; element-wise product features model the dimension-wise interactions between concepts, enabling the discovery of deeper association patterns. These features are concatenated and input into the prediction layer to predict the preconditions of concepts. ,
[0087] in and For prediction layer parameters, It is the sigmoid activation function. Representing concepts yes Predicted probability of prerequisites.
[0088] Document-level prerequisite relation prediction uses the same approach, performing a nonlinear transformation through the fused embedding representation of the document: ,
[0089] in Document The fusion and embedding. For document pairs The predicted probability is calculated as follows: ,
[0090] Document-level prediction and concept-level prediction maintain symmetry in their prediction function structures. This not only unifies the model design but, more importantly, enables the two tasks to form consistent constraints at the feature learning level through the same interactive feature modeling approach, which is beneficial for joint optimization.
[0091] Both tasks use binary cross-entropy loss as the loss function. Let the concept-level training sample set be... ,in If the true label is represented, then the concept-level prediction loss is: ,
[0092] Similarly, document-level prediction loss Can be trained on document-level sample sets The calculation yielded the result. The binary cross-entropy loss function is suitable for binary classification tasks, and its gradient properties can effectively drive the model to polarize the predicted probabilities of positive and negative samples.
[0093] Ultimately, the overall loss function of the model is a weighted sum of the losses from the two tasks: ,
[0094] in Used to balance the relative weights of concept-level and document-level prediction tasks during training. When When the model degenerates into single-task learning, it only optimizes concept-level predictions; when At that time, only document-level predictions are optimized. By adjusting... The value of can control the degree of influence of the two tasks on model optimization, ensuring that the model focuses on the concept-level prediction task while making full use of the supervision information provided by the document-level prediction auxiliary task.
[0095] S503. In response to the convergence of the joint optimization based on the weighted total loss, output the directed prerequisite probabilities of the concept pair to be predicted.
[0096] More specifically, the primary task is concept-level prerequisite relation prediction, while the secondary task is document-level prerequisite relation prediction, with both tasks sharing all underlying network modules. Self-features, difference features, and element-wise product features are constructed for concept pairs / document pairs to explicitly model the directed asymmetry of prerequisite relations. Binary cross-entropy is used to calculate the concept-level loss Lc. BCE and document-level loss Ld BCE, the total loss is obtained by weighting the values by the balance coefficient α∈[0,1]: L=α Lc BCE+(1 α) Ld BCE. Trained using the Adam optimizer with a batch size of 256 and weight decay of 5×10⁻⁶. 4. A fixed random seed of 25 ensures reproducibility. Utilizing the richer annotation features of document-level tasks, auxiliary supervision signals are provided to alleviate the sparsity problem of concept-level relationships, thereby improving the model's generalization ability and robustness. The final output is the directed prior probability p(ci,cj) of the concept pair (ci,cj).
[0097] The above steps utilize auxiliary supervision to improve the performance of sparse master tasks, alleviate overfitting, explicitly model directed dependencies, and achieve stable prediction in extremely sparse scenarios.
[0098] refer to Figures 4 to 8To evaluate the effectiveness, robustness, and sparsity mitigation performance of the proposed SCPR-HM in the concept prerequisite relation prediction task, this chapter conducts a series of experiments, including baseline model comparison, significance testing, sparsity analysis, ablation experiments, and parameter sensitivity analysis. Three representative public datasets were selected for this experiment: MOOC, LectureBank, and UniversityCourse. These datasets all contain prerequisite relations between concepts, documents, and manually annotated concepts and documents, and are commonly used data sources for concept prerequisite relation prediction tasks.
[0099] This section fully demonstrates the specific implementation process and technical effects of the present invention under different datasets, different sparsity levels, and different module configurations through six typical embodiments. All embodiments are based on a unified software and hardware environment, with clear steps, well-defined parameters, and reproducible effects, fully proving the universality, stability, and superiority of the present invention in the task of predicting sparse concept prerequisite relationships.
[0100] Referring to Table 1, in one specific embodiment, concept prerequisite relationship prediction is performed based on a MOOC dataset: Table 1
[0101] This embodiment selects a publicly available MOOC dataset with a concept relation sparsity rate of approximately 2.1% for experimentation. This dataset contains 382 teaching documents, 406 concepts, and 1008 sets of labeled prerequisite relations, serving as a typical benchmark for sparse relation prediction in online education scenarios. The implementation process sequentially completes the following steps: constructing a document-concept association matrix, training a three-layer autoencoder to generate 64-dimensional dense features, constructing a multi-attribution hypergraph using K-means clustering, extracting structural embeddings using hypergraph convolution combined with multi-head attention, fusing text semantic encoding with multi-view gating, and performing multi-task joint prediction. The clustering number is set to 24, the hypergraph convolution layer number to 6, the attention head number to 8, and the multi-task balance coefficient to 0.7. Experimental results show that the model in this embodiment achieves an ACC of 0.9154, an AUC of 0.9648, and an F1 score of 0.9119. Compared to the current state-of-the-art methods, the F1 score is improved by 7.1%, fully demonstrating that this invention can achieve high-precision and high-reliability concept prerequisite relation prediction on typical sparse educational data such as online courses.
[0102] Referring to Table 2, in one specific embodiment, concept prerequisite relation prediction is based on the UniversityCourse dataset. Table 2
[0103] This embodiment uses the UniversityCourse dataset, which has a more complex knowledge structure and covers a wider variety of institutions and course types. It contains 654 course documents and 407 concepts, suitable for verifying the model's structural modeling and prediction capabilities under complex knowledge systems. While maintaining the overall technical solution, the number of clusters is set to 24, the number of hypergraph convolutional layers to 6, and the learning rate to 5e-4, completing the entire process of document-concept matrix construction, hypergraph generation, multi-view fusion, and multi-task learning. The final model achieves ACC=0.8861, AUC=0.9232, and F1=0.8856, representing a 5.3% improvement in the F1 score compared to existing mainstream methods. This demonstrates that the present invention can still efficiently capture high-order association information and maintain excellent predictive performance in scenarios like university courses, where knowledge levels are high and concept relationships are complex.
[0104] Referring to Table 3, in one specific embodiment, concept prerequisite relation prediction is performed based on the LectureBank dataset: Table 3
[0105] This embodiment uses the LectureBank dataset for robustness verification. Approximately 20% of the concepts in this dataset lack complete Wikipedia descriptions, effectively testing the model's stability in situations with missing textual information. While maintaining consistency with the overall process and core parameters, adaptation is only applied to the text encoding stage, still achieving hypergraph structure enhancement, multi-view feature fusion, and multi-task optimization. Experimental results show that the model's ACC is 0.8182, AUC is 0.8768, and F1 is 0.8308, maintaining stable output even with missing text for some concepts. This demonstrates that this invention, with hypergraph structure enhancement as its core, can significantly reduce dependence on external textual knowledge bases and possesses stronger adaptability to real-world scenarios.
[0106] Referring to Tables 4 to 6, the results of the significance test are shown below: Table 4
[0107] Table 5
[0108] Table 6
[0109] To verify the statistical reliability and robustness of the model performance, this application conducted significance tests on three datasets—MOOC, UniversityCourse, and LectureBank—with five different random seeds, comparing SCPR-HM with the optimal baseline model, GKROM. The results show that on the MOOC dataset, SCPR-HM significantly outperforms GKROM in terms of mean ACC, AUC, and F1 scores, with p-values less than 0.01 for all metrics, and AUC reaching a highly significant level of p < 0.001. On the UniversityCourse dataset, SCPR-HM also achieved significant improvements in ACC and F1 scores (p < 0.01) and AUC reached a significant level of p < 0.05. However, on the LectureBank dataset, the differences between the two models in all metrics did not reach a statistically significant level (p > 0.05). In summary, SCPR-HM's performance improvement on datasets with complete data and sufficient textual information is statistically significant, and the results are stable and highly reproducible across multiple experiments.
[0110] refer to Figure 4 In one specific embodiment, concept prerequisite relation prediction in an extremely sparse scenario: This embodiment sets five levels of positive-to-negative sample ratios (1:2, 1:3, 1:4, 1:5, and 1:6) to progressively increase data sparsity, simulating the extremely scarce annotation environment in real educational knowledge graphs. While maintaining the model structure and key parameters unchanged, training and testing are performed for each sparsity level. Results show that as data sparsity increases, the performance decline of the model in this invention is far less than that of existing traditional methods. Even under the extreme sparsity condition of 1:6, the model still maintains a high F1 score, and the prediction results are stable and reliable. This fully verifies that the technical approach of this invention, which supplements implicit associations through hypergraph construction and enhances supervision signals through multi-task transfer, can fundamentally improve the robustness and generalization ability of the model in extremely sparse scenarios.
[0111] refer to Figure 5 In one specific embodiment, a core module ablation control experiment was conducted: This embodiment employs the controlled variable method to conduct ablation experiments, sequentially removing the core innovative modules of the invention to verify the necessity and contribution of each module. Experimental results show that removing the K-means clustering-based hypergraph construction module reduces the model's F1 score by at least 5.2%, indicating that the hypergraph structure is crucial for alleviating data sparsity and enhancing structural expressiveness. Removing the multi-view gating fusion module reduces the model's F1 score by at least 2.8%, demonstrating that adaptive fusion significantly improves feature representation quality. Removing the multi-task learning module reduces the model's F1 score by at least 3.5%, indicating that document-level auxiliary tasks effectively enhance supervision signals and improve model accuracy. Overall, the ablation results demonstrate that the three modules—hypergraph construction, multi-view gating fusion, and multi-task learning—work synergistically and are indispensable, jointly supporting the invention's superior performance in sparse scenarios.
[0112] refer to Figures 6 to 8 Parameter sensitivity analysis experiments: To verify the robustness and optimal configuration of key hyperparameters of the SCPR-HM model, this application conducted univariate parameter sensitivity tests on three datasets: MOOC, LectureBank, and UniversityCourse, targeting the number of hypergraph convolutional layers, learning rate, number of relation clusters, and distance metric. Experimental results show that model parameters need to be adapted to the characteristics of the dataset: on the MOOC dataset, the performance improves with increasing depth; on the LectureBank dataset, 6 layers are optimal, but overpropagation can easily introduce noise; the UniversityCourse dataset performs best at medium depth; the learning rate is 5×10⁻⁶. - ³、1×10 - ³、5×10 -4 The system achieves optimal results on three datasets. The number of clusters exhibits a pattern of "too few clusters are insufficient, too many are redundant," with MOOC and LectureBank showing the best results at 24 and 28 clusters, respectively. Regarding distance metrics, cosine distance is more suitable for MOOC and UniversityCourse datasets, while Euclidean distance performs better for LectureBank due to missing textual information. Overall, SCPR-HM demonstrates stable performance within a reasonable parameter range, and the framework exhibits good practicality and robustness.
[0113] It is understandable that this invention uses the principle of structural completion to solve the core problem of sparse annotation of concept prerequisite relationships at the data level. Compared with the shortcomings of traditional methods that rely solely on explicit sparse pairwise relationships and thus cause information propagation breaks, this invention combines K-means clustering with hypergraph construction to connect semantically similar concepts through hyperedges, actively supplementing the implicit higher-order associations between concepts, and enabling originally isolated concept nodes to obtain effective connectivity paths. This solves the problem of insufficient representation learning caused by the sparsity of knowledge networks from the underlying structure.
[0114] This invention enhances the completeness and discriminative power of concept representation based on the principle of multi-view complementarity. Structural embedding provides knowledge network topology and concept group affiliation information, while text embedding represents the semantic connotation of the concept itself. Furthermore, a multi-view gating fusion mechanism dynamically selects more reliable information sources and adaptively allocates weights to achieve efficient complementarity between structural and semantic information. This effectively overcomes the shortcomings of single-view information and makes concept representation more comprehensive and accurate.
[0115] This invention utilizes the principle of multi-task transfer to compensate for the scarcity of supervision signals. Relying on the inclusion and dependency transitive relationships between concepts and documents, it uses document-level prerequisite relationships, which have lower annotation costs and richer sample sizes, as auxiliary supervision signals to regularize and guide concept representation learning. This enables positive knowledge transfer from weak supervision to strong supervision, allowing the model to still converge stably and predict efficiently even under extremely sparse conditions where the annotation relationship accounts for less than 3%.
[0116] Example 2 refer to Figure 9 In a second aspect, the present invention provides a sparse concept prerequisite relation prediction device 1 based on hypergraph learning, comprising: an acquisition module 11, configured to acquire co-occurrence data of concept nodes and associated resource carriers in a target object set; a construction module 12, configured to perform feature mapping and cluster analysis on the co-occurrence data, and construct a concept hypergraph for representing high-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis; an extraction module 13, configured to extract structural features of the concept nodes based on the concept hypergraph, and acquire textual semantic features of the concept nodes; a generation module 14, configured to fuse the structural features and the textual semantic features through an adaptive gating mechanism to generate a concept fusion representation; and an output module 15, configured to perform multi-task joint learning based on the concept fusion representation, with concept-level prerequisite relation prediction as the main task and resource carrier-level prerequisite relation prediction as the auxiliary task, and output the prerequisite relation prediction results between the concept nodes.
[0117] Furthermore, the construction module 12 includes: an extraction unit, used to convert the co-occurrence data into an association matrix and input it into an unsupervised feature mapping module to extract dense concept features; a generation unit, used to divide the dense concept features using a clustering algorithm to generate multiple semantic clusters; and an allocation unit, used to use a multi-attribution strategy to allocate each concept node to at least one semantic cluster and map each semantic cluster as a hyperedge to construct the concept hypergraph.
[0118] In one specific embodiment of the present invention, it includes: a document-concept association matrix construction module: converting the original educational text into a structured co-occurrence matrix; Autoencoder feature dimensionality reduction module: performs nonlinear dimensionality reduction on high-dimensional sparse matrices to obtain dense concept features; K-means clustering hypergraph construction module: Clusters based on semantic similarity to generate concept hypergraphs and hyperedges; Multi-view embedding generation and gated fusion module: generates structural embeddings and text embeddings respectively and then adaptively fuses them; Multi-task learning and prediction module: Jointly optimize the prediction of concept-level and document-level prerequisite relationships.
[0119] Example 3 refer to Figure 10 A third aspect of the present invention provides an electronic device comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of the first aspect of the present invention.
[0120] Electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An input / output (I / O) interface 505 is also connected to bus 504.
[0121] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, hard disks; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 10 Each box shown can represent a device or multiple devices as needed.
[0122] Specifically, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by a processing device 501, it performs the functions defined in the methods of embodiments of this disclosure. It should be noted that the computer-readable medium described in embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more computer programs, which, when executed by the electronic device, cause the electronic device to: Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, C++, and Python—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A sparse concept prerequisite relation prediction method based on hypergraph learning, characterized in that, include: Obtain co-occurrence data of concept nodes and associated resource carriers in the target object set; Feature mapping and cluster analysis are performed on the co-occurrence data, and a concept hypergraph is constructed based on the semantic clusters generated by the cluster analysis to represent higher-order associations of multiple concepts. Based on the concept hypergraph, the structural features of the concept nodes are extracted, and the textual semantic features of the concept nodes are obtained; The structural features and the textual semantic features are fused using an adaptive gating mechanism to generate a concept fusion representation; With concept-level prerequisite relationship prediction as the main task and resource carrier-level prerequisite relationship prediction as the auxiliary task, multi-task joint learning is performed based on the concept fusion representation to output the prerequisite relationship prediction results between the concept nodes.
2. The sparse concept prerequisite relation prediction method based on hypergraph learning according to claim 1, characterized in that, The step of performing feature mapping and cluster analysis on the co-occurrence data, and constructing a concept hypergraph to represent higher-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis, includes: The co-occurrence data is converted into an association matrix and input into the unsupervised feature mapping module to extract dense concept features; Clustering algorithms are used to divide the dense concept features into multiple semantic clusters; A multi-attribution strategy is adopted to assign each concept node to at least one semantic cluster, and each semantic cluster is mapped to a hyperedge to construct the concept hypergraph.
3. The sparse concept prerequisite relation prediction method based on hypergraph learning according to claim 2, characterized in that, The step of converting the co-occurrence data into an association matrix and inputting it into an unsupervised feature mapping module to extract dense concept features includes: Construct an unsupervised representation network containing an encoder and a decoder; The correlation matrix is used as a sparse co-occurrence representation and input into the encoder. Feature dimensionality reduction is performed through multi-layer nonlinear transformation to output low-dimensional dense concept features. The dense concept features are input into the decoder to reconstruct the association matrix, and the reconstruction error between the reconstructed matrix and the original association matrix is calculated. In response to the convergence of network backpropagation aimed at minimizing the reconstruction error, the dense concept features of the encoder's final output are obtained.
4. The sparse concept prerequisite relation prediction method based on hypergraph learning according to claim 1, characterized in that, The extraction of structural features of concept nodes based on the concept hypergraph includes: Information is propagated bidirectionally along nodes and hyperedges in the conceptual hypergraph to capture local higher-order structural information; A global attention mechanism is introduced to establish a global dependency relationship across hyperedges based on the local high-order structural information in order to output the structural features.
5. The sparse concept prerequisite relation prediction method based on hypergraph learning according to claim 1, characterized in that, The step of fusing the structural features and the textual semantic features through an adaptive gating mechanism to generate a concept fusion representation includes: Project the structural features and the text semantic features into a unified dimensional space; Calculate the dimension-wise gated weight matrix based on the aligned features; Based on the gating weight matrix, the structural features and the text semantic features are adaptively weighted and fused to generate the concept fusion representation.
6. The sparse concept prerequisite relation prediction method based on hypergraph learning according to claim 1, characterized in that, The step of performing multi-task joint learning based on the concept fusion representation and outputting the prediction results of the prerequisite relationships between the concept nodes includes: For the concept node pair to be predicted, a composite feature for characterizing directed asymmetry is constructed based on the concept fusion representation. The composite feature includes self-feature, difference feature and element-product feature. Calculate the loss of the main task at the concept level and the loss of the auxiliary task at the resource carrier level separately, and calculate the weighted total loss according to the preset balance coefficient; In response to the convergence of the joint optimization based on the weighted total loss, the directed prerequisite probabilities of the concept pairs to be predicted are output.
7. A sparse concept prerequisite relation prediction device based on hypergraph learning, characterized in that, include: The acquisition module is used to acquire co-occurrence data of concept nodes and associated resource carriers in the target object set; The construction module is used to perform feature mapping and cluster analysis on the co-occurrence data, and to construct a concept hypergraph to represent higher-order associations of multiple concepts based on the semantic clusters generated by the cluster analysis. The extraction module is used to extract the structural features of the concept nodes based on the concept hypergraph and obtain the textual semantic features of the concept nodes; The generation module is used to fuse the structural features and the text semantic features through an adaptive gating mechanism to generate a concept fusion representation; The output module is used to perform multi-task joint learning based on the concept fusion representation, with concept-level prerequisite relationship prediction as the main task and resource carrier-level prerequisite relationship prediction as the auxiliary task, and output the prerequisite relationship prediction results between the concept nodes.
8. The sparse concept prerequisite relation prediction device based on hypergraph learning according to claim 7, characterized in that, The building module includes: The extraction unit is used to convert the co-occurrence data into an association matrix and input it into the unsupervised feature mapping module to extract dense concept features; The generation unit is used to divide the dense concept features using a clustering algorithm to generate multiple semantic clusters; An allocation unit is used to employ a multi-attribution strategy to assign each of the concept nodes to at least one of the semantic clusters, and to map each of the semantic clusters as a hyperedge to construct the concept hypergraph.
9. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the sparse concept prerequisite relation prediction method based on hypergraph learning as described in any one of claims 1 to 6.
10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the sparse concept prerequisite relation prediction method based on hypergraph learning as described in any one of claims 1 to 6.