A multimodal preference driven graph convolution combined optimization learning path generation method
By generating learning paths through multimodal feature fusion and graph convolution optimization, the adaptability and executability issues of learning path generation in online education are solved, achieving dynamic adjustment and efficient learning path planning.
Patent Information
- Application Number
- CN202511434701.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing online education learning path generation methods fail to effectively consider students' multiple preferences, changes in learning time, and the order of knowledge points, resulting in generated paths that are difficult to adapt to students' dynamic learning needs in practical applications, and the algorithms are costly and have poor execution.
A multimodal preference-driven graph convolutional ensemble optimization learning path generation method is adopted. Through multimodal feature fusion, graph convolutional feature update and integer programming, an optimized learning path that takes into account interest, time and knowledge coverage is generated, and dynamic reprogramming is supported.
It improves the adaptability and executability of the learning path, reduces algorithm costs, ensures that the path can be dynamically adjusted during the student's learning process to adapt to changes in time and goals, and improves learning outcomes.
Smart Images

Figure CN120894204B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of learning path generation, and particularly relates to a multi-modal preference driven graph convolution combined optimization learning path generation method. BACKGROUND
[0002] At present, the resource types of online education are various, which include video courses, test papers, study plans, test questions and the like, and there is often a prerequisite dependency relationship (i.e. based on the learning order of knowledge before and after) among these resources. Among them, students at different stages have obvious differentiated preferences for the presentation form, explanation method and difficulty of learning resources. When students face different learning goals (such as score improvement, leakage checking and supplement, and high school entrance examination rush), the available learning time of students also changes from time to time, which makes it more complex to generate a learning path that can take into account interest and the ability of students to complete.
[0003] For example, the disclosed technology with the publication number CN120386941A and the name of a personalized learning content recommendation method and system based on neighborhood and hypergraph cooperation, which constructs a hypergraph view and a neighbor graph view by obtaining user session sequence data; extracts global high-order relationship features based on the hypergraph view; extracts local co-occurrence relationship features based on the neighbor graph view, generates enhanced co-occurrence relationship features; concatenates the sequence position encoding of the user learning path and the learning content semantic embedding, and fuses the context information through a gated dynamic attention mechanism to generate a dynamic learning interest embedding representation. The focus of this technology is on personalized learning content recommendation based on interest, and the core goal is to pick out the most suitable single learning content without complex and multiple considerations of the order before and after. The algorithm goal of this technology is not to develop a learning route, and the content that students are most interested in.
[0004] In addition, in the disclosed technology with the publication number CN120355539A and the name of an online teaching optimization method and system based on emotion recognition, multi-modal sensors are used to collect emotion data, and then relevant data and algorithms are used to generate teaching strategies, and an improved genetic algorithm and a knowledge graph are used to optimize the learning route, and real-time adjustment is performed in cooperation with hierarchical teaching control. This technology does not consider the constraints and combinations of the order before and after the key knowledge points (out of class), the time invested in learning, the importance of the generated tasks, the completeness of the knowledge points and the coverage degree. At the same time, this technology does not construct a general linear feasible region model, and does not have a clear path-level cost function. Here, the completeness of the knowledge points requires that each knowledge point is at least involved once by the learning resources; and the coverage degree refers to how many knowledge points are covered on the basis of the completeness of the knowledge points.
[0005] In "CN120354883A, the name is based on hypergraph neural network and knowledge tracking adaptive learning path recommendation method", it determines the association between learning resources as the edge of the learning resource undirected graph; the characteristics of the learning resources are used as the embedding feature vectors of each node of the learning resource undirected graph; the embedding feature vectors are updated by using the graph neural network to obtain the resource embedding vector, which is used as the node feature of the learner hypergraph structure; the hypergraph neural network is used to iteratively aggregate the learner hypergraph structure to obtain the dynamic resource embedding and the learner behavior sequence, generate an initial recommendation list, further generate a candidate learning path set, and use the non-dominated sorting genetic algorithm II to generate a Pareto front solution set; based on the dynamic weight distribution strategy and the comprehensive utility function, the comprehensive score of each path in the solution set is calculated to determine the optimal learning path. The disadvantages of this technology are that the constraints are soft, the executability is discounted, and the key conditions such as time length, coverage, task number and quality are mainly adjusted by weight penalty, which belongs to "soft constraint". Therefore, the generated path looks reasonable, but it is not easy to grasp the result whether the algorithm can be completed on time and in quantity when the software platform is used. There is a lack of online re-mechanism, which is a key defect. Because students often learn halfway through in reality, due to changes in learning time, the need for temporary practice, or adjustments to stage goals, the plan needs to be adjusted, but there is no online re-mechanism in the scheme. The adaptability is weak in actual application. The hypergraph and multi-objective evolutionary algorithm have high computational overhead and require heavy parameter tuning. Different paths may be generated for the same input, and the cost of computing power is high when the project is implemented. Therefore, the lack of online re-mechanism needs to be explained: Generally, an offline fixed learning path is generated. Before learning begins, a recommended path is calculated by the hypergraph neural network and evolutionary algorithm. Once the path is determined, even if the actual learning situation, time arrangement, and target requirements of the student change, the algorithm will not automatically recalculate. This scheme belongs to a one-time offline static plan. However, in actual application, students often encounter temporary practice, reduced learning time, or target adjustment. If the path cannot be adjusted in real time according to these changes, it is easy to cause the path to be unable to complete and the learning effect to decline.
[0006] Therefore, it is urgent to propose a multi-modal preference driven graph convolution combination optimization learning path generation method with simple logic and adaptability. SUMMARY
[0007] To solve the above problems, the purpose of the present application is to provide a multi-modal preference driven graph convolution combination optimization learning path generation method. The technical solution adopted by the present application is as follows:
[0008] A multi-modal preference driven graph convolution combination optimization learning path generation method, comprising the following steps:
[0009] The multi-modal learning feature data is acquired, and multi-modal feature data fusion processing is performed to obtain a resource initial feature matrix.
[0010] A multi-relation adjacency matrix is constructed for the learning resource initial feature matrix, and normalization processing is performed to construct a normalized propagation matrix.
[0011] According to the normalized propagation matrix, the graph convolution feature is updated, and a learning gain score is obtained.
[0012] The comprehensive utility is calculated based on the learning gain score.
[0013] According to the score of the comprehensive utility, an integer programming is established under multiple constraint conditions to screen and obtain optimized candidate resources.
[0014] An optimal learning path is generated in the optimized candidate resources by using an edge cost function and a learning path target function; in the process of generating the optimal learning path, dynamic re-planning is adopted; the dynamic re-planning includes mastery degree updating, coverage demand decreasing, and time length budget updating.
[0015] Compared with the prior art, the present application has the following beneficial effects:
[0016] The present application can enhance the performance of the model when capturing the multi-dimensional dependence of the learning resources in the graph convolution propagation, and reduce the dilution problem in multi-layer propagation. Through normalization processing, the information of high nodes can be prevented from being excessively amplified and the information of low nodes can be prevented from being excessively weakened in the graph convolution multi-layer propagation, so that the stable propagation and balanced diffusion of information in the entire graph structure are maintained, and the stability and convergence speed of GCN learning are improved.
[0017] The present application updates the graph convolution feature and obtains a learning gain score. The graph convolution update with residual can introduce neighbor node information aggregation while preserving the input features, and realize feature update with residual connection. In addition, according to the score of the comprehensive utility, the present application establishes an integer programming under multiple constraint conditions to screen and obtain optimized candidate resources, which are based on multi-dimensional learning features and learning behavior data, dynamically calculate student fitness, and can generate student paths that are more suitable for the ability level and learning rhythm of different students.
[0018] In conclusion, the present application has the advantages of simple logic, reliable multi-modal fusion, and the like, and has high practical value and promotional value in the technical field of learning path generation. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as a limitation to the protection scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0020] Figure 1 The logic flow chart of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further illustrate the present application by combining with the drawings and embodiments, and the embodiments of the present application include but are not limited to the following embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0022] In the present embodiment, the term "and / or" is only used to describe the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone.
[0023] The terms "first" and "second" and the like in the specification and claims of the present embodiment are used to distinguish different objects, and are not used to describe the specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, and are not used to describe the specific order of the target objects.
[0024] In the embodiments of the present application, the words such as "exemplary" or "for example" are used to mean example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "exemplary" or "for example" are intended to present the relevant concept in a specific manner.
[0025] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.
[0026] As Figure 1As shown, the embodiment provides a multi-modal preference driven graph convolution combined optimization learning path generation method, which is based on multi-modal feature fusion, and then uses a multi-relation graph convolution network to update and extract the initial features after fusion, and finally uses combined optimization to select the path, realizes the cooperation of the two stages of offline global optimization and online dynamic adjustment, and constitutes a closed-loop personalized learning path planning that can achieve learning goals. Here, the multi-relation graph convolution feature extraction is combined with the initial features of the resources after constructing various adjacency relationships such as prerequisite relationships, similar relationships, etc., input into the graph convolution network for layer-by-layer propagation and neighbor aggregation, and the original information is maintained through the residual structure. After this process, the updated resource representation can be obtained, which can not only retain the characteristics of the node itself, but also integrate the context information under different relationships, providing a more reliable feature basis for subsequent learning gain calculation and path optimization. Specifically, it includes the following steps:
[0027] First, multi-modal feature fusion and initial feature matrix establishment: that is, obtaining multi-modal learning feature data and performing multi-modal feature data fusion processing to obtain a learning resource initial feature matrix. Here, the features of text, image, and behavior data of learning resources such as courses, test papers, study cases, and tests are fused into a unified vector.
[0028] Here, the expression of multi-modal feature data fusion processing is: (1) Here, the three types of features of image, text, and behavior are mapped to the same dimensional space to generate the initial vector of the learning resource , and the feature format is consistent when processing in the graph structure, making calculation more convenient. In formula (1), represents the initial feature vector of the th learning resource after multi-modal feature extraction and fusion mapping; represents the image feature vector of the th learning resource; represents the text feature vector of the th learning resource; represents the behavior feature vector of the th learning resource; represents the weight matrix adjusting the contribution degree of the mode, which combines the picture, text, and behavior features into a unified feature vector through linear combination; represents the bias vector, which adds an offset to fine-tune the fused features to ensure that the results are more in line with actual needs; is a positive integer.
[0029] Here, the expression of the learning resource initial feature matrix is: (2)
[0030] The processed result of formula (1) is converted into a unified input matrix for graph convolution. In the matrix, each row represents a feature vector of a learning resource, so that the similarity, prerequisite relationship and co-occurrence information between resources can be preserved during graph structure calculation. In formula (2), represents the initial feature vector of a learning resource after multi-modal feature extraction and fusion mapping; represents the initial feature vector of a learning resource after multi-modal feature extraction and fusion mapping; represents the initial feature vector of a learning resource after multi-modal feature extraction and fusion mapping; T represents vector transposition; represents the total number of nodes of the learning resource.
[0031] The second step is relationship fusion and normalization, which includes the establishment of fused adjacency matrix and normalized propagation matrix. Here, a multi-relationship adjacency matrix is constructed based on the initial feature matrix of the learning resource, and normalized processing is performed to construct a normalized propagation matrix. Here, the multi-relationship adjacency matrix (prerequisite, similarity, co-occurrence, etc.) is constructed, fused into a unified relationship graph, and then normalized to ensure stable feature propagation.
[0032] wherein the expression of the multi-relationship adjacency matrix is: (3)
[0033] Here, the prerequisite relationship, similarity relationship and co-occurrence relationship are fused into a unified adjacency matrix according to the weight, and the unit matrix is added to retain the node's own information. In graph convolution propagation, the performance of the model can be enhanced and the dilution problem in multi-layer propagation can be reduced when capturing the multi-dimensional dependence of the learning resource. In formula (3), represents the fused multi-relationship adjacency matrix, which is used as the adjacency weight input for graph convolution calculation; represents the prerequisite relationship adjacency matrix, and the element value represents the weight of the prerequisite dependence relationship between two resources; represents the similarity relationship adjacency matrix, and the element value represents the feature similarity weight between resources; represents the co-occurrence relationship adjacency matrix, and the element value represents the frequency weight of the common occurrence of resources in the learning path or learning record; represents the fusion weight coefficient corresponding to the prerequisite relationship adjacency matrix; represents the fusion weight coefficient corresponding to the similarity relationship adjacency matrix; represents the fusion weight coefficient corresponding to the co-occurrence relationship adjacency matrix; represents the self-loop weight coefficient of the unit matrix, which is used to adjust the proportion of self-feature preservation of each node in graph convolution; represents the unit matrix, and the main diagonal elements are always 1 and the other elements are 0, which is used to preserve the feature influence of each node itself in the fused adjacency matrix.
[0034] In addition, the expression of the normalized propagation matrix is: (4)
[0035] Here, the fused multi-relation adjacency matrix is symmetrically normalized to generate the normalized propagation matrix . Here, through the normalization process, the problem of over-amplification of high-degree node information and over-weakness of low-degree node information during graph convolution multi-layer propagation is prevented, thereby maintaining stable propagation and balanced diffusion of information in the entire graph structure, and improving the stability and convergence speed of GCN learning. In formula (4), where denotes the node degree matrix, which is a diagonal matrix, and the element denotes the degree of the th node; denotes the reciprocal of the square root of each non-zero element of the diagonal matrix, which is used to symmetrically normalize the adjacency matrix and balance the influence of different node degrees on the propagation process.
[0036] The third step is graph convolution update, which includes graph convolution feature update and learning gain scoring:
[0037] Here, a graph convolution update with residual is used, and its expression is: (5)
[0038] Here, formula (5) can be used to introduce neighbor node information aggregation while preserving input features, realizing feature update with residual connection. The residual term can prevent the feature from being over-smoothed in deep GCN, improve the convergence speed and generalization ability, and ensure the contribution of different nodes during information propagation based on the normalized propagation matrix , thereby preventing high-connectivity nodes from excessively affecting the update; the GCN weight matrix of the th layer of graph convolution controls the mapping from the input feature space to the next layer of feature space. In formula (5), denotes the learning resource node feature matrix of the +1th layer of graph convolution; denotes the learning resource node feature matrix of the th layer of graph convolution; denotes a nonlinear activation function; denotes the GCN weight matrix of the th layer of graph convolution.
[0039] In addition, the expression of the learning gain score is: (6)
[0040] Here, based on the final feature vectors of the student and the resource, the knowledge gain score that the student may obtain after using the resource is calculated. To balance the similarity between acquired resources and student features, methods such as inner product, weighted similarity, or MLP can be used, supporting non-linear interaction. Gain score serves as an important input indicator in learning path generation; a higher score indicates that the resource is more suitable for the current student's learning stage and ability level, while a lower score indicates suitability for the current ability level. Formula (6) can accept multimodal inputs (video features, text features, interactive behavior features, etc.) for personalized learning recommendations. In formula (6), Indicates the first The student on the first The learning gain score of each learning resource; The learning gain scoring function can be a differentiable function such as inner product, weighted cosine similarity, or MLP (Multilayer Perceptron), used to measure the matching degree between students and resource features; Indicates the first The final feature vector of each student is the output of multimodal preference modeling; Indicates the first The final feature vector of each learning resource is output by the multimodal resource representation and graph convolution update module.
[0041] The fourth step is personality and goal assessment, which includes learning resource preference score and goal alignment score:
[0042] Here, the expression for the learning resource preference score is: (7)
[0043] Here, formula (7) is the first Find the optimal learning path for each student It achieves this by maximizing the total learning gain ∑ along the path. To improve learning outcomes; here, through cost-penalty items. Controlling the rationality of the learning path prevents students from making incoherent resource choices or selecting excessive resources, thus avoiding an overburdened learning burden. The final path balances learning effectiveness with time and difficulty costs, and is the direct output of the personalized recommendation system. In formula (7), This represents any candidate learning path, consisting of a set of learning resources arranged in prerequisite order. composition; Indicates the first The set of all feasible learning paths for each student, satisfying prerequisite knowledge constraints, progressive difficulty requirements, and time / capacity limitations; This represents the cost penalty coefficient; a larger value indicates greater sensitivity to path cost. The path cost function takes into account: difficulty fluctuation (to avoid excessively large leaps), time cost (total resource duration), and resource redundancy (to avoid redundant learning).
[0044] Furthermore, the expression for the target alignment score is: ; (8)
[0045] Here, formula (8) represents the importance of the target to each knowledge point (the first... Target alignment score vector of learning resources ), based on the covering matrix Aggregated to the resource layer, the matching degree between each resource and the current learning objective is calculated. If a resource has a large number of target knowledge points with "high weight", its... It will be bigger. In the comprehensive utility score (formula (9), it is used as the "objective relevance item" and structural gain) and the students' own preferences They are weighted and fused together for subsequent candidate selection and path optimization. The target requirements are directly converted into resource layer scores to meet hard or soft requirements such as coverage or alignment. In formula (8), The total number of knowledge points; Indicates the first The learning resource for the first The coverage strength of each knowledge point; Indicates the first The target knowledge point weight vector for each knowledge point; This represents the weight vector of the target knowledge points; This represents the resource-level target alignment score vector.
[0046] Step 5, Calculate the overall utility:
[0047] The expression for this overall utility is: (9)
[0048] Here, structural gain score, target alignment score, and individual preference score are used to calculate the matching degree between each resource and the current learning objective. If a resource covers a large number of "high-weight" target knowledge points, its... The larger the value, the more suitable the resource is for the student's learning path. In formula (9), Indicates the first The first student recommended Overall priority of learning resources; This represents the structural gain weighting coefficient, used to control the importance of the graph convolution propagation result in the overall utility; This represents the target alignment weight coefficient, which is used to control the contribution of learning resources to the matching degree between learning targets; represents the individual preference weight coefficient, used to control the proportion of student individual preference score in the comprehensive utility; represents the structure gain score of the th student recommending the th learning resource, derived from residual graph convolution update, reflecting the importance of the resource in the knowledge graph structure and the association strength with the student's existing knowledge state; represents the structure gain score of the th student recommending the th learning resource, derived from residual graph convolution update, reflecting the importance of the resource in the knowledge graph structure and the association strength with the student's existing knowledge state;
[0049] Here, formula (5) outputs: , , . From formula (5), read read and substitute it into formula (9) to get ; where is the read vector, which linearly maps to scalar , and is a learnable parameter.
[0050] Step 6, consider multiple constraint conditions to establish integer programming, filter and obtain optimized candidate resources. Here, it includes learning path objective function, time constraint, knowledge point coverage constraint, task quantity constraint, average quality constraint, etc.
[0051] Specifically, the expression of the learning path objective function is: (10)
[0052] where represents the learning path resource sequence set, represented as a set of learning resource numbers organized in order; represents the weight coefficient of maximizing the total score of comprehensive utility, used to emphasize the overall quality of the learning path; represents the path length penalty coefficient, used to suppress excessively long learning paths and prevent ineffective or redundant learning content; represents the path length.
[0053] In addition, the expression of the time constraint is: (11)
[0054] where formula (11) requires the generated learning path to be within the student's learning time range. Here, the time range is controlled to prevent the recommended learning path from exceeding the student's available time, improving the executability. In addition, personalized matching: combined with the student's schedule, energy allocation, etc., dynamically set. Combined with the objective function: participate in solving as a hard constraint condition in the optimization process of formula (10), ensure that the final output path is both high quality and meets the time limit. In the solving stage, formula (11) will be one of the limiting conditions of the optimization algorithm (such as genetic algorithm, integer programming, branch and bound method, etc.), directly affecting the screening and generation of path feasible solutions. In formula (11), represents the predicted learning duration of the th learning resource; represents the maximum total learning duration limit that the th student can accept.
[0055] In addition, the expression of the knowledge point coverage constraint is: (12)
[0056] Here, it is ensured that the total coverage of the selected resources for each target knowledge point is not less than the preset threshold, thereby meeting the "planned completion, sufficient training" requirement of the syllabus or stage goal. The learning path objective function, duration constraint, knowledge point coverage constraint, task quantity constraint, and average quality constraint form a set of linear constraints that can be directly solved by a MILP solver. In formula (12), represents the selection decision variable of the th learning resource; represents the minimum coverage threshold of the th knowledge point.
[0057] The expression of the task quantity constraint of the present embodiment is: (13)
[0058] Here, formula (13) is used to control the size and learning intensity of the candidate set, ensuring that the path contains at least a sufficient number of tasks to achieve basic learning effects (lower bound ), and avoiding excessive tasks that lead to overload (upper bound ). In formula (13), represents the minimum allowed task quantity lower bound; represents the maximum allowed task quantity upper bound.
[0059] The expression of the average quality constraint of the present embodiment is: (14)
[0060] Here, it is ensured that the overall quality of the selected learning resources meets the minimum standard, preventing optimization from focusing only on quantity or individual high scores while ignoring overall quality. In formula (14), represents the quality score of the th learning resource; represents the average quality threshold, requiring the average quality of the selected resources to be not less than the threshold.
[0061] Step 7, Path Optimization Process, which generates the optimal learning path in the optimized candidate resources using the edge cost function and the learning path objective function. Specifically:
[0062] The expression of the edge cost function is: (15)
[0063] Here, the prerequisite penalty The mandatory guided path follows the prerequisite logic, difficulty jumps Reduce the difficulty of learning span, promote the smooth progression of difficulty, and switch the form Reduce frequent switching between different learning modes, affecting students' attention and increasing operating costs. In the path optimization objective, The penalty term of the adjacent step is accumulated to guide the heuristic (such as the LKH idea) or sequence decision-making in MILP to be smoother and more executable; it can be based on parameter tuning 、 、 Adapt to different teaching scenarios (such as intensive prerequisite and cutoff for sprint classes, and emphasize smooth progression for introductory classes). In formula (15), represents the edge cost from the th learning resource to the th learning resource; represents the prerequisite penalty weight, adjusting the proportion of the cost of violating the prerequisite relationship in the total edge cost; represents the prerequisite violation indication cost from the th learning resource to the th learning resource; represents the difficulty jump weight, controlling the impact strength of difficulty discontinuity on edge cost; represents the cognitive difficulty scalar of the th learning resource; represents the cognitive difficulty scalar of the th learning resource; represents the form switching weight, used to adjust the penalty strength of cross-resource form switching; represents the form switching indication quantity between the th learning resource and the th learning resource; represents the single switching base penalty amplitude, multiplied by to form the absolute magnitude of the switching cost.
[0064] Here, the expression of the learning path objective function is: (16)
[0065] Here, three goals of path optimization are achieved: (1) maximize the overall utility score of the path (personalization, goal alignment, structure gain); (2) minimize the sequence cost of adjacent learning steps (pre-requisite violation, difficulty step, form switching); (3) based on the soft deadline penalty, constrain the time axis, and complete the avoidable delay in advance. This embodiment optimizes the sequence quality under the premise of meeting the hard constraints of formulas (11) to (14) and the feasibility of pre-requisites, and outputs an executable and efficient final learning route. In formula (16), represents the path length of the th learning path resource sequence set; represents the path length of the th learning path resource sequence set; represents the overall utility score of selecting the th learning resource at the hth position; represents the edge cost from the th learning resource to the th learning resource; represents the cumulative time used in the first h steps of the path; represents the deadline penalty coefficient; represents the soft deadline time of the th learning resource; represents the expected learning duration of the th learning resource.
[0066] The eighth step is dynamic re-planning, which includes mastery update, coverage demand submission, and duration budget update.
[0067] Here, the mastery update expression is: (17)
[0068] When the student completes the learning resource , the mastery of the related knowledge points is improved, and the higher the coverage, the more the improvement. The updated will be used for dynamic re-planning to reduce the number of knowledge points that still need to be covered and adjust the priority of subsequent resources (such as reducing the weight of knowledge points that have been fully mastered). After updating, the is truncated to the interval to ensure numerical stability and interpretability. In formula (17), represents the current knowledge point mastery vector; represents the unit resource gain coefficient, which represents the proportion coefficient of the improvement of the mastery of the covered knowledge points after completing a resource; represents the resource coverage column vector.
[0069] The expression for the decreasing coverage demand of this embodiment is: (18)
[0070] When the student finishes the learning resource , according to the coverage strength of the knowledge points of the resource, the remaining coverage demand of the corresponding knowledge points is reduced , combined with formula (17), the learning state is updated from the "demand side" and the "ability side" synchronously. In formula (18), represents the remaining coverage demand threshold of the th knowledge point; represents the coverage degree of the th knowledge point by the learning resource.
[0071] Finally, the expression of the time budget update of the embodiment is:
[0072] (19)
[0073] After the student finishes the current resource learning, the remaining time is deducted synchronously. As a dynamic state quantity, it corresponds to the global upper limit of formula (11). When the time budget is close to 0, the adjustment strategy is triggered, and the resources with high priority or high target weight are preferentially executed, or the current stage task is terminated, to ensure that the overall path is executable and meets the synchronous station rules. In formula (19), represents the remaining time budget; represents the learning time of the completed learning resource.
[0074] The above embodiment is only a preferred embodiment of the present application, and is not a limitation on the protection scope of the present application. Any design principle adopted by the present application, and any changes made on the basis of non-creative labor shall belong to the protection scope of the present application.
Claims
1. A multi-modal preference driven graph convolutional combinatorial optimization learning path generation method, characterized in that, Comprising the following steps: The multi-modal learning feature data is acquired, and multi-modal feature data fusion processing is performed to obtain a learning resource initial feature matrix. An expression of the multi-modal feature data fusion processing is: ; wherein, represents an initial feature vector of a learning resource after multi-modal feature extraction and fusion mapping of the i th learning resource; represents an initial feature vector of a learning resource after multi-modal feature extraction and fusion mapping of the i th learning resource; represents an image feature vector of the i th learning resource; represents an image feature vector of the i th learning resource; represents a text feature vector of the i th learning resource; represents a text feature vector of the i th learning resource; represents a behavior feature vector of the i th learning resource; represents a behavior feature vector of the i th learning resource; represents a weight matrix for adjusting a modal contribution degree; represents a bias vector; is a positive integer. The expression of the learning resource initial feature matrix is: ; wherein, represents a learning resource initial feature vector after multi-modal feature extraction and fusion mapping of 1 learning resource; represents a learning resource initial feature vector after multi-modal feature extraction and fusion mapping of learning resources; T represents vector transposition; represents the total number of nodes of the learning resource; A multi-relation adjacency matrix is constructed for the initial characteristic matrix of the learning resource, and a normalized propagation matrix is constructed by normalizing the multi-relation adjacency matrix; an expression of the multi-relation adjacency matrix is: ; wherein, denotes the multi-relation adjacency matrix after fusion; denotes a prerequisite relation adjacency matrix; denotes a similarity relation adjacency matrix; denotes a co-occurrence relation adjacency matrix; denotes a fusion weight coefficient corresponding to the prerequisite relation adjacency matrix; denotes a fusion weight coefficient corresponding to the similarity relation adjacency matrix; denotes a fusion weight coefficient corresponding to the co-occurrence relation adjacency matrix; denotes a self-loop weight coefficient of the unit matrix; denotes a unit matrix; The normalized propagation matrix The expression for the normalized propagation matrix is: ; where, D represents the node degree matrix; D represents the inverse square root of each non-zero element of the diagonal matrix. According to the normalized propagation matrix, the graph convolution feature is updated, and a learning gain score is obtained; wherein the expression for updating the graph convolution feature is: ; wherein, represents the learning resource node feature matrix of the first layer graph convolution; represents the learning resource node feature matrix of the first layer graph convolution; represents a nonlinear activation function; represents the GCN weight matrix of the first layer graph convolution; The expression for the learning gain score is: ;in, Indicates the first The student on the first The learning gain score of each learning resource; This represents the learning gain scoring function; Indicates the first The final feature vector of each student; Indicates the first The final feature vector of each learning resource; Calculate the comprehensive utility based on the learning gain score; According to the score of the comprehensive utility, consider the establishment of integer programming under multiple constraint conditions to screen and obtain optimized candidate resources; Generate the optimal learning path in the optimized candidate resources by using the edge cost function and the learning path target function; in the generation process of the optimal learning path, dynamic re-planning is adopted; the dynamic re-planning comprises mastery update, coverage demand decrease and time length budget update.
2. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 1, characterized in that, Further comprising: For the Find the optimal learning path for each student Its expression is: ;in, This represents any candidate learning path; Indicates the first The set of all feasible learning paths for a given student; This represents the cost penalty coefficient; This represents the path cost function.
3. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 2, characterized in that, Further comprising: Find the first Target alignment score vector of learning resources Its expression is: ;in, This represents the total number of knowledge points. Indicates the first The learning resource for the first The coverage strength of each knowledge point; Indicates the first The target knowledge point weight vector for each knowledge point.
4. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 3, characterized in that, The overall utility is calculated based on the learning gain score, which is expressed as: ; wherein, represents the overall priority of the th student recommending the th learning resource; represents the structure gain weight coefficient; represents the goal alignment weight coefficient; represents the individual preference weight coefficient; represents the structure gain score of the th student recommending the th learning resource; represents the history preference degree of the th student recommending the th learning resource.
5. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 4, characterized in that, According to the score of the comprehensive utility, consider the establishment of integer programming under multiple constraint conditions to screen and obtain optimized candidate resources; the establishment of integer programming under multiple constraint conditions comprises a learning path target function, a time length constraint, a knowledge point coverage constraint, a task quantity constraint and an average quality constraint; An expression of the learning path objective function is: ; wherein, denotes a set of learning path resource sequences; denotes a weight coefficient of maximum total score of comprehensive utility; denotes a path length penalty coefficient; denotes a path length; The expression of the time length constraint is: ; wherein, represents the predicted learning time length of the th learning resource; represents the maximum total learning time length limit acceptable by the th student; The expression of the knowledge point coverage constraint is: ; wherein, represents the selection decision variable of the first learning resource; represents the minimum coverage threshold of the first knowledge point; The expression of the task quantity constraint is: ; wherein, represents a minimum task quantity lower bound allowed; represents a maximum task quantity upper bound allowed; The expression of the average quality constraint is: ; wherein, denotes the quality score of the th learning resource; denotes the average quality threshold.
6. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 5, characterized in that, The optimal learning path is generated from the optimized candidate resources using an edge cost function and a learning path objective function; the expression for the edge cost function is: ;in, Indicates from the first The learning resource to the first The marginal cost of a learning resource; Indicates the weight of the penalty before modification; Indicates from the first The learning resource to the first The cost of violating prerequisite instructions for a learning resource; Indicates the difficulty jump weight; Indicates the first The cognitive difficulty scalar of each learning resource; Indicates the first The cognitive difficulty scalar of each learning resource; Indicates the weight of form switching; Indicates the first The first learning resource and the first A quantity indicating the format switching of a learning resource; Indicates the base penalty for a single switch; The expression for the objective function of the learning path is: ; ;in, Indicates the first The path length of a set of learning path resource sequences; Indicates the first -1 is the path length of a set of learning path resource sequences; This indicates selecting the h-th position. The overall utility score of each learning resource; Indicates from the first Learning resources up to the The marginal cost of a learning resource; This indicates the cumulative time taken for the first h steps of the path; Indicates the cutoff penalty coefficient; Indicates the first The soft deadline for each learning resource; Indicates the first The estimated learning time for each learning resource.
7. The multi-modal preference driven graph convolutional combined optimization learning path generation method according to claim 6, characterized in that, The optimal learning path generation process adopts dynamic re-planning; the dynamic re-planning comprises mastery updating, coverage requirement decreasing and time length budget updating; an expression of the mastery updating is: Wherein, represents a current knowledge point mastery vector; represents a unit resource gain coefficient; represents a resource coverage column vector; The expression of the coverage requirement decreasing is: ; wherein, represents the remaining coverage requirement threshold of the th knowledge point; represents the coverage degree of the th learning resource to the th knowledge point; The expression of the time length budget update is: wherein, represents the remaining time length budget; represents the learning time length of the completed learning resource.
Citation Information
Patent Citations
Adaptive learning path recommendation method based on hypergraph neural network and knowledge tracking
CN120354883A
Online teaching optimization method and system based on emotion recognition
CN120355539A
Personalized learning content recommendation method and system based on neighborhood and hypergraph collaboration
CN120386941A
Decision optimization method fusing enhanced multi-modal learning and knowledge graph
CN120409642A
Multi-source heterogeneous learning path planning method based on personalized constraints
CN120409857A