A learning resource recommendation method and system based on user portrait
By employing implicit semantic matching, multi-objective optimization, and deep reinforcement learning techniques, an implicit semantic matching model between users and learning resources is constructed. This solves the problems of inaccurate user profiling and poor adaptability of recommendation systems to dynamic changes, enabling personalized, accurate, and diversified learning resource recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU LANFAN INFORMATION TECH CO LTD
- Filing Date
- 2024-09-10
- Publication Date
- 2026-06-02
AI Technical Summary
Existing learning resource recommendation systems struggle to capture subtle differences in user behavior over time, resulting in inaccurate user profiles and an inability to provide effective recommendations for new users or courses. Furthermore, traditional recommendation systems lack the ability to adapt to dynamic changes in users, leading to homogenized recommendation results and insufficient attention to niche resources.
By introducing implicit semantic matching, multi-objective optimization, and deep reinforcement learning techniques, an implicit semantic matching model between users and learning resources is constructed. Using factorization machine and collaborative filtering algorithm, combined with non-dominated sorting genetic algorithm and deep Q-network algorithm, the recommendation strategy is dynamically adjusted. With the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources, the optimal recommendation decision model is generated.
It improves the accuracy and dynamic adaptability of learning resource recommendations, meets users' personalized and dynamically changing learning needs, enhances the accuracy and diversity of recommendations, and strengthens the dynamic adaptability of the recommendation system.
Smart Images

Figure CN120256713B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of resource recommendation technology, and in particular to a learning resource recommendation method and system based on user profiles. Background Technology
[0002] In today's digital age, personalized learning has become an important trend in the education field. By analyzing user behavior and preferences, personalized learning resources are provided to users. However, users' interests and needs are constantly changing, while existing recommendation systems often rely on static user data, such as age, gender, and education level. This data is difficult to capture the subtle differences in users over time. While user behavior data can provide a more dynamic perspective, the collection and analysis of this data often lacks depth and breadth, resulting in inaccurate user profiles that fail to truly reflect the user's current learning status and needs. This is especially true for new users or new courses, where there is often a lack of sufficient user interaction data to build accurate user profiles. This makes it difficult for recommendation systems to provide effective recommendations for new users or new learning resources, leading to the cold start problem.
[0003] Secondly, traditional recommender systems mainly rely on single recommendation algorithms such as content-based recommendation or collaborative filtering. While these algorithms can provide personalized learning resources to some extent, they have significant limitations. These algorithms are not only inefficient in handling large-scale datasets, but also tend to recommend popular or highly rated learning resources, leading to homogenized recommendation results and ignoring users' need for diversity. At the same time, this trend exacerbates the uneven distribution of learning resources, causing some niche or low-rated but educationally valuable learning resources to receive insufficient attention. Furthermore, users' learning needs and abilities are constantly changing, and traditional recommender systems often lack the ability to adapt to these dynamic changes. Users need different types of learning resources at different learning stages, but traditional recommender systems often fail to capture these changes in a timely manner and find it difficult to adjust recommendation strategies in real time to adapt to these changes in users.
[0004] In summary, while traditional learning resource recommendation methods have made some progress in personalized recommendations, they still fall short in terms of recommendation diversity and dynamic adaptability. To address these issues, there is an urgent need to provide a learning resource recommendation method based on user profiles to improve the accuracy and diversity of recommendation results. Summary of the Invention
[0005] The purpose of this invention is to provide a learning resource recommendation method and system based on user profiles. By introducing techniques such as latent semantic matching, multi-objective optimization, and deep reinforcement learning, the recommendation strategy is dynamically adjusted to improve the accuracy and dynamic adaptability of learning resource recommendations.
[0006] To address the above technical problems, this invention provides a learning resource recommendation method and system based on user profiles.
[0007] In a first aspect, the present invention provides a learning resource recommendation method based on user profiles, the method comprising the following steps:
[0008] Based on the collected user profile information and learning resource information, establish a user profile matrix and a learning resource feature matrix;
[0009] Based on the user profile matrix and the learning resource feature matrix, a latent semantic matching model between users and learning resources is constructed using factorization machine and collaborative filtering algorithm;
[0010] The parameters of the latent semantic matching model are learned by alternating least squares method to generate latent semantic association features between users and learning resources.
[0011] A multi-objective learning resource recommendation optimization model is established with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources.
[0012] The user-learning resource latent semantic association features are used as individual codes, and the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model is searched using a non-dominated sorting genetic algorithm.
[0013] We construct a multidimensional learning state representation based on user profiles and learning resources, use the deep Q-network algorithm to estimate the long-term cumulative returns of different learning resources, and autonomously learn to obtain the optimal recommendation decision model.
[0014] Based on the Pareto optimal recommendation solution set and the optimal recommendation decision model, a learning resource recommendation scheme is generated.
[0015] In a further implementation, the step of establishing a user profile matrix and a learning resource feature matrix based on the collected user profile information and learning resource information includes:
[0016] Collect user profile information and learning resource information, and use the maximum relevance minimum redundancy algorithm to extract features from the user profile information and the learning resource information respectively to obtain the corresponding user profile feature subset and learning resource feature subset;
[0017] Principal component analysis was performed on the user profile feature subset and the learning resource feature subset respectively to obtain the corresponding user profile key feature vector and learning resource key feature vector.
[0018] The user profile key feature vectors are weighted and averaged according to the pre-acquired feature importance weights to obtain the user profile baseline feature values;
[0019] Based on the baseline feature values of the user profile, the key feature vectors of the user profile are standardized to establish a user profile matrix;
[0020] The learning resource baseline feature values are obtained by aggregating the key feature vectors of the learning resource using the moving average method to perform learning time series aggregation.
[0021] The key feature vectors of the learning resources are standardized based on the baseline feature values of the learning resources to establish a feature matrix of the learning resources.
[0022] In a further implementation, the step of constructing a latent semantic matching model between users and learning resources using factorization machines and collaborative filtering algorithms based on the user profile matrix and the learning resource feature matrix includes:
[0023] The user profile matrix and the learning resource feature matrix are mapped to the same latent semantic space to obtain the latent vector representation of the user profile and the latent vector representation of the learning resource.
[0024] The factorization machine is used to capture the feature interaction information between the latent vector representation of user profiles and the latent vector representation of learning resources;
[0025] The feature interaction information is decomposed using a collaborative filtering algorithm to obtain the user implicit feature vector and the learning resource implicit feature vector.
[0026] The feature interaction information, the user implicit feature vector, and the learning resource implicit feature vector are fused to obtain the fused feature;
[0027] Based on the fusion features, a latent semantic matching model between users and learning resources is constructed using a deep autoencoder attention network.
[0028] In a further implementation, the step of generating user-learning resource latent semantic association features by learning the parameters of the latent semantic matching model using alternating least squares includes:
[0029] The parameters in the latent semantic matching model are randomly initialized; the parameters of the latent semantic matching model include the user latent semantic space representation and the learning resource latent semantic space representation.
[0030] Using the alternating least squares method, the user latent semantic space representation and the learning resource latent semantic space representation are updated alternately by minimizing the reconstruction error to obtain user latent semantic features and learning resource latent semantic features;
[0031] Based on the latent semantic features of the user and the latent semantic features of the learning resources, the potential association between users and learning resources is captured, and latent semantic association features between users and learning resources are generated.
[0032] In a further implementation, the multi-objective learning resource recommendation optimization model is specifically as follows:
[0033] ;
[0034] in,
[0035] ;
[0036] ;
[0037] ;
[0038] In the formula, For users to learn time functions; For the cumulative learning resource quality function; is a weight parameter between 0 and 1; S is the number of learning resources; a and b are fitting parameters; The difficulty of learning resource i; For recommendation variables, ; The learning time for learning resource i; For users' familiarity with learning resources; , , , , These are the weighting coefficients; To broaden the knowledge coverage of learning resource i; The relevance of learning resources to the user's current learning progress and needs; The contribution of user interest in learning resource i to the cumulative quality of learning resources; The frequency of user interaction with learning resource i; The maximum frequency of interaction between the user and the learning resource i; The diversity gain after adding learning resource i to the learning resource set S; Let cosine similarity be the similarity between learning resource i and learning resource j. .
[0039] In a further implementation, the constraints of the multi-objective learning resource recommendation optimization model include learning time constraints and learning resource difficulty constraints, wherein the learning time constraints are specifically as follows:
[0040] ;
[0041] The specific difficulty constraints of the learning resources are as follows:
[0042] ;
[0043] In the formula, B represents the user's learning saturation.
[0044] In a further implementation, the step of constructing a multidimensional learning state representation based on user profiles and learning resources includes:
[0045] Obtain interaction data between user profile information and learning resource information, evaluate the connection strength between users and learning resources based on the interaction data, and construct a user-learning resource adjacency matrix;
[0046] Define multidimensional attributes for each learning resource based on the learning resource information, and construct a multidimensional attribute matrix for the learning resources.
[0047] The user-learning resource adjacency matrix and the learning resource multidimensional attribute matrix are concatenated to form a learning feature concatenation matrix.
[0048] The concatenated matrix of the learned features is input into the graph neural network to learn the feature representations of users and learning resources, thereby obtaining user feature vectors and learning resource feature vectors.
[0049] The learning resource information is sequentially categorized and multi-dimensional information entropy is extracted to obtain multi-dimensional learning resource representation data.
[0050] The user feature vector, the learning resource feature vector, and the multidimensional learning resource representation data are weighted and fused to obtain a multidimensional learning state representation.
[0051] In a further implementation, the deep Q-network algorithm uses the multidimensional learning state representation as its state space.
[0052] In a further implementation, the step of generating a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model includes:
[0053] Iterate through each candidate solution in the Pareto optimal recommendation solution set and calculate the matching degree between the candidate solution and the user's current progress requirement;
[0054] Based on the matching degree, the recommended candidate set with the highest matching degree with the user's current progress needs is selected;
[0055] Based on the optimal recommendation decision model, obtain the Q value of each learning resource in the recommendation candidate set;
[0056] Based on the Q-value of each learning resource, the learning resource with the highest Q-value is selected from the recommendation candidate set to generate a learning resource recommendation scheme.
[0057] Secondly, the present invention provides a learning resource recommendation system based on user profiles, the system comprising:
[0058] The data analysis module is used to build user profile matrices and learning resource feature matrices based on the collected user profile information and learning resource information.
[0059] The latent semantic analysis module is used to construct a latent semantic matching model between users and learning resources based on the user profile matrix and the learning resource feature matrix, using factorization machine and collaborative filtering algorithm.
[0060] The parameter learning module is used to learn the parameters of the latent semantic matching model through alternating least squares method and generate user-learning resource latent semantic association features.
[0061] The multi-objective optimization module is used to establish a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources; and to use the implicit semantic association features between users and learning resources as individual codes and a non-dominated sorting genetic algorithm to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model.
[0062] The recommendation decision module is used to construct a multi-dimensional learning state representation based on user profiles and learning resources, estimate the long-term cumulative returns of different learning resources using a deep Q-network algorithm, and autonomously learn to obtain the optimal recommendation decision model; and generate a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model.
[0063] This invention provides a learning resource recommendation method and system based on user profiles. The method establishes a user profile matrix and a learning resource feature matrix based on collected user profile and learning resource information, and constructs a latent semantic matching model between users and learning resources using factorization machines and collaborative filtering algorithms. It then learns the parameters of the latent semantic matching model using alternating least squares to generate latent semantic association features between users and learning resources. A multi-objective learning resource recommendation optimization model is established with the optimization objectives of minimizing user learning time and maximizing the cumulative quality of learning resources. The latent semantic association features between users and learning resources are used as individual codes, and a non-dominated sorting genetic algorithm is used to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model. A multi-dimensional learning state representation based on user profiles and learning resources is constructed, and a deep Q-network algorithm is used to estimate the long-term cumulative returns of different learning resources, autonomously learning to obtain the optimal recommendation decision model. Finally, a learning resource recommendation scheme is generated based on the Pareto optimal recommendation solution set and the optimal recommendation decision model. Compared with existing technologies, this method utilizes a latent semantic matching model to deeply mine the potential associations between users and learning resources, combines a multi-objective optimization algorithm to balance recommendation accuracy and diversity, and dynamically adjusts the recommendation strategy through deep reinforcement learning to meet users' personalized and dynamically changing learning needs, thereby improving the personalization, accuracy, and dynamic adaptability of learning resource recommendations. Attached Figure Description
[0064] Figure 1 This is a schematic diagram of the learning resource recommendation method based on user profiles provided in an embodiment of the present invention;
[0065] Figure 2 This is a block diagram of a learning resource recommendation system based on user profiles provided in an embodiment of the present invention. Detailed Implementation
[0066] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0067] refer to Figure 1 This invention provides a learning resource recommendation method based on user profiles, such as... Figure 1 As shown, the method includes the following steps:
[0068] S1. Based on the collected user profile information and learning resource information, establish a user profile matrix and a learning resource feature matrix.
[0069] In this embodiment, the step of establishing a user profile matrix and a learning resource feature matrix based on the collected user profile information and learning resource information includes:
[0070] Collect user profile information and learning resource information, and use the maximum relevance minimum redundancy algorithm to extract features from the user profile information and the learning resource information respectively to obtain the corresponding user profile feature subset and learning resource feature subset;
[0071] Principal component analysis was performed on the user profile feature subset and the learning resource feature subset respectively to obtain the corresponding user profile key feature vector and learning resource key feature vector.
[0072] The user profile key feature vectors are weighted and averaged according to the pre-acquired feature importance weights to obtain the user profile baseline feature values;
[0073] Based on the baseline feature values of the user profile, the key feature vectors of the user profile are standardized to establish a user profile matrix;
[0074] The learning resource baseline feature values are obtained by aggregating the key feature vectors of the learning resource using the moving average method to perform learning time series aggregation.
[0075] The key feature vectors of the learning resources are standardized based on the baseline feature values of the learning resources to establish a feature matrix of the learning resources.
[0076] Specifically, this embodiment collects user profile information and learning resource information respectively, and performs data cleaning on the user profile information and learning resource information to handle missing values, outliers, and duplicate data, providing a data foundation for the subsequent establishment of user profile matrix and learning resource feature matrix; wherein, the user profile information includes, but is not limited to, basic user information (such as age, gender, educational background), user behavior data (such as browsing history, learning progress, interaction records, knowledge mastery level), and user preference data (such as learning resources of interest and learning habits); the learning resource information includes, but is not limited to, learning resource metadata and learning resource performance data. Learning resource metadata may include topics, difficulty levels, resource types, authors, and the scope of application of knowledge points, etc., and learning resource performance data may include the completion rate, positive review rate, and user rating of learning resources, etc.
[0077] This embodiment employs the Maximum Relevance Minimum Redundancy (mRMR) algorithm to select the most representative and least redundant features from the user profile information and the learning resource information, respectively, to obtain user profile feature subsets and learning resource feature subsets. Then, Principal Component Analysis (PCA) is used to reduce the dimensionality of both subsets, preserving the main changes in the datasets and extracting key feature vectors for user profiles and learning resources. Next, for each key feature in the user profile feature vector, this embodiment calculates its feature importance weight in the entire feature vector based on the different contributions of different features to the user profile, using statistical methods such as Pearson correlation coefficient or mutual information. The user profile key feature vector is then weighted using these weights to obtain weighted feature values. This weighted weighting more accurately reflects the actual impact of features. Then, the following calculations are performed... The average of the weighted feature values is used as the baseline feature value for the user profile. It should be noted that this embodiment takes into account that the key feature vectors of learning resources will vary in different time periods. For example, the regular updates of knowledge points or content of learning resources, the impact of different times or events on the access volume of learning resources, and the quality of learning resources will change over time. Therefore, this embodiment uses the moving average method to aggregate the key feature vectors of learning resources and extract the average performance or trend of learning resources in different time periods to obtain the baseline feature value of learning resources. This allows for a more accurate evaluation of the quality and effectiveness of learning resources and provides support for subsequent decision-making. At the same time, considering that different features have different dimensions and numerical ranges, this embodiment standardizes the key feature vectors of user profiles and key feature vectors of learning resources according to their respective baseline feature values to transform the feature values to the same scale, thereby constructing the user profile matrix and the learning resource feature matrix.
[0078] S2. Based on the user profile matrix and the learning resource feature matrix, construct a latent semantic matching model between users and learning resources using factorization machine and collaborative filtering algorithm.
[0079] In this embodiment, the step of constructing a latent semantic matching model between users and learning resources using factorization machine and collaborative filtering algorithm based on the user profile matrix and the learning resource feature matrix includes:
[0080] The user profile matrix and the learning resource feature matrix are mapped to the same latent semantic space to obtain the latent vector representation of the user profile and the latent vector representation of the learning resource.
[0081] The factorization machine is used to capture the feature interaction information between the latent vector representation of user profiles and the latent vector representation of learning resources;
[0082] The feature interaction information is decomposed using a collaborative filtering algorithm to obtain the user implicit feature vector and the learning resource implicit feature vector.
[0083] The feature interaction information, the user implicit feature vector, and the learning resource implicit feature vector are fused to obtain the fused feature;
[0084] Based on the fusion features, a latent semantic matching model between users and learning resources is constructed using a deep autoencoder attention network.
[0085] This embodiment merges the latent vector representations of user profiles and learning resources as input to Factorization Machines (FM). The FM captures the implicit feature interactions between users and learning resources, such as clicks and learning duration. Simultaneously, a collaborative filtering algorithm is used to decompose the feature interaction information, obtaining implicit feature vectors for users and learning resources. Then, combining the FM and collaborative filtering algorithms, a latent semantic matching model is constructed. This embodiment fuses the feature interaction information learned by the FM with the implicit feature vectors obtained by collaborative filtering. This fusion process can be achieved through feature concatenation or weighted summation to obtain fused features. Based on these fused features, a latent semantic matching model capable of expressing the deep-seated relationship between users and learning resources is constructed.
[0086] Specifically, this embodiment, based on fusion features, employs a deep autoencoder attention network to construct a latent semantic matching model between users and learning resources. Unlike traditional recommendation methods, this embodiment combines factorization machines, collaborative filtering algorithms, and deep autoencoders to capture the deep-level relationships between users and learning resources. The deep autoencoder attention network combines a deep autoencoder, a multi-scale attention mechanism, and matrix factorization. The deep autoencoder attention network includes a matrix factorization initialization layer, a deep autoencoder, and a multi-scale attention mechanism introduced within the deep autoencoder. The deep autoencoder includes an encoder, a bottleneck layer, and a decoder. The encoder includes multiple neural network layers and a multi-scale attention mechanism placed before each neural network layer. The matrix factorization initialization layer is used to generate preliminary implicit features for users and learning resources using matrix factorization techniques. The fusion feature vector is weighted and fused with the initial implicit feature vector to obtain the enhanced fusion feature. In this embodiment, a multi-scale attention mechanism is introduced into the encoder. The role of the multi-scale attention mechanism is to dynamically adjust the weights so that the model can capture the complex relationships between different users and learning resources more flexibly and accurately. In this embodiment, the first layer of the multi-scale attention mechanism is used to adjust the weights of the enhanced fusion feature, and the enhanced fusion feature obtained after weighting by the multi-scale attention mechanism is nonlinearly transformed by a multi-layer neural network. According to the above steps, feature information is further extracted and compressed by the nonlinear transformation layer by layer of the multi-layer neural network, thereby capturing the core implicit semantic information output by the multi-layer neural network through the bottleneck layer, generating implicit semantic feature vectors, and then decoding the implicit semantic feature vectors by the decoder.
[0087] S3. The parameters of the latent semantic matching model are learned by alternating least squares method to generate user-learning resource latent semantic association features.
[0088] In this embodiment, the step of generating user-learning resource latent semantic association features by learning the parameters of the latent semantic matching model through alternating least squares method includes:
[0089] The parameters in the latent semantic matching model are randomly initialized; the parameters of the latent semantic matching model include the user latent semantic space representation and the learning resource latent semantic space representation.
[0090] Using the alternating least squares method, the user latent semantic space representation and the learning resource latent semantic space representation are updated alternately by minimizing the reconstruction error to obtain user latent semantic features and learning resource latent semantic features;
[0091] Based on the latent semantic features of the user and the latent semantic features of the learning resources, the potential association between users and learning resources is captured, and latent semantic association features between users and learning resources are generated.
[0092] This embodiment employs Alternating Least Squares (ALS) to learn the parameters of the latent semantic matching model. In the latent semantic matching model, for the user latent semantic space representation and the learning resource latent semantic space representation, ALS iteratively minimizes the reconstruction error by alternately fixing one latent semantic space representation (either the user latent semantic space representation or the learning resource latent semantic space representation) and optimizing the other latent semantic space representation, solving for the optimal solution under the current fixed parameters until a preset convergence condition is reached. Then, the newly solved parameters are fixed, and the previously fixed parameters are optimized. This process is repeated alternately until the model parameters converge. In this embodiment, the preset convergence condition is set as the change in reconstruction error being less than a preset reconstruction error threshold or reaching a preset number of iterations. This embodiment optimizes the model parameters using ALS to capture the latent semantic association between users and learning resources, obtaining user-learning resource latent semantic association features.
[0093] S4. Establish a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources.
[0094] S5. Using the latent semantic association features of the user-learning resources as individual codes, a non-dominated sorting genetic algorithm is used to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model.
[0095] This embodiment establishes a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources. A non-dominated sorting genetic algorithm (NSGA-II) is used to handle the multi-objective optimization problem, enabling the multi-objective learning resource recommendation optimization model to find an optimal balance among multiple optimization objectives, that is, simultaneously minimizing user learning time and maximizing the quality of accumulated learning resources. This yields a Pareto optimal recommendation solution set, which represents a comprehensive recommendation strategy that achieves a balance among multiple objectives such as user learning time, accumulated learning resource quality, and diversity. This provides the recommendation system with greater flexibility and adaptability, and can meet the needs of a wider range of users and scenarios. In this embodiment, the multi-objective learning resource recommendation optimization model is specifically as follows:
[0096] ;
[0097] in,
[0098] ;
[0099] ;
[0100] ;
[0101] In the formula, For users to learn time functions; For the cumulative learning resource quality function; is a weight parameter between 0 and 1; S is the number of learning resources; a and b are fitting parameters; The difficulty of learning resource i; For recommendation variables, If a user is recommended learning resource i, then =1; The learning time for learning resource i; For users' familiarity with learning resources; , , , , These are the weighting coefficients; To broaden the knowledge coverage of learning resource i; The relevance of learning resources to the user's current learning progress and needs; The contribution of user interest in learning resource i to the cumulative quality of learning resources; The frequency of user interaction with learning resource i; The maximum frequency of interaction between the user and the learning resource i; The diversity gain after adding learning resource i to the learning resource set S; Let cosine similarity be the similarity between learning resource i and learning resource j. .
[0102] Meanwhile, the constraints of the multi-objective learning resource recommendation optimization model may include learning time constraints and learning resource difficulty constraints, wherein the learning time constraints are specifically as follows:
[0103] ;
[0104] The specific constraints on the difficulty of learning resources are as follows:
[0105] ;
[0106] In the formula, B represents the user's learning saturation, which reflects the user's ability to accept new learning resources.
[0107] This embodiment also introduces learning resource attributes such as the importance of learning resources and the coherence of learning paths to construct individual evaluation functions. Based on these individual evaluation functions, a non-dominated sorting genetic algorithm is used to solve the multi-objective learning resource optimization model, searching for the Pareto optimal solution set. Specifically, the individual evaluation function is as follows:
[0108] ;
[0109] ;
[0110] ;
[0111] in,
[0112] ;
[0113] ;
[0114] In the formula, For individual evaluation functions; To maximize the overall benefits of learning resource i; To ensure the coherence of the learning path from learning resource i to learning resource j; Weighting coefficient for user learning time; For the availability of learning resource i; The importance of learning resources; A nonlinear adjustment parameter for the importance of learning resources; This refers to the conversion degree weighting coefficient; The degree of transformation from learning resource i to learning resource j; For learning resource j to be a prerequisite for learning resource i; The difficulty gap between learning resource i and learning resource j; This is a non-linear adjustment parameter for the difficulty of learning resources.
[0115] S6. Construct a multidimensional learning state representation based on user profiles and learning resources, use the deep Q-network algorithm to estimate the long-term cumulative returns of different learning resources, and obtain the optimal recommendation decision model through autonomous learning.
[0116] In this embodiment, the step of constructing a multidimensional learning state representation based on user profiles and learning resources includes:
[0117] Obtain interaction data between user profile information and learning resource information, evaluate the connection strength between users and learning resources based on the interaction data, and construct a user-learning resource adjacency matrix;
[0118] Define multidimensional attributes for each learning resource based on the learning resource information, and construct a multidimensional attribute matrix for the learning resources.
[0119] The user-learning resource adjacency matrix and the learning resource multidimensional attribute matrix are concatenated to form a learning feature concatenation matrix.
[0120] The concatenated matrix of the learned features is input into the graph neural network to learn the feature representations of users and learning resources, thereby obtaining user feature vectors and learning resource feature vectors.
[0121] The learning resource information is sequentially categorized and multi-dimensional information entropy is extracted to obtain multi-dimensional learning resource representation data.
[0122] The user feature vector, the learning resource feature vector, and the multidimensional learning resource representation data are weighted and fused to obtain a multidimensional learning state representation.
[0123] This embodiment combines the user-learning resource adjacency matrix with the learning resource multidimensional attribute matrix to form a learning feature concatenation matrix. The learning feature concatenation matrix is then input into the graph neural network to learn the feature representations of users and learning resources. This fully utilizes the inherent structure and features of the data to learn more comprehensive and accurate feature representations of users and learning resources, resulting in user feature vectors and learning resource feature vectors.
[0124] This embodiment captures and analyzes the interaction data between user profile information and learning resource information to evaluate the connection strength between them. Based on this connection strength, a user-learning resource adjacency matrix is constructed. In this embodiment, the user-learning resource adjacency matrix reflects the user's level of engagement with different learning resources. Then, to more comprehensively characterize the features of learning resources, this embodiment defines multi-dimensional attributes for each learning resource and constructs a learning resource multi-dimensional attribute matrix. The user-learning resource adjacency matrix and the learning resource multi-dimensional attribute matrix are combined to form a learning feature concatenation matrix. This learning feature concatenation matrix not only includes the user's level of engagement with learning resources but also incorporates the learning resource multi-dimensional attribute matrix, providing comprehensive data support for subsequent graph neural networks. Next, this embodiment uses a graph neural network (GNN) with the learning feature concatenation matrix as input to learn and extract deep-level feature representations of users and learning resources, namely user feature vectors and learning resource feature vectors. By introducing GNN, this embodiment effectively captures the complex relationships and structural characteristics in the learning feature concatenation matrix, improving the accuracy and efficiency of feature learning.
[0125] This embodiment constructs a multidimensional learning state representation based on user feature vectors, learning resource feature vectors, and multidimensional learning resource representation data. This multidimensional learning state representation is used as the state space of a deep Q-network algorithm, and the recommendation action of recommending a specific learning resource is used as the action space of the deep Q-network algorithm. The deep Q-network algorithm estimates the long-term cumulative reward of different recommendation actions and continuously updates its strategy to maximize the long-term cumulative reward, outputting the optimal recommendation decision strategy for the current state. In summary, the deep Q-network algorithm interacts with the environment using the multidimensional learning state representation as a state vector, selecting the recommendation action that maximizes the long-term cumulative reward based on the current state, thus obtaining the optimal recommendation decision strategy. This embodiment combines graph neural networks and deep Q-network algorithms to achieve accurate estimation and optimization of the long-term cumulative reward of user learning resources. This enables the optimal recommendation decision model to intelligently adjust the recommendation strategy according to the dynamic changes in the learning state, and by maximizing the user's long-term learning benefits, it can adaptively optimize the decision-making learning resource recommendation scheme.
[0126] S7. Generate a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model.
[0127] In this embodiment, the step of generating a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model includes:
[0128] Iterate through each candidate solution in the Pareto optimal recommendation solution set and calculate the matching degree between the candidate solution and the user's current progress requirement;
[0129] Based on the matching degree, the recommended candidate set with the highest matching degree with the user's current progress needs is selected;
[0130] Based on the optimal recommendation decision model, obtain the Q value of each learning resource in the recommendation candidate set;
[0131] Based on the Q-value of each learning resource, the learning resource with the highest Q-value is selected from the recommendation candidate set to generate a learning resource recommendation scheme.
[0132] Specifically, the Pareto optimal recommendation solution set contains multiple candidate solutions. In this embodiment, by traversing the entire Pareto optimal recommendation solution set, a matching degree is calculated for each candidate solution to quantify its suitability for the user's current learning progress needs. Then, based on the matching degree, the candidate solutions with the highest matching degree to the user's current progress needs are further selected to form a recommendation candidate set. These recommendation candidate sets can meet the user's current actual learning needs to the greatest extent possible. Next, this embodiment uses an optimal recommendation decision model to calculate the Q-value for each learning resource in the selected recommendation candidate set. The Q-value reflects... The learning resources are designed to meet the user's current needs and have potential value in optimizing the future learning path. Finally, each learning resource is sorted according to its Q-value, and the learning resource with the highest Q-value is selected. The learning resource with the highest Q-value best meets the user's current and future learning needs. A learning resource recommendation scheme is generated based on the selected learning resource with the highest Q-value. Through the above steps, this embodiment effectively combines the optimal recommendation decision model with the Pareto optimal recommendation solution set. It not only considers the balance of multiple objectives, but also can dynamically adjust according to the user's real-time feedback, efficiently selecting the most suitable recommendation scheme for the user's current progress needs from the Pareto optimal recommendation solution set.
[0133] In summary, the user profile-based learning resource recommendation method provided in this embodiment integrates latent semantic matching, multi-objective optimization, and deep reinforcement learning techniques. It utilizes the latent semantic matching model to deeply mine the potential associations between users and learning resources, combines multi-objective optimization algorithms to balance recommendation accuracy and diversity, and dynamically adapts to changes in learning state through deep reinforcement learning. This improves the accuracy and diversity of learning resource recommendations, enhances the dynamic adaptability of the recommendation system, and achieves more accurate, diverse, and intelligent personalized learning resource recommendations.
[0134] This invention provides a learning resource recommendation method based on user profiles. The method establishes a user profile matrix and a learning resource feature matrix based on collected user profile and learning resource information, and constructs a latent semantic matching model between users and learning resources using factorization machines and collaborative filtering algorithms. It then learns the parameters of the latent semantic matching model using alternating least squares to generate latent semantic association features between users and learning resources. A multi-objective learning resource recommendation optimization model is established with the optimization objectives of minimizing user learning time and maximizing the cumulative quality of learning resources. The latent semantic association features between users and learning resources are used as individual codes, and a non-dominated sorting genetic algorithm is used to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model. A multi-dimensional learning state representation based on user profiles and learning resources is constructed, and a deep Q-network algorithm is used to estimate the long-term cumulative returns of different learning resources, autonomously learning to obtain the optimal recommendation decision model. Finally, a learning resource recommendation scheme is generated based on the Pareto optimal recommendation solution set and the optimal recommendation decision model. Compared with existing technologies, the method proposed in this embodiment addresses the problems of insufficient accuracy, homogeneity of recommended content, and poor adaptability to dynamic user needs in traditional recommendation systems by introducing techniques such as latent semantic matching, multi-objective optimization, and deep reinforcement learning. By using a latent semantic matching model to mine the potential associations between users and learning resources, the accuracy and relevance of recommendations are significantly improved. At the same time, by combining multi-objective optimization and deep reinforcement learning techniques, dynamic learning and adaptive recommendations are achieved, enhancing the dynamic adaptability of the system and meeting the diverse and dynamically changing learning needs of users.
[0135] It should be noted that the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0136] In one embodiment, such as Figure 2 As shown, this embodiment of the invention provides a learning resource recommendation system based on user profiles, the system comprising:
[0137] Data analysis module 101 is used to establish a user profile matrix and a learning resource feature matrix based on the collected user profile information and learning resource information;
[0138] The latent semantic analysis module 102 is used to construct a latent semantic matching model between users and learning resources based on the user profile matrix and the learning resource feature matrix, using factorization machine and collaborative filtering algorithm.
[0139] The parameter learning module 103 is used to learn the parameters of the latent semantic matching model through alternating least squares method and generate user-learning resource latent semantic association features.
[0140] The multi-objective optimization module 104 is used to establish a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources; and to use the user-learning resource latent semantic association features as individual codes and use a non-dominated sorting genetic algorithm to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model.
[0141] The recommendation decision module 105 is used to construct a multi-dimensional learning state representation based on user profiles and learning resources, estimate the long-term cumulative returns of different learning resources using a deep Q-network algorithm, and autonomously learn to obtain the optimal recommendation decision model; and generate a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model.
[0142] For specific limitations regarding a user profile-based learning resource recommendation system, please refer to the above-described limitations regarding a user profile-based learning resource recommendation method, which will not be repeated here. Those skilled in the art will recognize that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] This invention provides a learning resource recommendation system based on user profiles. The system establishes a user profile matrix and a learning resource feature matrix through a data analysis module; a latent semantic analysis module constructs a latent semantic matching model between users and learning resources using factorization machines and collaborative filtering algorithms; a parameter learning module learns the parameters of the latent semantic matching model using alternating least squares to generate latent semantic association features between users and learning resources; a multi-objective optimization module establishes a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the cumulative quality of learning resources; using the latent semantic association features between users and learning resources as individual codes, a non-dominated sorting genetic algorithm is used to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model; a recommendation decision module constructs a multi-dimensional learning state representation based on user profiles and learning resources, uses a deep Q-network algorithm to estimate the long-term cumulative returns of different learning resources, and autonomously learns to obtain the optimal recommendation decision model; and generates learning resource recommendation schemes based on the Pareto optimal recommendation solution set and the optimal recommendation decision model. Compared with existing technologies, the system proposed in this embodiment addresses the problems of insufficient accuracy in personalized recommendations, homogeneous recommended content, and poor adaptability to dynamic user needs in traditional recommendation systems by introducing techniques such as latent semantic matching, multi-objective optimization, and deep reinforcement learning. By using a latent semantic matching model to mine the potential associations between users and learning resources, the accuracy and relevance of recommendations are significantly improved. At the same time, by combining multi-objective optimization and deep reinforcement learning techniques, dynamic learning and adaptive recommendations are achieved, enhancing the dynamic adaptability of the system and meeting the diverse and dynamically changing learning needs of users.
[0144] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A learning resource recommendation method based on user profiles, characterized in that, Includes the following steps: Based on the collected user profile information and learning resource information, establish a user profile matrix and a learning resource feature matrix; Based on the user profile matrix and the learning resource feature matrix, a latent semantic matching model between users and learning resources is constructed using factorization machine and collaborative filtering algorithm; The parameters of the latent semantic matching model are learned by alternating least squares method to generate latent semantic association features between users and learning resources. A multi-objective learning resource recommendation optimization model is established with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources. The user-learning resource latent semantic association features are used as individual codes, and the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model is searched using a non-dominated sorting genetic algorithm. A multidimensional learning state representation based on user profiles and learning resources is constructed, and the multidimensional learning state representation is used as the state space of the deep Q-network algorithm. The deep Q-network algorithm is used to estimate the long-term cumulative returns of different learning resources, and the optimal recommendation decision model is obtained through autonomous learning. Based on the Pareto optimal recommendation solution set and the optimal recommendation decision model, a learning resource recommendation scheme is generated; The step of constructing a multidimensional learning state representation based on user profiles and learning resources includes: Obtain interaction data between user profile information and learning resource information, evaluate the connection strength between users and learning resources based on the interaction data, and construct a user-learning resource adjacency matrix; Define multidimensional attributes for each learning resource based on the learning resource information, and construct a multidimensional attribute matrix for the learning resources. The user-learning resource adjacency matrix and the learning resource multidimensional attribute matrix are concatenated to form a learning feature concatenation matrix. The concatenated matrix of the learned features is input into the graph neural network to learn the feature representations of users and learning resources, thereby obtaining user feature vectors and learning resource feature vectors. The learning resource information is sequentially categorized and multi-dimensional information entropy is extracted to obtain multi-dimensional learning resource representation data. The user feature vector, the learning resource feature vector, and the multidimensional learning resource representation data are weighted and fused to obtain a multidimensional learning state representation. The step of generating a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model includes: Iterate through each candidate solution in the Pareto optimal recommendation solution set and calculate the matching degree between the candidate solution and the user's current progress requirement; Based on the matching degree, the recommended candidate set with the highest matching degree with the user's current progress needs is selected; Based on the optimal recommendation decision model, obtain the Q value of each learning resource in the recommendation candidate set; Based on the Q-value of each learning resource, the learning resource with the highest Q-value is selected from the recommendation candidate set to generate a learning resource recommendation scheme.
2. The learning resource recommendation method based on user profiles as described in claim 1, characterized in that, The steps of establishing the user profile matrix and the learning resource feature matrix based on the collected user profile information and learning resource information include: Collect user profile information and learning resource information, and use the maximum relevance minimum redundancy algorithm to extract features from the user profile information and the learning resource information respectively to obtain the corresponding user profile feature subset and learning resource feature subset; Principal component analysis was performed on the user profile feature subset and the learning resource feature subset respectively to obtain the corresponding user profile key feature vector and learning resource key feature vector. The user profile key feature vectors are weighted and averaged according to the pre-acquired feature importance weights to obtain the user profile baseline feature values; Based on the baseline feature values of the user profile, the key feature vectors of the user profile are standardized to establish a user profile matrix; The learning resource baseline feature values are obtained by aggregating the key feature vectors of the learning resource using the moving average method to perform learning time series aggregation. The key feature vectors of the learning resources are standardized based on the baseline feature values of the learning resources to establish a feature matrix of the learning resources.
3. The learning resource recommendation method based on user profiles as described in claim 1, characterized in that, The step of constructing a latent semantic matching model between users and learning resources based on the user profile matrix and the learning resource feature matrix, using factorization machine and collaborative filtering algorithm, includes: The user profile matrix and the learning resource feature matrix are mapped to the same latent semantic space to obtain the latent vector representation of the user profile and the latent vector representation of the learning resource. The feature interaction information between the latent vector representation of the user profile and the latent vector representation of the learning resources is captured by the factorization machine. The feature interaction information is decomposed using a collaborative filtering algorithm to obtain the user implicit feature vector and the learning resource implicit feature vector. The feature interaction information, the user implicit feature vector, and the learning resource implicit feature vector are fused to obtain the fused feature; Based on the fusion features, a latent semantic matching model between users and learning resources is constructed using a deep autoencoder attention network.
4. The learning resource recommendation method based on user profiles as described in claim 1, characterized in that, The step of generating user-learning resource latent semantic association features by learning the parameters of the latent semantic matching model through alternating least squares method includes: The parameters in the latent semantic matching model are randomly initialized; the parameters of the latent semantic matching model include the user latent semantic space representation and the learning resource latent semantic space representation. Using the alternating least squares method, the user latent semantic space representation and the learning resource latent semantic space representation are updated alternately by minimizing the reconstruction error to obtain user latent semantic features and learning resource latent semantic features; Based on the latent semantic features of the user and the latent semantic features of the learning resources, the potential association between users and learning resources is captured, and latent semantic association features between users and learning resources are generated.
5. The learning resource recommendation method based on user profiles as described in claim 1, characterized in that, The multi-objective learning resource recommendation optimization model is specifically as follows: ; in, ; ; ; In the formula, For users to learn time functions; For the cumulative learning resource quality function; is a weight parameter between 0 and 1; S is the number of learning resources; a and b are fitting parameters; The difficulty of learning resource i; For recommendation variables, ; The learning time for learning resource i; For users' familiarity with learning resources; , , , , These are the weighting coefficients; To broaden the knowledge coverage of learning resource i; The relevance of learning resources to the user's current learning progress and needs; The contribution of user interest in learning resource i to the cumulative quality of learning resources; The frequency of user interaction with learning resource i; The maximum frequency of interaction between the user and the learning resource i; The diversity gain after adding learning resource i to the learning resource set S; Let cosine similarity be the similarity between learning resource i and learning resource j. .
6. The learning resource recommendation method based on user profiles as described in claim 5, characterized in that, The constraints of the multi-objective learning resource recommendation optimization model include learning time constraints and learning resource difficulty constraints. The learning time constraints are specifically as follows: ; The specific difficulty constraints of the learning resources are as follows: ; In the formula, B represents the user's learning saturation.
7. A learning resource recommendation system based on user profiles, characterized in that, The system includes: The data analysis module is used to build user profile matrices and learning resource feature matrices based on the collected user profile information and learning resource information. The latent semantic analysis module is used to construct a latent semantic matching model between users and learning resources based on the user profile matrix and the learning resource feature matrix, using factorization machine and collaborative filtering algorithm. The parameter learning module is used to learn the parameters of the latent semantic matching model through alternating least squares method and generate user-learning resource latent semantic association features. The multi-objective optimization module is used to establish a multi-objective learning resource recommendation optimization model with the optimization objectives of minimizing user learning time and maximizing the quality of accumulated learning resources; and to use the implicit semantic association features between users and learning resources as individual codes and a non-dominated sorting genetic algorithm to search for the Pareto optimal recommendation solution set of the multi-objective learning resource recommendation optimization model. The recommendation decision module is used to construct a multi-dimensional learning state representation based on user profiles and learning resources, and use the multi-dimensional learning state representation as the state space of the deep Q-network algorithm. The deep Q-network algorithm is used to estimate the long-term cumulative returns of different learning resources and autonomously learn to obtain the optimal recommendation decision model. The module also generates a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model. Specifically, the construction of a multidimensional learning state representation based on user profiles and learning resources includes: Obtain interaction data between user profile information and learning resource information, evaluate the connection strength between users and learning resources based on the interaction data, and construct a user-learning resource adjacency matrix; Define multidimensional attributes for each learning resource based on the learning resource information, and construct a multidimensional attribute matrix for the learning resources. The user-learning resource adjacency matrix and the learning resource multidimensional attribute matrix are concatenated to form a learning feature concatenation matrix. The concatenated matrix of the learned features is input into the graph neural network to learn the feature representations of users and learning resources, thereby obtaining user feature vectors and learning resource feature vectors. The learning resource information is sequentially categorized and multi-dimensional information entropy is extracted to obtain multi-dimensional learning resource representation data. The user feature vector, the learning resource feature vector, and the multidimensional learning resource representation data are weighted and fused to obtain a multidimensional learning state representation. The step of generating a learning resource recommendation scheme based on the Pareto optimal recommendation solution set and the optimal recommendation decision model specifically includes: Iterate through each candidate solution in the Pareto optimal recommendation solution set and calculate the matching degree between the candidate solution and the user's current progress requirement; Based on the matching degree, the recommended candidate set with the highest matching degree with the user's current progress needs is selected; Based on the optimal recommendation decision model, obtain the Q value of each learning resource in the recommendation candidate set; Based on the Q-value of each learning resource, the learning resource with the highest Q-value is selected from the recommendation candidate set to generate a learning resource recommendation scheme.