Method and device for determining recommended learning resources
By selecting appropriate recommendation algorithms based on user characteristic data, the problem that the existing recommendation system is not accurate enough when recommending learning resources is solved, and a more accurate and personalized recommendation effect is achieved.
Patent Information
- Application Number
- CN202110405603.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-04-15
AI Technical Summary
The existing recommendation system is not accurate enough when recommending learning resources to meet users' personalized needs.
By determining the characteristic data of the target user, determining whether it meets preset conditions, and then selecting appropriate recommendation algorithms (such as clustering algorithms, association rules algorithms or collaborative filtering algorithms) to recommend learning resources.
It improves the accuracy of recommended learning resources and can better meet users' personalized needs.
Smart Images

Figure CN113094584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for determining recommended learning resources. Background Art
[0002] Internet education is a new form of education that combines Internet technology with the field of education. Faced with a vast amount of learning resource information, users often cannot accurately locate the resources they need. The recommendation system is a personalized information recommendation system that recommends information and products of interest to users based on their information needs and interests. However, the learning resource recommendation methods in existing recommendation systems are often not accurate enough and cannot meet the needs of users. Summary of the invention
[0003] In view of this, an embodiment of the present invention provides a method and device for determining recommended learning resources, which can more accurately recommend relevant learning resources to users.
[0004] In a first aspect, an embodiment of the present invention provides a method for determining recommended learning resources, including:
[0005] Determine the characteristic data of target users;
[0006] Determining whether the characteristic data meets a first preset condition;
[0007] Determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition;
[0008] The target recommendation algorithm and the characteristic data of the target user are used to determine the recommended learning resources for the target user.
[0009] Optionally, the feature data includes: user interest tags;
[0010] The determining whether the characteristic data meets the first preset condition includes:
[0011] Determine whether the number of the user interest tags is less than a first number threshold.
[0012] Optionally, determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition includes:
[0013] If the first judgment result indicates that the number of the user interest tags is less than a first number threshold, determining the clustering algorithm as the target recommendation algorithm;
[0014] The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes:
[0015] Determine a user group of the target user by using a clustering algorithm and the characteristic data of the target user;
[0016] Determine recommended learning resources for the target user according to the user group.
[0017] Optionally, determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition includes:
[0018] If the first judgment result indicates that the number of the user interest tags is greater than or equal to a first number threshold, determining the association rule algorithm as the target recommendation algorithm;
[0019] The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes:
[0020] Determine similar users of the target user by using an association rule algorithm and the characteristic data of the target user;
[0021] Determine recommended learning resources for the target user based on the similar users.
[0022] Optionally, determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition includes:
[0023] If the first judgment result indicates that the number of the user interest tags is greater than or equal to a first number threshold, then determining the collaborative filtering algorithm as the target recommendation algorithm;
[0024] The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes:
[0025] The recommended learning resources for the target user are determined by using a collaborative filtering algorithm and the characteristic data of the target user.
[0026] Optionally, determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition includes:
[0027] If the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, then judging whether the number of users in the system meets the second preset condition, the second preset condition includes: the number of users in the system is less than the second number threshold;
[0028] A target recommendation algorithm is determined according to a second judgment result corresponding to the second preset condition.
[0029] Optionally, determining a target recommendation algorithm according to a second judgment result corresponding to the second preset condition includes:
[0030] If the second judgment result indicates that the number of users in the system is greater than or equal to the second number threshold, determining the association rule algorithm as the target recommendation algorithm;
[0031] The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes:
[0032] Determine similar users of the target user by using an association rule algorithm and the characteristic data of the target user;
[0033] Determine recommended learning resources for the target user based on the similar users.
[0034] Optionally, determining a target recommendation algorithm according to a second judgment result corresponding to the second preset condition includes:
[0035] If the second judgment result indicates that the number of users in the system is less than the second number threshold, determining the collaborative filtering algorithm as the target recommendation algorithm;
[0036] The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes:
[0037] The recommended learning resources for the target user are determined by using a collaborative filtering algorithm and the characteristic data of the target user.
[0038] Optionally, it also includes:
[0039] Acquire a first training set, where the first training set includes a plurality of learning resources;
[0040] Obtaining the first behavior data of the test user on the training set;
[0041] Determining a first recommendation list for the test user according to the first behavior data;
[0042] Determine a first behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the first training set;
[0043] The accuracy of the method for determining the recommended learning resources is determined according to the first recommendation list and the first behavior list.
[0044] Optionally, it also includes:
[0045] Acquire a second training set, where the second training set includes a plurality of learning resources;
[0046] Obtaining second behavior data of the test user on the training set;
[0047] Determining a second recommendation list for the test user according to the second behavior data;
[0048] Determine a second behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the second training set;
[0049] The recall rate of the method for determining the recommended learning resources is determined according to the second recommendation list and the second behavior list.
[0050] Optionally, determining the characteristic data of the target user includes:
[0051] Obtaining a log file of the target user;
[0052] Preprocessing the log file, wherein the preprocessing includes one of the following: word segmentation, feature word recognition, and invalid character filtering;
[0053] The processed log files are statistically analyzed to determine the characteristic data of the target user.
[0054] Optionally, the characteristic data includes: explicit data and invisible data;
[0055] The explicit data includes at least one of the following: resource comments, resource likes, resource forwarding, resource downloads, user rating feedback, downloaded resources, question records, course resource searches, number of interactions with courses, duration of each interaction, and system online time;
[0056] The invisible data includes at least one of the following: browsing history, reading time, viewing record, and search log.
[0057] In a second aspect, an embodiment of the present invention provides a device for determining recommended learning resources, including:
[0058] A data determination module, used to determine characteristic data of target users;
[0059] A condition judgment module, used to judge whether the characteristic data meets a first preset condition;
[0060] An algorithm determination module, used to determine a target recommendation algorithm according to a first judgment result corresponding to the first preset condition;
[0061] The resource determination module is used to determine the recommended learning resources for the target user by using the target recommendation algorithm and the characteristic data of the target user.
[0062] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0063] one or more processors;
[0064] a storage device for storing one or more programs,
[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above embodiments.
[0066] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the above embodiments.
[0067] An embodiment of the above invention has the following advantages or beneficial effects: determining a target recommendation algorithm based on whether the characteristic data of the target user meets the first preset condition, and then determining the recommended learning resources for the target user using the target recommendation algorithm and the characteristic data of the target user. The embodiment of the present invention does not adopt a fixed recommendation algorithm, but determines the target recommendation algorithm based on the characteristic data of the user, so that the recommended learning resources obtained are more accurate and can better meet the needs of the user.
[0068] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific implementation examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The accompanying drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention.
[0070] Figure 1 is an exemplary system architecture diagram to which embodiments of the present invention may be applied;
[0071] Figure 2 is a schematic diagram of a process of a method for determining recommended learning resources provided by an embodiment of the present invention;
[0072] Figure 3 is a schematic diagram of a process of another method for determining recommended learning resources provided by an embodiment of the present invention;
[0073] Figure 4 is a schematic diagram of a process of another method for determining recommended learning resources provided by an embodiment of the present invention;
[0074] Figure 5 is a structural schematic diagram of a device for determining recommended learning resources provided by an embodiment of the present invention;
[0075] Figure 6 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0076] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0077] Figure 1 An exemplary system architecture 100 is shown to which the method for determining recommended learning resources or the apparatus for determining recommended learning resources according to the embodiments of the present invention can be applied.
[0078] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0079] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, office applications, etc. (only examples).
[0080] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0081] The server 105 may be a server that provides various services, such as a backend management server (for example only) that supports online education websites browsed by users using the terminal devices 101, 102, and 103. The server 105 may determine a target recommendation algorithm based on whether the characteristic data of the target user meets the first preset condition; and then use the target recommendation algorithm to determine the recommended learning resources for the target user.
[0082] It should be noted that the method for determining recommended learning resources provided in the embodiment of the present invention is generally executed by the server 105 , and accordingly, the device for determining the method for determining recommended learning resources is generally set in the server 105 .
[0083] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0084] Figure 2 is a schematic diagram of a process of determining a recommended learning resource provided by an embodiment of the present invention. Figure 2 As shown, the method includes:
[0085] Step 201: Determine characteristic data of a target user.
[0086] Collect relevant feature data of target users to generate user profiles for recommendation tasks. These feature data include user attributes, user behavior data, or resources accessed by users. By building user profiles, the recommendation system recommends learning resources to users.
[0087] Depending on the different types of input, feature data can be explicit data or implicit data. For example, the most direct explicit feedback can be the user directly inputting the content of interest. Implicit feedback can be indirectly inferring user preferences by observing user behavior, and mixed feedback can be obtained by combining explicit and implicit feedback.
[0088] Explicit data refers to data actively input by users, such as comments, likes, reposts, downloads, etc. Implicit data refers to users' browsing history, reading time, viewing records, search logs, etc. The backend will create a data set for each user who uses the product / visits the site.
[0089] Specifically, explicit data may include: resource comments, resource likes, resource forwarding, resource downloads, user rating feedback, downloaded resources, question records, course resource searches, number of interactions with courses, duration of each interaction, system online time, etc. Invisible data may include: browsing history, reading time, viewing records, search logs, etc.
[0090] Taking the online learning platform as an example, a user portrait is a collection of personal information associated with a specific user. This information includes the user's cognitive skills, intelligence level, learning style, hobbies, and interactive behaviors. User portraits are usually used for information retrieval when building user models. In other words, user portraits roughly reflect the user model. To make a successful recommendation system, it depends largely on its ability to represent user interests. To obtain accurate recommendation results, an accurate user model is essential.
[0091] Since each user has different preferences for products, the feature data sets collected for each user are also completely different. As time goes by, more and more user feature data is collected, and through a series of data analysis, the recommendation results will become more and more accurate.
[0092] User feature data can be represented by a data matrix, and the data matrix is standardized using the translation standard deviation transformation and the maximum and minimum transformation. Standardization of the data matrix can effectively avoid numerical problems, because large data may cause differences in orders of magnitude and have a greater impact on the final data classification results; at the same time, standardization can also balance the impact of user feature data on K-means clustering and association rule analysis, achieving better classification and recommendation results.
[0093] Step 202: Determine whether the characteristic data meets a first preset condition.
[0094] The first preset condition can be set according to specific needs and actual application environment. For example, the first preset condition can be: whether the number of user interest tags is less than the first number threshold, whether the number of users is less than the second number threshold, whether the number of learning resources is less than the third number threshold, etc.
[0095] Step 203: Determine a target recommendation algorithm according to a first judgment result corresponding to the first preset condition.
[0096] The target recommendation algorithm can be set according to specific needs and actual application environment. For example, the target recommendation algorithm can be a content-based recommendation algorithm, a collaborative filtering recommendation algorithm, a knowledge-based recommendation algorithm, etc.
[0097] Step 204: Determine the recommended learning resources for the target user by using the target recommendation algorithm and the characteristic data of the target user.
[0098] Learning resources are resources used to help users learn. Learning resources may include: videos, audios, courseware, e-books, question banks, mock examination papers, etc.
[0099] In the embodiment of the present invention, the target recommendation algorithm is determined according to whether the characteristic data of the target user meets the first preset condition; and the target recommendation algorithm and the characteristic data of the target user are used to determine the recommended learning resources for the target user. The embodiment of the present invention does not adopt a fixed recommendation algorithm, but determines the target recommendation algorithm according to the characteristic data of the user, and the recommended learning resources obtained are more accurate. Therefore, the problem of not being able to recommend more accurate learning resources to users can be solved.
[0100] After the resource is recommended, the results can be calculated by using the precision and recall metrics to effectively evaluate the recommendation effect.
[0101] In one embodiment of the present invention, the method further includes: obtaining a first training set, the first training set including multiple learning resources; obtaining first behavior data of a test user on the first training set; determining a first recommendation list of the test user based on the first behavior data; determining a first behavior list of the test user using a target recommendation algorithm, feature data of the test user and the first training set; determining the accuracy of the method for determining recommended learning resources based on the first recommendation list and the first behavior list. The accuracy is calculated as shown in the following formula:
[0102]
[0103] Wherein, Precision represents accuracy, R(u) is the first behavior list of the test user on the first test set, the first behavior list is predicted using the method of the embodiment of the present invention, and T(u) is the first recommendation list made based on the test user's behavior on the first training set.
[0104] In one embodiment of the present invention, the method further includes: obtaining a second training set, the second training set including a plurality of learning resources; obtaining second behavior data of the test user on the second training set; determining a second recommendation list of the test user based on the second behavior data; determining a second behavior list of the test user using a target recommendation algorithm, feature data of the test user and the second training set; determining a recall rate of the method for determining recommended learning resources based on the second recommendation list and the second behavior list. The recall rate is calculated as shown in the following formula:
[0105]
[0106] Wherein, Recall represents the recall rate, R(u) is the second behavior list of the test user on the second test set, the second behavior list is predicted using the method of the embodiment of the present invention, and T(u) is the second recommendation list made based on the test user's behavior on the second training set.
[0107] In one embodiment of the present invention, determining the characteristic data of the target user includes: obtaining a log file of the target user; preprocessing the log file, the preprocessing including one of the following: word segmentation, feature word recognition, and invalid character filtering; and performing statistical analysis on the processed log file to determine the characteristic data of the target user. Through the above method, the characteristic data of the target user can be accurately and completely extracted.
[0108] In view of the information overload problem existing in the online education platform of the Internet, the embodiment of the present invention adopts the K-means clustering algorithm to cluster the platform users, thus solving the cold start problem in the recommendation system; through the user's course resource learning record, the collaborative filtering algorithm is adopted, and the Apriori algorithm in the association rule mining technology is used to discover users with similar interests, and a personalized resource recommendation list for the target user is generated based on the course resources that other users are interested in, thus avoiding dependence on the user's explicit data and effectively analyzing the user's implicit feature data. The method of the embodiment of the present invention can provide users with personalized course learning resource recommendations, thus meeting the personalized needs of users. Figure 3 is a schematic diagram of a process of determining a recommended learning resource provided by an embodiment of the present invention. Figure 3 As shown, the method includes:
[0109] Step 301: Determine characteristic data of a target user, the characteristic data including: user interest tags.
[0110] Step 302: Determine whether the number of user interest tags is less than a first number threshold.
[0111] If the first judgment result indicates that the number of user interest tags is less than the first number threshold, execute step 303; if the first judgment result indicates that the number of user interest tags is greater than or equal to the first number threshold, execute step 305 or step 307.
[0112] Step 303: determine the clustering algorithm as the target recommendation algorithm, and use the clustering algorithm and the characteristic data of the target user to determine the user group of the target user.
[0113] The recommendation system has a cold start problem. For new users or users whose interest tags do not meet the number requirements, the mean clustering needs to be performed on the information of existing users because the interest tags of new users cannot be obtained. The mean clustering algorithm is used to cluster the user personal information data of the platform. Users are divided into different clusters, and each cluster represents a group of users with similar interests. The core idea of the K-means clustering algorithm is to first randomly select the initial cluster center, then calculate the Euclidean distance of each sample point to the initial cluster center, and assign them to the class represented by the cluster center with the greatest similarity according to the criterion of the closest distance; finally, the mean of all sample points in each cluster is calculated, and the cluster center is updated until the target criterion function converges.
[0114] Step 304: Determine recommended learning resources for target users based on the user group.
[0115] The recommended learning resources for a user may be determined by the following method.
[0116] Method 1: Extract the first several interest tags of the user group to which the target user belongs, and assign the tags to the target user in the cluster. Through the extracted interest tags, the course resources recommendation for the user's early use of the learning platform is completed.
[0117] Method 2: Determine the learning resources that the users in the user group to which the target user belongs pay attention to or browse more, and recommend these learning resources to the user as recommended learning resources.
[0118] Step 305: Determine the association rule algorithm as the target recommendation algorithm, and use the association rule algorithm and the characteristic data of the target user to determine similar users of the target user.
[0119] Step 306: Determine recommended learning resources for the target user based on similar users.
[0120] In order to recommend relevant resources to users, the Apriori algorithm in the association rule technology is used to mine the association rules between users. The goals of the algorithm analysis include two items: discovering frequent item sets and association rules. First, frequent item sets need to be found, and then association rules can be obtained. The two input parameters of the Apriori algorithm are the minimum support (min_support) and the data set. The algorithm first generates a list of item sets for all individual items, and then scans the data records to see which item sets meet the minimum support requirements. The sets that do not meet the minimum support are removed. Then, the remaining sets are combined to generate an item set containing two elements. Next, the transaction records are rescanned to remove the item sets that do not meet the minimum support. This process is repeated until all item sets are removed.
[0121] In order to complete the subsequent user-based collaborative filtering analysis process, the Apriori algorithm commonly used in association rule mining technology is used to mine association rules between users. The definition of association rule support is shown in the formula below:
[0122] support(A→B)=P(A∪B)
[0123] Among them, P(A∪B represents the probability that a transaction contains the union of sets A and B. The confidence of the association rule is defined as follows:
[0124]
[0125] In order to reduce the computational complexity of generating frequent itemsets, it is necessary to use the support to prune the candidate itemsets. This process needs to follow the following two laws:
[0126] If a set is a frequent itemset, then all its subsets are frequent itemsets. For example, if a set {A, B} is a frequent itemset, that is, the number of times A and B appear in a record at the same time is greater than or equal to the minimum support (min_support), then the number of times its subsets {A} and {B} appear must be greater than or equal to the minimum support, that is, its subsets are all frequent itemsets.
[0127] If a set is not a frequent itemset, then all its supersets are not frequent itemsets. For example: Assume that the set {A} is not a frequent itemset, that is, the number of occurrences of A is less than the minimum support, then the number of occurrences of any of its supersets such as {A, B} must be less than the minimum support, so its superset must not be a frequent itemset either.
[0128] After using association rules to mine associated users with strong association rules with the target user, the target user's score for the target resource can be predicted directly based on the associated user's score for the target resource. The target user has not generated the target behavior for the target resource, and the target behavior may include browsing, learning, collecting, following, evaluating, downloading, etc. Then select several learning resources with the highest predicted scores for the target user. Specifically, determine the score of the associated user for the target resource, and then use the similarity value between the target user and the associated user as the weight, calculate the weighted sum of the scores of each associated user for the target resource, and use the weighted sum as the target user's predicted score for the target resource.
[0129] In one embodiment of the present invention, a method for determining recommended learning resources for a target user comprises the following steps:
[0130] S01: Establish a user rating matrix, which includes the user's ratings of learning resources. The input data of the collaborative filtering algorithm is usually represented as a user-rating matrix. The rating value can be the number of times the user browses, learns, collects, etc., and can also use a displayed rating, such as the user's direct rating of the resource course or the number of times the user learns as the rating.
[0131] S02: Determine at least one nearest neighbor of the target user. This mainly completes the search for the nearest neighbor of the target user, and by calculating the similarity between the target user and other users, calculates the set of "nearest neighbors" that are most similar to the target user. This process is completed in two steps: first, the similarity between users is obtained by using the cosine value similarity calculation method, and second, the "nearest neighbor" is selected according to the following method. The following qualified users can be set as "nearest neighbors", including: selecting users with a similarity greater than a set threshold; selecting the first several users with the largest similarity; selecting several users with a similarity greater than a predetermined threshold, etc.
[0132] S03: Execute the above step 305 and use the association rule algorithm to generate a similar user result set of the target user. Use the association rule algorithm to mine users with strong association rules with the target user and add them to the neighbor set obtained in S02 to generate a similar user result set of the target user.
[0133] S04: Generate relevant recommended resource results, the calculation method is as follows: take the similarity value between the target user and the neighbor user in the similar user result set generated in S03 as the weight, and then determine the target user's predicted score for the target resource based on the neighbor user's score for the target resource. For example, take a weighted average of the difference between the neighbor user's score for the resource and all the scores of this neighbor user. Take a weighted average of the difference between the neighbor user's score for the target resource and the average (or median) of the neighbor user's score, or calculate the average of the weighted sum of the scores of each associated user for the target resource, and take the average of the weighted sum as the target user's predicted score for the target resource. The above method predicts the target user's score for the unrated resources, and then selects several items with the highest predicted scores to recommend to the target user.
[0134] Step 307: determine the collaborative filtering algorithm as the target recommendation algorithm, and determine the recommended learning resources for the target user by using the collaborative filtering algorithm and the characteristic data of the target user.
[0135] Method 1: Determine the recommended learning resources for the target user based on the similar user result set of the target user. The basic idea of this method is to search for the nearest neighbors of the target user based on the similarity between the users' course learning resources, and then generate recommendations to the target user based on the scores of the nearest neighbors. This method is mainly divided into the following steps:
[0136] S11: Establish a user rating matrix, which includes the user's ratings of learning resources. The input data of the collaborative filtering algorithm is usually represented as a user-rating matrix. The rating value can be the number of users' browsing times, learning times, collection times, etc., and can also use displayed ratings, such as the user's direct rating of the resource course or the number of learning times as the rating.
[0137] S12: Determine at least one nearest neighbor of the target user. This mainly completes the search for the nearest neighbor of the target user, and by calculating the similarity between the target user and other users, calculates the set of "nearest neighbors" that are most similar to the target user. This process is completed in two steps: first, the similarity between users is obtained by using the cosine value similarity calculation method, and second, the "nearest neighbor" is selected according to the following method. The following qualified users can be set as "nearest neighbors", including: selecting users with a similarity greater than a set threshold; selecting the first several users with the largest similarity; selecting several users with a similarity greater than a predetermined threshold, etc.
[0138] S13: Use an association rule algorithm to generate a similar user result set of the target user. For example, use the Apriori algorithm to mine users with strong association rules with the target user and add them to the neighbor set obtained in S02 to generate a similar user result set of the target user.
[0139] S14: Generate relevant recommended resource results, the calculation method is as follows: take the similarity value between the target user and the neighbor user in the similar user result set generated in S03 as the weight, and then determine the target user's predicted score for the target resource based on the neighbor user's score for the target resource. For example, take a weighted average of the difference between the neighbor user's score for the resource and all the scores of this neighbor user. Take a weighted average of the difference between the neighbor user's score for the target resource and the average (or median) of all the scores of this neighbor user, or calculate the average of the weighted sum of the scores of each associated user for the target resource, and take the average of the weighted sum as the target user's predicted score for the target resource. The above method predicts the target user's score for the unrated resources, and then selects several learning resources with the highest predicted scores to recommend to the target user.
[0140] S15: Use collaborative filtering algorithms, such as the ALS algorithm, to iteratively calculate the personalized resource recommendation results for users. The specific process is as follows: (1) Decompose the sparse rating matrix into the product of the user feature vector matrix and the resource item feature vector matrix; (2) Alternately use the least squares method to gradually calculate the user-resource item feature vectors to minimize the mean square error; (3) Use the user-resource item features to predict a user's rating for a resource item and generate personalized recommended learning resources.
[0141] In one embodiment of the present invention, a recommended learning resource set for a final target user is generated based on the recommended learning resource set generated by S14 and S15. Specifically, the union or intersection of the two learning resource sets can be extracted, or a number of learning resources can be extracted from the two learning resource sets respectively as the recommended learning resource set for the final target user. The embodiment of the present invention does not limit how to generate the recommended learning resource set for the final target user based on the recommended learning resource set generated by S14 and S15.
[0142] Method 2: Use the ALS algorithm as the target recommendation algorithm to determine the recommended learning resources for the target user. Specifically, first build a user-resource-rating data matrix, and predict each user's preference for resources through matrix decomposition to complete the recommendation of personalized course resources.
[0143] The ALS algorithm belongs to User-Item collaborative filtering, that is, a hybrid collaborative filtering algorithm. It considers both users and items at the same time. The relationship between users and resource items can be abstracted as the following triple:<User,Item,Rating> . Where User represents the user, Item represents the resource item, and Rating is the user's rating of the resource item, which represents the user's preference for the resource item. Assuming that the system contains m users and n resource items, the rating matrix R is defined as m×n , where the element r ui It represents the rating of the u-th user on the i-th resource item. Since it is impossible for a user to rate all resource items, the R matrix is a sparse matrix.
[0144] Based on the above data characteristics, we can assume that there are several correlation dimensions (age, region, position, etc.) between platform users and resource items. We only need to map the rating matrix R to these characteristic dimensions to obtain the following formula:
[0145]
[0146] Among them, k represents the dimension of the latent factor, and its value is generally 20 to 200. X is the user feature vector matrix, and Y is the resource item feature vector matrix. In order to make the low-rank matrices X and Y as close to R as possible, it is necessary to minimize the following square error function formula:
[0147]
[0148] Among them, r ui represents the rating of the u-th user on the i-th resource item, x u is the feature vector of the u-th user, y i is the feature vector of the i-th resource item. The specific process is as follows: decompose the sparse rating matrix into the product of the user feature vector matrix X and the resource item feature vector matrix Y; alternately use the least squares method to gradually calculate the user-resource item feature vector to minimize the mean square error; predict a user's rating of a resource item through the user-resource item features to generate personalized recommended resources.
[0149] The method of the embodiment of the present invention can achieve the following beneficial effects: a clustering algorithm is used to realize the generation of recommended resources for new users, solving the cold start problem of the recommendation system; an association rule analysis technology is used to mine the strong association rules between resource items, and generate relevant resource recommendation results for users; the ALS algorithm in collaborative filtering is used to reduce the computational complexity of the collaborative filtering recommendation algorithm, utilize the implicit feedback characteristics of users, avoid the data sparsity existing in traditional recommendation technologies, and improve the accuracy and real-time performance of resource recommendation results.
[0150] Figure 4 is a schematic diagram of a process of determining a recommended learning resource provided by an embodiment of the present invention. Figure 4 As shown, the method includes:
[0151] Step 401: Determine characteristic data of a target user, the characteristic data including: user interest tags.
[0152] Step 402: Determine whether the number of user interest tags is less than a first number threshold.
[0153] If the first judgment result indicates that the number of user interest tags is less than the first number threshold, step 403 is executed; if the first judgment result indicates that the number of user interest tags is greater than or equal to the first number threshold, step 404 is executed.
[0154] Step 403: determine the clustering algorithm as the target recommendation algorithm; and determine the recommended learning resources for the target user using the clustering algorithm and the characteristic data of the target user.
[0155] Step 404: Determine whether the number of users in the system is less than a second number threshold.
[0156] If the second judgment result indicates that the number of users in the system is greater than or equal to the second number threshold, step 405 is executed; if the second judgment result indicates that the number of users in the system is less than the second number threshold, step 407 is executed.
[0157] Step 405: determine the association rule algorithm as the target recommendation algorithm, and use the association rule algorithm and the characteristic data of the target user to determine similar users of the target user.
[0158] Step 406: Determine recommended learning resources for the target user based on similar users.
[0159] Step 407: determine the collaborative filtering algorithm as the target recommendation algorithm, and determine the recommended learning resources for the target user by using the collaborative filtering algorithm and the characteristic data of the target user.
[0160] In the embodiment of the present invention, the target recommendation algorithm is determined according to the number of user interest tags and the number of users in the system. When the number of users in the system is large, the associated users of the target user are determined by association rules, and the attribute difference between the determined associated users and the target user is small. Therefore, the recommended learning resources corresponding to the target user can be predicted more accurately through the association rule algorithm.
[0161] When the number of users in the system is small, if the association rules determine the associated users with the target user, the attributes of the determined associated users may be quite different from those of the target user, and the corresponding recommended learning resources for the target user may not be accurately predicted through the associated users. Therefore, when the number of users in the system is small, the collaborative filtering algorithm is used to more accurately predict the corresponding recommended learning resources for the target user.
[0162] Figure 5 FIG. 1 is a schematic diagram of a structure of a device for determining recommended learning resources provided by an embodiment of the present invention. Figure 5 As shown, the method includes:
[0163] Data determination module 501, used to determine characteristic data of target users;
[0164] The condition judgment module 502 is used to judge whether the characteristic data meets the first preset condition;
[0165] An algorithm determination module 503, configured to determine a target recommendation algorithm according to a first judgment result corresponding to a first preset condition;
[0166] The resource determination module 504 is used to determine the recommended learning resources for the target user by using the target recommendation algorithm and the characteristic data of the target user.
[0167] Optionally, the feature data includes: user interest tags;
[0168] The condition judgment module 502 is specifically used for:
[0169] Determine whether the number of user interest tags is less than a first number threshold.
[0170] Optionally, the algorithm determination module 503 is specifically used to:
[0171] If the first judgment result indicates that the number of the user interest tags is less than a first number threshold, the clustering algorithm is determined as the target recommendation algorithm;
[0172] The resource determination module 504 is specifically used for:
[0173] Determine the user group of the target user by using clustering algorithms and the characteristic data of the target user;
[0174] Based on the user group, determine the recommended learning resources for target users.
[0175] Optionally, the condition judgment module 502 is specifically used for:
[0176] If the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, the association rule algorithm is determined as the target recommendation algorithm;
[0177] The resource determination module 504 is specifically used for:
[0178] Use association rule algorithms and target user’s feature data to identify similar users to the target user;
[0179] Determine recommended learning resources for target users based on similar users.
[0180] Optionally, the algorithm determination module 503 is specifically used to:
[0181] If the first judgment result indicates that the number of the user's interest tags is greater than or equal to a first number threshold, then determining the collaborative filtering algorithm as the target recommendation algorithm;
[0182] The resource determination module 504 is specifically used for:
[0183] The collaborative filtering algorithm and the characteristic data of the target users are used to determine the recommended learning resources for the target users.
[0184] Optionally, the algorithm determination module 503 is specifically used to:
[0185] If the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, then judging whether the number of users in the system meets the second preset condition, the second preset condition includes: the number of users in the system is less than the second number threshold;
[0186] A target recommendation algorithm is determined according to a second judgment result corresponding to the second preset condition.
[0187] Optionally, the algorithm determination module 503 is specifically used to:
[0188] If the second judgment result indicates that the number of users in the system is greater than or equal to a second number threshold, determining the association rule algorithm as the target recommendation algorithm;
[0189] The resource determination module 504 is specifically used for:
[0190] Use association rule algorithms and target user’s feature data to identify similar users to the target user;
[0191] Determine recommended learning resources for target users based on similar users.
[0192] Optionally, the algorithm determination module 503 is specifically used to:
[0193] If the second judgment result indicates that the number of users in the system is less than a second number threshold, determining the collaborative filtering algorithm as the target recommendation algorithm;
[0194] The resource determination module 504 is specifically used for:
[0195] The collaborative filtering algorithm and the characteristic data of the target users are used to determine the recommended learning resources for the target users.
[0196] Optionally, it also includes:
[0197] An accuracy determination module 505 is used to obtain a first training set, where the first training set includes a plurality of learning resources;
[0198] Get the first behavior data of the test user in the training set;
[0199] Determine a first recommendation list for the test user according to the first behavior data;
[0200] Determine a first behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the first training set;
[0201] The accuracy of the method for determining the recommended learning resources is determined according to the first recommendation list and the first behavior list.
[0202] Optionally, it also includes:
[0203] A recall rate determination module 506 is used to obtain a second training set, where the second training set includes a plurality of learning resources;
[0204] Obtain the second behavior data of the test user in the training set;
[0205] Determine a second recommendation list for the test user according to the second behavior data;
[0206] Determine a second behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the second training set;
[0207] The recall rate of the method for determining the recommended learning resources is determined according to the second recommendation list and the second behavior list.
[0208] Optionally, the data determination module 501 is specifically used for:
[0209] Get the target user's log file;
[0210] Preprocess the log files, including one of the following: word segmentation, feature word recognition, and invalid character filtering;
[0211] Perform statistical analysis on the processed log files to determine the characteristic data of the target users.
[0212] Optionally, the characteristic data includes: explicit data and invisible data;
[0213] Explicit data includes at least one of the following: resource comments, resource likes, resource forwarding, resource downloads, user rating feedback, downloaded resources, question record, course resource search, number of interactions with the course, duration of each interaction, and system online time;
[0214] Invisible data includes at least one of the following: browsing history, reading time, viewing history, and search log.
[0215] An embodiment of the present invention provides an electronic device, including:
[0216] one or more processors;
[0217] a storage device for storing one or more programs,
[0218] When one or more programs are executed by one or more processors, the one or more processors implement the method of any of the above embodiments.
[0219] Reference below Figure 6 , which shows a schematic diagram of the structure of a computer system 600 of a terminal device suitable for implementing an embodiment of the present invention. Figure 6 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0220] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0221] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.
[0222] In particular, according to the embodiments disclosed in the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present invention are executed.
[0223] It should be noted that the computer-readable medium shown in the present invention may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0224] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0225] The modules involved in the embodiments of the present invention may be implemented by software or hardware. The modules described may also be set in a processor. For example, they may be described as: a processor includes a data determination module, a condition judgment module, an algorithm determination module, and a resource determination module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the data determination module may also be described as a "module for determining characteristic data of a target user."
[0226] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes:
[0227] Determine the characteristic data of target users;
[0228] Determining whether the characteristic data meets a first preset condition;
[0229] Determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition;
[0230] The target recommendation algorithm and the characteristic data of the target user are used to determine the recommended learning resources for the target user.
[0231] According to the technical solution of the embodiment of the present invention, the target recommendation algorithm is determined according to whether the characteristic data of the target user meets the first preset condition; and the recommended learning resources for the target user are determined by using the target recommendation algorithm and the characteristic data of the target user. The embodiment of the present invention does not adopt a fixed recommendation algorithm, but determines the target recommendation algorithm according to the characteristic data of the user, and the recommended learning resources obtained are more accurate and can better meet the needs of the user.
[0232] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for determining recommended learning resources, characterized in that: include: Determine the characteristic data of target users; The characteristic data includes: user interest tags; Determining whether the characteristic data meets a first preset condition includes: determining whether the number of the user interest tags is less than a first number threshold; Determining a target recommendation algorithm according to a first judgment result corresponding to the first preset condition, including: if the first judgment result indicates that the number of the user interest tags is less than a first number threshold, determining a clustering algorithm as the target recommendation algorithm; or if the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, determining an association rule algorithm as the target recommendation algorithm; or if the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, determining a collaborative filtering algorithm as the target recommendation algorithm; Determine the recommended learning resources for the target user by using the target recommendation algorithm and the characteristic data of the target user, including: determine the user group of the target user by using a clustering algorithm and the characteristic data of the target user, and determine the recommended learning resources for the target user based on the user group; or determine similar users of the target user by using an association rule algorithm and the characteristic data of the target user, and determine the recommended learning resources for the target user based on the similar users; or determine the recommended learning resources for the target user by using a collaborative filtering algorithm and the characteristic data of the target user.
2. The method according to claim 1, characterized in that The determining of the target recommendation algorithm according to the first judgment result corresponding to the first preset condition includes: If the first judgment result indicates that the number of the user interest tags is greater than or equal to the first number threshold, then judging whether the number of users in the system meets the second preset condition, the second preset condition includes: the number of users in the system is less than the second number threshold; A target recommendation algorithm is determined according to a second judgment result corresponding to the second preset condition.
3. The method according to claim 2, characterized in that The determining of the target recommendation algorithm according to the second judgment result corresponding to the second preset condition includes: If the second judgment result indicates that the number of users in the system is greater than or equal to the second number threshold, determining the association rule algorithm as the target recommendation algorithm; The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes: Determine similar users of the target user by using an association rule algorithm and the characteristic data of the target user; Determine recommended learning resources for the target user based on the similar users.
4. The method according to claim 2, characterized in that: The determining of the target recommendation algorithm according to the second judgment result corresponding to the second preset condition includes: If the second judgment result indicates that the number of users in the system is less than the second number threshold, determining the collaborative filtering algorithm as the target recommendation algorithm; The step of using the target recommendation algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user includes: The recommended learning resources for the target user are determined by using a collaborative filtering algorithm and the characteristic data of the target user.
5. The method according to claim 1, characterized in that Also includes: Acquire a first training set, where the first training set includes a plurality of learning resources; Obtaining first behavior data of a test user on the first training set; Determining a first recommendation list for the test user according to the first behavior data; Determine a first behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the first training set; The accuracy of the method for determining the recommended learning resources is determined according to the first recommendation list and the first behavior list.
6. The method according to claim 1, characterized in that Also includes: Acquire a second training set, where the second training set includes a plurality of learning resources; Acquire second behavior data of the test user on the second training set; Determining a second recommendation list for the test user according to the second behavior data; Determine a second behavior list of the test user by using the target recommendation algorithm, the feature data of the test user and the second training set; The recall rate of the method for determining the recommended learning resources is determined according to the second recommendation list and the second behavior list.
7. The method according to claim 1, characterized in that The determining of characteristic data of the target user includes: Obtaining a log file of the target user; Preprocessing the log file, wherein the preprocessing includes one of the following: word segmentation, feature word recognition, and invalid character filtering; The processed log files are statistically analyzed to determine the characteristic data of the target user.
8. The method according to claim 1, characterized in that The characteristic data includes: explicit data and invisible data; The explicit data includes at least one of the following: resource comments, resource likes, resource forwarding, resource downloads, user rating feedback, downloaded resources, question records, course resource searches, number of interactions with courses, duration of each interaction, and system online time; The invisible data includes at least one of the following: browsing history, reading time, viewing record, and search log.
9. A device for determining recommended learning resources, characterized in that: include: A data determination module, used to determine characteristic data of target users; The characteristic data includes: user interest tags; A condition judgment module, used to judge whether the characteristic data meets a first preset condition, including: judging whether the number of the user interest tags is less than a first number threshold; An algorithm determination module, used to determine a target recommendation algorithm according to a first judgment result corresponding to the first preset condition, including: if the first judgment result indicates that the number of user interest tags is less than a first number threshold, a clustering algorithm is determined as the target recommendation algorithm; or if the first judgment result indicates that the number of user interest tags is greater than or equal to the first number threshold, an association rule algorithm is determined as the target recommendation algorithm; or if the first judgment result indicates that the number of user interest tags is greater than or equal to the first number threshold, a collaborative filtering algorithm is determined as the target recommendation algorithm; A resource determination module is used to determine the recommended learning resources for the target user by using the target recommendation algorithm and the characteristic data of the target user, including: using a clustering algorithm and the characteristic data of the target user to determine the user group of the target user, and determining the recommended learning resources for the target user based on the user group; or using an association rule algorithm and the characteristic data of the target user to determine similar users of the target user, and determining the recommended learning resources for the target user based on the similar users; or using a collaborative filtering algorithm and the characteristic data of the target user to determine the recommended learning resources for the target user.
10. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
11. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Accurate recommendation system and method based on parallel processing of collaborative filtering and association rules
CN107220365A
Online user group classification method and device based on clustering and association rules
CN110532429A