Learning resource recommendation method and system based on big data

By adopting the combination of dual neighbor sets and near-neighbor affinity in the learning resource recommendation method, the problems of high-dimensional, nonlinear correlation and sparse data processing are solved, and more accurate student group division and learning resource recommendation are achieved.

CN120086450AActive Publication Date: 2025-06-03TAISHAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510587447.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-03
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing learning resource recommendation methods are difficult to deal with high-dimensional, nonlinear correlation and highly sparse student learning behavior data, resulting in rough student population division and low accuracy in learning resource recommendation.

Method used

The dual neighborhood set is used to calculate the cumulative local density of the neighborhood, and combine it with the relative distance to obtain the decision value, determine the initial clustering center, divide the data according to the number of data in the intersection of neighbors, introduce the affinity of the nearest neighbor to calculate the comprehensive correlation intensity, complete the clustering of free data, and optimize the cluster structure according to the fusion degree between clusters to obtain the student clustering model.

Benefits of technology

It can more accurately identify student data points that are relatively important in the learning behavior characteristic space, capture students' diverse learning needs, significantly improve the rationality and pertinence of learning resource recommendations, and improve the accuracy of learning resource recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086450A_ABST
    Figure CN120086450A_ABST
Patent Text Reader

Abstract

The invention discloses a learning resource recommendation method and system based on big data, and belongs to the technical field of resource recommendation, and the method comprises the steps of learning data collection, learning data preprocessing, student clustering, student clustering model parameter optimization and learning resource recommendation list generation. According to the scheme, the neighborhood cumulative local density is calculated based on a dual neighborhood set, data division is performed according to the number of data in a neighbor intersection, clustering of condensed data is completed, neighbor affinity is introduced to calculate the comprehensive association strength, clustering of free data is completed, and a cluster structure is optimized according to the inter-cluster fusion degree; and calculating the quality of each exploration position based on the fitness value, constructing an optimal set, calculating the influence of each exploration position on each dimension, updating each exploration position based on the upper and lower limits of the parameter search space and the influence, and finding the optimal student clustering model parameters, thereby accurately providing appropriate learning resources for different students, and improving the learning efficiency. And the accuracy and effectiveness of learning resource recommendation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of resource recommendation, and specifically refers to a learning resource recommendation method and system based on big data. Background Art

[0002] The learning resource recommendation method is to use big data analysis and artificial intelligence technology to integrate students' multi-dimensional learning data, deeply explore the potential rules in massive learning data, and intelligently recommend personalized learning resources, thereby improving learning efficiency and overall learning outcomes. However, the existing learning resource recommendation methods have the problem of difficulty in handling high-dimensional, nonlinearly associated and highly sparse student learning behavior data, and cannot effectively capture its internal structure, resulting in rough division of student groups and low accuracy of learning resource recommendation; the existing learning resource recommendation methods have the problem of large differences between different students' learning behavior data, and it is difficult to effectively adapt to all students' learning behavior data, resulting in unstable division of student groups and inability to accurately recommend learning resources. Summary of the invention

[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a learning resource recommendation method and system based on big data. In view of the problems that the existing learning resource recommendation methods are difficult to handle high-dimensional, nonlinearly associated and highly sparse student learning behavior data, and cannot effectively capture their internal structure, resulting in rough division of student groups and low accuracy of learning resource recommendation, this scheme calculates the neighborhood cumulative local density based on the dual neighborhood set, and combines it with the relative distance to obtain the decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerated data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain the student clustering model, which can more accurately identify the students in The relatively important student data points in the learning behavior feature space can more accurately capture the diverse learning needs of students and improve the accuracy of learning resource recommendations. In view of the large differences between different students' learning behavior data in the existing learning resource recommendation methods, it is difficult to effectively adapt to all students' learning behavior data, resulting in unstable student group division and inability to accurately recommend learning resources. This solution calculates the quality of each exploration position based on the fitness value, constructs a preferred set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits and influence of the parameter search space, finds the optimal student clustering model parameters, makes it more suitable for the learning data characteristics of different students, improves the stability of student group division, and thus accurately provides suitable learning resources for different students.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides a learning resource recommendation method based on big data, which comprises the following steps:

[0005] Step S1: Learning data collection;

[0006] Step S2: Preprocessing of learning data;

[0007] Step S3: Student clustering;

[0008] Step S4: Optimization of student clustering model parameters;

[0009] Step S5: Generating a recommended list of learning resources.

[0010] Further, in step S1, the learning data collection is to collect historical student learning data, real-time student learning data, and learning resource metadata.

[0011] Further, in step S2, the preprocessing of learning data is to perform data cleaning, data standardization, and construction of data sets on the collected historical student learning data, real-time student learning data, and learning resource metadata, to obtain a student learning data set and a learning resource metadata set.

[0012] Further, in step S3, the student clustering specifically includes the following steps:

[0013] Step S31: Selecting an initial clustering center; including the following steps:

[0014] Step S311: Calculating the neighborhood cumulative local density; calculating the Euclidean distance between any two data in the student learning data set, finding the K nearest neighbors of each data and constructing a neighborhood set, and calculating the neighborhood cumulative local density of each data based on the dual neighborhood set; the formula used is as follows:

[0015] ;

[0016] where x i , x j and x m are the i-th, j-th, and m-th data in the student learning data set respectively, i, j, and m are data indices, ρ i is the neighborhood cumulative local density of x i , is the Euclidean distance between x j and x m , and are the neighborhood sets of x i and x j respectively, and K is the number of nearest neighbors;

[0017] Step S312: Determine; calculate the relative distance and decision value of each data in the student learning data set, select the top N data with the largest decision value from the student learning data set as the initial cluster centers, and construct an initial cluster center set H; where N is the number of initial cluster centers;

[0018] Step S32: Data clustering; includes the following steps:

[0019] Step S321: Data partitioning; based on the data in the student learning data set except the initial cluster centers, construct a data set to be assigned, and calculate the neighbor intersection and the number of data in the neighbor intersection between each data in the data set to be assigned and each initial cluster center and the number of data in the neighbor intersection , if there exists at least one initial cluster center such that the number of data in the neighbor intersection between x a and x m is , then classify x a as cohesive data; otherwise, classify x a as free data; where x a is the a-th data in the data set to be assigned, x m is the m-th initial cluster center, a and m are data indices and initial cluster center indices respectively, is the neighbor intersection between x a and x m , is the number of data in;

[0020] Step S322: Cohesive data clustering; calculate the number of data in the neighbor intersection between each cohesive data and each initial cluster center, and assign each cohesive data x v to the cluster where the initial cluster center x belongs that satisfies; where, m is the number of data in the neighbor intersection between the neighborhood set of x and the neighborhood set of x v , x m is the v-th cohesive data, x v is the m-th initial cluster center, v is the data index, m is the maximum value function; is the maximum value function;

[0021] Step S323: Free data clustering; based on the initial cluster centers and cohesive data, construct an assigned data set A, introduce the neighbor affinity, calculate the comprehensive association strength between each free data and each data in the assigned data set, and assign each free data x o to the x that satisfies inb in the cluster it belongs to; the formula used is as follows:

[0022] ;

[0023] ;

[0024] In the formula, x b is the b-th data in the allocated data set, x o is the o-th free data, b and o are data indices, and are the neighborhood sets of x o and x b respectively, x c is the c-th data, c is a data index, , and are the neighbor affinities between x o and x b , x o and x c and x b and x c respectively, is the comprehensive association strength between x o and x b , is and the number of data in the neighbor intersection between them, is the Euclidean distance between x o and x b ;

[0025] Step S33: Cluster structure optimization; calculate the inter-cluster fusion degree between every two clusters, and merge the two clusters with the maximum inter-cluster fusion degree; the formula used is as follows:

[0026] ;

[0027] In the formula, C n1 and C n2 are the n1-th and n2-th clusters respectively, is the inter-cluster fusion degree between C n1 and C n2 , x p and x q are the p-th and q-th data respectively, S n1 and S n2 are the number of data in C n1 and C n2 respectively, is the distance between x q and x pThe pairwise affinity between them, where p and q are data indices, and n1 and n2 are cluster indices;

[0028] Step S34: Model determination; Repeat step S33 until the number of clusters is equal to B, at which point the student clustering is completed, obtaining the student clustering model and outputting the clustering result; where B is the number of true clusters.

[0029] Furthermore, in step S4, the optimization of the student clustering model parameters specifically includes the following steps:

[0030] Step S41: Initialization; Establish a parameter search space for the number of nearest neighbors K, the number of initial cluster centers N, and the number of true clusters B in the student clustering model. Randomly initialize Z 1 exploration positions within the parameter search space. Use each exploration position to represent a set of student clustering model parameters, and take the average of all silhouette coefficients of the data in the clustering result obtained based on the exploration position as the fitness value corresponding to the exploration position; where Z 1 is the number of exploration positions;

[0031] Step S42: Calculate quality; Normalize the fitness value of each exploration position, and then determine the quality of each exploration position according to the normalized fitness value;

[0032] Step S43: Calculate influence; Construct a preferred set based on the top exploration positions with larger quality, and calculate the influence of each exploration position in each dimension based on the preferred set; The formula used is as follows:

[0033] ;

[0034] where E t is the preferred set at the t-th search, is the z-th exploration position at the t-th search, is the h-th exploration position in E t at the t-th search, and are respectively and the values in the r-th dimension, is the influence of the z-th exploration position at the t-th search in the r-th dimension, r is the dimension index, h is the exploration position index, s r is a random value within the range [0, 1] corresponding to the r-th dimension, is the L2 norm, is the quality of, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search number index, and T is the maximum number of searches. is an exponential function, is floor function, is a smoothing term;

[0035] Step S44: Update the exploration positions; update each exploration position based on the upper and lower bounds of the parameter search space and the influence; the formula used is as follows:

[0036] ;

[0037] In the formula, is the value of the z-th exploration position at the (t + 1)-th search in the r-th dimension, UB r and LB r are respectively the upper and lower bounds of the parameter search space in the r-th dimension;

[0038] Step S45: Parameter determination; preset a fitness threshold, update the fitness values of the exploration positions. When there is an exploration position whose fitness value is higher than the fitness threshold, the student clustering model parameters corresponding to this exploration position are the optimal parameters; otherwise, if the maximum search number is reached, return to Step S41 to re-initialize; otherwise, increment the search number by 1 and return to Step S42 to continue the search.

[0039] Further, in Step S5, the generating of the learning resource recommendation list is to input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing. Based on the output clustering result, determine the cluster D to which the real-time student learning data belongs. Construct a to-be-recommended resource set based on the resource IDs included in all historical student learning data in cluster D. Combine with the learning resource metadata set to obtain the recommendation value corresponding to each resource ID in the to-be-recommended resource set, and sort the resource IDs in the to-be-recommended resource set in descending order according to the recommendation value to obtain the learning resource recommendation list.

[0040] The learning resource recommendation system based on big data provided by the present invention includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module, and a generating learning resource recommendation list module;

[0041] The learning data collection module collects historical student learning data, real-time student learning data, and learning resource metadata, and sends the data to the learning data preprocessing module;

[0042] The learning data preprocessing module performs data cleaning, data standardization, and constructing data set processing on the collected data to obtain a student learning data set and a learning resource metadata set, and sends the data to the student clustering module;

[0043] The student clustering module calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerative data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, and sends the data to the student clustering model parameter optimization module;

[0044] The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, and sends the data to the generation of learning resource recommendation list module;

[0045] The generation of learning resource recommendation list module inputs the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructs a set of resources to be recommended based on the output clustering results, calculates the recommendation value of each resource ID, and generates a learning resource recommendation list.

[0046] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0047] (1) Aiming at the problem that the existing learning resource recommendation methods are difficult to process the student learning behavior data with high dimensionality, non-linear correlation and high sparsity, and cannot effectively capture its internal structure, resulting in rough division of the student group and low accuracy of learning resource recommendation. This scheme calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerative data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, which can more accurately identify the relatively important student data points in the learning behavior feature space, distinguish different types of students, more accurately capture the diverse learning needs of students, significantly improve the rationality and pertinence of learning resource recommendation, and improve the accuracy of learning resource recommendation.

[0048] (2)In view of the problem that there are significant differences among the learning behavior data of different students in the existing learning resource recommendation methods, making it difficult to effectively adapt to the learning behavior data of all students, resulting in unstable student group division and inaccurate learning resource recommendation, this solution uses exploration positions to represent a set of student clustering model parameters, calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, makes it more suitable for the learning data characteristics of different students, improves the stability of student group division, and can recommend resources based on a more accurate student group division, thereby accurately providing appropriate learning resources for different students and improving the accuracy and effectiveness of learning resource recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of the learning resource recommendation method based on big data provided by the present invention;

[0050] Figure 2 is a schematic diagram of the learning resource recommendation system based on big data provided by the present invention;

[0051] Figure 3 is a schematic flowchart of step S3;

[0052] Figure 4 is a schematic flowchart of step S4.

[0053] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0056] Example 1, refer toFigure 1 , the learning resource recommendation method based on big data provided by the present invention includes the following steps:

[0057] Step S1: Learning data collection; collecting historical student learning data, real-time student learning data, and learning resource metadata;

[0058] Step S2: Preprocessing of learning data; performing data cleaning, data standardization, and constructing data set processing on the collected data to obtain a student learning data set and a learning resource metadata set;

[0059] Step S3: Student clustering; calculating the neighborhood cumulative local density based on the double neighborhood set, combining it with the relative distance to obtain a decision value, determining the initial clustering center, partitioning the data according to the number of data in the neighbor intersection and completing the clustering of agglomerative data, introducing the near neighbor affinity to calculate the comprehensive association strength, completing the clustering of free data, and optimizing the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model;

[0060] Step S4: Optimization of student clustering model parameters; calculating the quality of each exploration position based on the fitness value, constructing an optimal set and calculating the influence of each exploration position on each dimension, updating each exploration position based on the upper and lower limits of the parameter search space and the influence, and finding the optimal student clustering model parameters;

[0061] Step S5: Generating a learning resource recommendation list; inputting the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructing a resource set to be recommended based on the output clustering result, calculating the recommendation value of each resource ID, and generating a learning resource recommendation list.

[0062] Example 2, refer to Figure 1 , based on the above example, in step S1, both the historical student learning data and the real-time student learning data include the browsing behavior, interaction behavior, and learning behavior of students on the learning platform;

[0063] The browsing behavior includes the resource ID accessed, access timestamp, access duration, page scroll depth, and number of repeated accesses;

[0064] The interaction behavior includes favorite records, rating data, sharing records, and download records;

[0065] The learning behavior includes resource learning progress, quiz scores, and types of wrong questions;

[0066] The learning resource metadata includes resource ID, number of times favorited, average rating, number of times accessed, and number of days since the resource was published.

[0067] Example 3, refer to Figure 1, based on the above embodiment, in step S2, data cleaning, data standardization, and data set construction processing are performed on the collected historical student learning data, real-time student learning data, and learning resource metadata to obtain a student learning data set and a learning resource metadata set;

[0068] The data cleaning includes missing value processing and outlier processing. Missing value processing is to fill missing data with the mean value, and outlier processing is to identify and delete outliers using the box plot method;

[0069] The data standardization is to map the data to a unified range using the maximum-minimum scaling method;

[0070] The construction of the data set includes constructing a student learning data set and a learning resource metadata set. The student learning data set is constructed from the historical student learning data and real-time student learning data after data cleaning and data standardization, and the learning resource metadata set is constructed from the learning resource metadata after data cleaning and data standardization.

[0071] Embodiment 4, refer to Figure 1 and Figure 3 , based on the above embodiment, in step S3, student clustering specifically includes the following steps:

[0072] Step S31: Select the initial clustering center; determine the density of each data in its neighborhood in the student learning data set by calculating the neighborhood cumulative local density, so as to find the relatively dense areas in the data distribution; select the initial clustering center by calculating the relative distance and decision value, which can comprehensively consider the local density and relative position relationship of the data, so that the initial clustering center is distributed in the relatively dense and representative areas of the data, and can better divide the student groups with different learning modes, thereby recommending learning resources that are more in line with their learning habits and progress for each group, improving the accuracy and pertinence of the recommendation; including the following steps:

[0073] Step S311: Calculate the neighborhood cumulative local density; calculate the Euclidean distance between any two data in the student learning data set, find the K nearest neighbors of each data and construct a neighborhood set, and calculate the neighborhood cumulative local density of each data based on the double neighborhood set. The value range of K is [5, 50]; the formula used is as follows:

[0074] ;

[0075] In the formula, x i , x j and x m are the i-th, j-th, and m-th data in the student learning data set respectively, i, j, and m are data indices, and ρ i is xi The neighborhood cumulative local density of is x j and x m The Euclidean distance between and are x i and x j The neighborhood sets of , K is the number of nearest neighbors;

[0076] Step S312: Determine; Calculate the relative distance and the decision value of each data in the student learning data set, select the top N data with the largest decision value from the student learning data set as the initial clustering centers, and construct the initial clustering center set H. The value range of N is [B, 2B], and the value range of B is [2, ; where N is the number of initial clustering centers, δ i and γ i are the relative distance and decision value of x i respectively, ρ j is the neighborhood cumulative local density of x j , ρ max is the maximum neighborhood cumulative local density among all data in the student learning data set, is x i and x j The Euclidean distance between , B is the number of true clusters, and M is the number of data in the student learning data set;

[0077] Step S32: Data clustering; Reasonably divide the student learning data into different types for more accurate clustering and resource recommendation. By calculating the neighbor intersection and the number of data in the intersection between each data and the initial clustering centers, divide the data to be assigned into cohesive data and free data, which can analyze the relationship structure of the student learning data more carefully; According to the number of data in the neighbor intersection between the cohesive data and the initial clustering centers, assign it to the most suitable cluster, which helps to construct a tight student cluster, so that students in the same cluster have similar learning behaviors and needs; According to the comprehensive association strength between the free data and the assigned data, assign the free data to the cluster where the assigned data with the largest comprehensive association strength belongs, which can make the free students integrate into the most similar student group; It includes the following steps:

[0078] Step S321: Data division; Based on the data in the student learning data set except the initial clustering centers, construct the data set to be assigned, and calculate the neighbor intersection and the number of data in the neighbor intersection for each data in the data set to be assigned and each initial clustering center. If there is at least one initial clustering center such that xa The number of data in the neighbor intersection between x m and x , then x a is classified as cohesive data; otherwise, x a is classified as free data; where x a is the a-th data in the data set to be assigned, and x m is the m-th initial clustering center, and a and m are the data index and the initial clustering center index respectively, is the neighbor intersection between x a and x m , and are the neighborhood sets of x a and x m respectively, is the number of data in ;

[0079] Step S322: Cohesive data clustering; Calculate the number of data in the neighbor intersection between each cohesive data and each initial clustering center, and assign each cohesive data x v to the cluster where the initial clustering center x belongs; where m is the number of data in the neighbor intersection between the neighborhood set of x and the neighborhood set of x v , x m is the v-th cohesive data, x v is the m-th initial clustering center, and v is the data index, m is the maximum function;

[0080] Step S323: Free data clustering; Based on the initial clustering center and the cohesive data, construct the assigned data set A, introduce the neighbor affinity, calculate the comprehensive association strength between each free data and each data in the assigned data set, and assign each free data x o to the cluster where x b belongs; The formula used is as follows: b ;

[0081] ;

[0082] ;

[0083] In the formula, x b is the b-th data in the assigned data set, x o is the o-th free data, and b and o are the data indices, and They are the neighborhood sets of x o and x b respectively, where x c is the c-th data and c is the data index. , and are the neighbor affinities between x o and x b , x o and x c respectively, and x b and x c respectively. is the comprehensive association strength between x o and x b . is and respectively, which is the number of data in the neighbor intersection between them. is the Euclidean distance between x o and x b .

[0084] Step S33: Cluster structure optimization; in the learning resource recommendation environment, there may be some unreasonable cluster structures in the initial clustering results, which need to be optimized to improve the clustering quality. Calculate the inter-cluster fusion degree between every two clusters, and merge the two clusters with the largest inter-cluster fusion degree to make the clustering structure more compact and reasonable. The optimized cluster structure can reduce the resource recommendation deviation caused by unreasonable clustering. The formula used is as follows:

[0085] ;

[0086] In the formula, C n1 and C n2 are the n1-th and n2-th clusters respectively, is the inter-cluster fusion degree between C n1 and C n2 , x p and x q are the p-th and q-th data respectively, S n1 and S n2 are the number of data in C n1 and C n2 respectively, is the neighbor affinity between x q and x p . p and q are data indices, and n1 and n2 are cluster indices.

[0087] Step S34: Model determination; Repeat step S33 until the number of clusters is equal to B, at which point the student clustering is completed, obtaining the student clustering model and outputting the clustering results, so as to perform precise learning resource recommendation for the characteristics of different clustered student groups, improve the matching degree between resources and student needs, and thus enhance the learning effect; where B is the number of true clusters.

[0088] By performing the above operations, for the problem that the existing learning resource recommendation methods are difficult to handle the high-dimensional, non-linear correlation and highly sparse student learning behavior data, and cannot effectively capture its internal structure, resulting in rough division of student groups and low accuracy of learning resource recommendation, this solution calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain the decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the near-neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain the student clustering model, which can more accurately identify the relatively important student data points in the learning behavior feature space, distinguish different types of students, capture the diverse learning needs of students more accurately, significantly improve the rationality and pertinence of learning resource recommendation, and improve the accuracy of learning resource recommendation.

[0089] Example 5, refer to Figure 1 and Figure 4 , based on the above example, in step S4, the optimization of the student clustering model parameters specifically includes the following steps:

[0090] Step S41: Initialization; By randomly initializing multiple groups of parameters, it provides diverse starting points for subsequent optimization, avoids falling into local optima, combines the silhouette coefficient to evaluate the clustering quality, ensures that the initial parameters have a certain degree of discrimination, and lays a foundation for subsequent optimization; establish a parameter search space for the number of nearest neighbors K, the number of initial clustering centers N, and the number of true clusters B in the student clustering model, randomly initialize Z 1 exploration positions within the parameter search space, use each exploration position to represent a set of student clustering model parameters, and take the average value of all data silhouette coefficients in the clustering results obtained based on the exploration positions as the fitness value corresponding to the exploration position; where Z 1 is the number of exploration positions;

[0091] Step S42: Calculate quality; Normalize the fitness value of each exploration position, and then determine the quality of each exploration position according to the normalized fitness value, reflecting the compactness and separation degree of the clustering results. A high-quality parameter combination can better distinguish the learning behavior patterns of students, so as to recommend more suitable resources for each group; the formula used is as follows:

[0092] ;

[0093] ;

[0094] In the formula, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search number index, is the normalized fitness value of the z-th exploration position at the t-th search, is the fitness value of the z-th exploration position at the t-th search, and are the maximum fitness value and the minimum fitness value at the t-th search respectively;

[0095] Step S43: Calculate the influence; During the parameter optimization process, blind search needs to be avoided. Based on the first exploration positions with larger quality, a preferred set is constructed. Based on the preferred set, the influence of each exploration position in each dimension is calculated to adjust the search towards a better direction and accelerate convergence; The formula used is as follows:

[0096] ;

[0097] where, E t is the preferred set at the t-th search, is the z-th exploration position at the t-th search, is the h-th exploration position in E t at the t-th search, and are respectively and values in the r-th dimension, is the influence of the z-th exploration position at the t-th search in the r-th dimension, r is the dimension index, h is the exploration position index, s r is a random value in the range of [0, 1] corresponding to the r-th dimension, is the L2 norm, is quality, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search number index, T is the maximum search number, is the exponential function, is the floor function, is the smoothing term, , the smoothing term is used to avoid the denominator being 0;

[0098] Step S44: Update the exploration position; Based on the upper and lower limits of the parameter search space and the influence, each exploration position is updated to dynamically update the parameter combination to adapt to the optimization requirements at different stages, thereby improving the recommendation accuracy; The formula used is as follows:

[0099] ;

[0100] wherein, is the value of the z-th exploration position at the (t + 1)-th search in the r-th dimension, UB r and LB r are respectively the upper and lower limits of the parameter search space in the r-th dimension;

[0101] Step S45: Parameter determination; preset a fitness threshold, update the fitness value of the exploration positions. When there is an exploration position whose fitness value is higher than the fitness threshold, the student clustering model parameters corresponding to this exploration position are the optimal parameters; otherwise, if the maximum search number is reached, return to Step S41 to re-initialize; otherwise, increment the search number by 1 and return to Step S42 to continue the search; the finally determined parameters enable the clustering result to stably reflect the distribution of the real student population, thereby supporting long-term accurate recommendation.

[0102] By performing the above operations, for the problem in the existing learning resource recommendation method that there are large differences between different students' learning behavior data, it is difficult to effectively adapt to all students' learning behavior data, resulting in unstable student population division and inaccurate learning resource recommendation, this solution uses the exploration positions to represent a set of student clustering model parameters, calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, and finds the optimal student clustering model parameters, making it more suitable for the learning data characteristics of different students, improving the stability of student population division, being able to recommend resources based on a more accurate student population division, thereby accurately providing suitable learning resources for different students and improving the accuracy and effectiveness of learning resource recommendation.

[0103] Embodiment Six, refer to Figure 1 , this embodiment is based on the above embodiment. In Step S5, generating the learning resource recommendation list is to input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing. Based on the output clustering result, determine the cluster D to which the real-time student learning data belongs, construct a to-be-recommended resource set based on the resource IDs included in all historical student learning data in cluster D, combine with the learning resource metadata set, comprehensively consider the number of times collected, average rating, number of times accessed, and the number of days since the resource was published, obtain the recommendation value corresponding to each resource ID in the to-be-recommended resource set, and sort the resource IDs in the to-be-recommended resource set in descending order according to the size of the recommendation value to obtain the learning resource recommendation list. The formula used is as follows:

[0104] ;

[0105] Among them, P g is the recommended value corresponding to the g-th resource ID in the set of resources to be recommended, where g is the resource ID index, , , and are respectively the number of times of being collected, the average score, the number of times of being accessed, and the number of days since the resource was published corresponding to the g-th resource ID in the set of learning resource metadata in the set of resources to be recommended.

[0106] Example 7. Refer to Figure 2 . Based on the above example, the learning resource recommendation system provided by the present invention includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module, and a module for generating a learning resource recommendation list;

[0107] The learning data collection module collects historical student learning data, real-time student learning data, and learning resource metadata, and sends the data to the learning data preprocessing module;

[0108] The learning data preprocessing module performs data cleaning, data standardization, and construction of a data set on the collected data to obtain a student learning data set and a learning resource metadata set, and sends the data to the student clustering module;

[0109] The student clustering module calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerative data, introduces the near-neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, and sends the data to the student clustering model parameter optimization module;

[0110] The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs a preferred set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, and sends the data to the module for generating a learning resource recommendation list;

[0111] The module for generating a learning resource recommendation list inputs the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructs a set of resources to be recommended based on the output clustering result, calculates the recommended value of each resource ID, and generates a learning resource recommendation list.

[0112] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0113] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

[0114] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and design similar structural forms and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A learning resource recommendation method based on big data, characterized by: The method comprises the following steps: Step S1: Learning data collection: collecting historical student learning data, real-time student learning data and learning resource metadata; Step S2: learning data preprocessing; Step S3: student clustering; Step S4: Optimizing parameters of the student clustering model; Step S5: Generate a learning resource recommendation list; input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, construct a resource set to be recommended based on the output clustering results, calculate the recommendation value of each resource ID, and generate a learning resource recommendation list; Step S3 includes step S321: data partitioning; constructing a data set to be allocated based on the data of the student learning data set except the initial cluster center, and calculating the neighbor intersection between each data in the data set to be allocated and each initial cluster center The number of data in the intersection with neighbors , if there is at least one initial cluster center such that x a With x m The number of data in the intersection of neighbors , then x a Divide into agglomerated data; otherwise, a Divided into free data; where x a is the ath data in the data set to be allocated, x m is the mth initial cluster center, a and m are the data index and the initial cluster center index respectively, is x a With x m The neighbor intersection between yes The number of data in , K is the number of nearest neighbors.

2. The method for recommending learning resources based on big data according to claim 1, characterized in that: In step S3, the student clustering specifically includes the following steps: Step S31: Selecting an initial cluster center; Step S32: data clustering; Step S33: cluster structure optimization; calculate the inter-cluster fusion degree between every two clusters, and merge the two clusters with the largest inter-cluster fusion degree; the formula used is as follows: ; In the formula, C n1 and C n2 are the n1th and n2th clusters respectively, It is C n1 and C n2 The inter-cluster fusion degree, x p and x q are the pth and qth data respectively, S n1 and S n2 C n1 and C n2 The number of data in is x q With x p The neighbor affinity between them, p and q are data indexes, n1 and n2 are cluster indexes; Step S34: Model determination; repeat step S33 until the number of clusters is equal to B, the student clustering is completed, the student clustering model is obtained, and the clustering results are output; where B is the actual number of clusters.

3. The method for recommending learning resources based on big data according to claim 2, characterized in that: In step S31, the selection of the initial cluster center specifically includes the following steps: Step S311: Calculate the neighborhood cumulative local density; calculate the Euclidean distance between any two data in the student learning data set, find the K nearest neighbors of each data and construct a neighborhood set, and calculate the neighborhood cumulative local density of each data based on the double neighborhood set; the formula used is as follows: ; In the formula, x i 、x j and x m are the i-th, j-th and m-th data in the student learning data set, i, j and m are data indexes, ρ i is x i The neighborhood cumulative local density of is x j With x m The Euclidean distance between and They are x i and x j The neighborhood set of ; Step S312: Determine; calculate the relative distance and decision value of each data in the student learning data set, select the first N data with the largest decision value from the student learning data set as the initial clustering centers, and construct an initial clustering center set H; where N is the number of initial clustering centers.

4. The method for recommending learning resources based on big data according to claim 2, characterized in that: In step S32, the data clustering specifically includes the following steps: Step S321: data division; Step S322: clustering the agglomerative data; calculating the number of data in the neighbor intersection between each agglomerative data and each initial cluster center, and clustering each agglomerative data x v Assign to meet The initial cluster center x m The cluster to which it belongs; among them, is x v The neighborhood set of x m The number of data in the neighbor intersection between the neighborhood sets of v is the vth condensed data, x m is the mth initial cluster center, v is the data index, is the maximum value function; Step S323: clustering of free data; constructing the allocated data set A based on the initial cluster center and the condensed data, introducing the neighbor affinity, calculating the comprehensive correlation strength between each free data and each data in the allocated data set, and clustering each free data x o Assign to meet x b The cluster to which it belongs; the formula used is as follows: ; ; In the formula, x b is the bth data in the assigned data set, x o is the oth free data, b and o are data indexes, and They are x o and x b The neighborhood set of x c is the cth data, c is the data index, , and They are x o With x b 、x o With x c and x b With x c The neighbor affinity between is x o and x b The comprehensive correlation strength between yes and The number of data in the intersection of neighbors, is x o With x b The Euclidean distance between .

5. The method for recommending learning resources based on big data according to claim 1, characterized in that: In step S4, the student clustering model parameter optimization specifically includes the following steps: Step S41: Initialization; establish a parameter search space for the number of nearest neighbors K, the number of initial cluster centers N, and the number of real clusters B in the student clustering model, randomly initialize Z1 exploration positions in the parameter search space, use each exploration position to represent a set of student clustering model parameters, and use the average value of all data silhouette coefficients in the clustering results obtained based on the exploration position as the fitness value of the corresponding exploration position; where Z1 is the number of exploration positions; Step S42: Calculate the quality; normalize the fitness value of each exploration position, and then determine the quality of each exploration position according to the normalized fitness value; Step S43: Calculate influence; Step S44: Update the exploration position; update each exploration position based on the upper and lower limits and influence of the parameter search space; the formula used is as follows: ; In the formula, and are the values ​​of the zth exploration position in the rth dimension during the t+1th and tth searches, respectively, and UB r and LB r are the upper and lower limits of the parameter search space in the rth dimension, r is the dimension index, T is the maximum number of searches, is the influence of the zth exploration position in the rth dimension during the tth search, is the quality of the zth explored position in the tth search, z is the explored position index, and t is the search number index; Step S45: parameter determination; pre-set the fitness threshold, update the fitness value of the exploration position, and when there is an exploration position with a fitness value higher than the fitness threshold, the student clustering model parameters corresponding to the exploration position are the optimal parameters; otherwise, if the maximum number of searches is reached, return to step S41 for reinitialization; otherwise, add 1 to the number of searches and return to step S42 to continue searching.

6. The method for recommending learning resources based on big data according to claim 5, characterized in that: In step S43, the influence is calculated based on the previous The optimal set is constructed from the exploration positions, and then the influence of each exploration position on each dimension is calculated based on the optimal set; the formula used is as follows: ; Among them, E t is the preferred set for the tth search, is the zth explored position during the tth search, is the tth search time E t The hth exploration position in and They are and The value in the rth dimension, is the influence of the zth exploration position in the rth dimension during the tth search, h is the exploration position index, and s r is a random value in the range [0, 1] corresponding to the rth dimension, is the L2 paradigm, yes Quality, is an exponential function, is the smoothing term, It is rounded down.

7. The method for recommending learning resources based on big data according to claim 1, characterized in that: In step S5, the generation of a recommended learning resource list is to input the student learning data set into a student clustering model constructed based on optimal parameters for clustering processing, determine the cluster D to which the real-time student learning data belongs based on the output clustering result, and construct a resource set to be recommended based on the resource ID contained in all historical student learning data in cluster D. Combined with the learning resource metadata set, the recommendation value corresponding to each resource ID in the resource set to be recommended is obtained, and the resource IDs in the resource set to be recommended are arranged in descending order according to the size of the recommendation value to obtain a recommended learning resource list.

8. The method for recommending learning resources based on big data according to claim 1, characterized in that: In step S1, the learning data collection is to collect historical student learning data, real-time student learning data and learning resource metadata.

9. The method for recommending learning resources based on big data according to claim 1, characterized in that: In step S2, the learning data preprocessing is to clean, standardize and construct data sets of the collected historical student learning data, real-time student learning data and learning resource metadata to obtain student learning data sets and learning resource metadata sets.

10. A learning resource recommendation system based on big data, used to implement the learning resource recommendation method based on big data as described in any one of claims 1 to 9, characterized in that: It includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module and a learning resource recommendation list generation module; The learning data collection module collects historical student learning data, real-time student learning data and learning resource metadata, and sends the data to the learning data preprocessing module; The learning data preprocessing module performs data cleaning, data standardization and data set construction on the collected data to obtain a student learning data set and a learning resource metadata set, and sends the data to the student clustering module; The student clustering module calculates the neighborhood cumulative local density based on the dual neighborhood set, and combines it with the relative distance to obtain the decision value, determines the initial cluster center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerated data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain the student clustering model, and sends the data to the student clustering model parameter optimization module; The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs a preferred set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits and influence of the parameter search space, finds the optimal student clustering model parameters, and sends the data to the module for generating a learning resource recommendation list; The module for generating a recommended learning resource list inputs a student learning data set into a student clustering model constructed based on optimal parameters for clustering processing, constructs a set of resources to be recommended based on the output clustering results, calculates the recommendation value of each resource ID, and generates a recommended learning resource list.

Citation Information

Patent Citations

  • Intelligent classroom teaching activity recommendation method and system

    CN111339386A

  • Personalized learning-oriented online education resource recommendation method

    CN114238431A

  • Pharmacological learning recommendation method based on big data

    CN118132837A

  • Personalized learning resource recommendation system and device

    CN118260487A

  • KR20230155630A