Learning Resource Recommendation Method and System Based on Big Data

Through the big data-based learning resource recommendation method, the dual neighborhood collection and near-neighborhood affinity are used to optimize student clustering, which solves the problems of high-dimensionality and sparseness, and achieves more accurate student group division and personalized resource recommendation, improving the accuracy and stability of learning resource recommendation.

CN120086450BActive Publication Date: 2025-07-08TAISHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510587447.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-08
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing learning resource recommendation methods are difficult to deal with high-dimensional, nonlinear correlation and highly sparse student learning behavior data, resulting in rough student group division, low accuracy in recommendation of learning resource, and large differences in learning behavior data of different students, unstable group division, and inaccurate learning resources, and inaccurate learning resources.

Method used

Using the learning resource recommendation method based on big data, the cumulative local density and relative distance of neighborhoods are calculated through dual neighbor sets, the initial clustering center is determined, the clustering of condensed and free data is carried out, and the comprehensive correlation intensity is calculated by introducing the nearest neighbor affinity to optimize the cluster structure. At the same time, the exploration location is updated based on the fitness value and influence, and the optimal student clustering model parameters are found.

Benefits of technology

It significantly improves the accuracy and accuracy of learning resource recommendations, can more accurately capture students' diverse learning needs, improve the rationality and pertinence of resource recommendations, and ensures stable division and personalized recommendations of different student groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086450B_ABST
    Figure CN120086450B_ABST
Patent Text Reader

Abstract

The present invention discloses a learning resource recommendation method and system based on big data, belonging to the technical field of resource recommendation. The method includes: learning data collection, learning data preprocessing, student clustering, optimization of student clustering model parameters, and generation of a learning resource recommendation list. This solution calculates the neighborhood cumulative local density based on a dual neighborhood set, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the near-neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree; calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, and finds the optimal student clustering model parameters, so as to accurately provide suitable learning resources for different students, improving the accuracy and effectiveness of learning resource recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of resource recommendation, and specifically refers to a learning resource recommendation method and system based on big data. Background Art

[0002] The learning resource recommendation method utilizes big data analysis and artificial intelligence technologies. By integrating multi-dimensional learning data of students, it deeply explores the potential laws in the massive learning data and intelligently recommends personalized learning resources, thereby improving learning efficiency and enhancing the overall learning outcome. However, in the existing learning resource recommendation methods, there are problems such as difficulty in processing high-dimensional, non-linearly correlated, and highly sparse student learning behavior data, inability to effectively capture its internal structure, resulting in rough student group division and low accuracy of learning resource recommendation; in the existing learning resource recommendation methods, there are significant differences between different students' learning behavior data, making it difficult to effectively adapt to all students' learning behavior data, resulting in unstable student group division and inability to accurately recommend learning resources. Summary of the Invention

[0003] In view of the above situation, to overcome the defects of the existing technology, the present invention provides a learning resource recommendation method and system based on big data. Regarding the problem in the existing learning resource recommendation methods that it is difficult to process high-dimensional, non-linearly correlated, and highly sparse student learning behavior data, unable to effectively capture its internal structure, resulting in rough student group division and low accuracy of learning resource recommendation, this solution calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the neighbor affinity to calculate the comprehensive correlation strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, which can more accurately identify the relatively important student data points in the learning behavior feature space, more precisely capture the diverse learning needs of students, and improve the accuracy of learning resource recommendation; regarding the problem in the existing learning resource recommendation methods that there are significant differences between different students' learning behavior data, making it difficult to effectively adapt to all students' learning behavior data, resulting in unstable student group division and inability to accurately recommend learning resources, this solution calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, and finds the optimal student clustering model parameters to make it more suitable for different student learning data characteristics, enhance the stability of student group division, and thus accurately provide suitable learning resources for different students.

[0004] The technical solution adopted by the present invention is as follows: The learning resource recommendation method based on big data provided by the present invention includes the following steps:

[0005] Step S1: Learning data collection;

[0006] Step S2: Preprocessing of learning data;

[0007] Step S3: Student clustering;

[0008] Step S4: Optimization of student clustering model parameters;

[0009] Step S5: Generating a recommended list of learning resources.

[0010] Further, in step S1, the learning data collection is to collect historical student learning data, real-time student learning data, and learning resource metadata.

[0011] Further, in step S2, the preprocessing of learning data is to perform data cleaning, data standardization, and construction of data sets on the collected historical student learning data, real-time student learning data, and learning resource metadata, to obtain a student learning data set and a learning resource metadata set.

[0012] Further, in step S3, the student clustering specifically includes the following steps:

[0013] Step S31: Selecting an initial clustering center; including the following steps:

[0014] Step S311: Calculating the neighborhood cumulative local density; calculating the Euclidean distance between any two data in the student learning data set, finding the K nearest neighbors of each data and constructing a neighborhood set, and calculating the neighborhood cumulative local density of each data based on the dual neighborhood set; the formula used is as follows:

[0015] ;

[0016] where x i , x j and x m are the i-th, j-th, and m-th data in the student learning data set respectively, i, j, and m are data indices, ρ i is the neighborhood cumulative local density of x i , is the Euclidean distance between x j and x m , and are the neighborhood sets of x i and x j respectively, and K is the number of nearest neighbors;

[0017] Step S312: Determine; calculate the relative distance and decision value of each data in the student learning data set, select the top N data with the largest decision value from the student learning data set as the initial clustering centers, and construct an initial clustering center set H; where N is the number of initial clustering centers;

[0018] Step S32: Data clustering; includes the following steps:

[0019] Step S321: Data partitioning; construct an unassigned data set based on the data in the student learning data set except the initial clustering centers, and calculate the neighbor intersection and the number of data in the neighbor intersection between each data in the unassigned data set and each initial clustering center and the number of data in the neighbor intersection , if there exists at least one initial clustering center such that the number of data in the neighbor intersection between x a and x m is , then classify x a as cohesive data; otherwise, classify x a as free data; where x a is the a-th data in the unassigned data set, x m is the m-th initial clustering center, a and m are data indexes and initial clustering center indexes respectively, is the neighbor intersection between x a and x m , is the number of data in;

[0020] Step S322: Cohesive data clustering; calculate the number of data in the neighbor intersection between each cohesive data and each initial clustering center, and assign each cohesive data x v to the cluster where the initial clustering center x belongs that satisfies; where m is the number of data in the neighbor intersection between the neighborhood set of x and the neighborhood set of x v , x m is the v-th cohesive data, x v is the m-th initial clustering center, v is the data index, m is the maximum value function; is the maximum value function;

[0021] Step S323: Free data clustering; construct an assigned data set A based on the initial clustering centers and cohesive data, introduce the near-neighbor affinity, calculate the comprehensive association strength between each free data and each data in the assigned data set, and assign each free data x o to the x that satisfies inb in the cluster to which it belongs; the formula used is as follows:

[0022] ;

[0023] ;

[0024] In the formula, x b is the b-th data in the allocated data set, and x o is the o-th free data. b and o are data indices, and are the neighborhood sets of x o and x b respectively. x c is the c-th data, and c is a data index, , and are the nearest neighbor affinities between x o and x b , x o and x c , and x b and x c respectively, is the comprehensive association strength between x o and x b , is and the number of data in the neighbor intersection between them, is the Euclidean distance between x o and x b ;

[0025] Step S33: Cluster structure optimization; calculate the inter-cluster fusion degree between every two clusters, and merge the two clusters with the largest inter-cluster fusion degree; the formula used is as follows:

[0026] ;

[0027] In the formula, C n1 and C n2 are the n1-th and n2-th clusters respectively, is the inter-cluster fusion degree between C n1 and C n2 . x p and x q are the p-th and q-th data respectively. S n1 and S n2 are the number of data in C n1 and C n2 respectively, is the distance between x q and x pThe nearest neighbor affinity between them, where p and q are data indices, and n1 and n2 are cluster indices;

[0028] Step S34: Model determination; Repeat step S33 until the number of clusters is equal to B, at which point the student clustering is complete, obtaining the student clustering model and outputting the clustering result; where B is the number of true clusters.

[0029] Furthermore, in step S4, the optimization of the student clustering model parameters specifically includes the following steps:

[0030] Step S41: Initialization; Establish a parameter search space for the number of nearest neighbors K, the number of initial cluster centers N, and the number of true clusters B in the student clustering model. Randomly initialize Z1 exploration positions within the parameter search space. Each exploration position represents a set of student clustering model parameters. Take the average of all silhouette coefficients in the clustering result obtained based on the exploration position as the fitness value corresponding to the exploration position; where Z1 is the number of exploration positions;

[0031] Step S42: Calculate quality; Normalize the fitness value of each exploration position, and then determine the quality of each exploration position according to the normalized fitness value.

[0032] Step S43: Calculate influence; Based on the top exploration positions with larger quality, construct a preferred set, and calculate the influence of each exploration position in each dimension based on the preferred set; The formula used is as follows:

[0033] ;

[0034] where E t is the preferred set at the t-th search, is the z-th exploration position at the t-th search, is the h-th exploration position in E t at the t-th search, and are respectively and the values at the r-th dimension, is the influence of the z-th exploration position at the t-th search in the r-th dimension, r is the dimension index, h is the exploration position index, s r is a random value within the range of [0, 1] corresponding to the r-th dimension, is the L2 norm, is 's quality, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search time index, T is the maximum number of searches, is the exponential function, is floor division, is the smoothing term;

[0035] Step S44: Update the exploration positions; update each exploration position based on the upper and lower bounds and influence of the parameter search space; the formula used is as follows:

[0036] ;

[0037] In the formula, is the value of the z-th exploration position at the (t + 1)-th search in the r-th dimension, UB r and LB r are respectively the upper and lower bounds of the parameter search space in the r-th dimension;

[0038] Step S45: Parameter determination; preset a fitness threshold, update the fitness values of the exploration positions. When there is an exploration position with a fitness value higher than the fitness threshold, the parameters of the student clustering model corresponding to this exploration position are the optimal parameters; otherwise, if the maximum search times are reached, return to Step S41 to re-initialize; otherwise, increment the search times by 1 and return to Step S42 to continue the search.

[0039] Further, in Step S5, the generation of the learning resource recommendation list is to input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing. Based on the output clustering result, determine the cluster D to which the real-time student learning data belongs. Construct a resource set to be recommended based on the resource IDs included in all historical student learning data in cluster D. Combine with the learning resource metadata set to obtain the recommendation value corresponding to each resource ID in the resource set to be recommended, and sort the resource IDs in the resource set to be recommended in descending order of the recommendation value to obtain the learning resource recommendation list.

[0040] The learning resource recommendation system based on big data provided by the present invention includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module, and a generation of learning resource recommendation list module;

[0041] The learning data collection module collects historical student learning data, real-time student learning data, and learning resource metadata, and sends the data to the learning data preprocessing module;

[0042] The learning data preprocessing module performs data cleaning, data standardization, and data set construction processing on the collected data to obtain a student learning data set and a learning resource metadata set, and sends the data to the student clustering module;

[0043] The student clustering module calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, and sends the data to the student clustering model parameter optimization module;

[0044] The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs a preferred set and calculates the influence of each exploration position on each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, and sends the data to the generate learning resource recommendation list module;

[0045] The generate learning resource recommendation list module inputs the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructs a resource set to be recommended based on the output clustering result, calculates the recommendation value of each resource ID, and generates a learning resource recommendation list.

[0046] The beneficial effects achieved by the present invention using the above solution are as follows:

[0047] (1) Aiming at the problem that the existing learning resource recommendation methods are difficult to process the student learning behavior data with high dimensions, non-linear associations and high sparsity, and cannot effectively capture its internal structure, resulting in rough division of student groups and low accuracy of learning resource recommendation. This solution calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, which can more accurately identify the relatively important student data points in the learning behavior feature space, distinguish different types of students, capture the diverse learning needs of students more accurately, significantly improve the rationality and pertinence of learning resource recommendation, and improve the accuracy of learning resource recommendation.

[0048] (2)In view of the problem that there are significant differences among the learning behavior data of different students in the existing learning resource recommendation methods, making it difficult to effectively adapt to all students' learning behavior data, resulting in unstable student group division and inaccurate learning resource recommendation, this solution uses exploration positions to represent a set of student clustering model parameters, calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, makes it more suitable for the learning data characteristics of different students, improves the stability of student group division, and can recommend resources based on a more accurate student group division, thereby accurately providing appropriate learning resources for different students and improving the accuracy and effectiveness of learning resource recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a schematic flowchart of the learning resource recommendation method based on big data provided by the present invention;

[0050] Figure 2 is a schematic diagram of the learning resource recommendation system based on big data provided by the present invention;

[0051] Figure 3 is a schematic flowchart of step S3;

[0052] Figure 4 is a schematic flowchart of step S4.

[0053] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0056] Embodiment 1. Refer toFigure 1 , the learning resource recommendation method based on big data provided by the present invention includes the following steps:

[0057] Step S1: Learning data collection; collecting historical student learning data, real-time student learning data, and learning resource metadata;

[0058] Step S2: Preprocessing of learning data; performing data cleaning, data standardization, and constructing a data set on the collected data to obtain a student learning data set and a learning resource metadata set;

[0059] Step S3: Student clustering; calculating the neighborhood cumulative local density based on the dual neighborhood set, combining it with the relative distance to obtain a decision value, determining the initial clustering center, partitioning the data according to the number of data in the neighbor intersection and completing the clustering of cohesive data, introducing the near-neighbor affinity to calculate the comprehensive association strength, completing the clustering of free data, and optimizing the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model;

[0060] Step S4: Optimization of student clustering model parameters; calculating the quality of each exploration position based on the fitness value, constructing an optimal set and calculating the influence of each exploration position in each dimension, updating each exploration position based on the upper and lower limits of the parameter search space and the influence, and finding the optimal student clustering model parameters;

[0061] Step S5: Generating a learning resource recommendation list; inputting the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructing a resource set to be recommended based on the output clustering result, calculating the recommendation value of each resource ID, and generating a learning resource recommendation list.

[0062] Example 2, refer to Figure 1 , based on the above example, in step S1, the historical student learning data and the real-time student learning data both include the browsing behavior, interaction behavior, and learning behavior of students on the learning platform;

[0063] The browsing behavior includes the resource ID accessed, access timestamp, access duration, page scroll depth, and number of repeated accesses;

[0064] The interaction behavior includes favorite records, rating data, sharing records, and download records;

[0065] The learning behavior includes resource learning progress, quiz scores, and types of wrong questions;

[0066] The learning resource metadata includes resource ID, number of times favorited, average rating, number of times accessed, and number of days since the resource was published.

[0067] Example 3, refer to Figure 1, based on the above embodiment, in step S2, data cleaning, data standardization, and data set construction processing are performed on the collected historical student learning data, real-time student learning data, and learning resource metadata to obtain a student learning data set and a learning resource metadata set;

[0068] The data cleaning includes missing value processing and outlier processing. Missing value processing is to fill missing data with the mean value, and outlier processing is to identify and remove outliers using the box plot method;

[0069] The data standardization is to map the data to a unified range using the maximum-minimum scaling method;

[0070] The construction of the data set includes constructing a student learning data set and a learning resource metadata set. The student learning data set is constructed from the historical student learning data and real-time student learning data after data cleaning and data standardization, and the learning resource metadata set is constructed from the learning resource metadata after data cleaning and data standardization.

[0071] Embodiment 4, refer to Figure 1 and Figure 3 , based on the above embodiment, in step S3, the student clustering specifically includes the following steps:

[0072] Step S31: Select the initial clustering center; determine the density of each data in its neighborhood in the student learning data set by calculating the neighborhood cumulative local density, so as to find the relatively dense regions in the data distribution; select the initial clustering center by calculating the relative distance and decision value, which can comprehensively consider the local density and relative position relationship of the data, so that the initial clustering centers are distributed in the relatively dense and representative regions of the data, and can better divide the student groups with different learning modes, thereby recommending learning resources that are more in line with their learning habits and progress for each group, improving the accuracy and pertinence of the recommendation; including the following steps:

[0073] Step S311: Calculate the neighborhood cumulative local density; calculate the Euclidean distance between any two data in the student learning data set, find the K nearest neighbors of each data and construct a neighborhood set, and calculate the neighborhood cumulative local density of each data based on the double neighborhood set. The value range of K is [5, 50]; the formula used is as follows:

[0074] ;

[0075] In the formula, x i , x j and x m are the i-th, j-th, and m-th data in the student learning data set respectively, i, j, and m are data indices, and ρ i is xi The neighborhood cumulative local density of is x j With x m The Euclidean distance between and They are x i and x j Neighborhood set of , K is the number of nearest neighbors;

[0076] Step S312: Determine; calculate the relative distance of each data in the student learning data set and decision value , select the first N data with the largest decision value from the student learning data set as the initial cluster center, and construct the initial cluster center set H, the value range of N is [B, 2B], the value range of B is [2, ]; where N is the number of initial cluster centers, δ i and γ i They are x i The relative distance and decision value of j is x j The neighborhood cumulative local density, ρ max is the maximum neighborhood cumulative local density among all data in the student learning data set, is x i With x j The Euclidean distance between them, B is the number of true clusters, and M is the number of data in the student learning data set;

[0077] Step S32: data clustering; reasonably divide the student learning data into different types in order to make more accurate clustering and resource recommendation; divide the data to be allocated into condensed data and free data by calculating the neighbor intersection and the number of data in the intersection between each data and the initial cluster center, so as to more carefully analyze the relationship structure of the student learning data; allocate the condensed data to the most appropriate cluster according to the number of data in the neighbor intersection between the condensed data and the initial cluster center, which is helpful to build a tight student clustering so that students in the same cluster have similar learning behaviors and needs; allocate the free data to the cluster to which the allocated data with the largest comprehensive correlation strength belongs according to the comprehensive correlation strength between the free data and the allocated data, so as to integrate the free students into the most similar student group; comprising the following steps:

[0078] Step S321: Data division; construct a data set to be allocated based on the data in the student learning data set except the initial cluster center, and calculate the neighbor intersection between each data in the data set to be allocated and each initial cluster center The number of data in the intersection with neighbors , if there is at least one initial cluster center such that xa The number of data in the neighbor intersection between x m and , then x a is classified as cohesive data; otherwise, x a is classified as free data; where x a is the ath data in the data set to be assigned, x m is the mth initial clustering center, a and m are data index and initial clustering center index respectively, is the neighbor intersection between x a and x m , and are the neighborhood sets of x a and x m respectively, is the number of data in ; the intersection operator is

[0079] Step S322: Cohesive data clustering; calculate the number of data in the neighbor intersection between each cohesive data and each initial clustering center, and assign each cohesive data x v to the cluster where the initial clustering center x belongs; where m ; is the number of data in the neighbor intersection between the neighborhood set of x v and the neighborhood set of x m , x v is the vth cohesive data, x m is the mth initial clustering center, v is the data index, is the maximum value function;

[0080] Step S323: Free data clustering; based on the initial clustering center and cohesive data, construct the assigned data set A, introduce the neighbor affinity, calculate the comprehensive association strength between each free data and each data in the assigned data set, and assign each free data x o to the cluster where x belongs; the formula used is as follows: b ;

[0081] ;

[0082] ;

[0083] In the formula, x b is the bth data in the assigned data set, x o is the oth free data, b and o are data indexes, and They are the neighborhood sets of x o and x b respectively, where x c is the c-th data and c is the data index. , and are the neighborhood affinities between x o and x b , x o and x c , and x b and x c respectively. is the comprehensive association strength between x o and x b . is and the number of data in the neighbor intersection between them. is the Euclidean distance between x o and x b .

[0084] Step S33: Cluster structure optimization; in the learning resource recommendation environment, there may be some unreasonable cluster structures in the initial clustering results, which need to be optimized to improve the clustering quality. Calculate the inter-cluster fusion degree between every two clusters, and merge the two clusters with the largest inter-cluster fusion degree to make the clustering structure more compact and reasonable. The optimized cluster structure can reduce the resource recommendation deviation caused by unreasonable clustering. The formula used is as follows:

[0085] ;

[0086] In the formula, C n1 and C n2 are the n1-th and n2-th clusters respectively, is the inter-cluster fusion degree between C n1 and C n2 , x p and x q are the p-th and q-th data respectively, S n1 and S n2 are the number of data in C n1 and C n2 respectively, is the neighborhood affinity between x q and x p . p and q are data indices, and n1 and n2 are cluster indices.

[0087] Step S34: Model determination; repeat step S33 until the number of clusters is equal to B, student clustering is completed, a student clustering model is obtained, and the clustering results are output so as to make accurate learning resource recommendations based on the characteristics of student groups in different clusters, improve the matching degree between resources and student needs, and thus improve learning effects; where B is the actual number of clusters.

[0088] By performing the above operations, the existing learning resource recommendation methods have the problem that they are difficult to handle high-dimensional, nonlinearly associated and highly sparse student learning behavior data, and cannot effectively capture their internal structure, resulting in rough division of student groups and low accuracy of learning resource recommendation. This scheme calculates the neighborhood cumulative local density based on the dual neighborhood set, and combines it with the relative distance to obtain the decision value, determines the initial cluster center, divides the data according to the number of data in the neighbor intersection and completes the clustering of the agglomerated data, introduces the neighbor affinity to calculate the comprehensive association strength, completes the clustering of the free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain the student clustering model, which can more accurately identify the relatively important student data points in the learning behavior feature space, distinguish different types of students, and more accurately capture the diverse learning needs of students, significantly improve the rationality and pertinence of learning resource recommendations, and improve the accuracy of learning resource recommendations.

[0089] Example 5, see Figure 1 and Figure 4 This embodiment is based on the above embodiment. In step S4, the optimization of student clustering model parameters specifically includes the following steps:

[0090] Step S41: Initialization; by randomly initializing multiple groups of parameters, a variety of starting points are provided for subsequent optimization to avoid falling into local optimum, and the clustering quality is evaluated in combination with the silhouette coefficient to ensure that the initial parameters have a certain degree of discrimination, laying the foundation for subsequent optimization; a parameter search space is established for the number of nearest neighbors K, the number of initial cluster centers N, and the number of real clusters B in the student clustering model, and Z1 exploration positions are randomly initialized in the parameter search space, each exploration position represents a set of student clustering model parameters, and the average value of the silhouette coefficients of all data in the clustering results obtained based on the exploration position is used as the fitness value of the corresponding exploration position; wherein Z1 is the number of exploration positions;

[0091] Step S42: Calculate the quality; normalize the fitness value of each exploration position, and then determine the quality of each exploration position based on the normalized fitness value, reflecting the compactness and separation of the clustering results. High-quality parameter combinations can better distinguish students' learning behavior patterns, thereby recommending more suitable resources for each group; the formula used is as follows:

[0092] ;

[0093] ;

[0094] In the formula, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search number index, is the normalized fitness value of the z-th exploration position at the t-th search, is the fitness value of the z-th exploration position at the t-th search, and are the maximum fitness value and the minimum fitness value at the t-th search respectively;

[0095] Step S43: Calculate the influence; During the parameter optimization process, blind search needs to be avoided. Based on the first exploration positions with larger quality, a preferred set is constructed, and based on the preferred set, the influence of each exploration position in each dimension is calculated to adjust the search towards a better direction and accelerate convergence; The formula used is as follows:

[0096] ;

[0097] where E t is the preferred set at the t-th search, is the z-th exploration position at the t-th search, is the h-th exploration position in E t at the t-th search, and are respectively and the values in the r-th dimension, is the influence of the z-th exploration position at the t-th search in the r-th dimension, r is the dimension index, h is the exploration position index, s r is a random value in the range of [0, 1] corresponding to the r-th dimension, is the L2 norm, is the quality of, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, t is the search number index, T is the maximum number of searches, is the exponential function, is the floor function, is the smoothing term, , and the smoothing term is used to avoid the denominator being 0;

[0098] Step S44: Update the exploration position; Based on the upper and lower limits of the parameter search space and the influence, each exploration position is updated to dynamically update the parameter combination to adapt to the optimization requirements in different stages, thereby improving the recommendation accuracy; The formula used is as follows:

[0099] ;

[0100] In the formula, is the value of the z-th exploration position at the (t + 1)-th search in the r-th dimension, UB r and LB r are the upper and lower limits of the parameter search space in the r-th dimension respectively;

[0101] Step S45: Parameter determination; preset a fitness threshold, update the fitness value of the exploration positions. When there is an exploration position whose fitness value is higher than the fitness threshold, the student clustering model parameters corresponding to this exploration position are the optimal parameters; otherwise, if the maximum search number is reached, return to Step S41 to re-initialize; otherwise, increment the search number by 1 and return to Step S42 to continue the search; the finally determined parameters enable the clustering result to stably reflect the distribution of the real student population, thereby supporting long-term accurate recommendation.

[0102] By performing the above operations, for the problem in the existing learning resource recommendation method that there are large differences between different students' learning behavior data, it is difficult to effectively adapt to all students' learning behavior data, resulting in unstable division of the student population and inability to accurately recommend learning resources. In this solution, the exploration positions represent a set of student clustering model parameters, calculate the quality of each exploration position based on the fitness value, construct an optimal set and calculate the influence of each exploration position in each dimension, update each exploration position based on the upper and lower limits of the parameter search space and the influence, find the optimal student clustering model parameters, make it more suitable for the characteristics of different students' learning data, improve the stability of the student population division, and be able to recommend resources according to a more accurate student population division, so as to accurately provide suitable learning resources for different students and improve the accuracy and effectiveness of learning resource recommendation.

[0103] Example Six, refer to Figure 1 , this example is based on the above example. In Step S5, generating the learning resource recommendation list is to input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing. Based on the output clustering result, determine the cluster D to which the real-time student learning data belongs, construct the resource set to be recommended based on the resource IDs included in all historical student learning data in cluster D, combine with the learning resource metadata set, comprehensively consider the number of times of being collected, average score, number of times of being accessed, and the number of days since the resource was published, obtain the recommendation value corresponding to each resource ID in the resource set to be recommended, sort the resource IDs in the resource set to be recommended in descending order according to the size of the recommendation value, and obtain the learning resource recommendation list. The formula used is as follows:

[0104] ;

[0105] Among them, P gis the recommended value corresponding to the g-th resource ID in the set of resources to be recommended, where g is the resource ID index. , , and are respectively the number of times collected, average score, number of times accessed, and number of days since the resource was published corresponding to the g-th resource ID in the set of learning resource metadata in the set of resources to be recommended.

[0106] Example Seven. Refer to Figure 2 . Based on the above example, the learning resource recommendation system provided by the present invention includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module, and a generation of learning resource recommendation list module.

[0107] The learning data collection module collects historical student learning data, real-time student learning data, and learning resource metadata, and sends the data to the learning data preprocessing module.

[0108] The learning data preprocessing module performs data cleaning, data standardization, and construction of data set processing on the collected data to obtain a student learning data set and a learning resource metadata set, and sends the data to the student clustering module.

[0109] The student clustering module calculates the neighborhood cumulative local density based on the double neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of cohesive data, introduces the calculation of the near-neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain a student clustering model, and sends the data to the student clustering model parameter optimization module.

[0110] The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs a preferred set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, and sends the data to the generation of learning resource recommendation list module.

[0111] The generation of learning resource recommendation list module inputs the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructs a set of resources to be recommended based on the output clustering result, calculates the recommended value of each resource ID, and generates a learning resource recommendation list.

[0112] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0113] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention.

[0114] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by this and, without departing from the gist of the present invention, design similar structural modes and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A learning resource recommendation method based on big data, characterized in that: The method includes the following steps: Step S1: Learning data collection; collecting historical student learning data, real-time student learning data, and learning resource metadata; Step S2: Preprocessing of learning data; Step S3: Student clustering; Step S4: Optimization of student clustering model parameters; Step S5: Generating a learning resource recommendation list; inputting the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructing a set of resources to be recommended based on the output clustering results, calculating the recommendation value for each resource ID, and generating a learning resource recommendation list; Step S3 includes step S321: data partitioning; constructing a data set to be assigned based on the data in the student learning data set except for the initial clustering centers, and calculating the neighbor intersections between each data in the data set to be assigned and each initial clustering center and the number of data in the neighbor intersections , if there exists at least one initial clustering center such that the number of data in the neighbor intersection between x a and x m is , then x a is classified as cohesive data; otherwise, x a is classified as free data; where x a is the ath data in the data set to be assigned, x m is the mth initial clustering center, a and m are data index and initial clustering center index respectively, is the neighbor intersection between x a and x m , is the number of data in it, and K is the number of nearest neighbors.

2. The learning resource recommendation method based on big data according to claim 1, characterized in that: In step S3, the student clustering specifically includes the following steps: Step S31: Selecting the initial clustering centers; Step S32: Data clustering; Step S33: Cluster structure optimization; calculating the inter-cluster fusion degree between every two clusters, and merging the two clusters with the largest inter-cluster fusion degree; the formula used is as follows: ; Where C n1 and C n2 are the n1th and n2th clusters respectively, is the inter-cluster fusion degree between C n1 and C n2 , x p and x q are the pth and qth data respectively, S n1 and S n2 are the numbers of data in C n1 and C n2 respectively, is the nearest neighbor affinity between x q and x p , p and q are data indices, and n1 and n2 are cluster indices; Step S34: Model determination; repeating step S33 until the number of clusters is equal to B, the student clustering is completed, a student clustering model is obtained, and the clustering results are output; where B is the number of true clusters.

3. The learning resource recommendation method based on big data according to claim 2, wherein: In step S31, the selection of the initial clustering centers specifically includes the following steps: Step S311: Calculating the neighborhood cumulative local density; calculating the Euclidean distance between any two data in the student learning data set, finding the K nearest neighbors of each data and constructing a neighborhood set, and calculating the neighborhood cumulative local density of each data based on the double neighborhood set; the formula used is as follows: ; where x i , x j and x m are the i-th, j-th, and m-th data in the student learning data set respectively, i, j, and m are data indices, and ρ i is the neighborhood cumulative local density of x i ; is the Euclidean distance between x j and x m ; and are the neighborhood sets of x i and x j respectively; Step S312: Determination; calculating the relative distance and decision value of each data in the student learning data set, selecting the top N data with the largest decision value from the student learning data set as the initial clustering centers, and constructing an initial clustering center set H; where N is the number of initial clustering centers.

4. The learning resource recommendation method based on big data according to claim 2, wherein: In step S32, the data clustering specifically includes the following steps: Step S321: Data partitioning; Step S322: Agglomerative data clustering; calculate the number of data in the neighbor intersection between each agglomerative data and each initial clustering center, and assign each agglomerative data x v to the cluster to which the initial clustering center x belongs; where m is the number of data in the neighbor intersection between the neighborhood set of x and the neighborhood set of x v , x m is the v-th agglomerative data, x v is the m-th initial clustering center, v is the data index, m and is the maximum value function. Step S323: Free-form data clustering; Based on the initial clustering centers and the agglomerative data, construct the assigned data set A, introduce the neighbor affinity, calculate the comprehensive association strength between each free-form data and each data in the assigned data set, and assign each free-form data x o to the cluster where the x b belongs; The formula used is as follows: ; ; where x b is the b-th data in the allocated data set, and x o is the o-th free data, where b and o are data indices, and are the neighborhood sets of x o and x b respectively, x c is the c-th data, where c is a data index, 、 and are the pairwise neighbor affinities between x o and x b , x o and x c , and x b and x c respectively, is the comprehensive association strength between x o and x b , is and the number of data in the neighbor intersection between them, is the Euclidean distance between x o and x b .

5. The learning resource recommendation method based on big data according to claim 1, characterized in that: In step S4, the optimization of the student clustering model parameters specifically includes the following steps: Step S41: Initialization; establishing a parameter search space for the number of nearest neighbors K, the number of initial clustering centers N, and the number of true clusters B in the student clustering model, randomly initializing Z1 exploration positions within the parameter search space, using each exploration position to represent a set of student clustering model parameters, and taking the average value of all the silhouette coefficients of the data in the clustering results obtained based on the exploration positions as the fitness value corresponding to the exploration position; where Z1 is the number of exploration positions; Step S42: Calculating the quality; normalizing the fitness value of each exploration position, and then determining the quality of each exploration position according to the normalized fitness value; Step S43: Calculating the influence; Step S44: Updating the exploration positions; updating each exploration position based on the upper and lower limits of the parameter search space and the influence; the formula used is as follows: ; Wherein, and are the values of the z-th exploration position at the (t + 1)-th and t-th searches in the r-th dimension respectively, UB r and LB r are the upper and lower limits of the parameter search space in the r-th dimension respectively, r is the dimension index, T is the maximum number of searches, is the influence of the z-th exploration position at the t-th search in the r-th dimension, is the quality of the z-th exploration position at the t-th search, z is the exploration position index, and t is the search number index; Step S45: Parameter determination; preset a fitness threshold, update the fitness value of the exploration position. When there is an exploration position whose fitness value is higher than the fitness threshold, the student clustering model parameters corresponding to this exploration position are the optimal parameters; otherwise, if the maximum search number is reached, return to step S41 to re-initialize; otherwise, increment the search number by 1 and return to step S42 to continue the search.

6. The learning resource recommendation method based on big data according to claim 5, characterized in that: In step S43, the calculated influence is based on the first exploration positions with larger mass to construct an optimal set, and then the influence of each exploration position in each dimension is calculated based on the optimal set; the formula used is as follows: ; Among them, E t is the preferred set at the t-th search, is the z-th exploration position at the t-th search, is the h-th exploration position in E t at the t-th search, and are respectively and values in the r-th dimension, is the influence of the z-th exploration position at the t-th search on the r-th dimension, h is the exploration position index, s r is a random value in the range [0, 1] corresponding to the r-th dimension, is the L2 norm, is quality, is the exponential function, is the smoothing term, is floor.

7. The learning resource recommendation method based on big data according to claim 1, wherein: In step S5, the generation of the learning resource recommendation list is to input the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing. Based on the output clustering result, determine the cluster D to which the real-time student learning data belongs. Construct a resource set to be recommended based on the resource IDs included in all historical student learning data in cluster D. Combine with the learning resource metadata set to obtain the recommendation value corresponding to each resource ID in the resource set to be recommended. Arrange the resource IDs in the resource set to be recommended in descending order according to the recommendation value to obtain the learning resource recommendation list.

8. The learning resource recommendation method based on big data according to claim 1, wherein: In step S1, the learning data collection is to collect historical student learning data, real-time student learning data, and learning resource metadata.

9. The learning resource recommendation method based on big data according to claim 1, wherein: In step S2, the learning data preprocessing is to perform data cleaning, data standardization, and data set construction processing on the collected historical student learning data, real-time student learning data, and learning resource metadata to obtain the student learning data set and the learning resource metadata set.

10. A learning resource recommendation system based on big data, which is used to implement the learning resource recommendation method based on big data as described in any one of claims 1-9, and is characterized in that: It includes a learning data collection module, a learning data preprocessing module, a student clustering module, a student clustering model parameter optimization module, and a generation of learning resource recommendation list module; The learning data collection module collects historical student learning data, real-time student learning data, and learning resource metadata, and sends the data to the learning data preprocessing module; The learning data preprocessing module performs data cleaning, data standardization, and data set construction processing on the collected data to obtain the student learning data set and the learning resource metadata set, and sends the data to the student clustering module; The student clustering module calculates the neighborhood cumulative local density based on the dual neighborhood set, combines it with the relative distance to obtain a decision value, determines the initial clustering center, divides the data according to the number of data in the neighbor intersection and completes the clustering of agglomerative data, introduces the near neighbor affinity to calculate the comprehensive association strength, completes the clustering of free data, and optimizes the cluster structure according to the inter-cluster fusion degree to obtain the student clustering model, and sends the data to the student clustering model parameter optimization module; The student clustering model parameter optimization module calculates the quality of each exploration position based on the fitness value, constructs an optimal set and calculates the influence of each exploration position in each dimension, updates each exploration position based on the upper and lower limits of the parameter search space and the influence, finds the optimal student clustering model parameters, and sends the data to the generation of learning resource recommendation list module; The generation of learning resource recommendation list module inputs the student learning data set into the student clustering model constructed based on the optimal parameters for clustering processing, constructs a resource set to be recommended based on the output clustering result, calculates the recommendation value of each resource ID, and generates the learning resource recommendation list.

Citation Information

Patent Citations

  • Intelligent classroom teaching activity recommendation method and system

    CN111339386A

  • Personalized learning-oriented online education resource recommendation method

    CN114238431A