Family education content recommendation method and system based on big data

The family education content recommendation method, which combines multi-source data processing and LSTM model with social network analysis, solves the problems of weak multi-source data fusion capability and lack of strategy feedback mechanism in the existing system, and improves the accuracy and real-time performance of personalized recommendations.

CN120823077AInactive Publication Date: 2025-10-21NANJING CHONGZHEN BIG DATA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510903264.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing family education content recommendation system has weak multi-source data fusion capabilities, coarse user behavior modeling granularity, insufficient utilization of social structure and lack of strategy feedback mechanism, resulting in insufficient timeliness and personalization of recommendation strategies.

Method used

By collecting multi-source data from family members for preprocessing, using the LSTM model to predict learning progress, combining social network analysis and reinforcement learning to optimize personalized recommendations, dynamically adjusting push content, and continuously optimizing recommended content based on user feedback.

Benefits of technology

It significantly improves the accuracy and real-time performance of personalized recommendations, can dynamically adapt to the long-term evolution of family education goals and learning habits, and improves the accuracy of learning progress prediction and the relevance of recommended content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823077A_ABST
    Figure CN120823077A_ABST
Patent Text Reader

Abstract

The invention discloses a family education content recommendation method and system based on big data, and relates to the technical field of big data, and the method comprises the steps: collecting family member multi-source data for preprocessing, carrying out the real-time analysis of the preprocessed data, updating a user file, setting a personalized recommendation target according to an analysis result, and carrying out the recommendation of family education content through the updated user file. The learning progress of a user is predicted by using an LSTM model, and a social relation and an interaction mode among family members are established through a social network analysis algorithm in combination with a robust algorithm and a Pearch ranking algorithm. According to the method, the family member social relation graph is constructed and optimized through social network analysis, the robust algorithm and the Pearch ranking algorithm, the social sub-groups are identified, family education content recommendation is dynamically adjusted in combination with the reinforcement learning algorithm, and the social adaptability and interest matching ability of personalized content recommendation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to a method and system for recommending family education content based on big data. Background Art

[0002] With the rapid development of information technology, the trend toward digitalization in family education is gaining momentum. Traditional offline family education methods are gradually being replaced by intelligent, data-driven online education platforms. Parents and children can access a wealth of educational resources through various channels, including mobile devices, online education platforms, and social media. This diverse information source and the widespread use of digital media have enabled family education to gradually shift from "experience-based guidance" to "data-enabled" learning, promoting the implementation of the concepts of "personalized learning" and "precision parenting." Against this backdrop, data-driven personalized recommendation technologies are beginning to be introduced into family education scenarios to improve the alignment of educational content with user needs. Technologies such as deep learning, social network analysis, time series modeling, and natural language processing are being gradually applied to learning behavior analysis, content preference modeling, and educational content recommendation tasks, forming a recommendation mechanism tailored to user profiles.

[0003] However, existing family education content recommendation systems generally suffer from several prominent problems. First, most recommendation algorithms rely on a single data source, such as learning records on learning platforms, ignoring important data on social networks that reflect family members' educational tendencies. Second, at the user modeling level, there is a lack of detailed modeling of time series changes in learning behavior, making it difficult to capture the dynamic evolution of family members' learning status, resulting in insufficient timeliness and personalization of recommendation strategies. In addition, mainstream recommendation models often use static or shallow feature extraction methods, which make it difficult to explore the potential correlations between high-order behavioral features, especially in terms of their limited ability to characterize parent-child interactions and group behavior patterns. Moreover, existing technologies generally lack a "feedback closed loop" mechanism, lacking in-depth utilization of user feedback behavior after recommendation, and are unable to dynamically adjust strategies to adapt to the long-term evolution of family education goals and learning habits. Therefore, existing family education content recommendation technologies generally suffer from technical problems such as weak multi-source data fusion capabilities, coarse granularity in user behavior modeling, insufficient utilization of social structures, and a lack of a strategy feedback mechanism. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a family education content recommendation method and system based on big data to solve the technical problems of weak multi-source data fusion capability, coarse granularity of user behavior modeling, insufficient utilization of social structure and lack of strategy feedback mechanism.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for recommending family education content based on big data, which comprises:

[0008] Collect multi-source data of family members for pre-processing, analyze the pre-processed data in real time and update user profiles, and set personalized recommendation goals based on the analysis results;

[0009] The multi-source data includes education platform data, social media data and family behavior data;

[0010] Through updated user profiles, the LSTM model is used to predict user learning progress. Social network analysis algorithms are combined with the Louvain algorithm and the Page ranking algorithm to establish social relationships and interaction patterns among family members. Reinforcement learning is used to dynamically optimize personalized recommendation goals.

[0011] Dynamically adjust the timing and frequency of content push based on the LSTM model prediction results, combine the intelligent push mechanism to push optimized personalized recommendation targets, and continuously optimize the recommended content based on user feedback.

[0012] As a preferred solution of the family education content recommendation method based on big data described in the present invention, wherein: the method of predicting the user learning progress by using the LSTM model through the updated user profile refers to sampling the real-time updated user profile to obtain the key learning behavior feature data of each family member, and converting it into a feature vector through feature engineering, and using a fixed sliding window length T to construct a time series training sample group {x1, x2, ..., x τ ,τ=T} to perform Z-score standardization and missing value filling, identify abnormal learning behavior data through the IQR method, delete the abnormal data from the time series samples and input them into the convolutional neural network, and extract local features by sliding the convolution kernel in the convolution layer on the time series.

[0013] Use the multi-scale output splicing fusion strategy to integrate the output sequence of the convolution kernel group G={G1,G2,...,G τ ,τ=T′}, and use time dimension average pooling to obtain the pooled sequence Use the additive attention scoring function to calculate the attention score q of the local behavior feature at each time step in the pooled sequence τ′ ;

[0014] Normalize the attention scores of all time steps to obtain the attention weight κ of each time step τ′ Perform element-wise weighted operations on the feature vector of each time step to obtain the final behavioral feature sequence The final behavior characteristics of each time step are input into the long short-term memory network for gating mechanism calculation to obtain the hidden state vector h τ′ Obtain hidden state sequence H={h1,h2,...,h τ′}, and an adaptive genetic algorithm is used to perform offline global optimization of the key hyperparameters and structural configuration of LSTM, and finally output the optimal individual parameter Θ * For the final model training, the final learning progress prediction value sequence Y={y1,y2,...,y z}.

[0015] As a preferred solution of the family education content recommendation method based on big data described in the present invention, the social network analysis algorithm is combined with the Louvain algorithm and the Page ranking algorithm to establish the social relationship and interaction pattern between family members, and the personalized recommendation target is dynamically optimized by reinforcement learning. A unique identifier is assigned to each family member, all educational content is numbered in chronological order, all interactive behaviors of each family member on each educational content are recorded, all interactive behaviors are converted into binary indicator variables, and a user-content interaction matrix γ∈R is constructed. u×N , through the user-content interaction matrix γ∈R u×N A weighted adjacency matrix A between members is constructed and imported into the graph analysis tool Gephi for visualization. An undirected weighted graph is constructed with family members as nodes and the interaction intensity between members as edge weights. Based on the interaction data between family members, the data is uniformly encoded as a feature vector for normalization. The similarity value Similarity(i,j) between nodes is calculated by weighted similarity. A threshold J is set. If the similarity value Similarity(i,j) is greater than the threshold J, they are considered to be members of the same community, forming a preliminary social group. The Louvain algorithm is used to optimize community division and calculate the modularity I of each community to identify the optimal community structure. The PageRank algorithm is used to calculate the relative importance of each node in the graph to evaluate the node's influence PR(i). The Jaccard similarity is used to calculate the common participation of nodes in interactive content to measure similarity JC(i,j). A threshold θ is set. When the Jaccard similarity JC(i,j) of node i and node j is greater than the threshold θ, they will be classified as the same social subgroup. The reinforcement learning algorithm is used to dynamically adjust the optimal recommended content Z by combining the social relationships in the social graph and the family member behavior data. * .

[0016] As a preferred solution of the family education content recommendation method based on big data described in the present invention, wherein: the timing and frequency of pushing content are dynamically adjusted according to the prediction results of the LSTM model, and the personalized recommendation target after pushing the optimization is combined with the intelligent push mechanism. The learning progress prediction sequence Y={y1,y2,...,y z}, analyze the learning rhythm and behavioral characteristics of family members, identify the learning bottlenecks and high-frequency active time periods faced by family members, build a multi-factor push priority model based on the probability of learning bottlenecks, learning progress status and individual activity, and use push priority to dynamically schedule push strategies, pushing the best recommended content that best matches the current status during high-priority time periods. * .

[0017] As a preferred solution of the family education content recommendation method based on big data described in the present invention, the continuous optimization of recommended content based on user feedback refers to collecting user-related behavior data when users interact, determining the reward R′ based on the feedback of user behavior, and updating the Q value according to the Q update formula of the Q-learning algorithm. Whenever the user interacts with the recommended content, the Q value will be updated once to help learn better recommendation strategies. After the Q value is updated, the content type with the highest Q value is selected as the recommended content by maximizing the Q value.

[0018] As a preferred solution of the family education content recommendation method based on big data described in the present invention, the preprocessing of multi-source data of family members refers to collecting education-related data of family members through multi-source channels, removing duplicates and filling missing values ​​in the collected data, and using natural language processing technology to segment text information extracted from social media data and education platforms, remove stop words and extract keywords, and perform Z-score standardization on numerical data that presents an approximately normal distribution.

[0019] As a preferred solution of the family education content recommendation method based on big data described in the present invention, wherein: the real-time analysis of pre-processed data and updating of user profiles, setting personalized recommendation goals based on the analysis results refers to constructing learning progress time series data based on the historical learning progress of family members, and predicting the learning trend and periodicity P through the autoregressive integral sliding average model. t , combined with K-means cluster analysis, user groups with similar learning habits or interest preferences are identified and transmitted to the user profile management office, and the digital profile of each family member is dynamically updated. Based on the updated profile of each family member and the results of real-time analysis, personalized recommendation goals are set.

[0020] In a second aspect, the present invention provides a family education content recommendation system based on big data, comprising:

[0021] Multi-source data preprocessing module, used to collect online platform, social platform and behavioral data, remove duplicates and fill in, and unify feature scale processing;

[0022] User profile update module, used to update user profiles based on ARIMA and cluster analysis and set personalized recommendation target content;

[0023] The user learning prediction module is used to model the behavior sequence through CNN, LSTM and AGA, and output the predicted value of the learning progress of family members;

[0024] The social relationship modeling module is used to combine graph algorithms and similarity analysis to build social graphs and identify social groups and core nodes;

[0025] The intelligent push mechanism module is used to calculate push priority based on predicted progress and learning activity, and dynamically push the best personalized content.

[0026] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the family education content recommendation method based on big data as described in the first aspect of the present invention.

[0027] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the family education content recommendation method based on big data as described in the first aspect of the present invention.

[0028] The beneficial effects of the present invention are as follows: the present invention significantly improves the accuracy and real-time performance of personalized recommendations by combining the ARIMA model to predict learning trends and analyzing learning behavior characteristics through K-means clustering, extracts local features through convolution, focuses on key behaviors through the attention mechanism, and performs deep modeling through long and short-term memory networks. Combined with the genetic algorithm to globally optimize hyperparameters, the present invention significantly improves the accuracy of learning progress prediction and model robustness, continuously updates the Q value through Q-learning, and selects recommended content based on the maximum Q value strategy, forming a real-time feedback and adaptive recommendation content adjustment mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 This is a flowchart of the family education content recommendation method based on big data in Example 1.

[0031] Figure 2 This is a structural diagram of the family education content recommendation system based on big data in Example 1.

[0032] Figure 3 This is a flow chart of data collection and preprocessing in Example 1. DETAILED DESCRIPTION

[0033] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0034] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0035] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0036] Example 1, reference Figures 1 to 3 , which is the first embodiment of the present invention, provides a method for recommending family education content based on big data, comprising the following steps:

[0037] S1. Collect multi-source data of family members for pre-processing, analyze the pre-processed data in real time and update user profiles, and set personalized recommendation goals based on the analysis results;

[0038] Specifically, collecting multi-source data of family members for pre-processing means collecting education-related data of family members through multiple sources, mainly including:

[0039] Collect learning resource data (videos, textbooks, and exercises) used by family members from online education platforms;

[0040] Capture user-generated content (discussions on educational topics, questions and feedback from parents, and interest tags) from parenting forums and family education social platforms;

[0041] Obtain behavioral record data from family members' daily learning behaviors (video viewing history, article reading time, homework submission rate, content interaction frequency (likes, comments, etc.));

[0042] The collected data were deduplicated and missing values ​​filled, and natural language processing technology was used to segment text information extracted from social media data and educational platforms, remove stop words, and extract keywords. Numerical data with an approximately normal distribution (learning progress, learning time) were normalized with Z-scores, and data with clear upper and lower limits (scores, completion levels) were normalized to the interval [0,1] to achieve consistent scaling of input features.

[0043] By collecting education-related data from multiple channels (online education platforms, parenting forums, family member behavior data, etc.), the comprehensiveness and multi-dimensionality of the data are improved, avoiding the limitations of data sources in traditional education platforms. Through data cleaning and text processing, noise can be eliminated, the reliability of the data can be improved, and the accuracy of personalized analysis can be improved, making the analysis of educational content more refined, and able to extract valuable educational trends and the specific needs of family members. Through standardization and normalization processing, the consistency of the data and the compatibility of the model are improved, ensuring that the comparison of various indicators is more fair. Through precise data analysis, highly personalized educational recommendations can be achieved.

[0044] Furthermore, the pre-processed data is analyzed in real time and the user profile is updated. Based on the analysis results, personalized recommendation goals are set. This means that learning progress time series data is constructed based on the historical learning progress of family members, and learning trends and periodicity are predicted using the Autoregressive Integrated Moving Average (ARIMA) model:

[0045] P t =αP t-1 +βP t-2 +...+ε t

[0046] Among them, P t is the learning trend and periodicity at time point t, α, β are the model coefficients, and ε t is the error term;

[0047] Based on the pre-processed behavioral feature data (video viewing frequency, homework completion, interactive behavior), the learning behavior of family members is clustered through K-means cluster analysis to identify user groups with similar learning habits or interest preferences, and the learning trends and periodic P t The cluster analysis results are transmitted to the user profile management office to dynamically update the digital profile of each family member. The updated content includes learning progress status, changes in interest preferences, and current personalized learning needs. Based on the updated profile of each family member and the results of real-time analysis, personalized recommendation goals (children's education content, parents' education skills, and family interactive activities) are set.

[0048] By using the autoregressive integrated moving average (ARIMA) model to analyze the time series data of family members' learning progress, the prediction of learning progress becomes more scientific, avoiding judgment based solely on experience or intuition. By using K-means cluster analysis to cluster the learning behaviors of family members, it is possible to identify potential learning similarities between students, avoiding a "one-size-fits-all" education plan, and providing personalized and differentiated educational content for different groups, thereby enhancing the sense of participation and effectiveness of learning. The results of trend analysis and cluster analysis are transmitted to the user profile management system, and the digital profile of each family member is dynamically updated to ensure that the educational content always matches the latest needs of family members, enhancing the timeliness of education. Based on the updated profiles and real-time analysis results, personalized recommendation goals are set, which significantly improves the relevance of educational content and enables family members to focus on the resources or activities they need most.

[0049] S2. Using updated user profiles, we use the LSTM model to predict user learning progress. We use a social network analysis algorithm combined with the Louvain algorithm and the Page ranking algorithm to establish social relationships and interaction patterns among family members. We then use reinforcement learning to dynamically optimize personalized recommendation goals.

[0050] Specifically, the LSTM model is used to predict the user learning progress through the updated user profile. The real-time updated user profile is sampled to obtain the key learning behavior feature data of each family member (daily learning time, number of video views and average completion rate, homework completion status (submission rate and accuracy) and content interaction behavior (such as likes, comments, and collections)). The key learning behavior feature data is converted into feature vectors through feature engineering, and a time series training sample group X = {x1, x2, ..., x3} is constructed by fixing the sliding window length T. τ ,τ=T}, where x τ is the learning behavior characteristic of time step τ, T is the length of the time series, the data in the time series training sample group are Z-score standardized and missing values ​​are filled, and the abnormal learning behavior data are identified by the IQR (interquartile range) method. The abnormality judgment standard is:

[0051]

[0052] Among them, x τ is the learning behavior feature vector at time step τ, U1 and U3 are the first and third quartiles of the ξ-th feature, IQR is the interquartile range, IQR = U3-U1;

[0053] After deleting the abnormal data from the time series samples, the data is input into the convolutional neural network (CNN). The convolution kernel in the convolution layer slides on the time series to extract local features. The convolution operation calculation formula is:

[0054]

[0055] in, is the local behavior feature of the convolution output at time step τ, ω a is the weight of the convolution kernel, which is set by Xavier initialization. is the bias term, is the size of the convolution kernel, a is the index of the convolution kernel in the convolution operation, is the activation function, and the ReLU function is used to improve the nonlinear expression ability, x τ+a is the behavioral feature vector of day τ+a;

[0056] Use a multi-scale output splicing and fusion strategy to integrate the outputs of the convolution kernel group and enhance the network's ability to recognize changes in different periodic behaviors:

[0057]

[0058] Among them, G τ is the fused local behavior feature, and Concat(·) is the feature dimension concatenation operation;

[0059] The output sequence after fusion G={G1,G2,...,G τ ,τ=T′}, where T′ is the length of the fused sequence, T is the fixed sliding window length;

[0060] To reduce time step redundancy and improve the model's focus on important behavioral changes, average pooling in the time dimension is used to reduce the length of the feature map and compress the calculation:

[0061]

[0062] in, is the local behavior feature of the time step τ′ after pooling, p is the pooling window size, j is the index of the jth pooling segment, G τ is the fused local behavior feature vector;

[0063] Get the pooled sequence Among them, T′1 is the sequence length after pooling, The additive attention scoring function is used to calculate the attention score of the local behavior features at each time step in the pooled sequence:

[0064]

[0065] Among them, q τ′ is the attention score at time step τ′, is the local behavior feature of the time step τ′ after pooling, Ψ πis a learnable weight matrix, is a learnable bias term, v π is the attention score vector, and T is the transpose operation of the matrix;

[0066] Perform softmax normalization on the attention scores of all time steps to obtain the attention weight of each time step:

[0067]

[0068] Among them, κ τ′ is the attention weight at time step τ′, q k is the attention score at time step k;

[0069] By the attention weight κ τ′ , perform element-wise weighted operations on the feature vector of each time step to obtain the enhanced feature vector:

[0070]

[0071] in, is the weighted attention-enhanced feature vector, ⊙ is the element-wise multiplication;

[0072] The final behavior feature sequence is obtained based on the weighted attention-enhanced feature vector of each time step The final behavior features of each time step are input into the long short-term memory network (LSTM) for gating mechanism (forget gate, input gate, candidate state, output gate) calculation:

[0073]

[0074] Among them, f τ′ is the forget gate vector at time step τ′, is the concatenation of the previous hidden state and the current input, W f is the forget gate weight, set by backpropagation and optimizer, b f is the forget gate bias, is the Sigmoid activation function, h τ′-1 is the hidden state of the previous time step, is the attention-enhanced feature vector at time step τ;

[0075]

[0076] Among them, τ′ is the input gate vector at time step τ′, W o is the input gate weight, set by backpropagation and optimizer, b o is the input gate bias;

[0077]

[0078] Among them, c′ τ′ is a candidate memory state, W c is the candidate state weight, set by backpropagation and optimizer, b c is the candidate state bias;

[0079] c τ′ =f τ′ ⊙c τ′-1 +o τ′ ⊙c′ τ′

[0080] Among them, c τ′ is the unit state at the current moment;

[0081] h τ′ =o τ′ ⊙tanh(c τ′ )

[0082] Among them, h τ′ is the hidden state vector, i.e. the final output of this time step;

[0083] Hidden state vector h τ′ As one of the inputs of the next moment, it is passed to the upper layer of the LSTM model. Finally, the LSTM model outputs a set of hidden state sequences H = {h1,h2,...,h τ′}, and an adaptive genetic algorithm (AGA) is used to perform offline global optimization of the key hyperparameters and structural configuration of LSTM to further improve the model prediction accuracy and avoid falling into local optimality due to improper hyperparameter settings. The mean square error (MSE) on the validation set is selected as the fitness evaluation indicator, and the optimization goal is to minimize the validation loss function:

[0084]

[0085] Among them, Θ is a set of hyperparameter configurations of LSTM, It is the loss value calculated using MSE loss on the validation set;

[0086] Use real number coding to encode each set of hyperparameters to form individual chromosomes:

[0087] Γ=[ι,φ,χ,ψ,∈]

[0088] Among them, Γ is the chromosome individual, ι is the number of hidden layer units, φ is the learning rate, χ is the batch size, that is, the number of training samples per round, ψ is the Dropout probability, which controls overfitting and affects robustness, and ∈ is the Attention projection dimension, which controls the attention representation ability;

[0089] Assume that the initial population size is Ω, and randomly uniformly sample the first generation of individuals in each parameter interval to obtain Each individual (i.e., a set of parameters) will be trained on the training set and its performance will be evaluated on the validation set. The fitness function is defined as:

[0090]

[0091] Among them, Fit(Γ) is the fitness value of chromosome individual Γ, is the loss value on the validation set after the model is trained using the hyperparameters corresponding to the chromosome individual Γ. is a smoothing constant to prevent the denominator from being zero;

[0092] The roulette wheel selection method is used to replicate individuals according to fitness probability, and through the uniform crossover strategy, any two parent individuals exchange each gene position with equal probability:

[0093]

[0094] Among them, Γ new [r] is the new chromosome individual, Γ Ξ , Γ Π are the two selected parent individuals;

[0095] The probability of individual mutation in the sth generation is set by the convergence criterion as:

[0096]

[0097] in, is the mutation probability of individuals in the sth generation, Φ max is the maximum mutation rate, is the new chromosome individual Γ in the sth generation new The fitness of [r], is the minimum and average value of the fitness of the sth generation, Λ is a small constant to avoid division by 0;

[0098] Let the maximum iteration number be Each round retains the best individual to enter the next generation (elite strategy), when the maximum number of generations is reached Finally, the optimal individual parameters are output for final model training:

[0099]

[0100] Among them, Θ * is the optimal individual parameter;

[0101] The final learning progress prediction value sequence Y={y1,y2,...,y z}.

[0102] By extracting key behavioral data and constructing time series samples, we can capture all aspects of learning behavior and ensure that the data input to the model is comprehensive and representative. By using the IQR (interquartile range) method to identify abnormal learning behavior data and delete abnormal data, it helps to improve the quality of training data and ensure that the model will not produce deviations or misjudgments due to data anomalies. Through the extraction of local features by CNN, the model can focus on more fine-grained learning behavior information, improving the accuracy of prediction. Through pooling operations in the time dimension, unnecessary information interference is reduced, which helps to improve the model's perception of important features. The long short-term memory network (LSTM) improves the ability to capture long-term dependence on learning progress, further improving the prediction accuracy of learning progress. Through the hyperparameter optimization of the adaptive genetic algorithm (AGA), the optimal performance of the model during training and verification is guaranteed, thereby avoiding training failure or overfitting problems caused by inappropriate parameter configuration.

[0103] Furthermore, the social network analysis algorithm is combined with the Louvain algorithm and the Page ranking algorithm to establish the social relationship and interaction pattern among family members. The personalized recommendation target is dynamically optimized by reinforcement learning. A unique identifier is assigned to each family member (the family number is named in the form of the member relationship), all educational content is numbered in chronological order, and all interactive behaviors of each family member on each educational content (watching, liking, commenting, sharing) are recorded. All interactive behaviors are converted into binary indicator variables to construct the user-content interaction matrix γ∈R u×N , where u is the total number of family members, N is the total number of educational contents recorded, and the user-content interaction matrix γ∈R u×N Construct the weighted adjacency matrix A between members:

[0104]

[0105] Among them, A ij is an element in the adjacency matrix, representing the connection strength between family member (node) i and family member (node) j, γ ik and γ jk indicates that family member i and family member j have interacted with the kth educational content (watching, commenting, liking, sharing), and N is the total number of all recorded educational contents;

[0106] Import the adjacency matrix A into the graph analysis tool Gephi for visualization, and construct an undirected weighted graph with family members as nodes and interaction strength between members as edge weights. ij When >0, connect the edge and assign the edge weight to A ij, so that the social graph has an interactive structure between family members based on educational content. Based on the interactive data between family members (video viewing, homework completion, discussion participation), it is uniformly encoded into a feature vector and normalized to avoid result deviation caused by dimensionality differences. The similarity value between nodes is calculated by weighted similarity:

[0107]

[0108] Among them, Similarity(i,j) is the similarity value between node i and node j, w im and w jm are the behavioral feature values ​​of node i and node j in the mth dimension (video watching, homework completion, discussion participation), and n is the number of behavioral features;

[0109] The threshold J is set through statistical analysis. When the similarity value Similarity(i,j) is greater than the threshold J, they are considered to be members of the same community, forming a preliminary social group. The Louvain algorithm is used to optimize the community division and calculate the modularity I of each community:

[0110]

[0111] Where l is the total number of edges in the social graph, A ij is an element in the adjacency matrix, ζ i and ζ j are the degrees of node i and node j, i.e. the number of direct neighbors, δ(O i ,O j ) is the Kroneckerdelta function, which takes the value 1 if nodes i and j are in the same community, and 0 otherwise;

[0112] The Louvain algorithm is used to calculate and maximize the modularity I of the graph, identify the optimal community structure, maximize the interaction within each community, and minimize the interaction between different communities. The PageRank algorithm is used to calculate the relative importance of each node in the graph to evaluate the influence of the node:

[0113]

[0114] Among them, PR(i) is the influence of node i, M(i) is the set of neighbor nodes pointing to node i, L(j) is the out-degree of node j, which represents the number of neighbors of node j, and d is the damping factor, which is used to simulate the behavior of random browsers to prevent a node from having too much influence. is the weighted contribution of node j’s influence on node i;

[0115] All neighboring nodes pointing to node i are accumulated to obtain the contribution of each node to the target node (the node whose influence is calculated). This process is repeated until the influence of all nodes converges. To further quantify the overlap of content interests, the Jaccard similarity is used to calculate the common participation of nodes (between family members) in interactive content to measure similarity:

[0116]

[0117] Where JC(i, j) is the Jaccard similarity between nodes i and j, that is, the degree of similarity in educational content, B(i) and B(j) are the sets of educational content in which nodes i and j interact, respectively, |B(i)∩B(j)| is the number of educational contents in which nodes i and j participate together, and |B(i)∪B(j)| is the total number of all educational contents in which nodes i and j participate.

[0118] A threshold θ is set through fuzzy logic. When the Jaccard similarity JC(i, j) between nodes i and j is greater than the threshold θ, that is, the interaction frequency and interests between them are similar (there is a high degree of overlap in the interaction of educational content), they will be classified as the same social sub-group. Through the reinforcement learning (Q-learning) algorithm, the recommendation strategy is dynamically adjusted in combination with the social relationships in the social graph and the behavior data of family members, so as to provide personalized recommendation content that best meets the needs of each family member. The state space S (learning progress, position in the social graph, and interest preferences), the action space Z (recommendations for learning videos, family education tips articles, and parent-child interactive games or activities) and the reward function R (content click-through rate, dwell time, and user ratings) are defined. The reward function R is calculated as follows:

[0119] R=λ·C+μ·V+υ·F

[0120] Where C is the content click-through rate, V is the normalized user dwell time, F is the user feedback score, and λ, μ, and υ are weight coefficients, which are set through experiments.

[0121] Update the Q value:

[0122]

[0123] Among them, Q(S v ,Z v ) is in state S v Next take action Z v Q value, R v+1 is the user's reward, θ is the discount factor, η is the learning rate, In the new state S v+1 Next, the action with the largest Q value among all possible actions;

[0124] Through continuous learning and updating, the Q value corresponding to each recommended action is optimized, so that the optimal recommended action is selected in each state, thereby maximizing the user's feedback reward, and the optimal recommended content is selected by maximizing the Q value:

[0125]

[0126] Among them, Z * It is the best recommended content in the current state.

[0127] By constructing an interaction matrix, a comprehensive mapping of family members and educational content is achieved, providing high-quality input data for the subsequent personalized recommendation system. By constructing a weighted adjacency matrix and graphical visualization, the interactive relationships between family members are visualized, facilitating the rapid identification of family member roles (such as core nodes and edge nodes), which helps to accurately provide educational advice and resources. By calculating the behavioral similarity between nodes and optimizing the node partitioning in the graph using the Louvain algorithm, the relevance of recommended content is improved, making the allocation of educational resources more efficient and targeted. The PageRank algorithm is used to evaluate the influence of nodes, enabling more accurate identification of core family members and ensuring a more balanced allocation of educational resources and activities. Jaccard similarity is used to calculate the overlap of family members' interests in educational content, minimizing information overload and avoiding irrelevant content recommendations. This improves the relevance of educational content and member engagement, making recommendations more tailored to actual needs. Reinforcement learning (Q-learning) is used to optimize personalized recommendations, giving the recommendation system the ability to self-optimize and adjust to changes in family member behavior and fluctuations in interests, thereby enhancing its adaptability and intelligence.

[0128] S3. Dynamically adjust the timing and frequency of content push based on the LSTM model prediction results, combine the intelligent push mechanism to push optimized personalized recommendation targets, and continuously optimize the recommended content based on user feedback.

[0129] Specifically, the timing and frequency of content push are dynamically adjusted according to the prediction results of the LSTM model, and the personalized recommendation target after push optimization is pushed in combination with the intelligent push mechanism. The learning progress prediction sequence Y={y1,y2,...,y z}, analyze the learning rhythm and behavioral characteristics of family members, identify the learning bottlenecks and high-frequency active time periods faced by family members, build a multi-factor push priority model based on the probability of learning bottlenecks, learning progress status, and individual activity, and comprehensively evaluate the timing and frequency of recommended content push:

[0130]

[0131] Among them, D e is the push priority at time e, Bo(e) is the probability of facing a learning bottleneck at time e, Pr(e) is the value of the current learning progress, Us(e) is the learning activity of family members at time e, is the weight coefficient, which is set through adaptive adjustment;

[0132] Use push priority to dynamically schedule push strategies and push the best recommended content that best matches the current status during high priority periods. * .

[0133] By accurately predicting learning progress through the quantum long short-term memory network (LSTM), we can effectively understand the rhythm of family members in the learning process and identify potential bottlenecks or stagnation periods. By accurately identifying learning bottlenecks, the system can push appropriate tutoring or motivational content based on the students' real-time status, preventing students from stagnating in difficulties and improving learning outcomes. Through a multi-factor model, the push priority can intelligently reflect the learning status of family members, ensuring the efficiency and accuracy of recommended content. Through flexible push strategies, the system can provide family members with the most needed learning resources at the appropriate time, significantly improving the timeliness of educational intervention.

[0134] Furthermore, continuously optimizing recommended content based on user feedback means collecting user-related behavior data (click content, dwell time, user ratings) when users interact, determining the reward R′ based on user behavior feedback (a click indicates a reward of 1, and the longer the dwell time, the higher the reward), and updating the Q value according to the Q-learning algorithm's Q update formula. Whenever a user interacts with the recommended content, the Q value is updated once, helping the system learn better recommendation strategies. After the Q value is updated, the content type with the highest Q value is selected as the recommended content by maximizing the Q value, thereby ensuring that the recommended content can maximize the user's interests and needs.

[0135] By collecting behavioral data (click content, dwell time, user ratings) through user interaction, a data foundation is provided, enabling the recommendation system to make feedback based on the user's actual operations rather than preset rules, thereby enhancing the adaptability and dynamism of the recommendation logic. Setting a reward mechanism based on behavioral feedback can guide the system to continuously optimize the recommendation results through trial and error exploration, making the recommendation strategy more in line with the user's actual preferences. By maximizing the Q value and selecting the optimal content type for recommendation, the time cost for users to obtain content of interest is reduced, thereby improving the overall user experience and platform usage efficiency.

[0136] This embodiment also provides a family education content recommendation system based on big data, including:

[0137] Multi-source data preprocessing module, used to collect online platform, social platform and behavioral data, remove duplicates and fill in, and unify feature scale processing;

[0138] User profile update module, used to update user profiles based on ARIMA and cluster analysis and set personalized recommendation target content;

[0139] The user learning prediction module is used to model the behavior sequence through CNN, LSTM and AGA, and output the predicted value of the learning progress of family members;

[0140] The social relationship modeling module is used to combine graph algorithms and similarity analysis to build social graphs and identify social groups and core nodes;

[0141] The intelligent push mechanism module is used to calculate push priority based on predicted progress and learning activity, and dynamically push the best personalized content.

[0142] This embodiment also provides a computer device, which is suitable for the family education content recommendation method based on big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the family education content recommendation method based on big data proposed in the above embodiment.

[0143] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.

[0144] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for recommending family education content based on big data as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.

Claims

1. A family education content recommendation method based on big data, characterized by: include, Collect multi-source data of family members for pre-processing, analyze the pre-processed data in real time and update user profiles, and set personalized recommendation goals based on the analysis results; The multi-source data includes education platform data, social media data and family behavior data; Through updated user profiles, the LSTM model is used to predict user learning progress. Social network analysis algorithms are combined with the Louvain algorithm and the Page ranking algorithm to establish social relationships and interaction patterns among family members. Reinforcement learning is used to dynamically optimize personalized recommendation goals. Dynamically adjust the timing and frequency of content push based on the LSTM model prediction results, combine the intelligent push mechanism to push optimized personalized recommendation targets, and continuously optimize the recommended content based on user feedback.

2. The method for recommending family education content based on big data according to claim 1, characterized in that: The method of predicting user learning progress by using the LSTM model through the updated user profile refers to sampling the real-time updated user profile to obtain the key learning behavior feature data of each family member, converting it into a feature vector through feature engineering, and using a fixed sliding window length to construct a time series training sample group X = {x1, x2, ..., x τ ,τ=T} to perform Z-score standardization and missing value filling, identify abnormal learning behavior data through the IQR method, delete the abnormal data from the time series samples and input them into the convolutional neural network, and extract local features by sliding the convolution kernel in the convolution layer on the time series. Use the multi-scale output splicing fusion strategy to integrate the output sequence of the convolution kernel group G={G1,G2,...,G τ ,τ=T′}, and use time dimension average pooling to obtain the pooled sequence Use the additive attention scoring function to calculate the attention score q of the local behavior feature at each time step in the pooled sequence τ′ ; Normalize the attention scores of all time steps to obtain the attention weight κ of each time step τ′ Perform element-wise weighted operations on the feature vector of each time step to obtain the final behavioral feature sequence The final behavior characteristics of each time step are input into the long short-term memory network for gating mechanism calculation to obtain the hidden state vector h τ′ Obtain hidden state sequence H={h1,h2,...,h τ′ }, and an adaptive genetic algorithm is used to perform offline global optimization of the key hyperparameters and structural configuration of LSTM, and finally output the optimal individual parameter Θ * For the final model training, the final learning progress prediction value sequence Y={y1,y2,...,y z }.

3. The method for recommending family education content based on big data according to claim 2, wherein: The social network analysis algorithm is combined with the Louvain algorithm and the Page ranking algorithm to establish the social relationship and interaction pattern among family members, and the reinforcement learning is used to dynamically optimize the personalized recommendation target. A unique identifier is assigned to each family member, all educational contents are numbered in chronological order, all interactive behaviors of each family member on each educational content are recorded, all interactive behaviors are converted into binary indicator variables, and the user-content interaction matrix γ∈R is constructed. u×N , through the user-content interaction matrix γ∈R u×N A weighted adjacency matrix A between members is constructed and imported into the graph analysis tool Gephi for visualization. An undirected weighted graph is constructed with family members as nodes and the interaction intensity between members as edge weights. Based on the interaction data between family members, the data is uniformly encoded as a feature vector for normalization. The similarity value Similarity(i,j) between nodes is calculated by weighted similarity. A threshold J is set. When the similarity value Similarity(i,j) is greater than the threshold J, they are regarded as members of the same community, forming a preliminary social group. The Louvain algorithm is used to optimize community division and calculate the modularity I of each community to identify the optimal community structure. The PageRank algorithm is used to calculate the relative importance of each node in the graph to evaluate the node's influence PR(i). The Jaccard similarity is used to calculate the common participation of nodes in interactive content to measure similarity JC(i,j). A threshold θ is set. When the Jaccard similarity JC(i,j) of node i and node j is greater than the threshold θ, they are classified as the same social subgroup. The reinforcement learning algorithm is used to dynamically adjust the optimal recommended content Z by combining the social relationships in the social graph and the family member behavior data. * .

4. The method for recommending family education content based on big data according to claim 3, wherein: The timing and frequency of content push are dynamically adjusted according to the prediction results of the LSTM model, and the personalized recommendation target after push optimization is pushed in combination with the intelligent push mechanism. The learning progress prediction sequence Y={y1,y2,...,y z }, analyze the learning rhythm and behavioral characteristics of family members, identify the learning bottlenecks and high-frequency active time periods faced by family members, build a multi-factor push priority model based on the probability of learning bottlenecks, learning progress status and individual activity, and use push priority to dynamically schedule push strategies, pushing the best recommended content that best matches the current status during high-priority time periods. * .

5. The method for recommending family education content based on big data according to claim 4, characterized in that: Continuously optimizing recommended content based on user feedback refers to collecting user-related behavior data when users interact, determining the reward R′ based on the feedback of user behavior, and updating the Q value according to the Q update formula of the Q-learning algorithm. Every time a user interacts with the recommended content, the Q value is updated once, helping to learn a better recommendation strategy. After the Q value is updated, the content type with the highest Q value is selected as the recommended content by maximizing the Q value.

6. The method for recommending family education content based on big data according to claim 5, characterized in that: The preprocessing of multi-source data collected from family members refers to collecting education-related data from family members through multiple sources, removing duplicates and filling missing values ​​in the collected data, and using natural language processing technology to segment text information extracted from social media data and education platforms, remove stop words and extract keywords, and perform Z-score normalization on numerical data that presents an approximately normal distribution.

7. The method for recommending family education content based on big data according to claim 6, characterized in that: The said real-time analysis of pre-processed data and updating of user profiles, and setting personalized recommendation goals based on the analysis results refer to constructing learning progress time series data based on the historical learning progress of family members, and predicting the learning trend and periodicity through the autoregressive integral sliding average model. t , combined with K-means cluster analysis, user groups with similar learning habits or interest preferences are identified and transmitted to the user profile management office, and the digital profile of each family member is dynamically updated. Based on the updated profile of each family member and the results of real-time analysis, personalized recommendation goals are set.

8. A family education content recommendation system based on big data, based on the family education content recommendation method based on big data according to any one of claims 1 to 7, characterized in that: include, Multi-source data preprocessing module, used to collect online platform, social platform and behavioral data, remove duplicates and fill in, and unify feature scale processing; User profile update module, used to update user profiles based on ARIMA and cluster analysis and set personalized recommendation target content; The user learning prediction module is used to model the behavior sequence through CNN, LSTM and AGA, and output the predicted value of the learning progress of family members; The social relationship modeling module is used to combine graph algorithms and similarity analysis to build social graphs and identify social groups and core nodes; The intelligent push mechanism module is used to calculate push priority based on predicted progress and learning activity, and dynamically push the best personalized content.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for recommending family education content based on big data according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for recommending family education content based on big data according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intelligent music recommendation method based on emotion perception and acoustic characteristics

    CN121614635A