Big data and machine learning driven personalized training recommendation system

By combining data cleaning and deep learning of the personalized training recommendation system, the problems of incomplete data and historical bias are solved, more accurate recommendations and diversified learning experiences are achieved, and students' learning effects are improved.

CN120563281APending Publication Date: 2025-08-29CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510537041.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the existing personalized training recommendation system, data quality problems lead to insufficient accuracy in recommendation content, and the biases brought by historical data limit students' learning diversity, forming an information cocoon.

Method used

Through data cleaning, denoising, standardized processing, combined with deep neural networks and enhanced learning strategies, a hybrid recommendation algorithm and knowledge graph are used to expand multi-dimensional data and adjust the recommendation strategy in real time.

Benefits of technology

It improves the accuracy of the recommended content, expands the students' learning horizons, avoids the information cocoon effect, and improves the students' learning effect and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563281A_ABST
    Figure CN120563281A_ABST
Patent Text Reader

Abstract

The invention provides a personalized training recommendation system driven by big data and machine learning. The big data and machine learning driven personalized training recommendation system comprises a, a data acquisition module for collecting personal information, learning behaviors, interaction data and learning preference data of students, and b, a data processing module. According to the personalized training recommendation system driven by big data and machine learning, through data cleaning, denoising and standardization processing, the system ensures the integrity and accuracy of collected data, and the negative influence of data incompleteness, excessive noise or deviation on model training is avoided. Besides, advanced machine learning technologies such as a deep neural network and a reinforcement learning strategy are adopted, so that the overfitting problem can be avoided when the recommendation model processes historical data, the generalization ability of the model is optimized, and the accuracy of recommended content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of personalized training recommendation, and in particular to a personalized training recommendation system driven by big data and machine learning. Background Art

[0002] A personalized training recommendation system driven by big data and machine learning typically includes a data acquisition module, a data processing module, a model training module, a recommendation engine, and a feedback system. First, the data acquisition module is responsible for collecting students' personal information, learning behaviors, and interaction data. The multidimensional data collected through sensors or platforms supports subsequent analysis. Second, the data processing module cleans and preprocesses the collected data to ensure its accuracy and usability. Subsequently, the model training module uses machine learning algorithms (such as decision trees and neural networks) to learn from historical data, thereby predicting student needs and generating personalized recommendations. The recommendation engine intelligently pushes personalized learning resources or courses based on students' interests and needs. Finally, the feedback system collects student feedback to continuously optimize the recommendation algorithm.

[0003] While this system possesses the ability to provide personalized recommendations, it still suffers from some significant flaws. First, data quality issues can severely impact recommendation effectiveness. If the collected data is incomplete, noisy, or biased, the model's accuracy will decrease, resulting in inaccurate recommendations. Second, recommendation algorithms based on historical data can be biased, causing the system to only recommend content that students have already encountered, thereby limiting student learning diversity and creating an information cocoon. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a personalized training recommendation system driven by big data and machine learning, which solves the problem that the accuracy of the model will be reduced due to incomplete data, excessive noise or deviation, resulting in inaccurate recommendations; the recommendation algorithm of historical data will be biased, limiting the learning diversity of students and forming an information cocoon.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a personalized training recommendation system driven by big data and machine learning, comprising:

[0006] a. Data collection module, used to collect students' personal information, learning behavior, interaction data and learning preference data;

[0007] b. Data processing module, used to clean, denoise and standardize the student data to ensure data integrity and accuracy;

[0008] c. A feature extraction module for extracting representative feature data from the student data as input to the machine learning model;

[0009] d. A machine learning model training module, which trains historical student data using reinforcement learning strategies and deep neural networks to generate a personalized recommendation model. The training formula of the machine learning model is:

[0010]

[0011] Among them, θ represents the model parameters, including the weight matrix and bias term in the neural network;

[0012] x i Represents the characteristic data of the i-th student, including the student's learning history, behavior, and preference information;

[0013] y i Represents the learning feedback of the i-th student, usually the student's clicks, evaluations or completion status of the recommended course;

[0014] L(f θ (x i ),y i ) is the loss function, which is used to measure the difference between the model prediction results and the actual feedback;

[0015] λ is a regularization parameter used to control the complexity of the model and avoid overfitting;

[0016] It is an L2 regularization term, which is used to penalize excessive weights and improve the generalization ability of the model;

[0017] e. A recommendation engine module that recommends the most suitable training content or learning resources to students based on the personalized recommendation model;

[0018] f. Feedback collection module, used to collect students' feedback on recommended content and optimize and adjust the recommendation engine.

[0019] Preferably, the data acquisition module includes:

[0020] Student personal information collection unit, used to collect students' basic personal information;

[0021] Learning behavior data collection unit, used to collect students' learning history, learning time, and learning path behavior data;

[0022] The interactive data collection unit is used to collect detailed data on the interaction between students and the system, including but not limited to clicks, browsing, likes, and comments.

[0023] Preferably, the data processing module includes:

[0024] Data cleaning unit, used to remove noise, outliers and duplicate data from collected data;

[0025] Data standardization unit, used to standardize student data from different sources to ensure data consistency;

[0026] The missing value filling unit is used to deal with the missing value problem in the student data and use the interpolation method to fill the data.

[0027] Preferably, the machine learning model training module adopts a recommendation algorithm based on deep learning, and the recommendation algorithm uses a loss function to measure the difference between the recommendation result and the student feedback, and the loss function is a cross entropy loss function:

[0028]

[0029] in, is the recommendation probability predicted by the model, y i Provide practical feedback to students.

[0030] Preferably, the machine learning model includes at least one of a deep neural network, a convolutional neural network, or a long short-term memory network, which is used to process student data and generate personalized training content recommendations.

[0031] Preferably, the recommendation engine module recommends appropriate learning resources to students by combining a collaborative filtering-based algorithm and a content-based recommendation algorithm. The recommendation engine module further includes a knowledge graph-based recommendation mechanism to solve the problem of mining students' potential learning needs.

[0032] Preferably, the system uses data enhancement technology to improve data quality and diversity by multi-dimensionally expanding and synthesizing student data, thereby improving the accuracy of the recommendation algorithm.

[0033] Preferably, the recommendation engine module adopts a hybrid recommendation algorithm to conduct a comprehensive analysis of the student's interests, learning history, and learning progress, and provides the most suitable courses or training content for the student.

[0034] This invention provides a personalized training recommendation system driven by big data and machine learning. It has the following beneficial effects:

[0035] This personalized training recommendation system, powered by big data and machine learning, ensures the integrity and accuracy of collected data through data cleaning, denoising, and standardization, preventing the negative impact of incomplete data, excessive noise, or bias on model training. Furthermore, the use of advanced machine learning techniques such as deep neural networks and reinforcement learning strategies prevents overfitting of the recommendation model when processing historical data, optimizes the model's generalization capabilities, and thus improves the accuracy of recommended content. Through these technological innovations, the system has significantly improved data quality and recommendation accuracy.

[0036] This system effectively solves the bias problem that may be caused by historical data recommendation algorithms by introducing a hybrid recommendation algorithm based on content and collaborative filtering, and combining it with technologies such as knowledge graphs and data enhancement. Traditional recommendation systems are often limited by the information cocoon effect and only recommend content that students have been exposed to. However, the present invention can expand students' learning horizons and recommend more diverse and challenging learning resources through multi-dimensional data analysis. Combined with the feedback collection module, the system can adjust the recommendation strategy in real time and continuously optimize the students' learning path, thereby providing a more personalized and innovative learning experience, significantly improving students' learning results and satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] Example 1

[0040] like Figure 1 As shown, an embodiment of the present invention provides a personalized training recommendation system driven by big data and machine learning, including: a. a data collection module for collecting students' personal information, learning behavior, interaction data, and learning preference data. The data collection module includes:

[0041] Student personal information collection unit, used to collect students' basic personal information;

[0042] Learning behavior data collection unit, used to collect students' learning history, learning time, and learning path behavior data;

[0043] Interactive data collection unit, used to collect detailed data on students' interactions with the system, including but not limited to clicks, browsing, likes, and comments;

[0044] b. Data processing module, used to clean, denoise, and standardize student data to ensure data integrity and accuracy. The data processing module includes:

[0045] Data cleaning unit, used to remove noise, outliers and duplicate data from collected data;

[0046] Data standardization unit, used to standardize student data from different sources to ensure data consistency;

[0047] Missing value filling unit is used to deal with the missing value problem in the student data and use interpolation method to fill the data;

[0048] c. Feature extraction module, used to extract representative feature data from student data as input to the machine learning model;

[0049] d. The machine learning model training module trains historical student data using reinforcement learning strategies and deep neural networks to generate a personalized recommendation model. The training formula for the machine learning model is:

[0050]

[0051] Among them, θ represents the model parameters, including the weight matrix and bias term in the neural network, x i Represents the characteristic data of the i-th student, including the student's learning history, behavior, preference information, y i Represents the learning feedback of the i-th student, usually the student's clicks, evaluations or completion status of the recommended course;

[0052] L(f θ (x i ),y i ) is the loss function, which is used to measure the difference between the model prediction results and the actual feedback;

[0053] λ is a regularization parameter used to control the complexity of the model and avoid overfitting;

[0054] The L2 regularization term is used to penalize excessive weights and improve the generalization ability of the model. The machine learning model includes at least one of a deep neural network, a convolutional neural network, or a long short-term memory network, which is used to process student data and generate personalized training content recommendations;

[0055] The machine learning model training module adopts a recommendation algorithm based on deep learning. The recommendation algorithm uses a loss function to measure the difference between the recommendation results and the student feedback. The loss function is the cross-entropy loss function:

[0056]

[0057] in, is the recommendation probability predicted by the model, y i Provide practical feedback to students;

[0058] e. The recommendation engine module recommends the most suitable training content or learning resources to students based on a personalized recommendation model. The recommendation engine module recommends appropriate learning resources to students through a combination of collaborative filtering and content-based recommendation algorithms. The recommendation engine module further includes a knowledge graph-based recommendation mechanism to address the problem of mining students' potential learning needs. The recommendation engine module uses a hybrid recommendation algorithm to conduct a comprehensive analysis of students' interests, learning history, and learning progress to provide the most suitable courses or training content for students.

[0059] f. Feedback collection module, used to collect students' feedback on recommended content and optimize the recommendation engine. The system uses data enhancement technology to expand and synthesize student data in multiple dimensions to improve data quality and diversity, thereby improving the accuracy of the recommendation algorithm.

[0060] Experimental Examples

[0061] Purpose of the experiment:

[0062] Verify the effectiveness of the "big data and machine learning driven personalized training recommendation system" provided by the present invention in improving data quality, avoiding bias and improving recommendation accuracy when processing student data.

[0063] Experimental steps:

[0064] Data Collection:

[0065] The learning behavior data of 1,000 students were collected from an online learning platform, including their personal information (such as age and subject preferences), learning history (such as completed courses and learning duration), interaction data (such as clicks, likes, and comments), and learning feedback (such as students' ratings of recommended content).

[0066] There is noise, missing values ​​and abnormal data in the data.

[0067] Data preprocessing:

[0068] The collected data is processed using the data cleaning unit to remove duplicate data and outliers.

[0069] Fill in missing data, use interpolation method to fill in missing values ​​in learning behavior data, and standardize the characteristic data of different students to ensure the uniformity of the data.

[0070] Data augmentation technology is used to expand data and generate additional student behavior data to simulate the student's learning process, enhance data diversity, and improve recommendation accuracy.

[0071] Model training:

[0072] The cleaned data is used to train a machine learning model. The model uses a deep neural network (DNN) as the recommendation algorithm, and the training objective of the model is to minimize the loss function.

[0073] Recommendation engine implementation:

[0074] A hybrid recommendation algorithm (a combination of collaborative filtering and content-based recommendation algorithms) is used to generate a personalized recommendation list for each student. The collaborative filtering algorithm makes recommendations based on the student's historical behavior, while the content recommendation algorithm delivers personalized recommendations based on the student's interests and learning preferences.

[0075] Add a knowledge graph-based recommendation mechanism to the recommendation engine to explore students' potential learning needs, expand their learning horizons, and avoid the information cocoon effect.

[0076] Experimental verification:

[0077] Experimental group: The personalized training recommendation system of the present invention was used for recommendation.

[0078] Control group: Use traditional recommendation systems based on historical data and simple recommendation algorithms (such as collaborative filtering) for recommendations.

[0079] During the experiment, the recommendation effect of the system was evaluated. The main evaluation indicators included the relevance of the recommended content, student satisfaction, diversity of learning content, and learning completion rate.

[0080] Result evaluation:

[0081] After the experiment, analyze the differences between the experimental group and the control group:

[0082] Recommendation relevance: Students in the experimental group had higher click-through rates and feedback scores for the recommended content, indicating that the recommended content was more in line with the students' needs and interests.

[0083] Learning diversity: Students in the experimental group were exposed to more different types of learning content, which expanded the scope of learning and avoided the information cocoon effect.

[0084] Student satisfaction: Through the questionnaire survey, students in the experimental group expressed high satisfaction, and gave feedback on the personalization of recommended content and the optimization of learning paths.

[0085] Learning completion rate: The learning completion rate of the experimental group was significantly higher than that of the control group, indicating that the system can better stimulate students' learning interest and initiative.

[0086] Experimental results:

[0087] The experimental group significantly outperformed the control group in terms of recommendation accuracy and content diversity. The click-through rate of recommended content increased by 15%, and the learning completion rate increased by 12%. Furthermore, students were exposed to a wider range of content during their learning process, avoiding the information cocoon effect.

[0088] Through the system optimization feedback mechanism, the system can dynamically adjust according to the students' real-time learning data, continuously improving the accuracy of personalized recommendations and students' learning satisfaction.

[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A personalized training recommendation system driven by big data and machine learning, characterized by: include: a. Data collection module, used to collect students' personal information, learning behavior, interaction data and learning preference data; b. Data processing module, used to clean, denoise and standardize the student data to ensure data integrity and accuracy; c. A feature extraction module for extracting representative feature data from the student data as input to the machine learning model; d. A machine learning model training module, which trains historical student data using reinforcement learning strategies and deep neural networks to generate a personalized recommendation model. The training formula of the machine learning model is: Among them, θ represents the model parameters, including the weight matrix and bias term in the neural network; x i Represents the characteristic data of the i-th student, including the student's learning history, behavior, and preference information; y i Represents the learning feedback of the i-th student, usually the student's clicks, evaluations or completion status of the recommended course; L(f θ (x i )y i ) is the loss function, which is used to measure the difference between the model prediction results and the actual feedback; λ is a regularization parameter used to control the complexity of the model and avoid overfitting; It is an L2 regularization term, which is used to penalize excessive weights and improve the generalization ability of the model; e. A recommendation engine module that recommends the most suitable training content or learning resources to students based on the personalized recommendation model; f. Feedback collection module, used to collect students' feedback on recommended content and optimize and adjust the recommendation engine.

2. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The data acquisition module includes: Student personal information collection unit, used to collect students' basic personal information; Learning behavior data collection unit, used to collect students' learning history, learning time, and learning path behavior data; The interactive data collection unit is used to collect detailed data on the interaction between students and the system, including but not limited to clicks, browsing, likes, and comments.

3. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The data processing module includes: Data cleaning unit, used to remove noise, outliers and duplicate data from collected data; Data standardization unit, used to standardize student data from different sources to ensure data consistency; The missing value filling unit is used to deal with the missing value problem in the student data and use the interpolation method to fill the data.

4. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The machine learning model training module adopts a recommendation algorithm based on deep learning. The recommendation algorithm uses a loss function to measure the difference between the recommendation results and the student feedback. The loss function is the cross entropy loss function: in, is the recommendation probability predicted by the model, y i Provide practical feedback to students.

5. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The machine learning model includes at least one of a deep neural network, a convolutional neural network, or a long short-term memory network, which is used to process student data and generate personalized training content recommendations.

6. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The recommendation engine module recommends appropriate learning resources to students by combining collaborative filtering-based algorithms and content-based recommendation algorithms. The recommendation engine module further includes a knowledge graph-based recommendation mechanism to solve the problem of mining students' potential learning needs.

7. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The system adopts data enhancement technology to expand and synthesize student data in multiple dimensions.

8. The big data and machine learning-driven personalized training recommendation system according to claim 1, characterized in that: The recommendation engine module uses a hybrid recommendation algorithm to conduct a comprehensive analysis of students' interests, learning history, and learning progress.