Career Data Collection and Analysis System and Method Based on Big Data

Through a career data acquisition and analysis system based on big data, network crawling technology and deep neural networks are used to build a personalized exercise recommendation model, which solves the problem that traditional systems cannot be personalized and improves learning efficiency and accuracy.

CN119441610BActive Publication Date: 2025-07-25JIANGXI NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411509083.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-07-25
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

The traditional education system cannot personalize the individual differences of each student, resulting in students being unable to obtain exercises and learning content for their specific needs, increasing learning costs and small data collection, which cannot fully reflect learning behaviors and habits.

Method used

A career data acquisition and analysis system based on big data is obtained through network crawling technology, a test prediction model is constructed, unexpected coefficients are added and solved using Newton-Ravson iterative method, a cognitive level matrix and knowledge point matrix are established, and a deep neural network model is trained for exercise recommendations.

Benefits of technology

It improves the accuracy and learning efficiency of exercise recommendations, reduces the time for blindly doing questions, reduces the cost of learning, and improves the learning experience and fun.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119441610B_ABST
    Figure CN119441610B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data analysis, specifically a career data collection and analysis system and method based on big data, including a matrix construction unit, a model training unit, a model construction unit, and an exercise recommendation unit. It further includes: a data collection unit, which is used to obtain the career data of the target user on the online learning website through web crawler technology. The career data includes the knowledge points of the exercises done by the target user, and transmits the career data to the probability prediction unit and the matrix construction unit. By analyzing the career data of users, the present invention establishes a cognitive level matrix of users, identifies the strengths and weaknesses of users in knowledge points, so as to provide targeted exercise recommendations for users. Through the exercise recommendation model, the most suitable exercises for users can be recommended, helping users focus on the knowledge points that need improvement, thereby improving learning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and particularly to a career data collection and analysis system and method based on big data. Background Art

[0002] Career data includes the exercises done by the target user during the learning process and the knowledge points involved. These data can be used to track the learning progress of students, evaluate learning outcomes, identify learning difficulties, personalize learning paths, etc. By analyzing career data, educational institutions and educators can better understand the learning needs of students, provide targeted support and guidance, and thus promote the learning and development of students.

[0003] Traditional systems are usually based on fixed curriculum designs and cannot be adjusted individually for each student, which may result in students not being able to obtain exercises and learning content tailored to their specific needs. Moreover, traditional systems usually rely on the experience of teachers or teaching materials rather than data-based analysis, which causes the system to be unable to adjust in a timely manner to meet the changing needs of students. And due to the lack of personalization and intelligent recommendation in traditional systems, students may need to purchase additional materials or spend extra time on tutoring, increasing the learning cost. Also, traditional systems usually rely on paper-based assignments, exams, and classroom activities, and the data collected in this way is relatively small and cannot comprehensively reflect the learning behaviors, knowledge mastery, and learning habits of students. Summary of the Invention

[0004] The purpose of the present invention is to address the problems in the background art and propose a career data collection and analysis system and method based on big data.

[0005] The technical solution of the present invention: A career data collection and analysis system based on big data includes a matrix construction unit, a model training unit, a model construction unit, and an exercise recommendation unit, and further includes:

[0006] A data collection unit, which is used to obtain the career data of the target user on the online learning website through web crawler technology. The career data includes the knowledge points of the exercises done by the target user, and transmits the career data to the probability prediction unit and the matrix construction unit;

[0007] A probability prediction unit, which receives the career data transmitted by the data collection unit, constructs an exercise prediction model based on the career data, and adds an accident coefficient to the exercise prediction model to obtain a corrected exercise prediction model. The accident coefficient includes a carelessness coefficient and a guessing coefficient, and transmits the corrected exercise prediction model to the model solving unit;

[0008] The model solving unit receives the corrected exercise prediction model transmitted by the probability prediction unit, constructs the conditional distribution and marginal distribution of the answering results of the exercises done by the target user based on the corrected exercise prediction model, and uses the Newton-Raphson iteration method to solve each unknown parameter in the corrected exercise prediction model based on the conditional distribution and marginal distribution, so as to obtain the probability of the target user being correct on each exercise after correction, and transmits the probability of the target user being correct on each exercise to the cognitive matrix construction unit.

[0009] Preferably, the matrix construction unit receives the probability of the target user being correct on each exercise transmitted by the model solving unit and the career data transmitted by the data collection unit, constructs the cognitive level matrix of the target user based on the probability of the target user being correct on each exercise, matches the cognitive level matrix with the career data, so as to obtain the knowledge point matrix corresponding to the cognitive level matrix, and transmits the cognitive level matrix and the knowledge point matrix to the model training unit.

[0010] Preferably, the model construction unit is used to construct an exercise recommendation model based on a deep neural network. The exercise recommendation model includes an input layer, an embedding layer, a fusion layer, a hidden layer and a prediction layer, and transmits the exercise recommendation model to the model training unit.

[0011] Preferably, the model training unit receives the cognitive level matrix and the knowledge point matrix transmitted by the matrix construction unit and the exercise recommendation model transmitted by the model construction unit, and trains the exercise recommendation model through the cognitive level matrix and the knowledge point matrix, so as to obtain a trained exercise recommendation model, and transmits the trained exercise recommendation model to the exercise recommendation unit.

[0012] Preferably, the exercise recommendation unit receives the trained exercise recommendation model transmitted by the model training unit, inputs the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the trained exercise recommendation model, outputs the exercise recommendation probability through the trained exercise recommendation model, and compares the exercise recommendation probability with a preset threshold, so as to obtain the recommended exercises.

[0013] Preferably, obtaining the career data of the target user on the online learning website through web crawler technology includes the following steps:

[0014] A1. Locate the website address of the online learning website, and screen the target content from the online learning website based on the screening conditions provided by the target user;

[0015] A2. Obtain the response information of the online learning website server through the requests module in Python. The server forms a list of the target content in json format and returns it in the form of javascript;

[0016] A3. Use the Python regular expression module to match the content of the response information and extract the content of the response information to form a required data list;

[0017] A4. Based on the required data list, form a list of access links. Based on the list of access links, obtain the target content, and form the target content in the format of a Python dictionary to obtain career data. Create a career data list for all the career data.

[0018] Preferably, the exercise prediction model is as follows:

[0019]

[0020] where δ jk represents the predicted correct probability of the target user j on exercise k according to the knowledge point mastery situation, B kn represents the examination situation of exercise k on knowledge point n, C jn represents the mastery situation of the target user j on knowledge point n, and N represents the total number of knowledge points related to exercise k;

[0021] The corrected exercise prediction model is as follows:

[0022] P jk = β k (1 - γ k );

[0023] where P jk represents the probability of the target user j being correct on exercise k after correction, β k represents the probability of the target user j getting wrong due to carelessness on exercise k, and γ k represents the probability of the target user j guessing correctly on exercise k.

[0024] Preferably, the conditional distribution of the answering results of the exercises done by the target user is as follows:

[0025]

[0026] where P jk represents the probability of the target user j being correct on exercise k after correction, L1 represents the conditional distribution of the answering results of the exercises done by the target user, and m represents the total number of exercises done by the target user;

[0027] The marginal distribution of the answering results of the exercises done by the target user is as follows:

[0028]

[0029] Among them, P jk represents the probability that the target user j is correct on exercise k after correction, L2 represents the marginal distribution of the answering results of the exercises done by the target user, J represents the total number of the target users, and m represents the total number of exercises done by the target user.

[0030] Preferably, the method of using Newton-Raphson iteration based on the conditional distribution and the marginal distribution is used to solve each unknown parameter in the corrected exercise prediction model, including the following steps:

[0031] B1. Take the logarithm of the marginal distribution of the answering results of the exercises done by the target user. The marginal distribution of the answering results of the exercises done by the target user after taking the logarithm is as follows:

[0032]

[0033] B2. Based on the marginal distribution of the answering results of the exercises done by the target user after taking the logarithm, take the derivative of the probability that the target user j makes a careless mistake on exercise k and the probability that the target user j guesses correctly on exercise k. The marginal distribution of the answering results of the exercises done by the target user after taking the logarithm and taking the derivative is as follows:

[0034]

[0035] B3. Calculate the marginal distribution of the answering results of the exercises done by the target user after taking the logarithm and taking the derivative through the Newton-Raphson iteration method, so as to obtain the estimated values of the probability that the target user j makes a careless mistake on exercise k and the probability that the target user j guesses correctly on exercise k.

[0036] Preferably, inputting the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the trained exercise recommendation model, and outputting the exercise recommendation probability through the trained exercise recommendation model, including the following steps:

[0037] C1. Input the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the input layer. The input layer inputs the cognitive level matrix and the knowledge point matrix into the embedding layer, and the input layer performs one-hot encoding on the cognitive level matrix and the knowledge point matrix, and inputs the one-hot encoded cognitive level matrix and knowledge point matrix into the embedding layer;

[0038] C2. The embedding layer uses a fully connected layer to embed the cognitive level matrix and the knowledge point matrix, thereby obtaining a cognitive level preference vector and a knowledge point preference vector. Moreover, the embedding layer embeds the one-hot encoded cognitive level matrix and knowledge point matrix, thereby obtaining a cognitive level latent vector and a knowledge point latent vector. The cognitive level preference vector and the knowledge point preference vector are as follows:

[0039]

[0040] Among them, e u represents the cognitive level preference vector, s u represents the cognitive level matrix, W u represents the weight matrix of the embedding layer with respect to the cognitive level matrix, d i represents the knowledge point preference vector, s i represents the knowledge point matrix, W i represents the weight matrix of the embedding layer with respect to the knowledge point matrix;

[0041] The cognitive level latent vector and the knowledge point latent vector are as follows:

[0042]

[0043] Among them, p u represents the cognitive level latent vector, P T represents the parameter matrix of the embedding layer with respect to the one-hot encoded cognitive level matrix, represents the one-hot encoded cognitive level matrix, q i represents the knowledge point latent vector, Q T represents the parameter matrix of the embedding layer with respect to the one-hot encoded knowledge point matrix, V i I represents the one-hot encoded knowledge point matrix;

[0044] C3. The fusion layer performs feature fusion on the cognitive level preference vector and the knowledge point preference vector with the corresponding cognitive level latent vector and knowledge point latent vector, thereby obtaining a target user preference vector. The target user preference vector is as follows:

[0045]

[0046] Among them, f represents the target user preference vector, represents element-wise multiplication between two vectors;

[0047] C4. The hidden layer is a tower neural network structure composed of multiple fully connected layers. Through the hidden layer, multiple target user preference vectors are jointly encoded to capture the non-linear relationship between multiple target user preference vectors, and the ReLU activation function is used as the non-linear factor of the hidden layer;

[0048] C5. The prediction layer maps the output of the hidden layer to the exercise recommendation probability, and the expression of the exercise recommendation probability is as follows:

[0049]

[0050] where, represents the exercise recommendation probability, sigmoid() represents the activation function of the prediction layer, and x represents the specific value output by the hidden layer.

[0051] The technical solution of the present invention: A career data collection and analysis method based on big data, which is applicable to the career data collection and analysis system based on big data, includes the following steps:

[0052] S1. Collect career data from the online learning website of the target user through web crawler technology. The career data includes the exercises done by the target user and related knowledge points;

[0053] S2. Analyze the career data and construct an exercise prediction model, and add a carelessness coefficient and a guessing coefficient to the exercise prediction model to correct the exercise prediction model;

[0054] S3. Use the Newton-Raphson iteration method to solve the unknown parameters in the corrected exercise prediction model;

[0055] S4. Based on the unknown parameters solved in the corrected exercise prediction model, obtain the correct probability of the target user on each exercise, thereby construct the cognitive level matrix of the user, and match the cognitive level matrix with the career data to generate a knowledge point matrix corresponding to the cognitive level matrix;

[0056] S5. Use the cognitive level matrix and the knowledge point matrix to train the exercise recommendation model;

[0057] S6. Input the cognitive level matrix and the knowledge point matrix of the target user into the exercise recommendation model, and provide the finally recommended exercises according to the comparison between the exercise recommendation probability output by the exercise recommendation model and the preset threshold.

[0058] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:

[0059] 1. The present invention analyzes the user's career data, establishes a cognitive level matrix of the user, identifies the strengths and weaknesses of the user in knowledge points, and thus provides targeted exercise recommendations for the user. Through the exercise recommendation model, the most suitable exercises for the user can be recommended, helping the user focus on the knowledge points that need improvement, thereby improving learning efficiency. And by adopting big data technology and using methods such as web crawlers to collect and analyze the user's online learning data, the system makes decisions based on real data, improving the accuracy of exercise recommendations.

[0060] 2. The accuracy of exercise recommendations is continuously improved through the exercise prediction model and the deep neural network to better meet the learning needs of users. Moreover, through intelligent recommendations, users can reduce the time and energy spent on blindly doing exercises, focus their attention on the most important knowledge points, thereby reducing the overall learning cost. And personalized exercise recommendations can enhance the user's learning experience, increasing the fun and effectiveness of learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a schematic flowchart of the overall system in an embodiment proposed by the present invention;

[0062] Figure 2 It is a schematic flowchart of obtaining career data in an embodiment proposed by the present invention;

[0063] Figure 3 It is a schematic flowchart of solving each unknown parameter in the corrected exercise prediction model in an embodiment proposed by the present invention;

[0064] Figure 4 It is a schematic flowchart of the exercise recommendation model outputting the exercise recommendation probability in an embodiment proposed by the present invention;

[0065] Figure 5 It is a schematic flowchart of the overall method in an embodiment proposed by the present invention.

[0066] Reference numerals: 1, data acquisition unit; 2, probability prediction unit; 3, model solving unit; 4, matrix construction unit; 5, model training unit; 6, model construction unit; 7, exercise recommendation unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] Embodiment 1, as Figure 1 shown, the big data-based career data acquisition and analysis system proposed by the present invention includes a matrix construction unit 4, a model training unit 5, a model construction unit 6, and an exercise recommendation unit 7, and further includes:

[0068] Data acquisition unit 1, which is used to obtain the career data of target users on the online learning website through web crawler technology. The career data includes the knowledge points of the exercises done by the target users, and transmits the career data to the probability prediction unit 2 and the matrix construction unit 4;

[0069] Probability prediction unit 2, which receives the career data transmitted by the data acquisition unit 1, constructs an exercise prediction model based on the career data, and adds accident coefficients to the exercise prediction model to obtain a corrected exercise prediction model. The accident coefficients include a carelessness coefficient and a guessing coefficient, and transmits the corrected exercise prediction model to the model solving unit 3;

[0070] Model solving unit 3, which receives the corrected exercise prediction model transmitted by the probability prediction unit 2, constructs the conditional distribution and marginal distribution of the answering results of the exercises done by the target user based on the corrected exercise prediction model, and uses the Newton-Raphson iteration method to solve each unknown parameter in the corrected exercise prediction model based on the conditional distribution and marginal distribution, so as to obtain the probability of the target user being correct in each exercise after correction, and transmits the probability of the target user being correct in each exercise to the cognitive matrix construction unit 4.

[0071] In the present invention, the Newton-Raphson iteration method is a numerical method for solving equations and optimization problems. Its basic principle is to iteratively approximate the root or extreme value of the objective function through linear approximation. Specifically, it uses the function and its derivative to calculate the iteration step size and then approximates the solution. This method is widely used in numerical calculations to solve non-linear equations and optimization problems.

[0072] In an optional embodiment, the matrix construction unit 4 receives the probability of the target user being correct in each exercise transmitted by the model solving unit 3 and the career data transmitted by the data acquisition unit 1, constructs the cognitive level matrix of the target user based on the probability of the target user being correct in each exercise, matches the cognitive level matrix with the career data, so as to obtain a knowledge point matrix corresponding to the cognitive level matrix, and transmits the cognitive level matrix and the knowledge point matrix to the model training unit 5.

[0073] In an optional embodiment, the model construction unit 6 is used to construct an exercise recommendation model based on a deep neural network. The exercise recommendation model includes an input layer, an embedding layer, a fusion layer, a hidden layer and a prediction layer, and transmits the exercise recommendation model to the model training unit 5.

[0074] It should be noted that a deep neural network (DNN) is a type of artificial neural network, and its main feature is having multiple hidden layers. Compared with traditional shallow neural networks, DNNs can learn complex patterns and features in data through more complex structures and greater computing power.

[0075] In an optional embodiment, the model training unit 5 receives the cognitive level matrix and knowledge point matrix transmitted by the matrix construction unit 4 and the exercise recommendation model transmitted by the model construction unit 6, and trains the exercise recommendation model through the cognitive level matrix and knowledge point matrix, thereby obtaining a trained exercise recommendation model and transmitting the trained exercise recommendation model to the exercise recommendation unit 7.

[0076] In an optional embodiment, the exercise recommendation unit 7 receives the trained exercise recommendation model transmitted by the model training unit 5, inputs the cognitive level matrix and knowledge point matrix of the target user who needs exercise recommendation into the trained exercise recommendation model, outputs the exercise recommendation probability through the trained exercise recommendation model, and compares the exercise recommendation probability with a preset threshold to obtain the recommended exercises.

[0077] Embodiment 2, as Figures 2 - 4 shown, the career data collection and analysis system based on big data proposed by the present invention, compared with Embodiment 1, this embodiment further includes:

[0078] Obtain the career data of the target user on the online learning website through web crawler technology, including the following steps:

[0079] A1. Locate the website address of the online learning website and screen the target content from the online learning website based on the screening conditions provided by the target user;

[0080] A2. Obtain the response information of the online learning website server through the requests module of Python. The server forms a list of the target content in json format and returns it in the form of javascript;

[0081] A3. Use the Python regular expression module to match the content of the response information and extract the content of the response information to form the required data list;

[0082] A4. Based on the required data list, form a list of access links, obtain the target content based on the list of access links, and form the target content in the format of a Python dictionary to obtain career data, and establish a career data list for all career data.

[0083] In this embodiment, web crawler technology refers to using automated scripts or programs to browse and extract information from websites. This process usually involves simulating the behavior of human users on the Internet, retrieving web page content, and extracting useful data according to predefined rules. The application scope of web crawler technology is very wide, including search engine indexing, data collection, market research, competitive analysis, price tracking, content monitoring, academic research, etc.

[0084] In an alternative embodiment, the exercise prediction model is as follows:

[0085]

[0086] where, δ jk represents the predicted correct probability of target user j on exercise k according to the mastery of knowledge points, B kn represents the examination situation of exercise k on knowledge point n, C jn represents the mastery situation of target user j on knowledge point n, and N represents the total number of knowledge points related to exercise k;

[0087] The corrected exercise prediction model is as follows:

[0088] P jk = β k (1 - γ k );

[0089] where, P jk represents the corrected probability of target user j being correct on exercise k, β k represents the probability of target user j getting wrong due to carelessness on exercise k, and γ k represents the probability of target user j guessing correctly on exercise k.

[0090] In an alternative embodiment, the conditional distribution of the answering results of the exercises done by the target user is as follows:

[0091]

[0092] where, P jk represents the corrected probability of target user j being correct on exercise k, L1 represents the conditional distribution of the answering results of the exercises done by the target user, and m represents the total number of exercises done by the target user;

[0093] The marginal distribution of the answering results of the exercises done by the target user is as follows:

[0094]

[0095] where, P jkIt represents the probability that the corrected target user j answers question k correctly. L2 represents the marginal distribution of the answering results of the questions done by the target user, J represents the total number of target users, and m represents the total number of questions done by the target user.

[0096] In an alternative embodiment, the Newton-Raphson iteration method is used based on the conditional distribution and the marginal distribution to solve for each unknown parameter in the corrected exercise prediction model, including the following steps:

[0097] B1. Take the logarithm of the marginal distribution of the answering results of the questions done by the target user. The marginal distribution of the answering results of the questions done by the target user after taking the logarithm is as follows:

[0098]

[0099] B2. Take the derivatives of the probability that the target user j answers question k carelessly wrong and the probability that the target user j guesses correctly based on the marginal distribution of the answering results of the questions done by the target user after taking the logarithm. The marginal distribution of the answering results of the questions done by the target user after taking the logarithm and taking the derivatives is as follows:

[0100]

[0101] B3. Calculate the marginal distribution of the answering results of the questions done by the target user after taking the logarithm and taking the derivatives through the Newton-Raphson iteration method, so as to obtain the estimated values of the probability that the target user j answers question k carelessly wrong and the probability that the target user j guesses correctly.

[0102] In an alternative embodiment, the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation are input into the trained exercise recommendation model, and the exercise recommendation probability is output through the trained exercise recommendation model, including the following steps:

[0103] C1. Input the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the input layer. The input layer inputs the cognitive level matrix and the knowledge point matrix into the embedding layer, and the input layer performs one-hot encoding on the cognitive level matrix and the knowledge point matrix, and inputs the one-hot encoded cognitive level matrix and knowledge point matrix into the embedding layer;

[0104] C2. The embedding layer uses a fully connected layer to embed the cognitive level matrix and the knowledge point matrix, so as to obtain a cognitive level preference vector and a knowledge point preference vector. And the embedding layer embeds the one-hot encoded cognitive level matrix and knowledge point matrix, so as to obtain a cognitive level latent vector and a knowledge point latent vector. The cognitive level preference vector and the knowledge point preference vector are as follows:

[0105]

[0106] Among them, e u represents the cognitive level preference vector, s u represents the cognitive level matrix, W u represents the weight matrix of the cognitive level matrix in the embedding layer, d i represents the knowledge point preference vector, s i represents the knowledge point matrix, W i represents the weight matrix of the knowledge point matrix in the embedding layer;

[0107] The cognitive level latent vector and the knowledge point latent vector are as follows:

[0108]

[0109] Among them, p u represents the cognitive level latent vector, P T represents the parameter matrix of the one-hot encoded cognitive level matrix in the embedding layer, represents the one-hot encoded cognitive level matrix, q i represents the knowledge point latent vector, Q T represents the parameter matrix of the one-hot encoded knowledge point matrix in the embedding layer, V i I represents the one-hot encoded knowledge point matrix;

[0110] C3. The fusion layer fuses the features of the cognitive level preference vector and the knowledge point preference vector with the corresponding cognitive level latent vector and knowledge point latent vector to obtain the target user preference vector, and the target user preference vector is as follows:

[0111]

[0112] Among them, f represents the target user preference vector, represents element-wise multiplication between two vectors;

[0113] C4. The hidden layer is a tower neural network structure composed of multiple fully connected layers. Through the hidden layer, multiple target user preference vectors are jointly encoded, and the non-linear relationship between multiple target user preference vectors is captured, and the ReLU activation function is used as the non-linear factor of the hidden layer;

[0114] C5. The prediction layer maps the output of the hidden layer to the exercise recommendation probability, and the exercise recommendation probability expression is as follows:

[0115]

[0116] Among them, represents the exercise recommendation probability, sigmoid() represents the activation function of the prediction layer, and x represents the specific value output by the hidden layer.

[0117] It should be noted that the tower neural network structure is a deep learning network structure, usually used in application scenarios such as recommendation systems, ranking tasks, click-through rate prediction, etc. Its characteristic is that the input passes through a series of fully connected layers that gradually shrink (from wide to narrow), forming a tower-like structure, hence the name. The design concept of the tower structure is to gradually capture more complex non-linear relationships between data in a deep neural network.

[0118] Example 3, as Figure 5 shown, the method for collecting and analyzing career data based on big data proposed by the present invention is applicable to the system for collecting and analyzing career data based on big data, and includes the following steps:

[0119] S1. Collect career data from the online learning websites of target users through web crawler technology. The career data includes the exercises done by the target users and the relevant knowledge points;

[0120] S2. Analyze the career data and construct an exercise prediction model, and add a carelessness coefficient and a guessing coefficient to the exercise prediction model to correct the exercise prediction model;

[0121] S3. Use the Newton-Raphson iteration method to solve the unknown parameters in the corrected exercise prediction model;

[0122] S4. Based on the unknown parameters solved in the corrected exercise prediction model, obtain the correct probability of the target user on each exercise, thereby construct the cognitive level matrix of the user, and match the cognitive level matrix with the career data to generate a knowledge point matrix corresponding to the cognitive level matrix;

[0123] S5. Use the cognitive level matrix and the knowledge point matrix to train the exercise recommendation model;

[0124] S6. Input the cognitive level matrix and the knowledge point matrix of the target user into the exercise recommendation model, and provide the finally recommended exercises according to the comparison between the exercise recommendation probability output by the exercise recommendation model and the preset threshold.

[0125] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.

Claims

1. A career data collection and analysis system based on big data, including a matrix construction unit (4), a model training unit (5), a model construction unit (6), and an exercise recommendation unit (7), characterized in that: A data collection unit (1), which is used to obtain career data of target users on online learning websites through web crawler technology. The career data includes the knowledge points of the exercises done by the target users, and transmits the career data to a probability prediction unit (2) and a matrix construction unit (4); A probability prediction unit (2), which receives the career data transmitted by the data collection unit (1), constructs an exercise prediction model based on the career data, and adds an accident coefficient to the exercise prediction model to obtain a corrected exercise prediction model. The accident coefficient includes a carelessness coefficient and a guessing coefficient, and transmits the corrected exercise prediction model to a model solving unit (3); The exercise prediction model is as follows: ; wherein, represents the predicted correct probability of the target user j on exercise k according to the knowledge point mastery situation, represents the examination situation of exercise k on knowledge point n, represents the mastery situation of the target user j on knowledge point n, represents the total number of knowledge points related to exercise k; The corrected exercise prediction model is as follows: ; wherein, represents the probability that the corrected target user j answers exercise k correctly, represents the probability that the target user j carelessly answers exercise k wrongly, represents the probability that the target user j guesses correctly for exercise k; The conditional distribution of the answering results of the exercises done by the target users is as follows: ; Among them, represents the probability that the corrected target user j is correct on exercise k, represents the conditional distribution of the answering results of the exercises done by the target user, represents the total number of exercises done by the target user; The marginal distribution of the answering results of the exercises done by the target users is as follows: ; Among them, represents the probability that the corrected target user j is correct on exercise k, represents the marginal distribution of the answering results of the exercises done by the target user, represents the total number of the target users, represents the total number of exercises done by the target user; Using the Newton-Raphson iteration method based on the conditional distribution and the marginal distribution to solve each unknown parameter in the corrected exercise prediction model, including the following steps: B1. Take the logarithm of the marginal distribution of the answering results of the exercises done by the target users. The marginal distribution of the answering results of the exercises done by the target users after taking the logarithm is as follows: ; B2. Derive the probability of the target user j carelessly getting a wrong answer on exercise k and the probability of the target user j guessing correctly on exercise k based on the marginal distribution of the answering results of the exercises done by the target users after taking the logarithm. The marginal distribution of the answering results of the exercises done by the target users after taking the logarithm after derivation is as follows: ; B3. Calculate the marginal distribution of the answering results of the exercises done by the target users after taking the logarithm after derivation through the Newton-Raphson iteration method, so as to obtain the estimated values of the probability of the target user j carelessly getting a wrong answer on exercise k and the probability of the target user j guessing correctly on exercise k; A model solving unit (3), which receives the corrected exercise prediction model transmitted by the probability prediction unit (2), constructs the conditional distribution and the marginal distribution of the answering results of the exercises done by the target users based on the corrected exercise prediction model, and uses the Newton-Raphson iteration method based on the conditional distribution and the marginal distribution to solve each unknown parameter in the corrected exercise prediction model, so as to obtain the corrected probability of the target user getting a correct answer on each exercise, and transmits the probability of the target user getting a correct answer on each exercise to a cognitive matrix construction unit (4).

2. The career data collection and analysis system based on big data according to claim 1, characterized in that The matrix construction unit (4) receives the probability of the target user being correct on each exercise transmitted by the model solving unit (3) and the career data transmitted by the data acquisition unit (1), constructs the cognitive level matrix of the target user based on the probability of the target user being correct on each exercise, matches the cognitive level matrix with the career data, thereby obtaining a knowledge point matrix corresponding to the cognitive level matrix, and transmits the cognitive level matrix and the knowledge point matrix to the model training unit (5); The model construction unit (6) is used to construct an exercise recommendation model based on a deep neural network. The exercise recommendation model includes an input layer, an embedding layer, a fusion layer, a hidden layer, and a prediction layer, and transmits the exercise recommendation model to the model training unit (5).

3. The career data collection and analysis system based on big data according to claim 2, wherein The model training unit (5) receives the cognitive level matrix and the knowledge point matrix transmitted by the matrix construction unit (4) and the exercise recommendation model transmitted by the model construction unit (6), and trains the exercise recommendation model through the cognitive level matrix and the knowledge point matrix, thereby obtaining a trained exercise recommendation model, and transmits the trained exercise recommendation model to the exercise recommendation unit (7); The exercise recommendation unit (7) receives the trained exercise recommendation model transmitted by the model training unit (5), inputs the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the trained exercise recommendation model, outputs an exercise recommendation probability through the trained exercise recommendation model, and compares the exercise recommendation probability with a preset threshold, thereby obtaining the recommended exercises.

4. The career data collection and analysis system based on big data according to claim 1, characterized in that Obtaining the career data of the target user through web crawler technology includes the following steps: A1. Locate the website address of the online learning website, and screen the target content from the online learning website based on the screening conditions provided by the target user; A2. Obtain the response information of the online learning website server through the requests module of Python. The server forms a list of the target content in json format and returns it in the form of javascript; A3. Use the Python regular expression module to match the response information content and extract the response information content to form a required data list; A4. Based on the required data list, form an access link list, obtain the target content based on the access link list, and form the target content in the format of a Python dictionary, thereby obtaining career data, and establish a career data list for all the career data.

5. The career data collection and analysis system based on big data according to claim 3, characterized in that, Inputting the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the trained exercise recommendation model, and outputting an exercise recommendation probability through the trained exercise recommendation model includes the following steps: C1. Input the cognitive level matrix and the knowledge point matrix of the target user who needs exercise recommendation into the input layer. The input layer inputs the cognitive level matrix and the knowledge point matrix into the embedding layer. Moreover, the input layer performs one-hot encoding on the cognitive level matrix and the knowledge point matrix, and inputs the one-hot encoded cognitive level matrix and knowledge point matrix into the embedding layer; C2. The embedding layer uses a fully connected layer to embed the cognitive level matrix and the knowledge point matrix, so as to obtain a cognitive level preference vector and a knowledge point preference vector. And the embedding layer embeds the one-hot encoded cognitive level matrix and knowledge point matrix, so as to obtain a cognitive level latent vector and a knowledge point latent vector. The cognitive level preference vector and the knowledge point preference vector are as follows: ; Among them, represents the cognitive level preference vector, represents the cognitive level matrix, represents the weight matrix of the embedding layer with respect to the cognitive level matrix, represents the knowledge point preference vector, represents the knowledge point matrix, represents the weight matrix of the embedding layer with respect to the knowledge point matrix; The cognitive level latent vector and the knowledge point latent vector are as follows: ; Among them, represents the potential vector of the cognitive level, represents the parameter matrix of the cognitive level matrix after one-hot encoding in the embedding layer, represents the cognitive level matrix after one-hot encoding, represents the potential vector of the knowledge point, represents the parameter matrix of the knowledge point matrix after one-hot encoding in the embedding layer, represents the knowledge point matrix after one-hot encoding; C3. The fusion layer performs feature fusion on the cognitive level preference vector and the knowledge point preference vector with the corresponding cognitive level latent vector and knowledge point latent vector, so as to obtain a target user preference vector. The target user preference vector is as follows: ; Among them, represents the target user preference vector, represents the multiplication of corresponding elements between two vectors; C4. The hidden layer is a tower neural network structure composed of multiple fully connected layers. Through the hidden layer, multiple target user preference vectors are jointly encoded, and the non-linear relationship between multiple target user preference vectors is captured. And the ReLU activation function is used as the non-linear factor of the hidden layer; C5. The prediction layer maps the output of the hidden layer to an exercise recommendation probability. The expression of the exercise recommendation probability is as follows: ; Among them, represents the exercise recommendation probability, represents the activation function of the prediction layer, represents the specific value output by the hidden layer.

6. A career data collection and analysis method based on big data, which is applicable to the career data collection and analysis system based on big data according to any one of claims 1-5, characterized in that, Including the following steps: S1. Collect career data from the online learning website of the target user through web crawler technology. The career data includes the exercises done by the target user and the relevant knowledge points; S2. Analyze the career data and construct an exercise prediction model. Add a carelessness coefficient and a guessing coefficient to the exercise prediction model to correct the exercise prediction model; S3. Use the Newton-Raphson iteration method to solve the unknown parameters in the corrected exercise prediction model; S4. Based on solving the unknown parameters in the corrected exercise prediction model, obtain the probability of the target user being correct on each exercise, thereby construct the cognitive level matrix of the user, and match the cognitive level matrix with the career data to generate a knowledge point matrix corresponding to the cognitive level matrix; S5. Use the cognitive level matrix and the knowledge point matrix to train the exercise recommendation model; S6. Input the cognitive level matrix and the knowledge point matrix of the target user into the exercise recommendation model, and provide the finally recommended exercises according to the comparison between the exercise recommendation probability output by the exercise recommendation model and the preset threshold.

Citation Information

Patent Citations

  • Mathematical question and answering process generation method thereof

    CN117251533A

  • Exercise recommendation method and system based on preorder relation between cognitive diagnosis and knowledge points

    CN118690784A