Online education platform data storage management method
By using the multi-head self-attention mechanism on the online education platform, improving the combined model of gated cycle units and multi-layer perceptrons, accurately identifying and storing content of users' interest, the problem of long waiting time in the existing technology is solved and the user experience is improved.
Patent Information
- Application Number
- CN202510294996.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The data storage mechanism of the existing online education platform cannot accurately identify and store content that users are interested in, resulting in long wait time for users to explore content that is interested, affecting the user experience.
The multi-head self-attention mechanism is adopted, the gated cycle unit and multi-layer perception machine are improved, and the content of users is accurately identified through the educational resource model, user interest model and project interest prediction model, and the content of users is stored and managed separately as hot data.
It realizes accurate identification and rapid access to content that is of interest to users, reduces user waiting time and improves user exploration experience.
Smart Images

Figure CN120216770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data storage, and particularly to a method for managing data storage in an online education platform. Background Art
[0002] "Internet + Education" has become a development model in the education industry. The popularization of online education platforms has led to a sharp increase in the number of users, generating a large amount of data including user personal information, learning behavior records, teaching videos, and course materials. An efficient data storage mechanism is crucial for ensuring the smooth operation of online education platforms and enhancing the user experience.
[0003] Currently, the MySQL data storage repository adopted by online education platforms uses the cold and hot data separation technology, which divides data into cold data and hot data according to different data access frequencies and stores them separately. Since the amount of hot data is small, this method can greatly improve the data access efficiency and the response speed of the system. However, it is not accurate enough to distinguish data only based on the data access frequency, which cannot meet the personalized access needs of users on online education platforms, and cannot allow users to quickly access the content they may be interested in. When users explore the content they are interested in, they are very likely to wait for a long time for the system to access and load the cold data in the cold table, resulting in a poor user exploration experience. They may cancel the access while waiting for the cold data to load, thus reducing the motivation of users to explore new content.
[0004] Therefore, how to achieve the storage management of data of interest to users on online education platforms is a technical problem to be solved currently. Summary of the Invention
[0005] For this reason, the present invention provides a method for managing data storage in an online education platform. Through the multi-head self-attention mechanism, the improved gated recurrent unit, and the multi-layer perceptron, the content of interest to users on the online education platform is accurately identified and stored and managed separately. As a result, when users access the content they are interested in, the platform needs to access and load a large amount of data from the database, resulting in a long waiting time and a poor user exploration experience.
[0006] To achieve the above object, the present invention proposes a method for managing data storage in an online education platform. An educational resource database for separating and storing cold data and hot data is set up on the online education platform, including:
[0007] Generating educational resource interest features from the user access resource data of the online education platform through an educational resource model, where the educational resource model is constructed based on the multi-head self-attention mechanism;
[0008] Generating user interest features from the user time series data of the online education platform through a user interest model, where the user interest model is constructed based on the improved gated recurrent unit;
[0009] Determine the predicted interest items based on the educational resource features and the user interest features through an item interest prediction model, where the item interest prediction model is established based on a multi-layer perceptron;
[0010] Store the data corresponding to the predicted interest items as hot data in the educational resource database.
[0011] Further, the educational resource model includes a multi-head self-attention mechanism, an attention mechanism, and a pooling layer. The process of generating educational resource interest features from the user access resource data of the online education platform through the educational resource model includes:
[0012] Generate a resource feature head vector from the user access resource data through the multi-head self-attention mechanism;
[0013] Generate a target weight from the resource feature head vector through the attention mechanism;
[0014] Generate the educational resource interest features from the target weight through the pooling layer.
[0015] Further, the educational resource model further includes an auxiliary loss term for learning the feature representation of the user access resource data, and the auxiliary loss term is constructed based on a logarithmic likelihood form and an activation function.
[0016] In the above solution, the multi-head self-attention mechanism is beneficial to capturing more and larger-scale correlation features, increasing the expressive power of the model. Using the multi-head self-attention mechanism to model various potential interests of users, learning the potential relationships between items at the same time, and making interest mining diversified.
[0017] Further, the user interest model includes a gated recurrent unit module and an attention module. The process of generating user interest features from user time-series data through the user interest model includes:
[0018] Input the user time-series data into the gated recurrent unit module. The gated recurrent unit module outputs the user interest features as the hidden state, and controls the update of the hidden state of the gated recurrent unit module through the attention module.
[0019] Further, the hidden state includes a previous hidden state, a current hidden state, and a transmitted hidden state. The process of controlling the update of the hidden state of the gated recurrent unit module through the attention module includes:
[0020] Generate an attention weight from the current hidden state and the target vector through the attention module;
[0021] Generate the transmitted hidden state based on the attention weights, the previous hidden state, and the current hidden state, and transmit the transmitted hidden state to the gated recurrent unit module at the next time step.
[0022] Further, construct an objective function based on the cross-entropy loss term, and the objective function is used for the optimal prediction of the gated recurrent unit module.
[0023] In the above solution, by adaptively learning the weights of each behavior through the attention mechanism in the improved gated recurrent unit, the latest changes in expressing user interests are realized better.
[0024] Further, the process of determining the predicted interest items for the educational resource features and the user interest features through the item interest prediction model includes:
[0025] Construct a loss function based on the negative logarithmic function and the access probability of the predicted interest items, where the loss function is used for the prediction of the item interest prediction model;
[0026] Input the educational resource features, the user interest features, and the target item data into the item interest prediction model. The item interest prediction model outputs the access probability of the predicted interest items, and selects the predicted interest items from the target item data according to the access probability of the predicted interest items.
[0027] Further, the item interest prediction model includes a fully connected layer, at least two layers of Dice activation functions, and a Sigmoid activation function. The process of inputting the educational resource features, the user interest features, and the target item data into the item interest prediction model and the item interest prediction model outputting the access probability of the predicted interest items includes:
[0028] Pass the educational resource features, the user interest features, and the target item data through the fully connected layer, at least two layers of Dice activation functions, and the Sigmoid activation function in sequence to output the access probability of the predicted interest items.
[0029] In the above solution, by using the Dice activation function in the multi-layer perceptron, the expression ability of low-frequency features is strengthened, the evolution of interests is further perceived, and thus the influence of noise input into the multi-layer perceptron is reduced.
[0030] Further, the user access resource data includes user information data, user learning behavior data, and target item data;
[0031] The user time series data includes user historical behavior time series data, the user information data, and the target item data.
[0032] Furthermore, the educational resource database includes a cold data table and at least two hot data tables. The process of storing the data corresponding to the predicted interest items as hot data in the educational resource database includes:
[0033] Classify the data corresponding to the predicted interest items according to their access frequencies and their association degrees with the current item set in use, and store them separately in multiple hot data tables;
[0034] Regularly convert the cold data and the hot data through a thread pool.
[0035] Compared with the prior art, the beneficial effects of the present invention are as follows
[0036] 1. Through the multi-head self-attention mechanism, the improved gated recurrent unit and the multi-layer perceptron, the content of interest to users on the online education platform is accurately identified and stored and managed separately. As a result, when users access the content they are interested in, the platform needs to access and load a large amount of data from the database, resulting in a long waiting time and a poor exploration experience for users.
[0037] 2. The multi-head self-attention mechanism is conducive to capturing larger and more extensive correlation features, increasing the expression ability of the model. Using the multi-head self-attention mechanism to model the multiple potential interests of users and the potential relationships between educational resources at the same time makes interest mining diversified.
[0038] 3. In the improved gated recurrent unit, the weights of each behavior are adaptively learned through the attention mechanism, realizing a better expression of the latest changes in user interests.
[0039] 4. By using the Dice activation function in the multi-layer perceptron, the expression ability of low-frequency features is strengthened, further perceiving the evolution of interests, thereby reducing the influence of noise input to the multi-layer perceptron. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic flowchart of the method for managing data storage in an online education platform according to an embodiment of the present invention;
[0041] Figure 2 It is a schematic detailed flowchart of the model of the method for managing data storage in an online education platform according to an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of the model structure of the method for managing data storage in an online education platform according to an embodiment of the present invention;
[0043] Figure 4 It is a schematic structural diagram of an embodiment of the method for managing data storage in an online education platform according to an embodiment of the present invention;
[0044] Figure 5Another implementation structural diagram of the data storage management method for the online education platform according to the embodiments of the present invention. Detailed implementation manners
[0045] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] The preferred implementation manners of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0047] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.
[0048] In addition, it should be noted that in the description of the present invention, unless otherwise clearly defined and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0049] As Figures 1 to 5 shown, the present invention provides a data storage management method for an online education platform. Through the multi-head self-attention mechanism, the improved gated recurrent unit, and the multi-layer perceptron, the content of interest to users on the online education platform is accurately identified and stored and managed separately. As a result, when users access the content they are interested in, the platform needs to access and load data from a large database, resulting in a long waiting time and a poor exploration experience for users.
[0050] As Figures 1 to 5 shown, this embodiment proposes a data storage management method for an online education platform. The online education platform is provided with an education resource database for separating and storing cold data and hot data, including:
[0051] Generating education resource interest features from the user access resource data of the online education platform through an education resource model, where the education resource model is constructed based on the multi-head self-attention mechanism;
[0052] Generate user interest features from the user time-series data of the online education platform through a user interest model, where the user interest model is constructed based on an improved gated recurrent unit;
[0053] Determine the predicted interest items from the educational resource features and the user interest features through an item interest prediction model, where the item interest prediction model is established based on a multi-layer perceptron;
[0054] Store the data corresponding to the predicted interest items as hot data in the educational resource database.
[0055] It can be understood that the educational resource model introducing the multi-head self-attention mechanism models the multi-layer latent interests of users, and with the help of the auxiliary loss, it helps in the expression and learning of interest features. In the user interest model, user time-series data is added, that is, the user access historical data is encoded and input in time series, enhancing the intensity of user personalized interest expression.
[0056] Therefore, by accurately identifying the educational resource / item data that users are interested in and isolating it for processing as hot data, it helps the online education platform to make targeted configurations according to the type requirements of the items, such as allocating different priorities, thread numbers, etc., enabling the platform resources to be more reasonably allocated to data processing tasks of different access possibility types, and further optimizing and improving the overall performance of the online education platform.
[0057] It can be understood that as Figure 4 and 5 shown, the process of identifying the data corresponding to the predicted interest items can be used for the construction of the distinction between the hot data table and the cold data table in the initial stage, and can also be used for the conversion of data between the hot data table and the cold data table executed regularly.
[0058] Furthermore, the educational resource model includes a multi-head self-attention mechanism, an attention mechanism, and a pooling layer. The process of generating educational resource interest features from the user access resource data of the online education platform through the educational resource model includes:
[0059] Generate resource feature head vectors from the user access resource data through the multi-head self-attention mechanism;
[0060] Generate target weights from the resource feature head vectors through the attention mechanism;
[0061] Generate the educational resource interest features from the target weights through the pooling layer.
[0062] It is understandable that the multi-head self-attention mechanism is beneficial to capturing correlation features in a larger and more extensive range, increasing the expressive power of the model, and has been proficient in various training tasks. Here, the multi-head self-attention mechanism is used to model various potential interests of users.
[0063] Specifically, the multi-head self-attention mechanism is as follows:
[0064]
[0065] In the formula, head i represents the i-th resource feature head vector, and W Q,i , W K,i , W V,i are the projection matrices of the i-th heads of Q (Query), K (Key), and V (Value) respectively, and they all belong to R, that is, W Q,i , W K,i , W V,i ∈R, R = Multihead(Q, K, V), and x1 represents the input user access resource data.
[0066] Furthermore, the educational resource model further includes an auxiliary loss term for learning the feature representation of the user access resource data. The auxiliary loss term is constructed based on the logarithmic likelihood form and activation function, specifically as follows:
[0067]
[0068] In the formula, Loss multihead is the auxiliary loss term, N is the number of samples in the training dataset, σ is the Sigmoid activation function, <,> is the inner product used to represent the similarity between two variables, and r j , e′ j+1 represent the educational resource interest features learned by the model at the j-th moment and the original educational resource vector representation input to the model at the j-th moment, that is, the positive example sample respectively. Therefore, the auxiliary loss term can help learn a feature representation that is more in line with the discriminative features between different resources.
[0069] It is understandable that feature interaction through the item interest prediction model refers to a process of generating new features by combining two or more features. Feature interaction mainly performs feature cross-multiplication through the inner product to generate new features. This method not only reflects the user's interest evaluation of the product combination but also can uncover the potential relationships hidden between features, thereby improving the prediction ability of the model.
[0070] In the above solution, the multi-head self-attention mechanism is beneficial to capturing correlation features in a larger and wider range, increasing the expressive ability of the model. The multi-head self-attention mechanism is used to model various potential interests of users, and the potential relationships between items are learned at the same time to diversify interest mining.
[0071] Furthermore, the user interest model includes a gated recurrent unit module and an attention module. The process of generating user interest features from user time-series data through the user interest model includes:
[0072] Input the user time-series data into the gated recurrent unit module. The gated recurrent unit module outputs the user interest features as hidden states, and the attention module controls the update of the hidden states by the gated recurrent unit module.
[0073] Furthermore, the hidden states include a previous hidden state, a current hidden state, and a transmitted hidden state. The process of controlling the update of the hidden states by the gated recurrent unit module through the attention module includes:
[0074] Generate attention weights by passing the current hidden state and the target vector through the attention module;
[0075] Generate the transmitted hidden state according to the attention weights, the previous hidden state, and the current hidden state, and transmit the transmitted hidden state to the gated recurrent unit module at the next time step.
[0076] Specifically, the gated recurrent unit module (GRU) of the user interest model is:
[0077]
[0078] z t =σ(W z e t +U z h t-1 +b z )
[0079]
[0080] r t =σ(W r e t +U r h t-1 +b r )
[0081] In the formula, h t 、h t-1 、h t′ represent the activation state at the t-th moment, the activation state at the (t - 1)-th moment, and the updated value of the activation state at the t-th moment, respectively, z t , e t , r t represent the update gate at the t-th moment, the input representation at the t-th moment, and the reset gate at the t-th moment, respectively. σ is the Sigmoid activation function, W z , U z , W h , U h , W r , U r are all learnable weight vectors, b z , b h , b r are all learnable bias terms.
[0082] Furthermore, construct an objective function based on the cross-entropy loss term, and the objective function is used for the optimization prediction of the gated recurrent unit module. Specifically:
[0083]
[0084] In the formula, Attention t represents the attention module, and softmax(e t ) represents that the softmax activation function controls whether to update the input representation e t at the t-th moment. h t represent the activation state at the t-th moment respectively, and W represents the weight vector.
[0085] In the above solution, by using the attention mechanism to adaptively learn the weights of each behavior in the improved gated recurrent unit, the latest changes in user interests are better expressed.
[0086] Furthermore, the process of determining the predicted interest items by the educational resource features and the user interest features through the item interest prediction model includes:
[0087] Construct a loss function based on the negative logarithm function and the predicted access probability of the interest items, where the loss function is used for the prediction of the item interest prediction model;
[0088] Input the educational resource features, the user interest features and the target item data into the item interest prediction model. The item interest prediction model outputs the predicted access probability of the interest items, and selects the predicted interest items from the target item data according to the predicted access probability of the interest items.
[0089] Specifically, the loss function is:
[0090]
[0091] In the formula, Loss target is the loss function, x represents the input educational resource features and user interest features, y represents the user's accessed interest items in the training set, 1 - y represents the user's non - accessed interest items in the training set, N is the number of samples in the dataset, and O(x) represents the predicted interest item access probability output by the model. Therefore, it increases the prediction accuracy of the predicted interest item access probability and the accessed interest items. It enables the representation learned by the model to be more similar to the accessed interest items and less similar to the non - accessed interest items.
[0092] Furthermore, the item interest prediction model includes a fully - connected layer, at least two layers of Dice activation functions, and a Sigmoid activation function. The process of inputting the educational resource features, the user interest features, and the target item data into the item interest prediction model and the model outputting the predicted interest item access probability includes:
[0093] Passing the educational resource features, the user interest features, and the target item data through the fully - connected layer, at least two layers of Dice activation functions, and the Sigmoid activation function in sequence to output the predicted interest item access probability.
[0094] Preferably, the item interest prediction model is sequentially provided with two fully - connected layers, two Dice activation functions, and one Sigmoid activation function. Therefore, the training stability of the hidden layer is optimized through the Dice activation function, complex features are captured by using the double - layer fully - connected layer, and finally the predicted interest item access probability is constructed probabilistically with Sigmoid, achieving a balance among gradient control, non - linear modeling, and prediction accuracy.
[0095] In the above - mentioned solution, the Dice activation function is used through the multi - layer perceptron to strengthen the expression ability of low - frequency features, further perceive the evolution of interest, and thus reduce the influence of noise input into the multi - layer perceptron.
[0096] Furthermore, the user - accessed resource data includes user information data, user learning behavior data, and target item data;
[0097] The user time - series data includes user historical behavior time - series data, the user information data, and the target item data.
[0098] Specifically, the user - accessed resource data is expressed as: x1 ∈ (X a , X b , X c ), where x1 represents the user - accessed resource data, and X a , X b , X crespectively represent user information data, user learning behavior data, and target project data.
[0099] Specifically, the user time series data is represented as: x ∈ (X d , X a , X c ), where x represents the user time series data, and X d , X a , X c respectively represent the user historical behavior time series data, the user information data, and the target project data.
[0100] Furthermore, the educational resource database includes a cold data table and at least two hot data tables. The process of storing the data corresponding to the predicted interest items as hot data in the educational resource database includes:
[0101] Classify the data corresponding to the predicted interest items according to their access frequencies and their association degrees with the current used item set, and store them in multiple said hot data tables respectively;
[0102] Regularly convert the cold data and the hot data through a thread pool.
[0103] Specifically, the access frequency of the data corresponding to the predicted interest items is calculated using time decay weighting to highlight the recent popularity:
[0104]
[0105] In the formula, CF(k) represents the access frequency of the data corresponding to the predicted interest item, c t (k) represents the number of accesses of the predicted interest item k at time t, T now , T first represent the current time of accessing the data corresponding to the predicted interest item and the first time of accessing the data corresponding to the predicted interest item, and α represents an adjustment coefficient, preferably 0.3. Therefore, the more the number of accesses or the closer the access time interval, the larger the value of the access frequency.
[0106] Specifically, the association degree between the data corresponding to the predicted interest items and the current used item set is calculated using cosine similarity:
[0107]
[0108] In the formula, CR(k) represents the association degree, λ represents a weight, preferably 0.5, co(k, U) represents the co-occurrence times of the current predicted interest item k and the current used item set U, max k co(k, U) represents the maximum value of the co-occurrence times of all predicted interest items and the current used item set U, cos(vk , v U ) represents the embedding vector v corresponding to the predicted interesting item k k and the average embedding vector v of the currently used item set U U of the cosine similarity.
[0109] Normalize the correlation degree and the access frequency respectively by the maximum value and the minimum value, and calculate the comprehensive heat value:
[0110]
[0111] In the formula, S(k) represents the comprehensive heat value, CF(k) represents the access frequency of the data corresponding to the predicted interesting item, CR(k) represents the correlation degree, and β represents the exponential weight, preferably 0.4.
[0112] According to the comprehensive heat value, classify and store the data corresponding to the predicted interesting item in the hot data table:
[0113]
[0114] In the formula, Bucket l represents the hot data table where the data corresponding to the predicted interesting item is stored, l represents the label number of the hot data table level, L represents the total number of label numbers of the hot data table level, S(k) represents the comprehensive heat value of the predicted interesting item k, and S max (k) represents the maximum value of the comprehensive heat value of the predicted interesting item k. Therefore, the classification storage of the data corresponding to the predicted interesting item can be realized. Preferably, L is preferably 3, that is, the hot data is divided into the hot data table 1 with high heat, the hot data table 2 with medium heat, and the hot data table n with low heat through the above formula.
[0115] The proposed hot and cold data separation scheme aims to improve the data access efficiency of the system. By placing a small amount of data frequently accessed by users in the hot table and storing a large amount of historical data less frequently accessed in the cold table, users can search in the hot table first when accessing data. Since the amount of data in the hot table is small, this method can greatly improve the data access efficiency.
[0116] As Figure 4 and 5 shown, the process of an online education platform loading hot data or cold data is as follows:
[0117] Data search: When the search function is called and executed, the system first searches the hot data table. Since the amount of data in the hot data table has been maintained at a certain scale, the search process is relatively fast. As Figure 4As shown, during the data search process, the database first accesses the hot data table 1, then sequentially accesses the hot data table 2 until the data is found in the hot data table n or the cold data table. The hot data tables 1 to n are serially connected. Therefore, more efficient storage mechanisms and access mechanisms can be adopted for the hot data tables to improve the overall performance of the database. For example, Figure 5 As shown, the database performs multi-threaded parallel access to the hot data tables 1 to n, which can further improve the access efficiency.
[0118] Data transfer: Use a thread pool to write the data transfer function into a single thread. During execution, call through the Executor and put it into the thread pool as a task. The transfer function includes: passing in the data that needs to change the operation time, copying the data to the hot data table, updating its operation time, and at the same time deleting the corresponding data in the cold data table. For example, Figure 4 As shown, transfer the cold data table and the hot data table 2 through the thread pool, first write to the hot data table 2 and then perform the transfer between the hot data tables. Figure 5 As shown, transfer the cold data table and the hot data table group as a whole through the thread pool, and perform dynamic allocation when writing to the hot data table group.
[0119] Gradually scan the hot data table and copy it regularly: Use the data volume of the check data table and all data Scheduler timed tasks to ensure the freshness of the data at the latest update time in the early morning.
[0120] For example, Figure 4 and 5 As shown, Thread Pool 1 and Thread Pool 2 can dynamically allocate thread resources according to the actual load.
[0121] To simulate and reproduce the real environment, data simulation and stress testing are performed on the database tables. The test plan is based on the multiple-choice question table in the online education management platform for performance testing. The test process is divided into the following 3 steps: Select the test table and execute the test statement: Use the statement for testing, gradually increase the data volume and then adjust it to the ten million level. Gradually increase the data volume from the one hundred thousand level to the ten million level and observe the change in query performance.
[0122] In this embodiment, through the multi-head self-attention mechanism, the improved gated recurrent unit, and the multi-layer perceptron, the content that users are interested in on the online education platform is accurately identified and stored and managed separately. As a result, when users access the content they are interested in, the platform needs to access and load data from a large database with a large amount of data, resulting in a long waiting time and a poor exploration experience for users. The multi-head self-attention mechanism is beneficial to capturing larger and more extensive correlation features, increasing the expressive ability of the model. Using the multi-head self-attention mechanism to model various potential interests of users and the potential relationships between educational resources at the same time makes interest mining diversified. In the improved gated recurrent unit, the weights of each behavior are adaptively learned through the attention mechanism, achieving a better expression of the latest changes in user interests. By using the Dice activation function in the multi-layer perceptron, the expressive ability of low-frequency features is strengthened, further perceiving the evolution of interests, thereby reducing the influence of noise input to the multi-layer perceptron.
[0123] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0124] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A data storage management method for an online education platform, characterized in that: The online education platform sets up an education resource database for separate storage of cold data and hot data, including: Generate educational resource interest features from the user access resource data of the online education platform through an educational resource model, wherein the educational resource model is constructed based on a multi-head self-attention mechanism; Generate user interest features from the user time series data of the online education platform through a user interest model, wherein the user interest model is constructed based on an improved gated recurrent unit; Determine predicted interest items by using the educational resource features and the user interest features through an item interest prediction model, wherein the item interest prediction model is established based on a multi-layer perceptron; The data corresponding to the predicted interest items are stored as hot data in the educational resource database.
2. The online education platform data storage management method according to claim 1, characterized in that: The educational resource model includes a multi-head self-attention mechanism, an attention mechanism and a pooling layer. The process of generating educational resource interest features from the user access resource data of the online education platform through the educational resource model includes: Generate a resource feature head vector from the user access resource data through the multi-head self-attention mechanism; Generate a target weight by using the attention mechanism for the resource feature head vector; The target weight is passed through the pooling layer to generate the educational resource interest feature.
3. The online education platform data storage management method according to claim 2, characterized in that: The educational resource model also includes an auxiliary loss term for learning the feature representation of the user access resource data, and the auxiliary loss term is constructed based on a log-likelihood form and an activation function.
4. The online education platform data storage management method according to claim 1, characterized in that: The user interest model includes a gated recurrent unit module and an attention module. The process of generating user interest features from user time series data through the user interest model includes: The user time series data is input into the gated recurrent unit module, the gated recurrent unit module outputs the user interest feature as a hidden state, and the gated recurrent unit module is controlled to update the hidden state through the attention module.
5. The online education platform data storage management method according to claim 4, characterized in that: The hidden state includes a previous hidden state, a current hidden state and a transferred hidden state, and the process of controlling the gated recurrent unit module to update the hidden state through the attention module includes: The current hidden state and the target vector are passed through the attention module to generate an attention weight; The transfer hidden state is generated according to the attention weight, the previous hidden state and the current hidden state, and the transfer hidden state is passed to the gated recurrent unit module of the next time step.
6. The online education platform data storage management method according to claim 4, characterized in that: An objective function based on a cross entropy loss term is constructed, and the objective function is used for optimizing prediction of the gated recurrent unit module.
7. The online education platform data storage management method according to claim 1, characterized in that: The process of determining the predicted interest items by using the item interest prediction model based on the educational resource characteristics and user interest characteristics includes: Constructing a loss function based on a negative logarithmic function and a predicted probability of visiting an item of interest, wherein the loss function is used for prediction of the item interest prediction model; The educational resource characteristics, the user interest characteristics and the target project data are input into the project interest prediction model, the project interest prediction model outputs the predicted interest project access probability, and selects the predicted interest project from the target project data based on the predicted interest project access probability.
8. The online education platform data storage management method according to claim 7, characterized in that: The project interest prediction model includes a fully connected layer, at least two layers of Dice activation function and Sigmoid activation function, and the educational resource features, the user interest features and the target project data are input into the project interest prediction model. The project interest prediction model outputs the predicted access probability of the interest project, and the process includes: The educational resource features, the user interest features and the target project data are sequentially passed through the fully connected layer, at least two layers of Dice activation function and Sigmoid activation function to output the predicted probability of accessing the interest project.
9. The online education platform data storage management method according to any one of claims 1 to 8, characterized in that: The user access resource data includes user information data, user learning behavior data and target project data; The user time series data includes user historical behavior time series data, the user information data and the target project data.
10. The online education platform data storage management method according to any one of claims 1 to 8, characterized in that: The educational resource database includes a cold data table and at least two hot data tables, and the process of storing the data corresponding to the predicted interest items as hot data in the educational resource database includes: Classifying the data corresponding to the predicted interest items according to their access frequencies and their relevance to the currently used item set, and storing them in a plurality of the hot data tables respectively; The cold data and the hot data are converted periodically through a thread pool.
Citation Information
Patent Citations
Intelligent education storage system based on cloud computing
CN113535715A
Education resource recommendation method and system based on artificial intelligence
CN116628339A
Course recommendation method and system fusing user interest change
CN117609615A
Financial product recommendation method and device, equipment and storage medium
CN117726448A
Data storage method and system based on legal knowledge service platform and storage medium
CN118964496A