Teaching Resource Recommendation Method Based on User Behavior and Course Duration

By combining SIM model and user behavior characteristics in the course teaching resource recommendation, we can handle the relationship between long and short-term behavior and course duration, and solve the problem that the impact of course duration in the existing technology is not considered, and achieve more accurate and personalized course recommendations.

CN119691284BActive Publication Date: 2025-06-20GUOKAI ONLINE EDUCATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510206788.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-20
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing technology fails to effectively consider the impact of course duration on user learning behavior in the recommendation of course teaching resources, resulting in limited accuracy of recommendations.

Method used

The click-through rate prediction model based on the SIM model is adopted, combining the long-term and short-term behavior characteristics of the user, and processing behavior characteristics of different lengths through weighted summing, the user's relative behavior characteristics of the course are generated, and the course is predicted with the embedded characteristics of the course, and the course is recommended.

Benefits of technology

It improves the accuracy and personalization of course recommendations, and improves the effectiveness of the recommendation system by fitting individual learning preferences and reflecting the complex rules of learning behavior and course duration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691284B_ABST
    Figure CN119691284B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a teaching resource recommendation method based on user behavior and course duration. The method includes: obtaining a specific user to be recommended and multiple courses; for each course, respectively perform the following operations: generating long-term behavior characteristics of the user by using the SIM model; extracting a first duration with the largest value from the differences between the occurrence times of multiple behaviors input to the accurate search unit of the SIM model and the current time; according to the relationship among the current course duration, the first duration, and a second duration covered by the short-term behavior characteristics of the user, performing weighted summation on the long-term behavior characteristics and the short-term behavior characteristics to obtain relative behavior characteristics of the user for the current course; predicting the click-through rate of the user for the current course by using the relative behavior characteristics and the embedding characteristics of the current course; and recommending multiple courses to the user according to the click-through rates of the courses. This embodiment improves the accuracy of course recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical fields of artificial intelligence and intelligent education, and in particular to a teaching resource recommendation method based on user behavior and course duration. Background Art

[0002] Teaching resources refer to various materials provided for the effective development of teaching, including textbooks, cases, videos, pictures, courseware, etc. In online learning, how to recommend suitable courses to users from a vast amount of teaching resources is an important aspect of improving learning efficiency and giving full play to the advantages of online learning.

[0003] In the prior art, patent application CN116561410A provides a method for recommending course teaching resources, and patent CN117726082A provides a method for recommending teaching resources. Neither of them takes into account the influence of course duration on users' learning behaviors, and the accuracy of the recommendation is limited. Summary of the Invention

[0004] Embodiments of the present invention provide a teaching resource recommendation method based on user behavior and course duration to solve the above technical problems.

[0005] In a first aspect, embodiments of the present invention provide a teaching resource recommendation method based on user behavior and course duration, including:

[0006] S110. Obtain a specific user to be recommended and multiple courses;

[0007] S120. Perform operations S1-1 to S1-3 for each course respectively:

[0008] S1-1. Generate long-term behavior characteristics of the user by using the SIM model; extract the first duration with the largest value from the differences between the occurrence times of multiple behaviors input to the accurate search unit of the SIM model and the current time;

[0009] S1-2. According to the relationship among the current course duration, the first duration, and the second duration covered by the short-term behavior characteristics of the user, perform weighted summation on the long-term behavior characteristics and the short-term behavior characteristics to obtain relative behavior characteristics of the user for the current course;

[0010] S1-3. Predict the click-through rate of the user for the current course by using the relative behavior characteristics and the embedding characteristics of the current course;

[0011] S130. Recommend multiple courses to the user according to the click-through rates of each course.

[0012] In a second aspect, embodiments of the present invention provide an electronic device, and the electronic device includes:

[0013] One or more processors;

[0014] A memory for storing one or more programs,

[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the teaching resource recommendation method based on user behavior and course duration described in any embodiment.

[0016] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the teaching resource recommendation method based on user behavior and course duration described in any embodiment.

[0017] In summary, the embodiment of the present invention provides a teaching resource recommendation method based on user behavior and course duration. Combining with the SIM click-through rate prediction model, it analyzes the student's past click behavior on different types of resources, predicts the probability of their current preference for various resources, and fits the individual learning preferences. At the same time, combining with the rules of video learning behavior and time, different weights are assigned to long-term user behavior and short-term user behavior for different course durations to reflect the complex rules between learning behavior and course duration, further improving the accuracy and personalization of course recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a flowchart of a teaching resource recommendation method based on user behavior and course duration provided by an embodiment of the present invention;

[0020] Figure 2 is a structural diagram of a click-through rate prediction model provided by an embodiment of the present invention;

[0021] Figure 3 is an architecture diagram of a teaching resource recommendation system provided by an embodiment of the present invention;

[0022] Figure 4 is a structural diagram of a two-tower model provided by an embodiment of the present invention;

[0023] Figure 5 is a structural diagram of a two-tower model after feedback provided by an embodiment of the present invention;

[0024] Figure 6 It is the structural diagram of another double - tower model after feedback provided by the embodiments of the present invention;

[0025] Figure 7 It is the schematic structural diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.

[0027] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0028] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0029] Figure 1 It is the flowchart of a teaching resource recommendation method based on user behavior and course duration provided by the embodiments of the present invention. This method is applicable to the situation of online learning and is executed by an electronic device. As Figure 1 shown, the method specifically includes the following steps:

[0030] S110. Obtain a specific user and multiple courses to be recommended.

[0031] The user and courses here can be users in an online learning platform and online courses. This embodiment will take a learning user as an example to illustrate the recommendation process. When recommending for multiple users, this method can be executed for each user respectively.

[0032] S120. Use a click-through rate prediction model including the SIM model to predict the click-through rate of each course for the specific user.

[0033] The click-through rate here can be the probability of course learning, referring to the probability that a specific user will learn each course within a future period of time. Specifically, in this embodiment, the click-through rate prediction model shown as follows is used for prediction. This model includes a course embedding model, a SIM (Sequential Interaction Model) model, a short-term embedding model, a weighted summation layer, and a prediction layer. Figure 2 Among them, the course embedding model is used to generate the embedding features of the course. Optionally, the basic information of the course resources (such as course name, instructor, course duration), attribute information (subject label, difficulty level, course format), and context features (such as learning period, environment where the current learning progress is located), etc. can be extracted and input into the course embedding model to obtain the embedding representation of the course.

[0034] The SIM model is used to generate the long-term behavior features of the user, and is used for course recommendation based on the user's historical behavior, which is particularly suitable for capturing the long-term behavior characteristics of the user. The user behavior in this embodiment is the learning behavior of the user, including the user. The basic principle of SIM is to model the user's behavior sequence, mine the temporal and internal correlations between behaviors, and estimate the click-through rate according to the user's historical behavior sequence to generate personalized recommendations.

[0035] Further, combined with, the input of the SIM model is the long-term behavior sequence of the user. Optionally, the courses learned by the user within a relatively long period of time with the current time as the end point are extracted, and the embedding features of each course are arranged in chronological order to obtain the long-term behavior sequence. It should be noted that the time in this embodiment represents a time point.

[0036] Further, combined with Figure 2 , the input of the SIM model is the long-term behavior sequence of the user. Optionally, the courses learned by the user within a relatively long period of time with the current time as the end point are extracted, and the embedding features of each course are arranged in chronological order to obtain the long-term behavior sequence. It should be noted that the time in this embodiment represents a time point.

[0037] The SIM model includes a GSU (General Search Unit, rough search unit) and an ESU (Exact Search Unit, precise search unit). After the user's long-term behavior sequence is input into the GSU, the GSU extracts the top K courses with the strongest correlation with candidate course A from the long-term behavior sequence within linear time, and arranges their embedding vectors into a new subsequence, where K is much smaller than the length M of the original sequence. At the same time, the difference between the time when the user learned the course and the current time (i.e., the time difference, positive value) is added to the embedding vector of each selected course. The added subsequence is input into the ESU, and the ESU outputs the long-term behavior feature of the user through an attention operation, which represents the long-term interest of the user. Since the courses filtered out in the GSU have a very weak correlation with course A, the corresponding components in the attention operation are very small and hardly affect the final calculation result. Therefore, the SIM model can extract the long-term behavior feature with a very small computational cost.

[0038] The short-term behavior embedding model is used to generate the short-term behavior feature of the user. This model can adopt structures such as DIN and DIEN, and the input is the short-term behavior sequence of the user. Optionally, the courses learned by the user within a short period of time ending at the current time are extracted, and the embedding vectors of each course are arranged in chronological order to obtain the short-term behavior sequence. After being processed by the short-term embedding model, the short-term behavior sequence outputs the short-term behavior feature of the user, which represents the short-term interest of the user.

[0039] Based on the above three basic models, the following operations are respectively performed on each course in the click-through rate prediction in this embodiment:

[0040] S1-1. Use the SIM model to generate the long-term behavior feature of the user, and extract the maximum duration value from the differences between the occurrence times of multiple behaviors input into the ESU of the SIM model and the current time. Combine Figure 2 , for a specific user and the current candidate course A, input the long-term behavior sequence of the user into the SIM model to obtain the long-term behavior feature of the user. At the same time, extract the duration from the learning time of each course to the current time from the top-K subsequences input into the ESU, and extract a maximum value as the duration corresponding to the long-term behavior feature.

[0041] S1-2. Compare the current course duration, the maximum duration extracted in S1-1, and the duration covered by the short-term behavior feature of the user, and perform weighted summation on the long-term behavior feature and the short-term behavior feature according to the relationship among the three durations to obtain the relative behavior feature of the user for the current course.

[0042] This operation corresponds to Figure 2The weighted fusion layer in it. Specifically, since the learning frequency and interest show a "decaying - stable" pattern in the learning of a series of courses, the click - through rate of the courses may show a decaying trend over time. Based on the above rules, in this embodiment, more attention is paid to long - term behavior characteristics when recommending long courses, and more attention is paid to short - term behavior characteristics when recommending short courses. To achieve this effect, in this embodiment, by controlling the dimensions of the SIM model and the short - term embedding model, the two behavior characteristic dimensions are made the same, and the two behavior characteristics are weighted and fused according to the duration of the candidate course, which is used to balance the influence of long - term behavior and short - term behavior when matching with the candidate course later.

[0043] Optionally, the weights of the two behavior characteristics can be set according to experience. For example, when the course duration is greater than 5 weeks, the weights of the long - term behavior characteristic and the short - term behavior characteristic are respectively determined as and , where ; when the course is less than 5 weeks, the weights of the long - term behavior characteristic and the short - term behavior characteristic are respectively determined as and , where, .

[0044] In addition, this embodiment also provides a method to determine different weights for each course according to the ratio of three durations, realizing the differential setting of weights with the courses. For the convenience of distinction and description, in this embodiment, the duration corresponding to the long - term behavior characteristic extracted by S1 - 1 is called the first duration, and the duration covered by the short - term behavior sequence is called the second duration. Usually, the first duration is greater than the second duration, but in some extreme cases, the second duration is larger. To take all cases into account, in this embodiment, the longer one of the first duration and the second duration is denoted as , the shorter one is denoted as , and the current course duration is denoted as t. Then the method for setting the weights is:

[0045] If the current course duration is less than the shorter duration , then the weight of the behavior characteristic corresponding to the shorter duration is set to 1, and the weight of the other behavior characteristic is set to 0. That is, when the current course duration is very short, it is completely driven by short - term interest.

[0046] If the current course duration is greater than the longer duration , then the weight of the behavior characteristic corresponding to the shorter duration is set to 0, and the weight of the other behavior characteristic is set to 1. That is, when the current course duration is very long, it is completely driven by long - term interest.

[0047] If the current course duration t is between the two durations and If it is between, set the weight of the behavior feature corresponding to the shorter duration to and set the weight of the behavior feature corresponding to the longer duration to . That is, the current course duration is driven by a combination of short-term and long-term interests. According to the coverage or proportion of the two behavior features by the course duration, the influence degrees of the two are calculated.

[0048] Regardless of which method is used, after weighted summation using the above weights, a new behavior feature can be obtained. This behavior feature is related to both the current course and the long-term and short-term behaviors of the user, reflecting the complex law between the user's long-term and short-term behaviors and the course length. In this embodiment, this feature is called the relative behavior feature of the current user with respect to the current course.

[0049] S1-3. Use the relative behavior feature and the embedding feature of the current course to predict the click-through rate of the user for the current course. Combining Figure 2 , after weighted fusion of the relative feature and the embedding feature of the current course, jointly input Figure 2 to the prediction layer in . This layer reduces the dimension through linear or nonlinear operations and finally outputs the probability that the specific user will learn the current course in the future.

[0050] S130. Recommend multiple courses to the user according to the click-through rates of each course.

[0051] After performing the above operations for all courses, the probability that the specific user will learn each course in the future can be obtained; several courses with the highest probabilities are recommended to the specific user.

[0052] After performing the operations of S110-S130 for each user, the course recommendation for each user can be completed.

[0053] The following describes an application scenario of the above method. Figure 3 is an architecture diagram of a teaching resource recommendation system provided by an embodiment of the present invention. As shown in Figure 1, the system includes:

[0054] Offline layer: Use a large model to be responsible for preprocessing a large amount of teaching resource data, including cleaning, annotation, classification, etc., to build a teaching resource knowledge base; at the same time, complete the model training task, train relevant models such as multi-channel recall and click-through rate prediction based on historical data, and generate various data reports for subsequent analysis and optimization. In this link, incorporate the SIM model into the training system, use historical teaching resource browsing, usage and other behavior data to train the SIM model to capture the long-term and short-term learning resource preference patterns of users, and establish a click-through rate prediction model to provide a basis for subsequent recommendations.

[0055] Near-line layer: Focus on real-time feature processing. On the one hand, it captures users' learning behavior features in real time, such as the current browsing resource type, learning duration, operation records, etc. On the other hand, it monitors the basic attributes of newly launched teaching resources in real time. On this basis, a multi-channel recall operation is carried out. Based on dual-tower models such as DSSM, thousands of teaching resources that users may be interested in are initially screened from the content pool and sent to the rough ranking layer. At the same time, it also undertakes the rough ranking task, using the dual-tower model to improve the combination efficiency of multi-channel recall and relieve the subsequent fine ranking pressure.

[0056] Online layer: Mainly processes users' real-time requests. It receives hundreds of resources after rough ranking from the near-line layer, uses the highly accurate SIM model in the fine ranking layer to estimate the user click-through rate, screens out dozens of resources that best meet the user's current needs, and then re-ranks them using the MMR algorithm in the re-ranking layer to balance accuracy, diversity, and importance before delivering them to the user to respond to the user's resource acquisition needs.

[0057] The teaching resource recommendation method based on user behavior and course duration provided in this embodiment can be applied to the fine ranking layer in the above system. Combining Figure 1 and Figure 2 , the teaching resource recommendation method of this recommendation system sequentially executes the following steps:

[0058] S10. The multi-channel recall layer uses multiple recall branches to provide a recall course set.

[0059] Optionally, the multi-channel recall strategies include:

[0060] Content-based recall strategy: Popular strategy recall, that is, sorting according to the popularity of teaching resources being browsed, downloaded, and used as a whole, and giving priority to pushing popular courses, etc.; Attribute strategy recall, accurately positioning according to attributes such as subject, knowledge point, and resource format; New upload strategy recall, timely pushing newly uploaded high-quality resources to learners who are concerned about the corresponding field.

[0061] Collaborative filtering recall: A strategy based on similar user behavior patterns, finding user groups with similar learning behaviors and preferences, and recommending resources that the target user is interested in; Knowledge-based association strategy, based on the relevance between resources, if a user browses a certain course, recommend the supporting exercises and analysis videos.

[0062] Embedding vector recall: In this branch, dual-tower models such as DSSM are introduced. The user features and teaching resource features are respectively constructed into independent tower structures to decouple their calculations, reduce the calculation complexity. At the same time, the dot product calculation requires less computing power to ensure real-time performance, and the user vector is generated in real time and the dot product calculation is completed.

[0063] Taking DSSM as an example, combining Figure 4, the two-tower model consists of a user tower on the left and a course tower on the right. The user tower incorporates basic user information (such as age, gender, location), statistical attributes of the affiliated learning group (major, subject classification), and the user's past learning behavior sequence (course resources browsed, learned, and favorited); the course tower covers basic course resource information (course name, instructor, course duration), attribute information (subject tags, difficulty level, course format), and context features (such as learning time period, environment where the current learning progress is located) can be placed in the user tower as appropriate. The matching layer uses inner product or Cosine similarity to measure the matching degree between the user and the course resources.

[0064] S20. The rough ranking layer uses another two-tower model to match each recalled course with the user, obtains multiple roughly ranked courses, and provides them to the fine ranking layer.

[0065] Rough ranking is the first checkpoint in the recommendation system. Its main purpose is to quickly screen out a batch of relatively potential candidates from a large number of candidates to reduce the pressure of subsequent calculations. In the rough ranking stage, methods based on rules or simple models are often used. The rough ranking layer can also use a two-tower model, which separately constructs independent tower structures for user features and teaching resource features, decouples the two calculations, reduces the computational complexity, and at the same time, the dot product calculation requires less computing power to ensure real-time performance, generates user vectors in real time, and completes the dot product calculation.

[0066] Finally, the rough ranking layer uses a fusion strategy combined with algorithm optimization to select hundreds of resources from thousands of recalled resources to enter the fine ranking layer, solve the problem of the combination efficiency of multi-channel recall, and avoid the performance bottleneck of fine ranking.

[0067] S30. The fine ranking layer takes multiple roughly ranked courses as the multiple courses to be recommended in S110, and executes the methods of S110 - S130 above to obtain multiple finely ranked courses.

[0068] Fine ranking means that on the basis of rough ranking, a more detailed and accurate ranking is performed on the selected candidates. In this stage, complex deep learning models or ensemble learning models are often used to pursue higher recommendation accuracy. Combining Figure 3 , in this embodiment, according to the click-through rate prediction index, 10 high-quality teaching resources that best match the user's needs and are most likely to be clicked and used are strictly selected from 300 resources sent by the rough ranking and sent to the re-ranking layer. Through the methods of S110 - S130, it can be ensured that the screening process fully considers the long-term and short-term dynamic changes of the user's interest behavior, and avoids the disconnection between the recommended resources and the user's current learning process. It should be noted that Figure 3 the quantities of each course in are only for illustration and are not absolutely limited.

[0069] S40. The re-ranking layer executes the MMR algorithm to re-rank the multiple finely ranked courses obtained in S130 to obtain the final recommended courses.

[0070] The rearrangement comprehensively considers more business rules, real-time user feedback, diversity and other factors, and adjusts and optimizes the results after fine ranking. In this embodiment, the MMR (Maximal Marginal Relevance) algorithm is mainly used to balance the relevance and diversity weights, which not only ensures that the pushed resources are closely related to the user's learning needs, but also breaks the information cocoon and improves the long-term use experience. For example, for students who are continuously learning mathematics, while pushing core knowledge point materials, interesting mathematics cases that expand thinking are also matched.

[0071] Further, in order to make full use of the advantages of the SIM model and the relativity feature in capturing the long-term and short-term interests of users, after the fine ranking layer in this embodiment executes S130, it can also feed back the behavioral characteristics of each user generated in the SIM model to the two-tower model in the multi-channel recall layer and / or the rough ranking layer, and fine-tune the two-tower model after feedback. Taking the two-tower model in the multi-channel recall layer as an example below, the method of feature feedback and model fine-tuning will be described. The two-tower model in the rough ranking layer is the same.

[0072] Specifically, regarding the feedback method, according to the different feedback levels, this embodiment provides two optional implementation manners:

[0073] The first optional implementation manner is to feed back the relative behavioral characteristics of each user for each course to the last layer of the two-tower model. Combining Figure 5 , the feedback process includes the following steps:

[0074] Step 1: The fine ranking layer provides the relative behavioral characteristics of each user for each course to the multi-channel recall layer, and the multi-channel recall layer adds an MLP (Multilayer Perceptron) between the last layer of the user tower and the matching layer in the two-tower model.

[0075] Step 2: When the two-tower model runs, the relative behavioral characteristics of the current user for the current course are concatenated with the embedding characteristics of the last layer of the user tower (i.e., the user embedding characteristics in Figure 5 ). The concatenated characteristics are input into the multi-layer perceptron for dimension transformation, and the output characteristics of the multi-layer perceptron are used to be jointly input into the matching layer with the embedding characteristics of the last layer of the course tower (i.e., the course embedding characteristics) for matching.

[0076] In this implementation manner, the relative behavioral characteristics are fed back to the last layer of the user tower, and the corresponding relative behavioral characteristics can be called according to the user and course to be matched before entering the matching layer, realizing the interaction between the user and the course. The MLP among them is used to transform the dimension of the concatenated characteristics to the same dimension as the Item Embedding for Cosine calculation.

[0077] The second optional implementation method feeds back the relative behavior characteristics of each user for each course to the first layer of the dual tower model. Combining Figure 6 , this feedback process includes the following steps:

[0078] Step 1: The fine ranking layer uses the relative behavior characteristics of each user for each course to decouple the independent components of each user and each course in the influence law between the long-term and short-term behaviors of the user and the course duration, and provides each independent component to the dual tower model of the multi-channel recall layer.

[0079] Step 2: When the dual tower model runs, it fuses the independent component of each user with its embedding feature in the first layer of the user tower, and fuses the independent component of each course with its embedding feature in the first layer of the course tower. The two fused features are respectively input into the subsequent processing layers of the user tower and the course tower. The fusion method is not limited, and it can be weighted fusion or dimensionality reduction fusion after splicing.

[0080] The reason for decoupling in this implementation method is that the user tower and the course tower in the dual tower model are independent of each other before the last layer of interaction and do not transmit information to each other. Therefore, the combination of the first layer requires decoupling the independent user component and course component from the relative behavior characteristics.

[0081] In a specific implementation method, a user decoupling layer can be pre-constructed. This layer can adopt a deep learning model to convert the relative behavior characteristics of each user for multiple courses into the same embedding space. By constraining the conversion results of the same user for multiple courses to tend to be consistent and maximizing the difference between the conversion results of different users, the training of the user decoupling layer can be completed.

[0082] Furthermore, in the training process, this model can adopt an encoder + decoder architecture, and both the encoder and the decoder adopt a fully convolutional network or a fully connected layer, etc. After inputting the relative behavior characteristic of user i for course j into the encoder, the encoded feature can be obtained; then, after inputting the encoded feature into the decoder, the decoded feature can be obtained. The difference between the decoded feature of the same relative behavior characteristic and the original feature can be constrained to be minimized in the loss function, and the encoded features of the relative behavior characteristics of the same user for multiple courses tend to be consistent, and the difference between the encoded features of the relative behavior characteristics of different users for each course is maximized to update the model parameters.

[0083] Exemplarily, the following loss function can be adopted:

[0084]

[0085] Among them, represents the loss function value, i, and respectively represent the user index, j, and respectively represent the course index, , and respectively represent the corresponding weights.

[0086] After training is completed, remove the decoder, and the trained encoder is the user decoupling layer. Inputting the relative behavior characteristics of any user for any course into the user decoupling layer can obtain the independent components of the relative behavior characteristics of the user, that is, the independent components in the complex law between the long-term and short-term behaviors of the user and the course duration.

[0087] Similarly, a course decoupling layer can be pre-constructed. This layer can adopt a deep learning model to convert the relative behavior characteristics of multiple users for each course into the same embedding space. By constraining the conversion results of multiple users for the same course to tend to be consistent and maximizing the differences in the conversion results between different courses, the training of the course decoupling layer is completed.

[0088] Furthermore, during the training process, the model can adopt an encoder + decoder architecture, and both the encoder and decoder adopt fully convolutional networks or fully connected layers, etc. Inputting the relative behavior characteristics of user i for course j into the encoder can obtain the encoded feature ; then inputting the encoded feature into the decoder can obtain the decoded feature . The difference between the decoded feature

[0089] of the same relative behavior characteristic and the original feature can be constrained to be minimized in the loss function, and the encoded features of the relative behavior characteristics of multiple users for the same course tend to be consistent, and the differences in the encoded features of the relative behavior characteristics of each user for different courses are maximized to update the model parameters.

[0090]

[0091] Exemplarily, the following loss function can be adopted: represents the loss function value, , and respectively represent the corresponding weights, and the meanings of the remaining variables are the same as .

[0092] After training is completed, the decoder is removed, and the trained encoder is the course decoupling layer. Inputting the relative behavior features of any user for any course into the course decoupling layer can obtain the independent components of the any course in the relative behavior features, that is, the independent components in the complex rules of the user's long-term and short-term behaviors and the course duration.

[0093] In the above two feedback methods, the first optional implementation method does not require decoupling, but it is necessary to separately identify the user and the course for each interaction in the matching layer; the second implementation method can provide a brand-new user tower, which can be reused in various scenarios after one decoupling. In practical applications, it can be flexibly selected according to needs.

[0094] After the feedback is completed, due to the change of the embedding vector, it is necessary to fine-tune the double-tower model after the feedback. This process can be carried out separately outside the user recommendation process, or gradually adjusted in the subsequent course recommendation process. Optionally, based on the principle of reinforcement learning, in this embodiment, the double-tower model to be fed back is used as the Actor network, the part of the entire recommendation system except the double-tower model is used as the Critic network, the user is used as the environment, the courses actually learned by the user are used as state variables, the courses recommended by the double-tower model are used as action variables, and the hit rate of the recommended courses actually learned by the user is used as the reward function, and the parameters of the Actor network are fine-tuned using the principle of reinforcement learning. With the accumulation of user data, the Actor network will be gradually optimized and can adjust the network performance in a timely manner as the user data changes.

[0095] It should be noted that the user data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0096] In summary, the embodiment of the present invention provides a teaching resource recommendation method based on user behavior and course duration, and applies it to a teaching resource recommendation system including a multi-channel recall layer, a rough ranking layer, a fine ranking layer, and a re-ranking layer, achieving the following beneficial effects:

[0097] 1. The recommendation systems in the current education field lack multiple recall paths, fail to solve the data sparsity problem, and at the same time, the understanding of user characteristics is still at the basic demographic characteristics stage, and the user characteristics of the recall strategy are less, affecting the personalization degree of the recommendation.

[0098] The recommendation method based on multi-channel recall and click-through rate prediction proposed in this embodiment is used to improve the recommendation effect. Resources are recalled from different perspectives through multi-channel recall, such as recall through links based on popular courses, user portraits, knowledge point associations, teacher reputations, learning stages, etc., ensuring the diversity of recommendation results, effectively alleviating the data sparsity problem. Even if the user interacts with the system less, appropriate recommendation results can be obtained based on multi-dimensional features, ensuring that the resource types are rich and diverse to meet diverse teaching scenarios.

[0099] 2. The current click-through rate prediction model has not fully characterized the cognitive dynamics and individual learning laws of learners. For example, in the existing SIM click-through rate prediction model, the non-linear laws of learners' learning behaviors and learning durations are often ignored, resulting in systematic biases in long-term and short-term behavior predictions, thus affecting the effect of resource recommendation.

[0100] This embodiment combines the SIM click-through rate prediction model, analyzes the students' past click behaviors on different types of resources, predicts the probability of their current preferences for various resources, and fits the individual learning preferences. More importantly, by combining the laws of video learning behaviors and time, different weights are assigned to long-term user behaviors and short-term user behaviors for different course durations to reflect the complex laws between learning behaviors and course durations, further improving the accuracy and personalization of click-through rate prediction.

[0101] 3. The current dual-tower model and SIM model operate independently, and there is also a lack of effective information feedback between recall, fine ranking, and rough ranking, and the information resources between layers are not fully utilized.

[0102] This embodiment integrates the user behavior features generated based on the SIM model and course duration into the dual-tower model of the recall and rough ranking strategies, further improving the accuracy of the dual-tower model and realizing the feedback of information between layers.

[0103] Figure 7 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more, Figure 7 taking one processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected through a bus or other means, Figure 7 taking connection through a bus as an example.

[0104] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the teaching resource recommendation method based on user behavior and course duration in the embodiments of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, to implement the above-mentioned teaching resource recommendation method based on user behavior and course duration.

[0105] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 61 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely set relative to the processor 60, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] The input device 62 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 63 may include a display device such as a display screen.

[0107] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the teaching resource recommendation method based on user behavior and course duration in any embodiment.

[0108] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0109] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0110] The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0111] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A teaching resource recommendation method based on user behavior and course duration, characterized in that: include: The multi-channel recall layer uses multiple recall branches to provide a collection of recalled courses, among which the Embedding vector branch uses a dual-tower model to provide some recalled courses. The coarse ranking layer uses another dual-tower model to match each recalled course with the user, obtains multiple coarse-ranked courses and provides them to the fine ranking layer. The fine sorting layer performs the operations of S110-S140: S110, obtaining a specific user and multiple courses to be recommended; S120. Perform operations from S1-1 to S1-3 for each course: S1-1, using the SIM model to generate the long-term behavior characteristics of the user; extracting the first duration with the largest value from the difference between the occurrence time of multiple behaviors input to the SIM model accurate search unit and the current time; S1-2. According to the relationship between the current course duration, the first duration, and the second duration covered by the short-term behavior characteristics of the user, the long-term behavior characteristics and the short-term behavior characteristics are weighted and summed to obtain the relative behavior characteristics of the user for the current course. The behavior characteristics are related to the current course, as well as the long-term behavior and short-term behavior of the user, and reflect the complex rules between the long-term and short-term behaviors of the user and the course length. Specifically, the current course is driven by a combination of short-term interests and long-term interests. According to the coverage or proportion of the current course duration on the two behavior characteristics, the influence of the two is calculated. S1-3, predicting the click rate of the user on the current course by using the relative behavior feature and the embedded feature of the current course; S130, recommending multiple courses to the user according to the click rate of each course; S140, feed back the relative behavioral characteristics of each user for each course to the dual-tower model of the multi-way recall layer and / or the coarse ranking layer, and fine-tune the dual-tower model after the feedback; specifically, pre-construct a user decoupling layer, convert the relative behavioral characteristics of each user for multiple courses into the same embedding space, and complete the training of the user decoupling layer by constraining the conversion results of the same user for multiple courses to be consistent and maximizing the difference in conversion results between different users; pre-construct a course decoupling layer, convert the relative behavioral characteristics of multiple users for each course into the same embedding space, and complete the training of the user decoupling layer by constraining multiple users for the same course The conversion results of different courses tend to be consistent and the difference in conversion results between different courses is maximized, thus completing the training of the course decoupling layer; the relative behavioral characteristics of any user for any course are respectively input into the trained user decoupling layer and the trained course decoupling layer, and the independent components of any user and any course in the influence law between the user's long-term and short-term behaviors and the course duration are respectively obtained; when the dual-tower model is running, the independent components of each user are fused with the embedded features of the first layer of the user tower, and the independent components of each course are fused with the embedded features of the first layer of the course tower, and the two fused features are respectively input into the subsequent processing layers of the user tower and the course tower.

2. The method according to claim 1, characterized in that The weighted summing of the long-term behavior feature and the short-term behavior feature according to the relationship between the current course duration, the first duration, and the second duration covered by the short-term behavior feature of the user includes: Determine the longer duration t of the first duration and the second duration L and shorter duration t S ; If the current course duration is less than or equal to the shorter duration, the weight of the behavior feature corresponding to the shorter duration is reset to 1, and the weight of the other behavior feature is reset to 0; If the current course duration is greater than or equal to the longer duration, the weight of the behavior feature corresponding to the shorter duration is reset to 0, and the weight of the other behavior feature is reset to 1; If the current course duration t is between the two durations, the weight of the behavior feature corresponding to the shorter duration is reset to The weight of the behavior feature corresponding to the longer duration is reset to 3. The method according to claim 1, characterized in that The fine-tuning of the double-tower model after the feedback includes: The twin-tower model is used as the Actor network, the part of the entire teaching resource recommendation system except the twin-tower model is used as the Critic network, the user is used as the environment, the courses actually studied by the user are used as the state variable, the courses recommended by the twin-tower model are used as the action variable, and the hit rate of the recommended courses actually studied by the user is used as the reward function. The parameters of the Actor network are fine-tuned using the principle of reinforcement learning.

4. The method according to claim 1, characterized in that: The teaching resource recommendation method further includes a re-ranking layer; accordingly, after the fine ranking layer executes S130, it also includes: The rearrangement layer executes the maximum margin correlation algorithm to rearrange the multiple courses obtained in S130 to obtain final recommended courses.

5. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the teaching resource recommendation method based on user behavior and course duration as described in any one of claims 1-4.

6. A computer-readable storage medium, characterized in that: A computer program is stored thereon, which, when executed by a processor, implements the teaching resource recommendation method based on user behavior and course duration as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Course teaching resource recommendation method

    CN116561410A

  • Teaching resource recommendation method and device, electronic equipment and readable storage medium

    CN117726082A

  • Hierarchical Attention deep learning model course recommendation method

    CN113435685A

  • Recommendation method and device of mobile chart and storage medium

    CN115481319A

  • Interaction data prediction method and device, electronic equipment and storage medium

    CN117615180A