An educational video playing method and system
Through the reinforcement learning network, the user interest summary is updated, combined with multi-dimensional behavioral analysis and similar user data, the problem of dynamic interest changes in educational video recommendations is solved, and the timeliness and accuracy of personalized video playback is improved.
Patent Information
- Application Number
- CN202410653846.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-05-24
AI Technical Summary
The existing educational video recommendation methods cannot capture the dynamic changes in users' interest during the learning process, and cannot consider the user's learning progress and phased needs, resulting in insufficient timeliness of recommendations and reducing the accuracy of personalized video playback.
Through the reinforcement learning network, the user interest summary is updated, combined with multi-dimensional user behavior analysis, dynamically adapt to user interest changes, and using interest fusion feature vectors and behavior data of similar users, calculate the interest matching degree and correction factor of the video, and recommend content suitable for the current learning stage.
It captures the dynamic interest changes of users in a timely and accurate manner, improves the timeliness and accuracy of personalized recommendations, and provides more accurate personalized video playback services.
Smart Images

Figure CN118400582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an educational video playing method and system. Background Art
[0002] It is not easy to quickly find educational videos of interest in a vast amount of educational videos. If users want to obtain high-quality educational videos, they often rely on the recommendations and sharing of surrounding classmates. However, everyone's interests vary, and the recommendations of surrounding classmates may not necessarily suit oneself. Therefore, an educational video playing method based on intelligent recommendation algorithms has emerged.
[0003] Intelligent educational video playing methods mainly adopt collaborative filtering recommendation algorithms to recommend educational videos based on the videos that users have played and the playing habits of similar users. User-based collaborative filtering can find users with similar viewing habits to the current user and recommend videos that they have watched but the current user has not. Item-based collaborative filtering can find videos similar to the videos that the current user has watched and recommend these similar videos.
[0004] However, in the prior art, whether it is user-based collaborative filtering or item-based collaborative filtering, it is often considered that users' interests are static and unchanging. In the actual learning process, users' interests will change as they learn deeper. For example, at the beginning of learning, users may be more concerned about basic content, and as their understanding deepens, they may turn to more advanced and professional content. The existing video recommendation and playing methods cannot capture the dynamic interest changes of users during the continuous learning process, cannot consider the progress and phased needs of users during the learning process, may recommend content that is not suitable for the current learning stage, resulting in insufficient timeliness of recommendations and reducing the accuracy of personalized video playing. Summary of the Invention
[0005] In order to solve the technical problems that the existing video recommendation and playing methods cannot capture the dynamic interest changes of users during the continuous learning process, cannot consider the progress and phased needs of users during the learning process, may recommend content that is not suitable for the current learning stage, resulting in insufficient timeliness of recommendations and reducing the accuracy of personalized video playing, the present invention provides an educational video playing method and system.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] An educational video playing method provided by an embodiment of the present invention includes:
[0009] S1: Obtain the historical video playing data of the current user;
[0010] S2: Update the user interest summary through a reinforcement learning network according to the historical video playback data;
[0011] S3: Determine the interest fusion feature vector of the current user according to the updated user interest summary;
[0012] S4: Calculate the interest similarity between other users and the current user according to the interest fusion feature vectors of other users and the current user, and determine multiple similar users with similar interests to the current user;
[0013] S5: Obtain the content to be learned by the current user;
[0014] S6: Determine multiple pending education videos for the content to be learned from an education video database according to the content to be learned;
[0015] S7: Determine the interest matching degree between each pending education video and the current user according to the similarity between the feature vector of each pending education video and the interest fusion feature vector of the current user;
[0016] S8: Determine the interest correction factor of each pending education video according to whether multiple similar users have watched each pending education video;
[0017] S9: Determine the comprehensive matching degree between each pending education video and the current user according to the interest matching degree between each pending education video and the current user and the interest correction factor;
[0018] S10: Play education videos for the current user in descending order of the comprehensive matching degree between each pending education video and the current user.
[0019] Second aspect:
[0020] An education video playback system provided by an embodiment of the present invention includes:
[0021] A processor;
[0022] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the education video playback method as described in the first aspect is implemented.
[0023] Third aspect:
[0024] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, the education video playback method as described in the first aspect is implemented.
[0025] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0026] In the present invention, the user interest summary is updated through the reinforcement learning network, the dynamic interest changes of the user in the continuous learning process are captured timely and accurately, the progress and stage-by-stage needs of the user in the learning process are considered, and the user's interest summary is fused into a feature vector, which can more comprehensively express the user's interest preferences, and then recommend content suitable for the current learning stage based on the interest fusion feature vector, thereby improving the timeliness of personalized recommendations and the accuracy of personalized video playback. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 A flowchart of an educational video playback method provided by an embodiment of the present invention;
[0029] Figure 2 A schematic diagram of the structure of an educational video playback system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0031] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0032] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0033] In the embodiments of the present invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are consistent.
[0034] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0035] Referring to the attached Figure 1 figures, a schematic flowchart of an educational video playback method provided by an embodiment of the present invention is shown.
[0036] An embodiment of the present invention provides an educational video playback method. This method can be implemented by an educational video playback device, which can be a terminal or a server. The processing flow of the educational video playback method can include the following steps:
[0037] S1: Obtain the historical video playback data of the current user.
[0038] S2: Update the user interest summary through a reinforcement learning network according to the historical video playback data.
[0039] Among them, the Reinforcement Learning Network (RLN) is a neural network trained using reinforcement learning algorithms, aiming to optimize strategies through trial-and-error and reward mechanisms to maximize cumulative rewards.
[0040] In a possible implementation manner, the environmental space of the reinforcement learning network is specifically: S=(s1, s2, s3, s4, s5, s6, s7), where s1 represents the cosine similarity between the feature vectors of each existing interest video in the user interest summary and the feature vector of the latest played educational video, s2 represents the playback behavior of the latest played educational video during playback, s3 represents the pause behavior of the latest played educational video during playback, s4 represents the fast-forward behavior of the latest played educational video during playback, s5 represents the rewind behavior of the latest played educational video during playback, s6 represents the download behavior of the latest played educational video during playback, and s7 represents the favorite behavior of the latest played educational video during playback.
[0041] It should be noted that introducing multi-dimensional behavioral data such as playback, pause, fast forward, rewind, download, and favorite can more comprehensively understand the user's interest level in the video. Different behaviors (such as pause and fast forward) may represent different interest signals, and through the reinforcement learning network, different weights can be assigned to these behaviors to refine the interest model.
[0042] The action space of the reinforcement learning network is specifically: A={0, 1}, where 0 means not to update the user interest summary, and 1 means to replace some existing interest videos with the latest played educational video to add the latest played educational video to the user interest summary and update the user interest summary.
[0043] The reward function of the reinforcement learning network is used to provide rewards or punishments for the reinforcement learning network according to the action execution results, specifically:
[0044] R(a,s) = μ1[log P(θ|s) - log P(θ′|s)] + μ2[Sim avg (x,x b ) - Sim avg (x,x′ b )]
[0045] Among them, R() represents the reward function of the reinforcement learning network, a represents the action of the reinforcement learning network, s represents the state of the reinforcement learning network, θ represents the updated user interest summary, θ′ represents the user interest summary before update, log P(θ|s) represents the logarithmic likelihood estimate of the updated user interest summary, log P(θ′|s) represents the logarithmic likelihood estimate of the user interest summary before update, x represents the latest played educational video, x b represents the existing interest videos in the updated user interest summary, x′ b represents the existing interest videos in the user interest summary before update, Sim avg (x,x b ) represents the average value of the cosine similarities between each existing interest video in the updated user interest summary and the latest played educational video, Sim avg (x,x′ b ) represents the average value of the cosine similarities between each existing interest video in the user interest summary before update and the latest played educational video, μ1 represents the weight coefficient of the logarithmic likelihood estimate term, and μ2 represents the weight coefficient of the cosine similarity term.
[0046] Among them, those skilled in the art can set the magnitudes of the weight coefficient μ1 of the logarithmic likelihood estimate term and the weight coefficient μ2 of the cosine similarity term according to actual situations, and the present invention does not make any limitations.
[0047] It should be noted that by combining the logarithmic likelihood estimate and the cosine similarity, the reward function can evaluate the update effect of the user interest model from two perspectives. The logarithmic likelihood estimate measures the probability matching degree of the model, while the cosine similarity measures the similarity between videos.
[0048] Specifically, the process of updating the user interest summary through the reinforcement learning network is as follows:
[0049] Whenever a user plays a new educational video, the system records the behavioral data related to this video, including behaviors such as play, pause, fast forward, rewind, download, and favorite. These data will be used as the input for the current state s. A reinforcement learning network (such as Q-learning or Deep Q-Network, DQN) is used to select an action a. The ε-greedy policy can also be used to decide whether to select the optimal action (exploitation) or randomly select an action (exploration) to balance exploration and exploitation. According to the selected action a, it is decided whether to update the user interest summary. 0 indicates that there is no need to update the user interest summary, and 1 indicates that part of the existing interest videos are replaced with the latest played educational video. Then, the reward value is calculated according to the reward function. The current state, action, reward, and next state are recorded, and according to the update rules of Q-learning or DQN, the current experience is used to update the Q value. For each new educational video play, the above process is repeated to continuously optimize the user interest summary.
[0050] In the present invention, the user interest summary is updated through a reinforcement learning network, combined with multi-dimensional user behavior analysis and a refined decision-making mechanism, which can not only dynamically adapt to the changes in user interests, but also improve the accuracy of the recommendation system and user satisfaction. By designing a reasonable reward function, this method can effectively optimize the user interest model, thereby providing a more accurate and personalized recommendation service.
[0051] Furthermore, the reinforcement learning network is trained by the gradient descent method. The gradient descent method is an iterative algorithm for optimizing the objective function, which is widely used in machine learning and deep learning to minimize (or maximize) the loss function. The following is a detailed introduction to the gradient descent method. Using the gradient descent method to train the reinforcement learning network is a very mature existing technology, which will not be elaborated in the present invention.
[0052] S3: Determine the interest fusion feature vector of the current user according to the updated user interest summary.
[0053] In a possible implementation manner, S3 is specifically: introducing an attention mechanism, performing feature fusion on the feature vectors of each interest video in the updated user interest summary to obtain the user's interest fusion feature vector:
[0054]
[0055] where Y represents the interest fusion feature vector, λ i represents the attention weight coefficient of the i-th interest video, Y i represents the feature vector of the i-th interest video, and n represents the total number of interest videos in the updated user interest summary.
[0056] Specifically, the attention weights can be extracted through the Transformer model. The specific calculation method of the attention weight coefficients is as follows:
[0057] λ i = h T ReLU[W(Y new ⊙ Y i ) + b]
[0058] where λ i represents the attention weight coefficient vector of the i-th interested video. The attention weight coefficient vector includes the attention weight coefficients of each existing interested video. h represents the projection vector, and h T represents the transpose of the projection vector. The projection vector is used to determine how to adjust the attention weights according to the existing interested videos and the target course in the updated user interest summary. ReLU() represents the activation function, Y new represents the feature vector of the latest played educational video, Y i represents the feature vector of the i-th existing interested video, ⊙ represents the element-wise product operation, W represents the weight matrix, and b represents the bias term.
[0059] In the present invention, through the attention weights, the system can adjust the importance of each interested video in the user interest model according to the specific feedback of the user on each video (such as viewing duration, interaction behavior, etc.). This dynamic adjustment of weights can more accurately capture and reflect the current interest state of the user, thereby enhancing the relevance and accuracy of recommendations.
[0060] S4: Fuse the feature vectors of other users with the feature vector of the current user's interest, calculate the interest similarity between other users and the current user, and determine multiple similar users with similar interests to the current user.
[0061] In a possible implementation manner, the interest similarity between the other user and the current user is specifically:
[0062]
[0063] where τ j represents the interest similarity between the j-th other user and the current user, Y j represents the interest fusion feature vector of the j-th other user, represents the transpose of the interest fusion feature vector of the j-th other user, Y represents the interest fusion feature vector of the current user, and || || represents the vector modulus operation.
[0064] In the present invention, by calculating the interest similarity between users, user groups with similar interests to the current user can be identified, and recommendations can be made based on the behaviors and preferences of these similar users.
[0065] In a possible implementation, the method for determining similar users is specifically as follows: when the interest similarity between other users and the current user is greater than a preset similarity, the other users are determined as similar users.
[0066] Among them, those skilled in the art can set the size of the preset similarity according to the actual situation, and the present invention does not make any limitations.
[0067] S5: Obtain the content to be learned by the current user.
[0068] S6: According to the content to be learned, determine multiple pending educational videos for the content to be learned from the educational video database.
[0069] In a possible implementation, S6 is specifically as follows: Select educational videos containing tags related to the content to be learned and / or with the name containing the content to be learned from the educational video database as the pending educational videos.
[0070] In the present invention, by matching tags and names, the system can more accurately screen out videos related to the current learning needs of users. Tags and names are direct descriptions of video content and can efficiently reflect the theme and main content of the video. At the same time, for the specific learning needs of users, more accurate content is recommended instead of general recommendations, which can avoid recommending irrelevant videos and improve the learning efficiency of users.
[0071] S7: Determine the interest matching degree between each pending educational video and the current user according to the similarity between the feature vector of each pending educational video and the interest fusion feature vector of the current user.
[0072] In a possible implementation, the interest matching degree between the pending educational video and the current user is specifically as follows:
[0073]
[0074] Among them, σ k represents the interest matching degree between the k-th pending educational video and the current user, Y k represents the feature vector of the k-th pending educational video, represents the transpose of the feature vector of the k-th pending educational video, Y represents the interest fusion feature vector of the current user, and || || represents the vector norm operation.
[0075] In the present invention, by calculating the cosine similarity of feature vectors, the correlation between the to-be-determined video and the user interest can be accurately measured, the subtle differences between the user interest and the video content can be effectively captured, the accuracy and relevance of the recommendation system can be significantly improved, and at the same time, the personalized recommendation ability and computational efficiency of the system can be enhanced, thereby providing a better user experience.
[0076] S8: Determine the interest correction factor of each to-be-determined educational video according to whether multiple similar users have watched each to-be-determined educational video.
[0077] In a possible implementation manner, the interest correction factor of the to-be-determined educational video is specifically:
[0078]
[0079] wherein, δ k represents the interest correction factor of the kth to-be-determined educational video, τ j represents the interest similarity between the jth similar user and the current user, r jk represents whether the jth similar user has watched the kth to-be-determined educational video. If so, r jk = 1, otherwise, r jk = 0, and J represents the total number of similar users.
[0080] In the present invention, by introducing the viewing behavior data of similar users, the system can more accurately judge the popularity and applicability of the to-be-determined video. This correction factor based on the behavior of similar users can improve the accuracy and personalization of the recommendation. Further, as the viewing behavior data of similar users changes, the interest correction factor will also be dynamically adjusted, enabling the recommendation system to timely reflect the latest interest trends of the user group, thereby maintaining the real-time and effectiveness of the recommendation.
[0081] S9: Determine the comprehensive matching degree between each to-be-determined educational video and the current user according to the interest matching degree between each to-be-determined educational video and the current user and the interest correction factor.
[0082] In a possible implementation manner, the comprehensive matching degree between the to-be-determined educational video and the current user is specifically:
[0083] ρ k = δ k σ k
[0084] wherein, ρ k represents the comprehensive matching degree between the kth to-be-determined educational video and the current user, σ k represents the interest matching degree between the kth to-be-determined educational video and the current user, and δ k represents the interest correction factor of the kth to-be-determined educational video.
[0085] In the present invention, by combining the interest matching degree (based on the similarity of feature vectors) and the interest correction factor (based on the behavior data of similar users), the relevance of each video to the current user can be evaluated more comprehensively. This combination method can accurately reflect the actual interests and potential preferences of users, thereby providing more personalized recommendations, which can significantly improve the accuracy, personalization, and user satisfaction of the recommendation system.
[0086] S10: Play educational videos for the current user in the order of the comprehensive matching degree between each pending educational video and the current user from high to low.
[0087] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:
[0088] In the present invention, the user interest summary is updated through a reinforcement learning network to timely and accurately capture the dynamic interest changes of users during the continuous learning process. Considering the progress and phased needs of users during the learning process, the user interest summary is fused into a feature vector, which can more comprehensively express the interest preferences of users. Furthermore, content suitable for the current learning stage is recommended according to the interest fusion feature vector, improving the timeliness of personalized recommendations and the accuracy of personalized video playback.
[0089] Refer to the attached Figure 2 illustrates a schematic structural diagram of an educational video playback system provided by the present invention.
[0090] The present invention also provides an educational video playback system 20, which is applied to the above-mentioned educational video playback method and includes:
[0091] A processor 201;
[0092] A memory 202, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor 201, the educational video playback method described in the method embodiment is implemented.
[0093] The educational video playback system 20 provided by the present invention can execute the above-mentioned educational video playback method and achieve the same or similar technical effects. To avoid repetition, the present invention will not be elaborated herein.
[0094] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:
[0095] In the present invention, the user interest summary is updated through a reinforcement learning network to timely and accurately capture the dynamic interest changes of the user during the continuous learning process. Considering the progress and phased requirements of the user during the learning process, the user interest summary is fused into a feature vector, which can more comprehensively express the user's interest preferences. Furthermore, content suitable for the current learning stage is recommended based on the interest fusion feature vector, improving the timeliness of personalized recommendation and the accuracy of personalized video playback.
[0096] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0097] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0098] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0099] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context before and after.
[0100] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0101] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0102] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0104] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0105] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, the functional units in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0107] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0108] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, it implements the educational video playing method as described in the method embodiment.
[0109] The computer-readable storage medium provided by the present invention can implement the steps and effects of the educational video playing method in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0110] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0111] In the present invention, the user interest summary is updated through a reinforcement learning network, timely and accurately capturing the dynamic interest changes of users during the continuous learning process, considering the progress and phased needs of users during the learning process, and fusing the user interest summary into a feature vector, which can more comprehensively express the interest preferences of users. Furthermore, content suitable for the current learning stage is recommended according to the interest fusion feature vector, improving the timeliness of personalized recommendation and the accuracy of personalized video playing.
[0112] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0113] The following points need to be explained:
[0114] (1) The drawings of the embodiment of the present invention only relate to the structures involved in the embodiment of the present invention, and other structures can refer to the general design.
[0115] (2) For clarity, in the drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.
[0116] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0117] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An educational video playing method, characterized in that, Including: S1: Obtain the historical video playback data of the current user; S2: According to the historical video playback data, update the user interest summary through a reinforcement learning network; S3: According to the updated user interest summary, determine the interest fusion feature vector of the current user; S4: According to the interest fusion feature vectors of other users and the interest fusion feature vector of the current user, calculate the interest similarity between other users and the current user, and determine multiple similar users with similar interests to the current user; S5: Obtain the content to be learned by the current user; S6: According to the content to be learned, determine multiple pending education videos for the content to be learned from the education video database; S7: According to the similarity between the feature vectors of each pending education video and the interest fusion feature vector of the current user, determine the interest matching degree between each pending education video and the current user; S8: According to whether multiple similar users have watched each pending education video, determine the interest correction factor of each pending education video; S9: According to the interest matching degree and interest correction factor between each pending education video and the current user, determine the comprehensive matching degree between each pending education video and the current user; S10: Play education videos for the current user in descending order of the comprehensive matching degree between each pending education video and the current user; Among them, the environmental space of the reinforcement learning network is specifically: S = (s1, s2, s3, s4, s5, s6, s7), where s1 represents the cosine similarity between the feature vectors of each existing interest video in the user interest summary and the feature vector of the latest played education video, s2 represents the playback behavior of the latest played education video during playback, s3 represents the pause behavior of the latest played education video during playback, s4 represents the fast-forward behavior of the latest played education video during playback, s5 represents the rewind behavior of the latest played education video during playback, s6 represents the download behavior of the latest played education video during playback, and s7 represents the favorite behavior of the latest played education video during playback; The action space of the reinforcement learning network is specifically: A = {0, 1}, where 0 means that there is no need to update the user interest summary, and 1 means replacing some existing interest videos with the latest played education video to add the latest played education video to the user interest summary and update the user interest summary; The reward function of the reinforcement learning network is used to provide rewards or punishments for the reinforcement learning network according to the action execution results, specifically: R(a, s) = μ1[log P(θ|s) - log P(θ′|s)] + μ2[Sim avg (x, x b ) - Sim avg (x, x′ b )] Among them, R() represents the reward function of the reinforcement learning network, a represents the action of the reinforcement learning network, s represents the state of the reinforcement learning network, θ represents the updated user interest summary, θ′ represents the user interest summary before update, logP(θ|s) represents the logarithmic likelihood estimation of the updated user interest summary, logP(θ′|s) represents the logarithmic likelihood estimation of the user interest summary before update, x represents the latest played educational video, x b represents the existing interest videos in the updated user interest summary, x′ b represents the existing interest videos in the user interest summary before update, Sim avg (x, x b ) represents the average value of the cosine similarities between each existing interest video in the updated user interest summary and the latest played educational video, Sim avg (x, x′ b ) represents the average value of the cosine similarities between each existing interest video in the user interest summary before update and the latest played educational video, μ1 represents the weight coefficient of the logarithmic likelihood estimation term, and μ2 represents the weight coefficient of the cosine similarity term; Among them, the interest matching degree between the pending education video and the current user is specifically: Among them, σ k represents the interest matching degree between the k-th to-be-determined educational video and the current user, Y k represents the feature vector of the k-th to-be-determined educational video, represents the transpose of the feature vector of the k-th to-be-determined educational video, Y represents the interest fusion feature vector of the current user, and |||| represents the modulus operation of the vector; Among them, the interest correction factor of the pending education video is specifically: Among them, δ k represents the interest correction factor of the k-th to-be-determined educational video, τ j represents the interest similarity between the j-th similar user and the current user, r jk represents whether the j-th similar user has watched the k-th to-be-determined educational video. If r jk = 1, otherwise, r jk = 0, and J represents the total number of similar users; Among them, the comprehensive matching degree between the pending education video and the current user is specifically: ρ k = δ k σ k Among them, ρ k represents the comprehensive matching degree between the k-th to-be-determined educational video and the current user, σ k represents the interest matching degree between the k-th to-be-determined educational video and the current user, δ k represents the interest correction factor of the k-th to-be-determined educational video.
2. The educational video playing method according to claim 1, wherein The S3 is specifically: Introduce an attention mechanism to perform feature fusion on the feature vectors of each interest video in the updated user interest summary to obtain the interest fusion feature vector of the user: Among them, Y represents the interest fusion feature vector, and λ i represents the attention weight coefficient of the i-th interest video, and Y i represents the feature vector of the i-th interest video, and n represents the total number of interest videos in the updated user interest summary.
3. The educational video playing method according to claim 1, wherein The interest similarity between other users and the current user is specifically: Among them, τ j represents the interest similarity between the j-th other user and the current user, and Y j represents the interest fusion feature vector of the j-th other user, represents the transpose of the interest fusion feature vector of the j-th other user, Y represents the interest fusion feature vector of the current user, and |||| represents the modulus operation of the vector.
4. The educational video playing method according to claim 3, wherein The specific method for determining similar users is as follows: When the interest similarity between other users and the current user is greater than a preset similarity, the other users are determined as similar users.
5. The educational video playing method according to claim 1, wherein The specific content of S6 is as follows: From the educational video database, select educational videos that contain tags related to the content to be learned and / or whose names contain the content to be learned as the to-be-determined educational videos.
6. An educational video playback system, characterized in that, It includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the educational video playing method described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multimedia content recommendation method and device
CN110781321A
Adaptive learning content recommendation method and system based on deep reinforcement learning
CN117009668A
Education resource recommendation method and system
CN117688241A