Data processing methods for reading recommendations
By inferring user interest states through time-varying structural causal models and variational Bayesian methods, combined with Hamiltonian Monte Carlo sampling and causal reinforcement learning, the problem of dynamic changes in user interests in the recommendation system is solved, personalized and real-time recommendation strategy optimization is achieved, and the adaptability and user experience of the recommendation system are improved.
Patent Information
- Application Number
- CN202510359561.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Existing recommendation systems are unable to track and capture the dynamic changes of user interests in real time, resulting in recommendation results that cannot adapt to user needs in a timely manner, lack of personalization and flexibility, and delayed utilization of feedback data and poor system adaptability.
The time-varying structural causal model and variational Bayesian method are used to infer user interest status. Combined with Hamiltonian Monte Carlo sampling and causal reinforcement learning, a personalized recommendation strategy is constructed to update user interest status in real time and optimize the recommendation strategy.
It achieves accurate capture and real-time response to user interests, improves the matching and diversity of recommended content, and enhances the system's adaptability and user experience.
Smart Images

Figure CN120296249B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a data processing method for reading recommendations. Background Art
[0002] Existing recommendation systems typically infer user interests based on static models. These systems predict user preferences using historical data or simple collaborative filtering algorithms, but struggle to track and capture dynamic changes in user interests in real time. Traditional systems often overlook the ever-changing nature of user interests, resulting in recommendations that fail to adapt to changing user needs. These methods often rely on past behavioral data and fail to accurately reflect users' current interests, leading to a mismatch between recommended content and actual user needs and diminished relevance.
[0003] Furthermore, existing technologies often employ fixed strategies for filtering and ranking recommended content, lacking sufficient personalization and flexibility. Recommendation systems typically base their recommendations on statistical features of user history or simple content tags, without fully considering the causal relationships and real-time changes in user interests. This templated or fixed rule-based recommendation approach often results in overly monotonous and repetitive recommendations, lacking novelty and diversity, impacting user experience and satisfaction. Especially given the diverse preferences of users, the system struggles to provide sufficiently diverse and relevant recommendations.
[0004] Finally, existing technologies have a certain lag in utilizing feedback data, and cannot update users' interest status in real time. User feedback data often accumulates over a long period of time before being reflected in the updated interest status. As a result, the system cannot respond to users' latest needs in a timely manner, resulting in unsatisfactory recommendation results. Users' immediate feedback data is often ignored or processed with delays, and the updating of recommended content is not flexible enough, resulting in poor system adaptability and an inability to quickly adjust strategies to cope with changing user behavior and preferences. Therefore, the present invention proposes a data processing method for reading recommendations to address the shortcomings of existing technologies. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the present invention provides a data processing method for reading recommendation, which solves the problems of low matching degree of recommended content, single recommendation and poor system adaptability.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data processing method for reading recommendations, comprising the following steps:
[0007] S1. Collect user behavior data, including user click behavior, dwell time, and characteristic information of reading content, and construct a user historical behavior time series;
[0008] S2. Based on the collected user behavior data, a time-varying structural causal model is established. The time-varying structural causal model is used to characterize the changing trend of the user's interest status and establish the interest status transfer relationship based on the user's historical behavior time series;
[0009] S3. Infer the user's interest state, calculate the posterior distribution of the user's interest state using the variational Bayes method, and update the interest state parameters by optimizing the lower bound of the evidence to obtain the optimized user interest state;
[0010] S4. Based on the optimized user interest state, the Hamiltonian Monte Carlo sampling method is used to sample the parameter space of the interest state to obtain a sample set, and the user interest state is further optimized based on the sample set to improve the accuracy and stability of the interest state estimation;
[0011] S5. Based on the optimized user interest state, a personalized recommendation strategy is constructed in combination with causal reinforcement learning methods. In the recommendation strategy, causal reasoning is used to calculate the causal impact of the recommended content on the user's interests, and reinforcement learning is used to optimize the long-term benefits of the recommendation strategy.
[0012] S6. Based on the personalized recommendation strategy, recommended content is generated and pushed to the user. The selection of recommended content is based on the interest status inference results and adjusted in combination with the user's historical reading preferences;
[0013] S7. Collect user feedback data on recommended content, and update user interest status based on the feedback data. On this basis, adjust the recommendation strategy to optimize subsequent recommendation effects.
[0014] The present invention also provides a data processing system for reading recommendation, comprising:
[0015] Data collection module, used to collect user behavior data and build user historical behavior time series;
[0016] The interest state inference module is used to infer the user's interest state based on the time-varying structural causal model and update the interest state parameters; the recommendation strategy optimization module is used to optimize the personalized recommendation strategy based on the user's interest state and calculate the causal impact of the recommended content on the user's interest;
[0017] The recommended content generation module is used to generate recommended content based on the optimized recommendation strategy and adjust it based on the user's historical reading preferences;
[0018] The feedback update module is used to collect user feedback data on recommended content and update user interest status based on the feedback data.
[0019] The present invention provides a data processing method for reading recommendations. It has the following beneficial effects:
[0020] 1. This invention uses a time-varying structural causal model to infer user interest states, enabling real-time dynamic capture of changing trends in user interests. This technical solution allows the system to more accurately identify the causal relationships behind user behavior, improving the accuracy of interest state inference. Compared to existing approaches that rely on static models, this invention addresses the inability of traditional methods to flexibly adapt to dynamic changes in user interests, ensuring a close match between recommended content and user interests.
[0021] 2. This invention introduces a personalized recommendation strategy optimization module to optimize the selection and sorting strategies for recommended content based on user interest status, making recommendation results more accurate and diverse. This technical solution significantly improves the overly simplistic and lacking diversity of traditional recommendation systems, enabling users to obtain recommendations that better meet their actual needs, thereby improving user experience and satisfaction.
[0022] 3. This invention utilizes a feedback update module to collect and analyze user behavior data in real time, forming a closed-loop update mechanism that enables the recommendation system to continuously adjust and optimize its recommendation strategies. Compared to existing solutions that lack rapid feedback processing, this invention effectively addresses the recommendation system's inability to quickly adapt to changing user interests, enhancing the system's adaptability and real-time responsiveness.
[0023] 4. The recommended content generation module of this invention combines user historical preferences with an optimized recommendation strategy to generate content, and introduces deduplication and diversity control mechanisms. This technical solution effectively avoids the problem of over-recommending similar content in traditional recommendation systems, improves the diversity and richness of recommendations, and enhances user interest and engagement in recommended content. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the method of the present invention;
[0025] Figure 2 This is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0027] See also Figure 1 , an embodiment of the present invention provides a data processing method for reading recommendation, comprising the following steps:
[0028] S1. Collect user behavior data, including user click behavior, dwell time, and characteristic information of reading content, and construct a user historical behavior time series;
[0029] In the technical solution of the present invention, the user's interest state is the basis for optimizing the recommendation strategy, and the accurate modeling of the interest state depends on comprehensive user behavior data. Therefore, before formally conducting interest modeling and causal reasoning, it is first necessary to collect data on the user's reading behavior to ensure that the subsequent calculation process has sufficient input data support. In general, the user's behavior data includes explicit behavior and implicit behavior, where explicit behavior mainly refers to the user's clicks, favorites and other clear interaction records, while implicit behavior involves scrolling, stay time and other features that do not directly express interest but have reference value. The present invention not only takes into account traditional user behavior data, but also introduces external factors such as device information and environmental characteristics to more comprehensively characterize the user's reading preferences. In one possible implementation method, the data acquisition module constructs the user's behavior history in a time series manner to provide input for subsequent causal modeling.
[0030] In this embodiment, the collection of user behavior data mainly includes but is not limited to the following categories:
[0031] Specifically, click behavior is an important indicator of user interest. The system records the number of clicks a user makes on the content they read, the order of the clicks, and the timestamps of the clicks. Click data not only reflects a user's interest in a particular article, but can also be used to analyze user preference trends for specific types of content. For example, if a user frequently clicks on articles on the same topic within a short period of time, it can be inferred that they have a high interest in that topic. Furthermore, click data can also be used to calculate short-term interest drift, that is, how a user's click preferences change over different time windows.
[0032] As an option, dwell time is also an important parameter for measuring user interest. In the present invention, the dwell time of the user on each article is recorded to assist in judging reading interest. Specifically, the system starts timing when the user opens an article, and records the dwell time when the user closes the page or switches to other pages. In general, a long dwell time usually means a higher reading interest, but in some cases, a long dwell time may also be due to the high difficulty of the article or the user repeatedly switching to read in a multitasking environment. Therefore, in one possible implementation method, the dwell time data will be corrected in combination with auxiliary features such as scrolling rate and page jump behavior to improve the accuracy of interest judgment.
[0033] In some embodiments, a user's collection behavior can directly reflect their deep interest in the content. In the present invention, the system records the number of times a user collects articles, the time of collection, and the characteristics of the collected content. For example, if a user collects a large number of articles in the same category in a short period of time, it indicates that their interest in this category may be relatively stable. In addition, collection behavior can also be used to distinguish short-term interests from long-term interests. In one possible implementation, if a user only collects a certain type of article in a certain period of time, and no longer reads similar content at other times, then this interest may be phased, rather than a long-term stable interest preference.
[0034] Generally speaking, scrolling behavior can provide a more fine-grained basis for judging interest than clicking. In the present invention, the system not only records the user's scrolling trajectory, but also includes information such as scrolling rate, rollback operation, and page browsing completion. Specifically, the scrolling rate can be used to judge the user's reading pattern. If the user slides the page quickly and stays for a short time, it may indicate that the user has less interest in the content. On the other hand, the rollback operation usually means that the user is interested in a certain content and wants to read it repeatedly. The system can analyze the user's focus on specific parts of the article through the number of rollbacks and the rollback area. In addition, in a possible implementation method, the browsing completion can be used to determine whether the user has read an article in its entirety, thereby assisting the analysis of click behavior and avoiding misjudgment of interest.
[0035] In the present invention, in order to further improve the accuracy of recommendations, data collection also includes device information. Device information can reflect the user's usage environment and affect their reading preferences in some cases. For example, users who use mobile devices are generally more inclined to short-term fragmented reading, while PC users may prefer to read longer articles. Therefore, in some embodiments, the device type (mobile phone, tablet, PC, etc.), operating system, network environment, etc. will be recorded and used as reference variables for recommendation optimization. Specifically, device information can be used to adjust the presentation of recommended content, for example, short content may be more recommended on the mobile side, while long articles may be recommended first on the PC side. In addition, in one possible implementation method, the system can analyze the user's reading habits on different devices in combination with the usage time pattern of the user's device, and further optimize the recommendation effect across devices.
[0036] In one possible implementation, in order to improve the validity of the data, the present invention uses a weighted fusion method to calculate the user's comprehensive reading interest index. Specifically, the user interest index I is obtained by weighted summation of multiple behavioral characteristics, and its calculation formula is as follows:
[0037] I = αC + βT + γF + δS;
[0038] Where C represents the number of clicks by the user; T represents the reading time of the user; F represents the number of times the user has saved an article; S represents the user's scrolling behavior; α, β, γ, and δ are normalized weight parameters, satisfying α + β + γ + δ = 1, and are adaptively adjusted based on the statistical characteristics of user behavior data.
[0039] In general, the weight parameters can be set based on statistical analysis of user historical data and adaptively adjusted through optimization methods. For example, in some scenarios, if it is found that the user's collection behavior has a strong predictive effect on their subsequent reading interest, the system can dynamically adjust the value of γ to increase the influence weight of the collection behavior. As an option, the weight parameters can also be optimized through machine learning methods, such as automatically adjusting the parameter values through the gradient descent method so that the calculated interest index best matches the user's actual reading preference.
[0040] In some embodiments, in order to ensure the real-time nature of the data, the system will divide the collected behavioral data into time windows. Specifically, the system can update the user's interest status at fixed time intervals (such as one day or one hour) to ensure that the recommendation strategy can adapt to the user's latest reading preferences. In addition, in one possible implementation, the data collection module can dynamically adjust the collection frequency according to the user's activity. For example, for highly active users, the data update cycle can be shortened, while for less active users, the collection interval can be appropriately extended to reduce computing overhead.
[0041] S2. Based on the collected user behavior data, a time-varying structural causal model is established. The time-varying structural causal model is used to characterize the changing trend of the user's interest status and establish the interest status transfer relationship based on the user's historical behavior time series;
[0042] In the technical solution of the present invention, the user's interest state changes dynamically, and the recommendation system needs to accurately model this change to ensure that the recommended content can adapt to the user's interest shift in a timely manner. In general, traditional collaborative filtering or deep learning-based recommendation methods mainly rely on historical data for modeling, which makes it difficult to capture the evolution of user interests over time. The present invention adopts a time-varying structural causal model (TS-CausalModel) to identify the main influencing factors of user interests through causal reasoning methods, and constructs interest state transfer relationships that change over time, so that the recommendation system can adjust the recommendation strategy in a dynamic environment. In one possible implementation method, the model combines time series modeling with causal reasoning to eliminate accidental correlations in the data, thereby improving the rationality and stability of recommendations.
[0043] In this embodiment, the establishment of the time-varying structural causal model includes the following key parts:
[0044] Specifically, the present invention defines the user's interest state as a time-dependent variable H t , represents the user’s interest distribution at time t. The change of interest status is affected by multiple factors, including the historical interest status H t-1 and the user's current behavior data X t To this end, the present invention constructs the following state transfer equation of interest:
[0045] P(H t |H t-1 ,X t )=f(H t-1 ,X t ;θ);
[0046] Among them, P(H t |H t-1 ,X t ) indicates that the interest state H at the previous moment is known t-1 and the current user behavior data X t In the case of user's current interest state H t The conditional probability distribution of H t Represents the user's interest state at time t; H t-1 represents the user's interest status at time t-1, reflecting the user's historical interest information; X t represents the user's behavioral data at time t, including features such as clicks, dwell time, favorites, and scrolling; θ represents the causal relationship parameter, which indicates the influence intensity of different influencing factors on the interest state transition.
[0047] As an option, the causal relationship parameter θ needs to be learned through a data-driven approach. Generally speaking, it is difficult to accurately characterize the causal relationship using maximum likelihood estimation directly. Therefore, the present invention adopts a structural learning method to determine the causal structure by calculating the conditional independence between different variables. In one possible implementation, a Bayesian structural learning method is used to calculate the posterior probability between variables:
[0048] P(G|D)∝P(D|G)P(G);
[0049] Among them, P(G|D) represents the posterior probability of the causal structure G given the user behavior data D; P(D|G) represents the likelihood probability of observing data D under the causal structure G; P(G) represents the prior distribution of the causal structure G, which is used to constrain the complexity of the causal structure.
[0050] Generally, it is difficult to directly solve the above Bayesian posterior distribution. Therefore, the present invention uses a Markov Chain Monte Carlo (MCMC) sampling method to sample possible causal structures and select the structure with the highest posterior probability as the final causal relationship model. In some embodiments, to improve computational efficiency, a gradient optimization method can be used to search for causal structures, thereby accelerating the learning process of causal relationships.
[0051] In one possible implementation, the time-varying structural causal model is used not only to predict interest states, but also to analyze the impact of different factors on interest changes. For example, if a user clicks on more science and technology articles in a certain period of time, but previously focused on sports content, the causal model can automatically identify this interest shift and adjust the recommendation strategy to gradually tilt it towards science and technology content. In addition, in some embodiments, the causal model can also be used to analyze the impact of environmental factors, such as the time of day, device type, network status, etc., to ensure that the recommendation strategy can adapt to different scenarios.
[0052] In order to further improve the prediction accuracy of the interest state, the present invention adopts the Bayesian Time-Varying Regression model to dynamically adjust the causal relationship parameter θ. Specifically, the update equation for setting the parameter is as follows:
[0053] θ t =θ t-1 +∈ t ;
[0054] Among them, θ t represents the causal relationship parameter at time t; θ t-1 Represents the causal relationship parameter at time t-1; ∈ t Represents the system noise term, which is used to characterize the dynamic changes of causal parameters and satisfies the Gaussian distribution σ 2 Represents the variance of the noise term, which controls the degree of random fluctuation of parameter updates.
[0055] In some embodiments, a particle filtering method may be used to t This method estimates the causal parameters, ensuring robustness during parameter updates. For example, if a user suddenly changes their reading interests within a short period of time, the system can quickly adjust the causal parameters without being overly influenced by long-term historical data. Furthermore, in one possible implementation, variational inference methods can be combined to estimate the posterior distribution of the interest state by optimizing the Evidence Lower Bound (ELBO), improving computational efficiency.
[0056] S3. Infer the user's interest state, calculate the posterior distribution of the user's interest state using the variational Bayes method, and update the interest state parameters by optimizing the lower bound of the evidence to obtain the optimized user interest state;
[0057] In the technical solution of the present invention, the user's interest state is modeled by a time-varying structural causal model, but the accurate inference of this state is crucial for recommendation optimization. Generally speaking, the inference of interest state involves complex probability calculations, and directly solving the posterior distribution P(H t |D) has a large amount of computation and is difficult to be applied efficiently in a large-scale data environment. Therefore, the present invention adopts the Variational Bayesian Inference (VBI) method to optimize the computational efficiency of the posterior distribution and improve the estimation accuracy of the state of interest. In one possible implementation, variational inference introduces the variational distribution Q(H t ) approximates the target posterior distribution P(H t |D) and optimizes it by maximizing the evidence lower bound (ELBO) to avoid directly calculating high-dimensional integrals, improve computational efficiency, and ensure the stability of the model.
[0058] In this embodiment, the inference process of user interest state mainly includes key steps such as interest state modeling, variational distribution construction, optimization target solution and parameter updating.
[0059] Specifically, the user's interest status H t is a potential random variable, and its true posterior distribution is difficult to solve directly. Therefore, the present invention adopts a parameterizable variational distribution Q(H t ; φ) to approximate the distribution, and optimize the variational parameter φ to make the two as close as possible. In general, the solution to the optimal variational distribution can be achieved by maximizing the following evidence lower bound (ELBO):
[0060]
[0061] in, Represents the lower bound of evidence, used to optimize the variational distribution; is the expectation of the log-likelihood term, calculated in the variational distribution Q(H t ), the expected log-likelihood value of user data D; D KL (Q(H t )∥P(H t )) is the Kullback-Leibler (KL) divergence, which is used to measure the variational distribution Q(H t ) and the true posterior distribution P(H t ) between the differences; H tis the user's interest state, which is inferred as a potential random variable; D is the user's historical behavior data set, including clicks, stay time, scrolling behavior, collection records, etc.; is a variational distribution, which is used to approximate the true posterior distribution P(H t |D);P(H t ) is the prior distribution of the state of interest; P(D|H t ) is the likelihood function, describing the given state of interest H t When , the probability of user behavior data D occurring is usually modeled as:
[0062]
[0063] Among them, d i is the i-th user behavior data sample, and N is the total number of data samples;
[0064] As an option, the variational distribution Q(H t ) adopts the form of multidimensional normal distribution to ensure the analyticity of the calculation while maintaining the continuity of the state of interest:
[0065]
[0066] in, The mean vector representing the interest state reflects the user's preference in each interest dimension; is the covariance matrix of the interest state, describing the correlation between the interest dimensions.
[0067] In some embodiments, in order to improve the computational efficiency of variational inference, the present invention adopts the mean-field variational inference (MFVI) method, that is, assuming the conditional independence between the state variables of interest in each dimension, thereby converting the covariance matrix Σ t Approximately a diagonal matrix:
[0068]
[0069] in, is the uncertainty measure of the i-th dimension of the state of interest, that is, the variance of the estimate in this dimension; d is the dimension of the state of interest; diag(.) means filling the diagonal elements into the matrix and setting the remaining positions to zero, thus forming a diagonal matrix.
[0070] In one possible implementation, the variational parameter (μ t ,Σ t ) is updated through gradient optimization, and the specific optimization goal is:
[0071]
[0072] Among them, θ* is the optimal parameter set; θ is all the trainable parameters of the model, including the causal parameter θ C and variational distribution parameter φ=(μ t ,Σ t ); is the evidence lower bound (ELBO).
[0073] In general, ELBO optimization can be performed using stochastic gradient variational inference (SGVI), estimating the gradient through Monte Carlo sampling, and using an adaptive learning rate optimization method (such as Adam) for parameter update. In some embodiments, in order to reduce the variance of the gradient estimate, the present invention adopts a reparameterization trick, that is, the random variable H is t Expressed as:
[0074]
[0075] in, Represents the user's interest state vector at time t; is the mean vector of interest states, which represents the estimation of user interest; represents the Cholesky decomposition of the covariance matrix; is the standard normal noise, It follows a multidimensional normal distribution with zero mean and identity covariance.
[0076] In one possible implementation, the variational inference process is used not only to estimate the state of interest, but also to calculate the uncertainty of the state of interest. Specifically, the present invention calculates the prediction variance Used to measure the reliability of the current interest state estimate:
[0077]
[0078] Where T represents the total number of samples used to estimate the variance, that is, from the variational distribution Q(H t ) the number of samples collected; Indicates the i-th time from the variational distribution Q(H t ) Sampled state samples of interest; Expresses the expectation of a state of interest.
[0079] Alternatively, if the prediction variance is high, indicating that the current interest state estimate is unstable, the system can add exploratory recommendations in the subsequent recommendation process to further collect user feedback and improve the accuracy of the interest state estimate. For example, during the cold start phase, due to the limited user behavior data, the estimated variance of the interest state is generally large. Therefore, the present invention adopts an uncertainty-driven exploration strategy in this phase to optimize the recommendation effect.
[0080] In addition, in some embodiments, in order to ensure the temporal continuity of the state of interest, the present invention introduces a temporal smoothing regularization term, that is, adding the following constraints in the optimization process:
[0081]
[0082] Where: λ is the smoothing coefficient, which controls the degree of interest change between time steps; T is the number of time steps; ||.|| 2 is the square of the Euclidean distance.
[0083] S4. Based on the optimized user interest state, the Hamiltonian Monte Carlo sampling method is used to sample the parameter space of the interest state to obtain a sample set, and the user interest state is further optimized based on the sample set to improve the accuracy and stability of the interest state estimation;
[0084] In the technical solution of the present invention, step S4 mainly focuses on improving the accuracy and timeliness of the model's prediction of the user's interest status by updating the user interest model in real time. Through this step, the system can adapt to the dynamic changes of user interests, improve the accuracy of interest prediction and the adaptability of the model. According to the interest status H inferred in step S3 t , and according to user feedback information F t Since HMC sampling can provide a more reliable estimate of the interest state distribution, the system uses the HMC sampling result as the starting point for gradient optimization to reduce the oscillation during parameter update and improve the convergence speed. In addition, during the optimization process, HMC sampling is combined with L2 regularization constraints to prevent model overfitting and make interest state prediction more robust. This embodiment proposes an efficient model update method. Generally speaking, as user behavior data continues to increase, the interest model needs to be dynamically adjusted to make the recommendation results more personalized and real-time.
[0085] Because HMC sampling incorporates gradient information, it enables efficient exploration within the high-dimensional parameter space of states of interest, resulting in more accurate estimation of states of interest and mitigating the slow convergence and local optima inherent in traditional MCMC methods. Furthermore, HMC sampling results can be used to initialize model parameters, making subsequent gradient optimization more stable, reducing oscillations, and improving optimization efficiency.
[0086] In this embodiment, the key steps for updating the user interest model include calculating a loss function, minimizing the loss, and updating the model parameters using an optimization algorithm. Specifically, the model's goal is to minimize the error in predicting interest states, and this error is reduced by adjusting the model parameters. To this end, the present invention employs an optimization method based on gradient descent, incorporating an L2 regularization term to prevent overfitting.
[0087] The model updating process of the present invention is based on the following loss function To optimize:
[0088]
[0089] Where T is the number of time steps; H t It represents the real user interest state at time step t, obtained through methods such as variational inference; is the state of interest estimated by the current model; Represents the prediction error of the state of interest, by calculating the state vector H t and the model prediction value , which measures the prediction performance of the model; λ is the regularization coefficient, which is used to control the complexity of the model; Θ is the trainable parameter set of the model, including the weights, biases and other training parameters of all network layers.
[0090] The first term of the loss function is the reconstruction error of the state of interest, which measures the difference between the state of interest predicted by the model and the true state of interest. The goal of this term is to make the model prediction as close to the true state of interest as possible. The second term λ∥Θ∥ 2 is the L2 regularization term, which prevents the model from overfitting by constraining the size of the model parameters.
[0091] To optimize the loss function, the present invention uses stochastic gradient descent (SGD) or its variants (e.g., the Adam optimization algorithm). Specifically, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters θ, and then the model parameters are updated based on the gradient. The optimized update rule can be expressed as:
[0092]
[0093] Among them, Θ t Represents the model parameters at time step t, including the weights and biases of all trainable network layers; Θ t+1 represents the updated model parameters at time step t+1; η represents the learning rate, which is a constant that controls the step size of each parameter update; Represents the loss function The gradient of the model parameters Θ represents the rate of change of the loss function at the current model parameter point and is used to adjust the model parameters to minimize the loss function.
[0094] Typically, the gradient is calculated by differentiating the loss function using the chain rule to obtain the gradient of each parameter, and using this gradient information to update the model parameters to minimize the prediction error.
[0095] In some embodiments, to further enhance the recommendation system's adaptability to changing user interests, the present invention combines a temporal weighting mechanism with an LSTM memory network to dynamically adjust interest states. This temporal weighting mechanism assigns higher weight to recent feedback, enabling more timely reflection of short-term interest changes while preventing rapid interest drift.
[0096] On the other hand, the LSTM memory network can model long-term user interest trends, ensuring that the system strikes a balance between short-term interest fluctuations and long-term interest evolution. Furthermore, the attention mechanism further optimizes interest state estimation, enabling the model to automatically focus on historical behaviors that have a significant impact on the current interest state, thereby improving prediction accuracy.
[0097] Considering that user interests have strong temporal characteristics, this paper introduces a temporal weighting mechanism to assign different weights to historical feedback of interest status, thereby highlighting the impact of recent feedback on interest status updates. For example, the following weighted loss function can be designed:
[0098]
[0099] Where: w t is the time weighting coefficient, which is usually inversely proportional to the time step t, giving lower weight to subsequent feedback as time goes by; is the reconstruction error of the state of interest, reflecting the prediction error of the model for the state of interest;
[0100] Weighting coefficient w t The choice of can be flexibly adjusted according to business needs. For example, in some scenarios, to better capture long-term interest change trends, we may give higher weights to earlier time steps, while in other cases we may focus more on recent interest changes.
[0101] Taking into account the temporal nature of interests, particularly the fact that changes in user interests depend not only on recent feedback but also on historical behavior, this paper employs memory networks or attention mechanisms to further optimize the updating of interest states. By introducing the attention mechanism, the system can dynamically assign different attention weights to historical interest states at each time step. This allows the model to not only accurately capture current interest changes but also leverage historical information to adjust and supplement interest states.
[0102] In this process, the attention weight α t The calculation of can be based on the correlation between historical interest states and current interest states:
[0103]
[0104] Among them, sec(H t-1 ,H t ) represents the historical interest state H t-1 With the current interest status H t The similarity between them is usually calculated by dot product or cosine similarity; exp(·) is an exponential function used to perform nonlinear transformation on the similarity score to enhance the influence of highly correlated interest states.
[0105] Through this mechanism, the model is able to focus on historical states that have a greater impact on the update of the current state of interest, thereby enhancing the model's responsiveness and accuracy to temporal changes.
[0106] To better adapt to long-term changes in user interests, the present invention can also introduce structures such as long-short-term memory networks (LSTMs) to help the model memorize long-term user interest patterns. For example, in some embodiments, a user's interest state is determined not only by their current behavior data but also by their behavior at multiple past moments. By introducing LSTM networks, the system can effectively capture these long-term dependencies, allowing model updates to better reflect the dynamic changes in user interests.
[0107] S5. Based on the optimized user interest state, a personalized recommendation strategy is constructed in combination with causal reinforcement learning methods. In the recommendation strategy, causal reasoning is used to calculate the causal impact of the recommended content on the user's interests, and reinforcement learning is used to optimize the long-term benefits of the recommendation strategy.
[0108] In the technical solution of the present invention, the key task of step S5 is to further adjust and optimize the interest state model through user feedback on the recommended content. Unlike the interest model update in the aforementioned step S4, step S5 focuses on adjusting the performance of the interest model in real time according to changes in real-time user feedback to ensure that the model can accurately reflect the user's latest preferences. Generally, as user behavior and feedback continue to accumulate, user interest preferences will change, and feedback information becomes an important basis for dynamically adjusting the interest model. Through this step, the present invention can ensure that the recommendation system always provides personalized recommendations based on the latest user interest status.
[0109] In this embodiment, the user interest status feedback and adjustment process involves the following key links: first, the system generates a feedback signal based on the user's interaction with the recommended content; second, based on the feedback signal, the parameters of the interest model are adjusted to optimize the estimation of the interest state; finally, the optimized interest model will be used for the next round of recommendations.
[0110] In step S5, the user's feedback signal F tIt is an important basis for adjusting interest status. Generally speaking, user feedback can be collected in many ways, such as user clicks, stay time, purchase behavior, likes, etc. In some scenarios, the system can further weight the feedback signal based on factors such as the frequency of user interaction with the content and the duration of interaction. Specifically, the feedback signal F t It can be represented as a vector where each element F t,i Indicates the user's preference for a certain category or feature at time step t. This feedback signal can be generated according to the following formula:
[0111] F t =f(C t ,I t );
[0112] Where: C t is the feature vector of the recommended content, such as the content category, label, recommendation time, etc.; I t is the user's interaction information, such as user clicks, favorites, etc.; f(·) is the function that generates the feedback signal, which can be implemented by weighted summation based on user behavior, machine learning model, etc.
[0113] According to the generated user feedback signal F t , the interest state adjustment process adopted by the present invention can be described by the following formula:
[0114]
[0115] in: is the adjusted interest state; H t is the original state of interest, usually obtained through methods such as variational Bayesian inference; ΔH t is the adjustment amount of the interest state, which reflects the impact of user feedback on the interest state, usually through the feedback signal F t Calculated; α is the adjustment coefficient, which controls the magnitude of the interest state adjustment. Generally, the adjustment coefficient α can be dynamically adjusted based on the strength of user feedback or the reliability of the feedback signal.
[0116] In some embodiments, the interest state adjustment amount ΔH t It can be calculated by the following method:
[0117] ΔH t =γ·f(F t ,H t );
[0118] Where: γ is the learning rate, which is used to adjust the influence of the feedback signal on the state of interest; f(F t ,H t) is a mapping function between the feedback signal and the current state of interest, which is usually modeled by a deep neural network or other regression model to map the feedback signal to the adjustment amount of the state of interest.
[0119] To prevent short-term fluctuations in user feedback from having a significant impact on interest status, the present invention introduces a feedback smoothing mechanism. Specifically, by weighting the adjustment of interest status by time, recent feedback information has a greater weight on the adjustment of interest status. This smoothing mechanism can be described by the following formula:
[0120]
[0121] in, represents the smoothed state of interest, the smoothed update result at time step t; represents the smoothed state of interest at the previous time step t-1; represents the interest state adjusted according to user feedback at time step t; the β smoothing coefficient, usually in the range of [0, 1], controls the weighted ratio between the historical interest state and the latest adjusted state. The larger the β, the stronger the dependence on the historical state; (1-β) represents the degree of influence of the current adjusted interest state on the smoothing process.
[0122] This smoothing mechanism ensures that the influence of historical behavior will not be excessively weakened during the adjustment of interest states, thereby avoiding excessive fluctuations in the interest model caused by short-term feedback changes.
[0123] Based on long-term user behavior, the feedback adjustment mechanism may need to be dynamically adjusted. In some embodiments, if the system identifies a significant change in a particular user's interests, the system can automatically increase the adjustment coefficient α to make the model more adaptable to the latest feedback. Specifically, the value of α can be dynamically adjusted based on the variance of historical feedback or the model's prediction error.
[0124] For example, when user feedback becomes more frequent or the categories of feedback content change significantly, the system will automatically adjust the feedback weighting strategy to give more weight to recent feedback. This adaptive adjustment strategy can be implemented through the following rules:
[0125] α=λ·Var(F t );
[0126] Among them, Var(F t ) represents the variance of the feedback signal at the current time step, which is used to measure the volatility of user feedback; λ is the adjustment factor, which is used to control the relationship between the variance and the adjustment coefficient.
[0127] Through the adjustment of the feedback information in step S5, the interest status is finally updated This adjustment process ensures that the system continuously adapts and dynamically optimizes to user interests, thereby providing more personalized recommendations.
[0128] S6. Based on the personalized recommendation strategy, recommended content is generated and pushed to the user. The selection of recommended content is based on the interest status inference results and adjusted in combination with the user's historical reading preferences;
[0129] In the technical solution of the present invention, the core purpose of step S6 is to generate personalized recommendation results based on the updated interest status. Through the previous step S5, the model has adjusted the interest status based on the user's feedback, making the estimation of the interest status more accurate and timely. Generally, as user behavior and feedback change dynamically, the interest status needs to be constantly adjusted, and step S6 further utilizes these adjusted interest states to generate recommendation results based on the user's personalized needs. Through this step, the present invention can ensure the accuracy and real-time nature of the recommendation results and optimize the overall effect of the recommendation system.
[0130] In this embodiment, interest state prediction and recommendation optimization involves two main steps: first, predicting the user's preferences based on their current interest state and generating a corresponding set of candidate recommendations; then, using an optimization algorithm, selecting and outputting the final recommended content. This process not only relies on the user's current interest state but also considers the influence of historical interest states, as well as the diversity and novelty of the recommended content.
[0131] In step S6, firstly, based on the adjusted interest state To predict the user's interest preferences in the future. Specifically, the prediction process is usually done through a mapping function The mapping function maps the smoothed interest state to the user's interest value for different recommended content. This process can be described by the following formula:
[0132]
[0133] in, Represents the predicted user interest value vector, reflecting the user's preference for various recommended content; is the prediction function, based on the smoothed state of interest Generate predicted user interests for different content.
[0134] In some embodiments, the mapping function f(·) can be trained using a deep neural network or other regression model, which can better capture complex nonlinear relationships and thus improve the accuracy of interest prediction. For the specific implementation of the mapping function, a possible neural network model can be used as follows:
[0135]
[0136] Among them, W1 and W2 are the weight matrices of the network respectively; b1 and b2 are bias terms; ReLU is the activation function used to increase the nonlinear characteristics of the model.
[0137] By predicting the interest value vector The present invention further generates a candidate set of recommended content. Generally, the recommended content is sorted according to the user's interest value for different categories or feature content, and several items with the highest interest value are selected as candidate recommendations. Assume that the content library C contains multiple recommended content c1, c2, ..., c m , the system can score the candidate content using the following formula:
[0138]
[0139] Among them, S(c i ) indicates content c i Recommendation score of c i is the content feature vector, such as content category, label, keyword, etc.
[0140] In some embodiments, the rating of recommended content not only depends on the interest status, but also can introduce optimization objectives such as diversity and deduplication. Specifically, a diversity constraint can be added to prevent the recommended content from being too similar, thereby improving the richness of the recommendation system. The optimization objective of the diversity constraint can be expressed as:
[0141]
[0142] in: represents the diversity loss function, which is used to optimize the diversity of recommended content and avoid recommending content that is too similar; |C| represents the size of the candidate recommendation content set, that is, the total number of recommended content in the content library; i, j represents the index of the content in the content set C, i and j traverse all content; c i and c j Represents the feature vectors of the i-th and j-th content in the candidate content set C; cosine sim (c i ,c j ) indicates content c i and c j The cosine similarity between them is used to measure the similarity between two contents.
[0143] By optimizing The system can increase the diversity of recommended content, thereby avoiding recommending too single content.
[0144] After generating candidate recommendation content, the present invention also uses an optimization algorithm to further improve the recommendation effect. Generally, in order to improve the quality of recommendation results, the optimization objective function will be comprehensively optimized by combining factors such as user feedback, content quality, and recommendation diversity. The optimization goal of the recommendation system can be expressed as:
[0145]
[0146] in, The loss function for the relevance between content and user interests is usually calculated based on the relevance between user behavior data and content. is the content diversity loss function; Optimize the loss function for ranking to ensure that the order of recommendation results is consistent with the user's potential interest order; λ div and λ rank are hyperparameters for adjusting diversity and ranking optimization weights, respectively.
[0147] In some embodiments, the ranking optimization loss function PairwiseRankingLoss can be used to optimize the ranking of content by comparing the attractiveness of different content to users.
[0148] Through this process, the final recommendation list generated by the present invention not only improves user satisfaction, but also enhances the long-term adaptability and accuracy of the recommendation system. Specifically, by optimizing the diversity, accuracy, and ranking of recommended content, the present invention ensures that the recommended content users receive is more personalized and relevant.
[0149] S7. Collect user feedback data on recommended content, and update user interest status based on the feedback data. Based on this, adjust the recommendation strategy to optimize subsequent recommendation effects;
[0150] In the technical solution of the present invention, step S7, a key subsequent step in the recommendation system process, is primarily intended to collect and process the presentation of recommendation results and user feedback, further establishing a closed-loop update mechanism for user interests. Generally, after receiving recommended content, users' interactive behaviors (such as clicks, browsing, staying, liking, forwarding, commenting, and adding to favorites) effectively reflect their true interests. Therefore, step S7, through the real-time collection of user behavior data, provides the necessary input data for subsequent interest status updates and recommendation system optimization.
[0151] Combined with the technical solution for generating and optimizing recommended content in the aforementioned step S6, step S7 constructs a dynamic user feedback data stream based on the actual interactive feedback between users and recommended content, ensuring that the system can accurately and timely capture the changing trends of user interests and provide real and effective feedback signals for subsequent steps.
[0152] In this embodiment, the collection of user feedback includes multi-dimensional monitoring of user interaction behaviors with recommended content. Specifically, the monitored user behaviors may include but are not limited to click behaviors, browsing time, collection behaviors, forwarding behaviors, comment behaviors, user ratings, and whether recommended content is ignored or skipped.
[0153] In one possible implementation, user behavior data is collected in real time through a log system or event tracking module and stored in a user behavior database. Subsequently, the system performs feature processing on the raw behavior data to extract feedback feature vectors that reflect the user's true interest tendency, which is denoted as:
[0154] F t =g(B t );
[0155] Among them, F t represents the feedback feature vector extracted at time step t; B t is the user’s original behavior data at time step t; g(·) is the feature extraction function, which can be implemented through statistical methods, feature engineering, or deep feature extraction models.
[0156] Specifically, the feature extraction function g(·) may adopt a neural network model or a decision tree-based model in some embodiments to achieve in-depth mining of implicit preference signals in the original behavior data. For example, it can be implemented as follows: t =MLP(B t );
[0157] Among them, MLP(·) is a multi-layer perceptron model, which is used to extract high-order behavioral features.
[0158] As an option, in some implementations, the feature vector F of user feedback t You can also check your current interest status Combined to form a joint feature representation for subsequent interest state updates, which can be specifically expressed as:
[0159]
[0160] Among them, U t represents the joint feature vector; φ(·) is the feature fusion function, which can be operations such as splicing, weighted summation, or attention mechanism.
[0161] In one possible implementation, the feature fusion function φ(·) is implemented using a weighted mechanism:
[0162]
[0163] Among them, α is the fusion weight coefficient, which ranges from [0, 1] and is used to balance the degree of fusion between user feedback features and interest status.
[0164] In addition, in this embodiment, the system also performs outlier detection and denoising on user feedback data to eliminate invalid behavior data that may be caused by misoperation or system anomalies, ensuring the accuracy and robustness of the subsequent interest update process. Generally, outlier detection can be implemented using the IQR method, the Z-score method, or a neural network-based anomaly detection model.
[0165] In some embodiments, to further enhance the guiding role of user feedback in updating interest status, the system may introduce a behavior weighting mechanism to assign different weight coefficients based on different types of user behaviors. For example:
[0166] The weight of the click behavior is denoted as w click ;
[0167] The weight of the collection behavior is denoted as w collect ;
[0168] The weight of the forwarding behavior is denoted as w share ;
[0169] The weight of comment behavior is denoted as w comment ;
[0170] The weight of the skipping behavior is denoted as w skip。
[0171] The system calculates the behavior contribution C by multiplying the weight coefficient of different behaviors by the behavior intensity. t :
[0172]
[0173] Among them, w i represents the weight of behavior i; s i is the intensity index of behavior i (such as number of clicks, duration of stay, etc.); n is the total number of behavior types.
[0174] Finally, the calculated behavior contribution C t As an important reference indicator of user feedback strength, it is input into the next step to guide the dynamic update process of user interest status.
[0175] In a possible implementation, the system can also adopt a time decay mechanism to calculate the contribution of historical behavior C t-k Assign a time attenuation factor γ, specifically:
[0176]
[0177] Among them, γ∈(0,1] is the time decay factor. The smaller γ is, the weaker the impact of historical behavior is. K is the length of the historical time window.
[0178] See also Figure 2 The present invention also provides a data processing system for reading recommendation, comprising:
[0179] The data collection module is used to collect multi-dimensional user behavior data and construct a time series of user historical behavior. The collected data includes behaviors such as clicks, browsing, staying, collecting, and comments. Combined with contextual information such as the time and scene of the behavior, a user behavior sequence data set is formed.
[0180] The interest state inference module is used to dynamically infer user interest states based on a time-varying structural causal model, and to update interest state parameters in real time, modeling the evolution of user interests over time and identifying key causal factors that affect interest changes.
[0181] The recommendation strategy optimization module is used to optimize personalized recommendation strategies based on user interest status, dynamically adjust the screening, sorting and distribution mechanisms of recommended content, and calculate the causal effect of recommended content on user interests, thereby optimizing the accuracy and diversity of recommendation strategies.
[0182] The recommended content generation module is used to generate recommended content based on the optimized recommendation strategy, adjust it based on the user's historical reading preferences, ensure that the recommendation results are highly matched with the user's interests, and introduce diversity and deduplication strategies to improve the richness of content recommendations.
[0183] The feedback update module is used to collect user feedback data on recommended content and analyze user interaction behavior in real time. The feedback update module uses the feedback results for the dynamic update of subsequent interest status, forming a closed-loop mechanism of recommendation-feedback-update and improving the system's adaptive ability.
[0184] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A data processing method for reading recommendations, characterized in that The following steps are involved: S1. Collect user behavior data, including user click behavior, dwell time, and characteristic information of reading content, and construct a user historical behavior time series; S2. Based on the collected user behavior data, a time-varying structural causal model is established. The time-varying structural causal model is used to characterize the changing trend of the user's interest status and establish the interest status transfer relationship based on the user's historical behavior time series; S3. Infer the user's interest state, calculate the posterior distribution of the user's interest state using the variational Bayes method, and update the interest state parameters by optimizing the lower bound of the evidence to obtain the optimized user interest state; S4. Based on the optimized user interest state, the Hamiltonian Monte Carlo sampling method is used to sample the parameter space of the interest state to obtain a sample set, and the user interest state is further optimized based on the sample set to improve the accuracy and stability of the interest state estimation; S5. Based on the optimized user interest state, a personalized recommendation strategy is constructed in combination with causal reinforcement learning methods. In the recommendation strategy, causal reasoning is used to calculate the causal impact of the recommended content on the user's interests, and reinforcement learning is used to optimize the long-term benefits of the recommendation strategy. S6. Based on the personalized recommendation strategy, recommended content is generated and pushed to the user. The selection of recommended content is based on the interest status inference results and adjusted in combination with the user's historical reading preferences; S7. Collect user feedback data on recommended content, and update user interest status based on the feedback data. On this basis, adjust the recommendation strategy to optimize subsequent recommendation effects.
2. The data processing method for reading recommendation according to claim 1, characterized in that: The user behavior data further includes device information, user reading time distribution, user collection behavior, and user scrolling behavior, and a weighted fusion method is used to calculate the user's comprehensive reading interest index. The comprehensive reading interest index calculation formula is: I = αC + βT + γF + δS; Among them, C represents the number of clicks of the user; T represents the reading time of the user; F represents the number of collections of the user; S represents the scrolling browsing behavior of the user; α, β, γ, δ are normalized weight parameters, satisfying α+β+γ+δ=1, and adaptive adjustment is performed according to the statistical characteristics of the user behavior data.
3. The data processing method for reading recommendation according to claim 1, characterized in that: The time-varying structural causal model optimizes the temporal dependencies of user interest states based on the attention mechanism, and is used to enhance the modeling capability of long-term dependencies.
4. The data processing method for reading recommendation according to claim 1, characterized in that: In the process of inferring the user's interest state, a hierarchical variational Bayesian method is used to jointly model the user's short-term interests and long-term interests and optimize the posterior distribution of the interest state.
5. The data processing method for reading recommendation according to claim 1, characterized in that: In the Hamiltonian Monte Carlo sampling process, an adaptive step size adjustment strategy is introduced to improve sampling efficiency and reduce computational cost. The step size adjustment formula is: ∈ t+1 =∈ t ·(1+η·(A t -A * )); Among them, ∈ t+1 Indicates the step size of the t+1th round of sampling; ∈ t represents the step size of the t-th round of sampling; η is the step size adjustment rate; A t Indicates the acceptance rate of the current sampling; A * is the preset target acceptance rate.
6. The data processing method for reading recommendation according to claim 1, characterized in that: The causal reinforcement learning method is based on a reverse backtracking reward estimation mechanism and combines the dynamic changes of interest states to optimize the recommendation strategy.
7. The data processing method for reading recommendation according to claim 1, characterized in that: In the process of selecting the recommended content, diversity constraints are further considered, and a user exploration factor is added to the optimization target of the recommendation strategy to avoid excessive recommendation of similar content.
8. The data processing method for reading recommendation according to claim 1, characterized in that: The recommended content is pushed in two ways: intelligent notification push and embedded content stream display, and the push frequency is dynamically adjusted based on user feedback.
9. The data processing method for reading recommendation according to claim 1, characterized in that: The feedback data is used to update the user's interest status and train the reinforcement learning model based on the optimization objective of the reward function to improve the long-term benefits of the personalized recommendation strategy.
10. A data processing system for reading recommendation, applied to the data processing method for reading recommendation according to any one of claims 1 to 9, characterized in that: include: Data collection module, used to collect user behavior data and build user historical behavior time series; The interest state inference module is used to infer the user's interest state based on the time-varying structural causal model and update the interest state parameters; The recommendation strategy optimization module is used to optimize the personalized recommendation strategy based on the user's interest status and calculate the causal impact of the recommended content on the user's interest; The recommended content generation module is used to generate recommended content based on the optimized recommendation strategy and adjust it based on the user's historical reading preferences; The feedback update module is used to collect user feedback data on recommended content and update user interest status based on the feedback data.
Citation Information
Patent Citations
Personalized learning recommendation system and method based on deep reinforcement learning
CN118628195A
Content recommendation method and system based on semantic recognition
CN119089398A