Data processing method for reading recommendation
The user interest state is inferred through the time-varying structural causal model and Bayesian method, combined with Hamilton Monte Carlo sampling and causal reinforcement learning, a personalized recommendation strategy is constructed, which solves the problem of dynamic changes in user interests in the existing recommendation system, realizes real-time matching and diversity of recommended content, and enhances the system's adaptability.
Patent Information
- Application Number
- CN202510359561.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing recommendation system cannot track the dynamic changes in user interests in real time, resulting in the mismatch of recommendation results with user needs, lack of personalization and flexibility, and the feedback data utilization is lagging, and the system is poor in adaptability.
The time-varying structure causal model and variational Bayesian method are used to infer user interest status, combine Hamilton Monte Carlo sampling and causal reinforcement learning, and build a personalized recommendation strategy, and optimize the strategy through feedback update modules to form a closed-loop mechanism.
Real-time dynamic capture of user interests is realized, the matching and diversity of recommended content is improved, and the adaptability and response speed of the system is enhanced.
Smart Images

Figure CN120296249A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and specifically to a data processing method for reading recommendations. Background Art
[0002] Existing recommendation systems usually infer user interests based on static models. These systems predict user preferences through historical data or simple collaborative filtering algorithms, but it is difficult to track and capture the dynamic changes of user interests in real time. The interest state of users changes all the time, but traditional systems often ignore this, resulting in the recommendation results being unable to adapt to the changes in user needs in a timely manner. Such methods usually rely on past behavioral data and fail to accurately reflect the current interests of users, resulting in a mismatch between the recommended content and the real needs of users and reducing the relevance of recommendations.
[0003] In addition, in the prior art, the screening and ranking of recommended content often adopt fixed strategies, lacking sufficient personalization and flexibility. Recommendation systems usually recommend based on the statistical features of users' historical behaviors or simple tags of content, without fully considering the causal relationships and real-time changes of user interests. This way of recommendation based on templates or fixed rules makes the recommended content often too single and repetitive, lacking novelty and diversity, affecting the user experience and satisfaction. Especially in the case of diverse user preferences, it is difficult for the system to provide sufficiently diverse and highly relevant recommended content.
[0004] Finally, there is a certain lag in the utilization of feedback data in the prior art, and the interest state of users cannot be updated in real time. User feedback data often takes a long time to accumulate before it can be reflected in the update of the interest state. In this way, the system cannot respond to the latest needs of users in a timely manner, resulting in unsatisfactory recommendation effects. The instant feedback data of users is usually ignored or processed with a delay, and the update of recommended content is not flexible enough, resulting in poor adaptability of the system and being unable to quickly adjust strategies to cope with the changing user behaviors and preferences; Therefore, the present invention proposes a data processing method for reading recommendations to solve the deficiencies of the prior art. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a data processing method for reading recommendations, which solves the problems of low matching degree of recommended content, singleization of recommendations, and poor adaptability of the system.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A data processing method for reading recommendations, including the following steps: S1. Collect user behavior data, where the user behavior data includes the click behavior, residence time, and characteristic information of the read content of the user, and construct a user historical behavior time series; S2. Based on the collected user behavior data, establish a time-varying structural causal model, which is used to depict the changing trend of the user's interest state, and establish an interest state transition relationship based on the user's historical behavior time series; S3. Infer the user's interest state, calculate the posterior distribution of the user's interest state using the variational Bayesian method, and update the interest state parameters by optimizing the evidence lower bound to obtain the optimized user interest state; S4. Based on the optimized user interest state, use the Hamiltonian Monte Carlo sampling method to sample the parameter space of the interest state to obtain a sample set, and further optimize the user interest state based on this sample set to improve the accuracy and stability of the interest state estimation; S5. Based on the optimized user interest state, combine the causal reinforcement learning method to construct a personalized recommendation strategy. In this recommendation strategy, use causal inference to calculate the causal impact of the recommended content on the user's interest, and optimize the long-term benefits of the recommendation strategy based on reinforcement learning; S6. Based on the personalized recommendation strategy, generate recommended content and push it to the user. The selection of the recommended content is based on the interest state inference result and is adjusted in combination with the user's historical reading preferences; S7. Collect the feedback data of the user on the recommended content, and update the user's interest state based on the feedback data. On this basis, adjust the recommendation strategy to optimize the subsequent recommendation effect.
[0007] The present invention also provides a data processing system for reading recommendations, including: A data collection module, which is used to collect user behavior data and construct a user historical behavior time series; An interest state inference module, which is used to infer the user's interest state based on the time-varying structural causal model and update the interest state parameters; A recommendation strategy optimization module, which is used to optimize the personalized recommendation strategy based on the user's interest state and calculate the causal impact of the recommended content on the user's interest; A recommended content generation module, which is used to generate recommended content according to the optimized recommendation strategy and adjust it in combination with the user's historical reading preferences; A feedback update module, which is used to collect the feedback data of the user on the recommended content and update the user's interest state based on the feedback data.
[0008] The present invention provides a data processing method for reading recommendations. It has the following beneficial effects: 1. The present invention uses a time-varying structure causal model to infer the user interest state, capable of dynamically capturing the trend of user interest changes in real time. Through this technical solution, the system can more accurately identify the causal relationships behind user behaviors, improving the accuracy of interest state inference. Compared with the existing solutions that rely on static models, the present invention solves the problem that traditional methods cannot flexibly adapt to the dynamic changes of user interests, ensuring a close match between the recommended content and user interests.
[0009] 2. The present invention optimizes the screening and sorting strategies of recommended content based on the user interest state by introducing a personalized recommendation strategy optimization module, making the recommendation results more accurate and diverse. This technical solution significantly improves the problem of being too single and lacking diversity in traditional recommendation systems. Users can obtain recommended content that better meets their real needs, thus enhancing the user experience and satisfaction.
[0010] 3. The present invention realizes the real-time collection and analysis of user behavior data through a feedback update module, forming a closed-loop update mechanism, enabling the recommendation system to continuously adjust and optimize the recommendation strategy. Compared with the existing solutions lacking fast feedback processing, the present invention effectively solves the deficiency that the recommendation system cannot quickly adapt to the changes of user interests, enhancing the adaptive ability and real-time response ability of the system.
[0011] 4. The recommended content generation module of the present invention generates content by combining the user's historical preferences and the optimized recommendation strategy, and introduces a duplicate removal and diversity control mechanism. This technical solution effectively avoids the problem of over-recommending similar content in traditional recommendation systems, improving the diversity and richness of recommendations, and enhancing the user's interest and participation in the recommended content. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the method flow chart of the present invention; Figure 2 is the system architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0014] Please refer to Figure 1 , the embodiments of the present invention provide a data processing method for reading recommendations, including the following steps: S1. Collect user behavior data, where the user behavior data includes the user's click behavior, stay time, and characteristic information of the content read, and construct a user historical behavior time series; In the technical solution of the present invention, the user's interest state is the basis for optimizing the recommendation strategy, and the accurate modeling of the interest state depends on comprehensive user behavior data. Therefore, before formally performing interest modeling and causal reasoning, it is first necessary to collect data on the user's reading behavior to ensure that there is sufficient input data support for the subsequent calculation process. Generally, the user's behavior data includes explicit behavior and implicit behavior, where the explicit behavior mainly refers to clear interaction records such as the user's clicks and favorites, while the implicit behavior involves features such as scrolling and stay duration that do not directly express interest but have reference value. The present invention not only considers traditional user behavior data but also introduces external factors such as device information and environmental characteristics to more comprehensively characterize the user's reading preferences. In a possible implementation manner, the data collection module constructs the user's behavior history in a time series manner to provide input for subsequent causal modeling.
[0015] In this embodiment, the collection of user behavior data mainly includes but is not limited to the following categories: Specifically, the click behavior is an important manifestation of the user's interest. The system records the number of clicks, click order, and click timestamps of the user on the reading content. The click data not only reflects the user's interest in a certain article but also can be used to analyze the user's preference trend for specific types of content. For example, if the user frequently clicks on articles of the same theme within a short period, it can be inferred that the user has a high interest in that theme. In addition, the click data can also be used to calculate short-term interest drift, that is, how the user's click preferences change under different time windows.
[0016] As an option, the stay time is also an important parameter for measuring the user's interest. In the present invention, the reading interest is assisted by recording the stay time of the user on each article. Specifically, the system starts timing when the user opens an article and records the stay duration when the user closes the page or switches to another page. Generally, a long stay time usually means a higher reading interest, but in some cases, a long stay may also be due to the high difficulty of the article or the user repeatedly switching reading in a multitasking environment. Therefore, in a possible implementation manner, the stay time data is corrected in combination with auxiliary features such as the scrolling rate and page jump behavior to improve the accuracy of interest judgment.
[0017] In some embodiments, the user's collection behavior can directly reflect their deep - seated interest in content. In the present invention, the system records the user's collection times, collection time, and the characteristics of the collected content. For example, if a user collects a large number of articles of the same category in a short period of time, it indicates that their interest in this category may be relatively stable. In addition, the collection behavior can also be used to distinguish short - term interests from long - term interests. In one possible implementation, if a user only intensively collects a certain type of article within a specific time period and no longer reads similar content at other times, then this interest may be stage - based rather than a long - term and stable interest preference.
[0018] Generally, the scrolling behavior can provide a more fine - grained basis for interest judgment than clicking. In the present invention, the system not only records the user's scrolling trajectory but also information such as the scrolling rate, roll - back operation, and page - viewing completion rate. Specifically, the scrolling rate can be used to judge the user's reading pattern. If the user quickly swipes the page and stays for a short time, it may indicate that their interest in this content is relatively low. On the other hand, the roll - back operation usually means that the user is interested in a certain section of the content and hopes to reread it. The system can analyze the user's focus on specific parts of the article through the number of roll - back times and the roll - back area. In addition, in one possible implementation, the page - viewing completion rate can be used to judge whether the user has read an article completely, thus assisting in the analysis of click behavior and avoiding misjudgment of interest.
[0019] In the present invention, in order to further improve the accuracy of recommendations, data collection also includes device information. Device information can reflect the user's usage environment and, in some cases, affect their reading preferences. For example, users using mobile devices usually prefer short - time fragmented reading, while PC users may be more willing to read longer articles. Therefore, in some embodiments, the device type (mobile phone, tablet, PC, etc.), operating system, network environment, etc. will be recorded and used as reference variables for recommendation optimization. Specifically, device information can be used to adjust the presentation method of recommended content. For example, on mobile devices, short content is more likely to be recommended, while on PC, long articles can be recommended first. In addition, in one possible implementation, the system can combine the usage time pattern of the user's device to analyze the user's reading habits on different devices and further optimize the cross - device recommendation effect.
[0020] In one possible implementation, in order to improve the effectiveness of data, the present invention uses a weighted fusion method to calculate the user's comprehensive reading interest index. Specifically, the user interest index I is obtained by weighted summation of multiple behavioral characteristics, and its calculation formula is as follows: I = αC + βT + γF + δS; Among them, C represents the number of clicks of the user; T represents the reading duration of the user; F represents the number of collections of the user; S represents the scrolling behavior of the user; α, β, γ, δ are normalized weight parameters, satisfying α + β + γ + δ = 1, and are adaptively adjusted according to the statistical characteristics of the user behavior data; Generally, the setting of the weight parameters can be statistically analyzed based on the user's historical data and adaptively adjusted through optimization methods. For example, in some scenarios, if it is found that the user's collection behavior has a strong predictive effect on their subsequent reading interest, the system can dynamically adjust the value of γ to increase the influence weight of the collection behavior. As an option, the weight parameters can also be optimized through machine learning methods, such as automatically adjusting the parameter values through the gradient descent method to make the calculated interest index match the user's actual reading preference most closely.
[0021] In some embodiments, to ensure the real-time nature of the data, the system will perform time window segmentation on the collected behavior data. Specifically, the system can update the user's interest status at fixed time intervals (such as one day, one hour), so as to ensure that the recommendation strategy can adapt to the user's latest reading preference. In addition, in a possible implementation manner, the data collection module can dynamically adjust the collection frequency according to the user's activity. For example, for highly active users, the data update cycle can be shortened, while for low-active users, the collection interval can be appropriately extended to reduce the computational overhead.
[0022] S2. Based on the collected user behavior data, establish a time-varying structural causal model, which is used to depict the change trend of the user's interest status and establish an interest status transition relationship based on the user's historical behavior time series; In the technical solution of the present invention, the user's interest status is dynamically changing, and the recommendation system needs to accurately model this change to ensure that the recommended content can timely adapt to the user's interest deviation. Generally, traditional collaborative filtering or deep learning-based recommendation methods mainly rely on historical data for modeling and are difficult to capture the evolution of the user's interest over time. The present invention adopts a time-varying structural causal model (TS-CausalModel), identifies the main influencing factors of the user's interest through causal inference methods, and constructs an interest status transition relationship that changes over time, enabling the recommendation system to adjust the recommendation strategy in a dynamic environment. In a possible implementation manner, this model combines time series modeling and causal inference to eliminate accidental correlations in the data, thereby improving the rationality and stability of the recommendation.
[0023] In this embodiment, the establishment of the time-varying structural causal model includes the following key parts: Specifically, the present invention defines the user's interest status as a time-dependent variable H t, representing the user's interest distribution at time t. The change of the interest state is affected by multiple factors, including the historical interest state H t-1 and the user's behavior data X at the current time t . For this reason, the present invention constructs the following interest state transition equation: P(H t |H t-1 ,X t ) = f(H t-1 ,X t ; θ); wherein, P(H t |H t-1 ,X t ) represents the conditional probability distribution of the user's current interest state H t-1 given the interest state H t at the previous time and the user's behavior data X t at the current time; H t represents the user's interest state at time t; H t-1 represents the user's interest state at time t - 1, reflecting the user's historical interest information; X t represents the user's behavior data at time t, including features such as clicks, dwell time, favorites, scrolling, etc.; θ represents the causal relationship parameter, indicating the influence intensity of different influencing factors on the interest state transition.
[0024] As an option, the causal relationship parameter θ needs to be learned by a data-driven method. Generally, it is difficult to accurately characterize the causal relationship directly using maximum likelihood estimation. Therefore, the present invention adopts a structure learning method to determine the causal structure by calculating the conditional independence between different variables. In a possible implementation manner, a Bayesian structure learning method is adopted to calculate the posterior probability between variables: P(G|D) ∝ P(D|G)P(G); wherein, P(G|D) represents the posterior probability of the causal structure G given the user's behavior data D; P(D|G) represents the likelihood probability of observing the data D under the causal structure G; P(G) represents the prior distribution of the causal structure G, used to constrain the complexity of the causal structure.
[0025] Generally, it is difficult to directly solve the above Bayesian posterior distribution. Therefore, the present invention adopts the Markov chain Monte Carlo (MCMC) sampling method to sample the possible causal structures, and selects the structure with the maximum posterior probability as the final causal relationship model. In some embodiments, in order to improve the calculation efficiency, a gradient optimization method can be adopted to search for the causal structure, thereby accelerating the learning process of the causal relationship.
[0026] In a possible implementation, the time-varying structural causal model is not only used to predict the interest state, but also to analyze the influence of different factors on the change of interest. For example, if a user clicks on a large number of technology articles within a certain period of time, while previously mainly focusing on sports content, the causal model can automatically identify this interest shift and adjust the recommendation strategy to gradually tilt towards technology content. In addition, in some embodiments, the causal model can also be used to analyze the influence of environmental factors, such as the time when the user is located, the device type, the network status, etc., to ensure that the recommendation strategy can adapt to different scenarios.
[0027] To further improve the prediction accuracy of the interest state, the present invention uses a Bayesian time-varying regression model (Bayesian Time-Varying Regression) to dynamically adjust the causal relationship parameter θ. Specifically, the update equation for setting the parameter is as follows: θ t = θ t-1 + ∈ t ; where, θ t represents the causal relationship parameter at time t; θ t-1 represents the causal relationship parameter at time t - 1; ∈ t represents the system noise term, which is used to characterize the dynamic change of the causal parameter and satisfies the Gaussian distribution σ 2 represents the variance of the noise term, which controls the random fluctuation degree of the parameter update.
[0028] In some embodiments, the particle filter method can be used to estimate θ t to ensure that the parameter update process has good robustness. For example, if a user suddenly changes their reading interest in a short period of time, the system can quickly adjust the causal parameter without being overly affected by long-term historical data. In addition, in a possible implementation, the variational inference method can be combined to estimate the posterior distribution of the interest state by optimizing the evidence lower bound (ELBO) to improve the calculation efficiency.
[0029] S3. Infer the user's interest state, use the variational Bayesian method to calculate the posterior distribution of the user's interest state, and update the interest state parameter by optimizing the evidence lower bound to obtain the optimized user interest state; In the technical solution of the present invention, the user's interest state is modeled by a time-varying structural causal model, but the accurate inference of this state is crucial for recommendation optimization. Generally, the inference of the interest state involves complex probability calculations, and directly solving the posterior distribution P(H t|D) has a large computational cost and is difficult to be efficiently applied in a large-scale data environment. Therefore, the present invention adopts the Variational Bayesian Inference (VBI) method to optimize the computational efficiency of the posterior distribution and improve the estimation accuracy of the interest state. In a possible implementation manner, variational inference approximates the target posterior distribution P(H t ) by introducing a variational distribution Q(H t |D), and optimizes it by maximizing the Evidence Lower Bound (ELBO) to avoid directly calculating high-dimensional integrals, improve the computational efficiency, and ensure the stability of the model.
[0030] In this embodiment, the inference process of the user interest state mainly includes key steps such as interest state modeling, variational distribution construction, optimization objective solution, and parameter update.
[0031] Specifically, the user's interest state H t is a latent random variable, and its true posterior distribution is difficult to directly solve. Therefore, the present invention adopts a parameterizable variational distribution Q(H t ; φ) to approximate this distribution, and optimizes the variational parameter φ to make the two as close as possible. Generally, the solution of the optimal variational distribution can be achieved by maximizing the following Evidence Lower Bound (ELBO): where, represents the Evidence Lower Bound, which is used to optimize the variational distribution; is the expectation of the log-likelihood term, calculating the expectation of the log-likelihood value of the user data D under the variational distribution Q(H t ); D KL (Q(H t )∥P(H t )) is the Kullback-Leibler (KL) divergence, which is used to measure the difference between the variational distribution Q(H t ) and the true posterior distribution P(H t ); H t is the user's interest state, which is inferred as a latent random variable; D is the set of the user's historical behavior data, including clicks, dwell time, scrolling behavior, favorite records, etc.; is the variational distribution, which is used to approximate the true posterior distribution P(H t |D); P(H t ) is the prior distribution of the interest state; P(D|H t ) is the likelihood function, which describes the probability of the user behavior data D occurring given the interest state H t , and is usually modeled as: where, di is the i-th user behavior data sample, and N is the total number of data samples; As an option, in the present invention, the variational distribution Q(H t ) adopts the form of a multi-dimensional normal distribution to ensure the analyticity of the calculation and maintain the continuity of the interest state: Among them, represents the mean vector of the interest state, reflecting the user's preferences in each interest dimension; is the covariance matrix of the interest state, describing the correlation between each interest dimension.
[0032] In some embodiments, in order to improve the computational efficiency of variational inference, the present invention adopts the mean-field variational (Mean-Field Variational Inference, MFVI) method, that is, assuming the conditional independence between the interest state variables of each dimension, so that the covariance matrix Σ t is approximated as a diagonal matrix: Among them, is the uncertainty measure of the i-th dimension of the interest state, that is, the estimated variance in this dimension; d is the dimension of the interest state; diag(.) means filling the diagonal elements into the matrix and the rest are zero, so as to form a diagonal matrix.
[0033] In a possible implementation manner, the variational parameters (μ t , Σ t ) are updated by gradient optimization, and the specific optimization objective is: Among them, θ * is the set of optimal parameters; θ is all trainable parameters of the model, including the causal relationship parameter θ C and the variational distribution parameter φ = (μ t , Σ t ); is the evidence lower bound (ELBO).
[0034] Generally, the optimization of ELBO can adopt stochastic gradient variational inference (SGVI), estimate the gradient through Monte Carlo sampling, and update the parameters using an adaptive learning rate optimization method (such as Adam). In some embodiments, in order to reduce the variance of gradient estimation, the present invention adopts the reparameterization trick, that is, representing the random variable H t as: Among them, Denote the user's interest state vector at time t; Is the mean vector of the interest state, representing the estimate of the user's interest; Denote the Cholesky decomposition of the covariance matrix; Is the standard normal noise, Subject to a multi-dimensional normal distribution with mean zero and covariance matrix being the identity matrix.
[0035] In a possible implementation, the variational inference process is not only used to estimate the interest state, but also to calculate the uncertainty of the interest state. Specifically, the present invention calculates the predictive variance To measure the reliability of the current interest state estimate: where T represents the total number of samples used to estimate the variance, that is, the number of samples sampled from the variational distribution Q(H t ); Denote the interest state sample obtained by the i-th sampling from the variational distribution Q(H t ); Denote the expectation of the interest state.
[0036] As an option, if the predictive variance is high, it indicates that the estimate of the current interest state is unstable. The system can increase the exploratory recommendation content in the subsequent recommendation process to further collect user feedback and improve the accuracy of the interest state estimate. For example, in the cold start phase, due to the lack of user behavior data, the estimation variance of the interest state is usually large. Therefore, the present invention adopts an exploration strategy driven by uncertainty at this stage to optimize the recommendation effect.
[0037] In addition, in some embodiments, in order to ensure the temporal continuity of the interest state, the present invention introduces a temporal smoothing regularization term, that is, adding the following constraint in the optimization process: where: λ represents the smoothing coefficient, controlling the degree of interest change between time steps; T is the number of time steps; ||.|| 2 Is the square of the Euclidean distance.
[0038] S4. Based on the optimized user interest state, use the Hamiltonian Monte Carlo sampling method to sample the parameter space of the interest state, obtain a sample set, and further optimize the user interest state based on this sample set to improve the accuracy and stability of the interest state estimate; In the technical solution of the present invention, step S4 mainly focuses on enhancing the accuracy and timeliness of the model for the user's interest state by updating the user interest model in real time. Through this step, the system can adapt to the dynamic changes of the user's interest, improve the accuracy of interest prediction and the adaptability of the model. According to the interest state H inferred in step S3 t , and according to the update of the user feedback information F t , since HMC sampling can provide a more reliable estimate of the interest state distribution, the system uses the HMC sampling result as the initial point for gradient optimization to reduce the oscillation during parameter update and improve the convergence speed. In addition, during the optimization process, HMC sampling is combined with L2 regularization constraints to prevent the model from overfitting and make the interest state prediction more robust. This embodiment proposes an efficient model update method. Generally, as the user behavior data continues to increase, the interest model needs to be dynamically adjusted to make the recommendation results more personalized and real-time.
[0039] Since HMC sampling combines gradient information and can efficiently explore in the high-dimensional interest state parameter space, it makes the interest state estimation more accurate and reduces the problems of slow convergence and easy entrapment in local optima of the traditional MCMC method. In addition, the result of HMC sampling can be used to initialize the model parameters, making the subsequent gradient optimization more stable, reducing oscillation, and improving the optimization efficiency.
[0040] In this embodiment, the key steps for updating the user interest model include: calculating the loss function, minimizing the loss, and updating the model parameters using an optimization algorithm. Specifically, the goal of the model is to minimize the interest state prediction error and reduce this error by adjusting the model parameters. For this purpose, the present invention adopts an optimization method based on gradient descent and combines an L2 regularization term on this basis to prevent the model from overfitting.
[0041] The model update process of the present invention is optimized based on the following loss function as follows: where T is the number of time steps; H t represents the true user interest state at time step t, which is obtained by methods such as variational inference; is the interest state estimated by the current model; represents the prediction error of the interest state, which is calculated by measuring the square of the Euclidean distance between the interest state vector H t and the model prediction value to measure the prediction performance of the model; λ is the regularization coefficient used to control the complexity of the model; Θ is the set of trainable parameters of the model, including the weights, biases, and other training parameters of all network layers.
[0042] The first term of this loss function is the reconstruction error of the interest state, which measures the difference between the interest state predicted by the model and the true interest state. The goal of this term is to make the model prediction as close as possible to the true interest state. And the second term λ∥Θ∥ 2 is the L2 regularization term, which prevents the model from overfitting by constraining the magnitude of the model parameters.
[0043] To optimize the loss function, the present invention employs Stochastic Gradient Descent (SGD) or its variants (such as the Adam optimization algorithm). Specifically, the backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters Θ, and then the model parameters are updated according to the gradient. The optimized update rule can be expressed as: where Θ t represents the model parameters at time step t, including the weights and biases of all trainable network layers; Θ t+1 represents the updated model parameters at time step t + 1; η represents the learning rate, which is a constant that controls the step size of each parameter update; represents the loss function is the gradient of the loss function with respect to the model parameters Θ, representing the rate of change of the loss function at the current model parameter point, and is used to adjust the model parameters to minimize the loss function.
[0044] Generally, the gradient is calculated by taking the derivative of the loss function using the chain rule to obtain the gradient of each parameter, and these gradient information are used to update the model parameters to minimize the prediction error.
[0045] In some embodiments, to further enhance the adaptability of the recommendation system to the changes in user interests, the present invention combines a temporal weighting mechanism with an LSTM memory network to achieve dynamic adjustment of the interest state. The temporal weighting mechanism assigns higher weights to recent feedback, enabling more timely reflection of short-term interest changes while avoiding excessive interest drift.
[0046] On the other hand, the LSTM memory network can model the long-term interest trends of users, ensuring that the system finds a balance between short-term interest fluctuations and long-term interest evolution. In addition, the attention mechanism can further optimize the interest state estimation, enabling the model to automatically focus on historical behaviors that have a greater impact on the current interest state and improving the prediction accuracy.
[0047] Considering that user interests have strong temporal characteristics, the present invention introduces a temporal weighting mechanism to assign different weights to the historical feedback of the interest state, thereby highlighting the impact of recent feedback on the update of the interest state. For example, a weighted loss function can be designed as follows: where: wt is the time-weighted coefficient, which is usually inversely proportional to the time step t. As time goes by, lower weights are given to subsequent feedbacks. is the reconstruction error of the interest state, reflecting the prediction error of the model for the interest state. weighting coefficient w t can be flexibly adjusted according to business requirements. For example, in some scenarios, in order to better capture the long-term interest change trend, higher weights may be assigned to earlier time steps, while in other cases, more attention may be paid to recent interest changes.
[0048] Further considering the temporal characteristics of interests, especially that the change of user interests not only depends on the recent feedback, but is also affected by historical behaviors, the present invention adopts a Memory Networks or an Attention Mechanism to further optimize the update of the interest state. By introducing the attention mechanism, the system can dynamically allocate different attention weights to historical interest states at each time step. In this way, the model can not only accurately capture the current interest change, but also use historical information to adjust and supplement the interest state.
[0049] During this process, the attention weight α t can be calculated based on the correlation between the historical interest state and the current interest state: where sec(H t-1 , H t ) represents the similarity between the historical interest state H t-1 and the current interest state H t , which is usually calculated by dot product or cosine similarity; exp(·) is the exponential function, used to perform a non-linear transformation on the similarity score to enhance the influence of highly correlated interest states.
[0050] Through this mechanism, the model can focus on those historical states that have a greater impact on the update of the current interest state, thereby enhancing the model's response ability and accuracy to temporal changes.
[0051] In order to better adapt to the long-term changes of user interests, the present invention can also introduce structures such as Long Short-Term Memory Networks (LSTM) to help the model remember long-term user interest patterns. For example, in some embodiments, the user's interest state is not only determined by the behavior data at the current moment, but is also affected by the behaviors at multiple past moments. By introducing the LSTM network, the system can effectively capture this long-term dependence relationship, enabling the update of the model to better reflect the dynamic changes of user interests.
[0052] S5. Based on the optimized user interest state, combined with the causal reinforcement learning method, construct a personalized recommendation strategy. In the said recommendation strategy, use causal inference to calculate the causal impact of the recommended content on the user's interest, and optimize the long-term benefits of the recommendation strategy based on reinforcement learning; In the technical solution of the present invention, the key task of step S5 is to further adjust and optimize the interest state model through the user's feedback on the recommended content. Different from the interest model update in the previous step S4, step S5 focuses on adjusting the performance of the interest model in real time according to the changes in real-time user feedback to ensure that the model can accurately reflect the latest preferences of the user. Generally, as the user's behavior and feedback accumulate, the user's interest preferences will change, and the feedback information becomes an important basis for dynamically adjusting the interest model. Through this step, the present invention can ensure that the recommendation system always provides personalized recommendations according to the latest user interest state.
[0053] In this embodiment, the user interest state feedback and adjustment process involves the following key links: First, the system generates a feedback signal according to the user's interaction with the recommended content; Second, based on this feedback signal, adjust the parameters of the interest model to optimize the estimation of the interest state; Finally, the optimized interest model will be used for the next round of recommendation.
[0054] In step S5, the user's feedback signal F t is an important basis for adjusting the interest state. Generally, user feedback can be collected in various ways, such as user clicks, dwell time, purchase behavior, likes, etc. In some scenarios, the system can also further weight and calculate the feedback signal according to factors such as the interaction frequency and interaction duration between the user and the content. Specifically, the feedback signal F t can be represented as a vector, where each element F t,i represents the preference intensity of the user for a certain category or a certain feature at time step t. This feedback signal can be generated according to the following formula: F t = f(C t , I t ); Where: C t is the content feature vector of the recommendation, such as the category, label, recommendation time, etc. of the content; I t is the user's interaction information, such as the user's click, collection, etc.; f(·) is the function for generating the feedback signal, and its specific implementation can be a weighted sum based on user behavior, a machine learning model, etc.
[0055] According to the generated user feedback signal F t , the interest state adjustment process adopted by the present invention can be described by the following formula: Wherein: is the adjusted interest state; H t is the original interest state, usually obtained by methods such as variational Bayesian inference; ΔH t is the adjustment amount of the interest state, which reflects the impact of user feedback on the interest state, usually calculated through the feedback signal F t ; α is the adjustment coefficient, which controls the amplitude of the interest state adjustment. Generally, the adjustment coefficient α can be dynamically adjusted according to the intensity of user feedback or the reliability of the feedback signal.
[0056] In some embodiments, the interest state adjustment amount ΔH t can be calculated by the following method: ΔH t = γ·f(F t , H t ); Wherein: γ is the learning rate, which is used to adjust the influence degree of the feedback signal on the interest state; f(F t , H t ) is the mapping function between the feedback signal and the current interest state, usually modeled by a deep neural network or other regression models, and is used to map the feedback signal to the interest state adjustment amount.
[0057] To avoid the excessive influence of short-term fluctuations in user feedback on the interest state, the present invention introduces a feedback smoothing mechanism. Specifically, by time-weighting the adjustment of the interest state, the recent feedback information has a greater weight on the adjustment of the interest state. This smoothing mechanism can be described by the following formula: Wherein, represents the smoothed interest state, which is the smoothed update result at time step t; represents the smoothed interest state at the previous time step t-1; represents the interest state adjusted according to user feedback at time step t; β is the smoothing coefficient, usually with a value range of [0, 1], which controls the weighted proportion between the historical interest state and the latest adjusted state. The larger β is, the stronger the dependence on the historical state; (1-β) represents the influence degree of the currently adjusted interest state on the smoothing process.
[0058] This smoothing mechanism ensures that in the process of adjusting the interest state, the influence of historical behaviors will not be overly weakened, thus avoiding excessive fluctuations in the interest model caused by short-term feedback changes.
[0059] According to the long-term behavior of users, the feedback adjustment mechanism may need to be dynamically adjusted. In some embodiments, if the system identifies that the interest of a specific user has changed significantly, the system can automatically increase the adjustment coefficient α to enable the model to have higher adaptability to the latest feedback. Specifically, the value of α can be dynamically adjusted according to the variance of historical feedback or the prediction error of the model.
[0060] For example, when the user's feedback becomes more frequent or the category of feedback content changes significantly, the system will automatically adjust the feedback weighting strategy to give more weight to recent feedback. This adaptive adjustment strategy can be implemented through the following rules: α = λ · Var(F t ); where, Var(F t ) represents the variance of the feedback signal at the current time step, which is used to measure the volatility of user feedback; λ represents the adjustment factor, which is used to control the relationship between the variance and the adjustment coefficient.
[0061] Through the adjustment of the feedback information in step S5, the finally updated interest state will be used as the input for the subsequent recommendation system. This adjustment process ensures the continuous adaptation and dynamic optimization of the system to the user's interests, so as to be able to provide more personalized recommendation content.
[0062] S6. Based on the personalized recommendation strategy, generate recommendation content and push it to the user. The selection of the recommendation content is based on the inference result of the interest state and is adjusted in combination with the user's historical reading preferences; In the technical solution of the present invention, the core purpose of step S6 is to generate personalized recommendation results according to the updated interest state. Through the previous step S5, the model has adjusted the interest state according to the user's feedback, making the estimation of the interest state more accurate and timely. Generally, as the user's behavior and feedback change dynamically, the interest state needs to be continuously adjusted, and step S6 generates recommendation results based on the user's personalized needs by further using these adjusted interest states. Through this step, the present invention can ensure the accuracy and real-time nature of the recommendation results and optimize the overall effect of the recommendation system.
[0063] In this embodiment, the prediction and recommendation optimization of the interest state include two main links: First, predict the user's preferences through the current interest state and generate a corresponding candidate recommendation set; then, select the final recommendation content through an optimization algorithm and output it. This process not only depends on the user's current interest state, but also considers the influence of the historical interest state, as well as the diversity and novelty of the recommendation content.
[0064] In step S6, first, it is necessary to be based on the adjusted interest state To predict the user's interest preferences over a period of time in the future. Specifically, the prediction process is usually carried out through a mapping function implemented, and this mapping function maps the smoothed interest state to the user's interest values for different recommended contents. This process can be described by the following formula: where, represents the predicted user interest value vector, reflecting the user's preference degree for various types of recommended contents; is the prediction function, which generates the predicted interests of the user for different contents according to the smoothed interest state
[0065] In some embodiments, the mapping function f(·) can be trained through a deep neural network or other regression models, which can better capture complex non-linear relationships, thereby improving the accuracy of interest prediction. For the specific implementation of the mapping function, a possible neural network model can be adopted as follows: where, W1 and W2 are the weight matrices of the network respectively; b1 and b2 are the bias terms; ReLU is the activation function, which is used to increase the non-linear characteristics of the model.
[0066] Through the predicted interest value vector The present invention further generates a candidate set of recommended contents. Usually, the recommended contents are sorted according to the user's interest values for different categories or feature contents, and several items with the highest interest values are selected as candidate recommendations. Assume that the content library C contains multiple recommended contents c1, c2,..., c m , the system can score the candidate contents through the following formula: where, S(c i ) represents the recommendation score of the content c i ; c i is the content feature vector, such as the category, label, keyword, etc. of the content.
[0067] In some embodiments, the scoring of the recommended contents not only depends on the interest state, but also can introduce optimization objectives such as diversity and deduplication. Specifically, a diversity constraint can be added to avoid the recommended contents being too similar, thereby improving the richness of the recommendation system. The optimization objective of the diversity constraint can be expressed as: where: Represents the diversity loss function, which is used to optimize the diversity of recommended content and avoid recommending overly similar content; |C| represents the size of the candidate recommended content set, that is, the total number of content available for recommendation in the content library; i and j represent the indices of the content in the content set C, and i and j traverse all content; c i and c j represent the feature vectors of the i-th and j-th content in the candidate content set C; cosine sim (c i , c j ) represents the cosine similarity between content c i and c j , which is used to measure the similarity between two pieces of content.
[0068] By optimizing the system can increase the diversity of recommended content, thereby avoiding overly single recommended content.
[0069] After generating candidate recommended content, the present invention also adopts an optimization algorithm to further improve the recommendation effect. Generally, in order to improve the quality of the recommendation results, the optimization objective function will comprehensively optimize by combining user feedback, content quality, and recommendation diversity and other factors. The optimization objective of the recommendation system can be expressed as: Among them, is the loss function of the relevance between content and user interests, which is usually calculated based on the relevance between user behavior data and content; is the content diversity loss function; is the ranking optimization loss function, which is used to ensure that the ranking of the recommendation results conforms to the order of the user's potential interests; λ div and λ rank are hyperparameters for adjusting the diversity and ranking optimization weights respectively.
[0070] In some embodiments, the ranking optimization loss function can adopt PairwiseRankingLoss to optimize the ranking of content by comparing the attractiveness of different content to users.
[0071] Through this process, the final recommended list generated by the present invention can not only improve user satisfaction, but also enhance the long-term adaptability and accuracy of the recommendation system. Specifically, by optimizing the diversity, accuracy, and ranking of recommended content, the present invention ensures that the recommended content obtained by users is more personalized and relevant.
[0072] S7. Collect user feedback data on the recommended content, update the user interest status based on the feedback data, and on this basis, adjust the recommendation strategy to optimize the subsequent recommendation effect; In the technical solution of the present invention, step S7, as a subsequent key step in the recommendation system process, mainly aims to collect and process the display of recommendation results and user feedback to further form a closed-loop update mechanism for user interests. Generally, after receiving the recommended content, the user's interaction behaviors (such as clicking, browsing, staying, liking, forwarding, commenting, collecting, etc.) can effectively reflect their true interest tendencies. Therefore, step S7 provides the necessary input data for subsequent interest state updates and recommendation system optimizations by collecting user behavior data in real time.
[0073] Combined with the technical solution for generating and optimizing recommended content in the foregoing step S6, step S7 constructs a dynamic user feedback data stream based on the actual interaction feedback between the user and the recommended content, ensuring that the system can accurately and timely capture the changing trends of user interests and provide real and effective feedback signals for subsequent steps.
[0074] In this embodiment, the collection of user feedback includes multi-dimensional monitoring of the user's interaction behaviors with the recommended content. Specifically, the monitored user behaviors may include, but are not limited to, click behaviors, browsing durations, collection behaviors, forwarding behaviors, comment behaviors, user ratings, and whether to ignore or skip the recommended content, etc.
[0075] In a possible implementation, user behavior data is collected in real time through a logging system or an event tracking module, and the behavior data is uniformly stored in a user behavior database; subsequently, the system performs feature extraction on the original behavior data to extract a feedback feature vector reflecting the user's true interest tendency, denoted as: F t =g(B t ); where, F t represents the feedback feature vector extracted at time step t; B t is the original behavior data of the user at time step t; g(·) is a feature extraction function, which can be implemented through statistical methods, feature engineering, or a deep feature extraction model.
[0076] Specifically, in some embodiments, the feature extraction function g(·) may adopt a neural network model or a decision tree-based model to achieve in-depth mining of the implicit preference signals in the original behavior data. For example, it can be implemented in the following way: F t =MLP(B t ); where, MLP(·) is a multi-layer perceptron model for extracting high-order behavior features.
[0077] As an option, in some implementations, the feature vector F t of the user feedback can also be combined with the current interest state Combine to form a joint feature representation for subsequent update of the interest state, which can be specifically expressed as: Among them, U t represents the joint feature vector; φ(·) is a feature fusion function, which can be operations such as concatenation, weighted summation, or attention mechanism.
[0078] In a possible implementation, the feature fusion function φ(·) is implemented using a weighted mechanism: Among them, α is the fusion weight coefficient, and its value range is [0, 1], which is used to balance the fusion degree of user feedback features and the interest state.
[0079] In addition, in this embodiment, the system also performs outlier detection and denoising processing on the user feedback data to eliminate invalid behavior data that may be caused by misoperations or system anomalies, ensuring the accuracy and robustness of the subsequent interest update process. Generally, outlier detection can be implemented using the IQR method, the Z-score method, or an outlier detection model based on a neural network.
[0080] In some embodiments, to further enhance the guiding role of user feedback in the interest state update, the system can introduce a behavior weight mechanism and assign different weight coefficients according to different types of user behaviors. For example: The weight of the click behavior is denoted as w click ; The weight of the favorite behavior is denoted as w collect ; The weight of the forward behavior is denoted as w share ; The weight of the comment behavior is denoted as w comment ; The weight of the skip behavior is denoted as w skip。 The system calculates the behavior contribution degree C t by multiplying the weight coefficient of different behaviors by the behavior intensity: Among them, w i represents the weight of behavior i; s i is the intensity index of behavior i (such as the number of clicks, the dwell time, etc.); n is the total number of behavior types.
[0081] Finally, the calculated behavior contribution degree C t is used as an important reference index for the user feedback intensity and is input into the next step to guide the dynamic update process of the user interest state.
[0082] In a possible implementation, the system may also adopt a time decay-based mechanism for the historical behavior contribution degree C t-k and assign a time decay factor γ, specifically: where γ ∈ (0, 1] is the time decay factor, and the smaller γ is, the weaker the influence of historical behavior; K is the length of the historical time window.
[0083] Please refer to Figure 2 , the present invention also provides a data processing system for reading recommendations, including: A data acquisition module, configured to collect multi-dimensional user behavior data, construct a user historical behavior time series. The collected data includes behaviors such as clicks, browsing, stays, favorites, and comments, and forms a user behavior sequence data set by combining context information such as the time and scenario of behavior occurrence.
[0084] An interest state inference module, configured to dynamically infer the user's interest state based on a time-varying structural causal model, and update the interest state parameters in real time, model the evolution process of the user's interest over time, and identify key causal factors affecting interest changes.
[0085] A recommendation strategy optimization module, configured to optimize the personalized recommendation strategy based on the user's interest state, dynamically adjust the screening, sorting, and distribution mechanisms of recommended content, and calculate the causal effect of recommended content on the user's interest, so as to optimize the accuracy and diversity of the recommendation strategy.
[0086] A recommended content generation module, configured to generate recommended content according to the optimized recommendation strategy, and adjust it in combination with the user's historical reading preferences to ensure a high degree of matching between the recommended result and the user's interest, and introduce diversity and deduplication strategies to improve the richness of content recommendations.
[0087] A feedback update module, configured to collect feedback data of the user on the recommended content, analyze the user's interaction behavior in real time. The feedback update module uses the feedback result for the dynamic update of the subsequent interest state, forming a closed-loop mechanism of recommendation-feedback-update to improve the adaptive ability of the system.
[0088] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data processing method for reading recommendations, characterized in that, It includes the following steps: S1. Collect user behavior data, where the user behavior data includes the user's click behavior, residence time, and characteristic information of the content read, and construct a user historical behavior time series; S2. Based on the collected user behavior data, establish a time-varying structural causal model, which is used to depict the change trend of the user interest state, and establish an interest state transition relationship based on the user historical behavior time series; S3. Infer the user interest state, calculate the posterior distribution of the user interest state using the variational Bayesian method, and update the interest state parameters by optimizing the evidence lower bound to obtain the optimized user interest state; S4. Based on the optimized user interest state, use the Hamiltonian Monte Carlo sampling method to sample the parameter space of the interest state to obtain a sample set, and further optimize the user interest state based on this sample set to improve the accuracy and stability of the interest state estimation; S5. Based on the optimized user interest state, combine the causal reinforcement learning method to construct a personalized recommendation strategy. In the recommendation strategy, use causal inference to calculate the causal impact of the recommended content on the user interest, and optimize the long-term benefit of the recommendation strategy based on reinforcement learning; S6. Based on the personalized recommendation strategy, generate recommended content and push it to the user. The selection of the recommended content is based on the interest state inference result and is adjusted in combination with the user's historical reading preference; S7. Collect the feedback data of the user on the recommended content, and update the user interest state based on the feedback data. On this basis, adjust the recommendation strategy to optimize the subsequent recommendation effect.
2. The data processing method for reading recommendations according to claim 1, wherein The user behavior data further includes device information, the user's reading duration distribution, the user's collection behavior, the user's scrolling browsing behavior, and use the weighted fusion method to calculate the user's comprehensive reading interest index. The formula for calculating the comprehensive reading interest index is: I = αC + βT + γF + δS; where C represents the number of clicks of the user; T represents the reading duration of the user; F represents the number of collections of the user; S represents the user's scrolling browsing behavior; α, β, γ, δ are normalized weight parameters, satisfying α + β + γ + δ = 1, and are adaptively adjusted according to the statistical characteristics of the user behavior data.
3. The data processing method for reading recommendation according to claim 1, characterized in that, The time-varying structural causal model optimizes the time series dependence relationship of the user interest state based on the attention mechanism, and is used to enhance the modeling ability of long-term dependence.
4. The data processing method for reading recommendations according to claim 1, wherein In the inference process of the user interest state, use the hierarchical variational Bayesian method, which is used to jointly model the user's short-term interest and long-term interest and optimize the posterior distribution of the interest state.
5. The data processing method for reading recommendations according to claim 1, characterized in that, In the Hamiltonian Monte Carlo sampling process, introduce an adaptive step size adjustment strategy to improve the sampling efficiency and reduce the calculation cost. The step size adjustment formula is: ∈ t+1 = ∈ t ·(1 + η·(A t - A * )); where, ∈ t+1 represents the step size of the (t + 1)-th round of sampling; ∈ t represents the step size of the t-th round of sampling; η is the step size adjustment rate; A t represents the acceptance rate of the current sampling; A * is the preset target acceptance rate.
6. The data processing method for reading recommendations according to claim 1, wherein The causal reinforcement learning method is based on the reward estimation mechanism of backward tracing, and combines the dynamic change of the interest state to optimize the recommendation strategy.
7. The data processing method for reading recommendations according to claim 1, wherein In the selection process of the recommended content, further consider the diversity constraint, and add a user exploration factor to the recommendation strategy optimization target to avoid over-recommending similar content.
8. The data processing method for reading recommendations according to claim 1, wherein The push methods of the recommended content include intelligent notification push and embedded content stream display, and the push frequency is dynamically adjusted based on user feedback.
9. The data processing method for reading recommendations according to claim 1, wherein The feedback data is used to update the user interest status, and the reinforcement learning model is trained based on the optimization objective of the reward function to improve the long-term benefits of the personalized recommendation strategy.
10. A data processing system for reading recommendations, applied to the data processing method for reading recommendations according to any one of claims 1-9, characterized in that, It includes: A data collection module, which is used to collect user behavior data and construct a time series of the user's historical behavior; An interest status inference module, which is used to infer the user interest status based on the time-varying structural causal model and update the interest status parameters; A recommendation strategy optimization module, which is used to optimize the personalized recommendation strategy based on the user interest status and calculate the causal impact of the recommended content on the user interest; A recommended content generation module, which is used to generate recommended content according to the optimized recommendation strategy and adjust it in combination with the user's historical reading preferences; A feedback update module, which is used to collect the feedback data of the user on the recommended content and update the user interest status based on the feedback data.
Citation Information
Patent Citations
Personalized learning recommendation system and method based on deep reinforcement learning
CN118628195A
Content recommendation method and system based on semantic recognition
CN119089398A
Video recommendation method and system based on mixed feedback and time sequence
CN119357428A
Sequential recommendation method based on long-term interest and short-term interest
WO2021139164A1
Cited By
Intelligent insurance business precise promotion and customer behavior analysis system and method
CN120951092A
E-commerce platform personalized commodity recommendation system based on user behavior sequence
CN122415194A