Information pushing method and device based on user behavior analysis, equipment and medium

By combining the target hierarchical spatiotemporal model with the target DQN model, the weights of user behavior characteristics and lifetime value characteristics are adjusted in real time, which solves the problem of low efficiency of information push and achieves the diversity of information push and improvement of user experience.

CN120639844AActive Publication Date: 2025-09-12QINGDAO CHENGYUN DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510913383.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

The information push efficiency in existing technologies is low and cannot respond to market changes or user behavior fluctuations in real time, resulting in a poor user experience.

Method used

An information push method based on user behavior analysis is adopted. Multimodal data sources are integrated through a target hierarchical spatiotemporal model to construct a target DQN model. The dynamic weights of user behavior instant features and user life cycle value features are adjusted in real time to determine the information push action.

Benefits of technology

It achieves the diversity and efficiency of information push, balances the CTR instant click rate and CLV customer lifetime value, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639844A_ABST
    Figure CN120639844A_ABST
Patent Text Reader

Abstract

The invention discloses an information pushing method and device based on user behavior analysis, equipment and a medium. The method comprises the following steps: acquiring a multi-modal data source for user behavior analysis, inputting the multi-modal data source into a target layered spatial-temporal model, and outputting a user state vector; fusing the user behavior instant feature and the user life cycle value feature to construct a target DQN model, and inputting the user state vector into the target DQN model; based on the user activeness and the market intensity, dynamic weights of the user behavior real-time features and the user life cycle value features are adjusted in real time, and the dynamic weights of the user behavior real-time features and the user life cycle value features are used for determining output of a target DQN model; and obtaining an output target function of the target DQN model, and determining at least one target information pushing action from the candidate information pushing actions based on the target function, thereby solving the technical problems of low information pushing efficiency and relatively low flexibility in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and more specifically, to a method, apparatus, device, and medium for pushing information based on user behavior analysis. Background Art

[0002] In existing apps, users receive push notifications as they interact with the application. Existing technologies often build user profiles based on historical data such as purchase history and browsing time, then employ collaborative filtering algorithms to generate recommendation lists. For example, they might push coupons for similar products to users who frequently purchase maternity and baby products. Alternatively, these methods pre-define user groups (such as those who haven't purchased in the past 30 days) and periodically push notifications. These methods ignore the impact of device performance and environmental characteristics on push efficiency, and due to long model update cycles, they fail to respond to market changes or fluctuations in user behavior in real time, resulting in a poor user experience.

[0003] In summary, the existing technology has technical problems of low information push efficiency and low flexibility. There is currently no effective solution to how to improve the efficiency of information push in complex and changing information flows. Summary of the Invention

[0004] The embodiments of the present application provide an information push method, apparatus, device, and medium based on user behavior analysis to at least solve the technical problem of improving the efficiency of information push during user interaction with an application.

[0005] According to one aspect of an embodiment of the present application, a method for pushing information based on user behavior analysis is provided, comprising: Responding to user behavior instructions, obtaining a multimodal data source for user behavior analysis, inputting the multimodal data source into a target hierarchical spatiotemporal model, and outputting a user state vector, wherein the target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, and is used to fuse multimodal data sources and process multiple different types of data; The target DQN model is constructed by integrating the user behavior instantaneous features and the user lifetime value features. The user state vector is input into the target DQN model, wherein the user feature vector serves as the state space of the target DQN model and the information push parameters serve as the action space of the target DQN model. Adjust the dynamic weights of user behavior features and user lifetime value features in real time based on user activity and market strength. The dynamic weights of these features are used to determine the output of the target DQN model. An output objective function of the target DQN model is obtained, and at least one target information push action is determined from candidate information push actions based on the objective function.

[0006] According to another aspect of the embodiment of the present application, there is also provided an information push device based on user behavior analysis, including: an acquisition unit, configured to respond to user behavior instructions, acquire a multimodal data source for user behavior analysis, input the multimodal data source into a target hierarchical spatiotemporal model, and output a user state vector, wherein the target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, and is configured to fuse multimodal data sources and process multiple different types of data; A fusion unit is configured to fuse the user behavior instantaneous features and the user lifetime value features to construct a target DQN model, and input the user state vector into the target DQN model, wherein the user feature vector serves as the state space of the target DQN model and the information push parameters serve as the action space of the target DQN model; An adjustment unit, configured to adjust in real time the dynamic weights of user behavior instantaneous features and user lifetime value features based on user activity and market strength, wherein the dynamic weights of the user instantaneous features and user lifetime value features are used to determine the output of the target DQN model; A determination unit is configured to obtain an output objective function of the target DQN model and determine at least one target information push action from candidate information push actions based on the objective function.

[0007] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned information push method based on user behavior analysis when running.

[0008] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned information push method based on user behavior analysis through the computer program.

[0009] In an embodiment of the present application, in response to user behavior instructions, a multimodal data source for user behavior analysis is obtained, the multimodal data source is input into a target hierarchical spatiotemporal model, and a user state vector is output, wherein the target hierarchical spatiotemporal model is divided into a multi-layer architecture by different time scales and modalities, and is used to fuse multimodal data sources and process a variety of different types of data; the instantaneous characteristics of user behavior and the characteristics of user lifetime value are integrated to construct a target DQN model, and the user state vector is input into the target DQN model, wherein the user feature vector is used as the state space of the target DQN model, and the information push parameter is used as the action space of the target DQN model; based on user activity The dynamic weights of the instantaneous features of user behavior and the features of user lifetime value are adjusted in real time based on the degree and market strength, wherein the dynamic weights of the instantaneous features of user behavior and the features of user lifetime value are used to determine the output of the target DQN model; the output objective function of the target DQN model is obtained, and at least one target information push action is determined from the candidate information push actions based on the objective function, thereby achieving the purpose of multi-source data fusion and improving the diversity of information push. The weight distribution is dynamically adjusted in real time based on user activity and market activity, thereby achieving a balance between the CTR instant click-through rate and the CLV customer lifetime value, and solving the technical problem of low information push efficiency in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a schematic diagram of a flow chart of an optional information push method based on user behavior analysis according to an embodiment of the present application; Figure 2 is a schematic diagram of an optional information push device based on user behavior analysis according to an embodiment of the present application; Figure 3 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0011] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0012] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0013] Most existing applications on the market have large user bases. As users interact with an app, their interactions—such as viewing, clicking, and purchasing—generate corresponding data, and users also receive push notifications. Existing technologies either build user profiles based on historical data such as purchase history and browsing time, and then employ collaborative filtering algorithms to generate recommendation lists. For example, they might push coupons for similar products to users who frequently purchase maternity and baby products, or they might trigger fixed coupon push notifications based on immediate actions, such as a page stay of more than 30 seconds. However, these strategies rely solely on a single behavioral dimension. Furthermore, existing push mechanisms typically use pre-defined user groups to periodically send fixed messages, or they filter targeted push notifications to high-gain users. This results in a reliance on historical data and infrequent updates. During data transmission, existing technologies typically employ a single data model, integrating only user behavior logs, such as clicks and add-to-cart items, while ignoring device performance, such as page load speed, and environmental characteristics, such as the impact of network latency on decision-making. Alternatively, they employ centralized computing architectures, requiring user behavior data to be transmitted back to a central server for processing. This poses technical challenges such as privacy risks and high response latency. Moreover, traditional deep learning models cannot explain the reasons for push notifications, resulting in low user trust and a poor user experience.

[0014] This embodiment realizes multi-source data fusion through a hierarchical spatiotemporal attention network and designs dual-objective deep learning. By constructing a target DQN model and adjusting the federated learning in real time based on user activity and market strength to dynamically adjust the weights of immediate conversion rate (CTR) and long-term user value (LTV), it can respond to market changes or user behavior fluctuations in real time, improve information push efficiency, and enhance user experience.

[0015] Alternatively, as an optional implementation, Figure 1 As shown, the information push method includes: S102, responding to the user behavior instruction, obtaining a multimodal data source for user behavior analysis, inputting the multimodal data source into a target hierarchical spatiotemporal model, and outputting a user state vector, wherein the target hierarchical spatiotemporal model is a multi-layered architecture divided by different time scales and modalities, for fusing multimodal data sources and processing multiple different types of data; S104, integrating the user behavior real-time features and the user lifetime value features to build a target DQN model, and inputting the user state vector into the target DQN model, wherein the user feature vector serves as the state space of the target DQN model, and the information push parameters serve as the action space of the target DQN model; S106, adjusting the dynamic weights of user behavior instantaneous features and user lifetime value features in real time based on user activity and market strength, wherein the dynamic weights of user instantaneous features and user lifetime value features are used to determine the output of the target DQN model; S108 , obtaining an output objective function Q of the target DQN model, and determining a target information push action from the candidate information push actions based on the objective function Q, wherein the target information push action is at least one of the candidate information push actions.

[0016] Optionally, in this embodiment, the user behavior instructions in S102 can be understood as instructions such as detecting the real-time click stream of the user terminal, the user's page dwell time, etc.; the acquisition of the multimodal data source for user behavior analysis can be, but is not limited to, collecting real-time behavior by the user terminal through the SDK, and the collected real-time behavior can be, but is not limited to, including page scrolling speed, click hot zone coordinates, GPU occupancy, GPS positioning, etc. Based on the above user behavior and environmental information, the data can be classified into multimodal data sources such as user behavior data, device and environment data, and historical preference data, among which user behavior data can be, but is not limited to, including real-time click stream, page dwell time, purchase path, etc., device and environment data can be, but is not limited to, including GPU occupancy, page loading speed, geographic location, and historical preference data can be, but is not limited to, including purchase cycles, category preferences, and other data aggregated across clients through a federated learning framework.

[0017] Optionally, in this embodiment, after determining the multimodal data source, the authors consider that any modal data source may be missing modalities due to factors such as device acquisition, resulting in data incompleteness. For example, the lack of GPS data on some clients may lead to missing device and environmental data, resulting in lower accuracy in the final result. This embodiment supplements or prioritizes the acquisition of missing modal data through Shapley value quantization and temporal feature fusion, thereby achieving the technical effect of improving data comprehensiveness.

[0018] Optionally, in this embodiment, a multimodal data source is input into a target hierarchical spatiotemporal model, and a user state vector is output, wherein the target hierarchical spatiotemporal model is divided into a multi-layer architecture by different time scales and modalities, and is used to fuse multimodal data sources and process a variety of different types of data. The target hierarchical spatiotemporal model can be, but is not limited to, a local end model based on the MHSTAN neural network, specifically including a short-term behavior layer for processing second-level operation sequences, a long-term periodic layer for analyzing historical periodic features, and cross-modal attention for integrating device performance. The outputs of each layer interact through cross-layer attention or feature splicing. For example, long-term periodic features are used as context inputs for the short-term behavior layer, and all layers are trained jointly rather than independently trained and then spliced. The target hierarchical spatiotemporal model outputs a user state vector ,in In mathematics, it represents an n-dimensional real vector space, where each component of each vector is a real number. represents the user's state vector at time t. It should be noted that the choice of 512 dimensions is an empirical one, striking a balance between expressiveness and computational costs, such as memory and communication overhead. In addition, high-dimensional space can more finely encode complex features, such as user behavior and image semantics.

[0019] Optionally, in this embodiment, in S104, a target DQN model is constructed based on a dual-objective machine learning strategy of user behavior instantaneous features (CTR) and user lifetime value features (LTV), and the user feature vector As the state space of the model, push information to the parameters As the action space of the model, the reward function of the model is dynamically adjusted. The formula for dynamic adjustment of the reward function is as follows: ; in, Indicates the user activity weight, which is used to control the user activity weight The impact of the intensity, for example, the business focuses on long-term retention, The value will be increased, so that highly active users can trigger the user behavior instant feature (CTR) weight more quickly; The weight of market strength is used to reflect the intervention of external marketing activities on strategy adjustment. For example, during the Double 11 promotion period, The value temporarily increases, and even highly active users need to maintain a high user behavior instant feature (CTR) weight to sprint the short-term total merchandise transaction volume (GMV).

[0020] Optionally, in this embodiment, in S106, the dynamic weights of the user behavior instant features and the user lifetime value features are adjusted in real time based on user activity and market strength. The dynamic weights of the user instant features and the user lifetime value features are used to determine the output of the target DQN model, which may include, but is not limited to, using an MLP neural network to analyze the user behavior instant features and an LSTM neural network to analyze the user lifetime value features. By adjusting the weights to adjust the output weights of the MLP neural network and the LSTM neural network, the objective function calculation value, i.e., the output of the target DQN model, is ultimately affected.

[0021] Optionally, in this embodiment, in S108, a target information push action is determined from candidate information push actions based on the objective function. For example, the candidate actions are pushing clothing coupons and pushing daily necessities discount coupons. The objective function of pushing clothing coupons is greater than the objective function of pushing daily necessities discount coupons. Then, the clothing coupon push action is selected as the target information push action, and there is a high probability that clothing coupon information is pushed to the user, and there is a low probability that daily necessities discount coupon information is attempted to be pushed.

[0022] Through this embodiment, in response to user behavior instructions, a multimodal data source for user behavior analysis is obtained, the multimodal data source is input into the target hierarchical spatiotemporal model, and a user state vector is output, wherein the target hierarchical spatiotemporal model is divided into a multi-layer architecture by different time scales and modalities, and is used to fuse multimodal data sources and process a variety of different types of data; the instantaneous characteristics of user behavior and the characteristics of user life cycle value are integrated to construct a target DQN model, and the user state vector is input into the target DQN model, wherein the user feature vector is used as the state space of the target DQN model, and the information push parameter is used as the action space of the target DQN model; based on user activity and Market strength adjusts the dynamic weights of user behavior instant features and user lifetime value features in real time, where the dynamic weights of user instant features and user lifetime value features are used to determine the output of the target DQN model; the output objective function of the target DQN model is obtained, and at least one target information push action is determined from the candidate information push actions based on the objective function, thereby achieving the purpose of multi-source data fusion and improving the diversity of information push. The weight distribution is dynamically adjusted in real time based on user activity and market activity, thereby achieving the purpose of balancing CTR instant click-through rate and CLV customer lifetime value, and achieving the technical effect of improving the efficiency of information push.

[0023] Optionally, as an optional implementation method, the dynamic weights of user behavior instant features and user lifetime value features are adjusted in real time based on user activity and market strength, including: S1, determining the user's activity level based on user behavior information, wherein the user behavior information includes user behavior frequency, number of user deep interactions, and number of user core actions; S2, obtains user activity and market strength; S3: When the user's activity level exceeds the target high activity threshold, the weight formula is triggered to adjust the dynamic weights of the user's immediate behavior features and the user's lifetime value features; S4, when the market activity is greater than the target high activity value, trigger the weight formula to adjust the dynamic weights of the user behavior immediate characteristics and the user life cycle value characteristics.

[0024] Optionally, in this embodiment, whether a user is highly active can be determined based on user behavior information such as user behavior frequency, number of user deep interactions, and number of user core actions. For example, a user who meets at least two of the following criteria is identified as a highly active user: (1) Behavior frequency threshold: Daily activity: logged in ≥5 days in the past 7 days, or average daily interaction times ≥10 times (such as clicks, favorites), Weekly activity: meeting the above daily activity standards in ≥4 weeks in the past 30 days; (2) Deep participation indicators: Content production: publishing ≥3 UGC contents per week (such as comments, notes), social interaction: average daily private messages / sharing behaviors ≥5 times; (3) Business core actions: average weekly add-to-cart times ≥8 times, or monthly repurchase rate ≥40%.

[0025] Optionally, in this embodiment, the market activity level can also be determined by the frequency of market activities and the strength of discounts. When the market activity level or the user activity level is greater than the target high activity threshold, a weight formula is triggered to adjust the dynamic weights of the user behavior immediate features and the user life cycle value features. The weight formula is the reward function:

[0026] As an optional embodiment, for example, the local model of each user device receives global parameters and calculates the local user behavior immediate feature CTR and user lifetime value LTV estimate based on user behavior such as clicks and add-to-cart and long-term values ​​such as historical ARPU and repurchase cycle. The user's activity level and market strength are obtained. If a user has added 10 items in the past week, the user's activity level is greater than the target high activity threshold, triggering the weight formula: ; Among them, the parameter setting (High activity Decline fast), (Weak impact of market activities); The final result is When (t)=0.3 (originally 0.7), the CTR weight dropped from 70% to 30%, while the LTV weight increased to 70%. This achieved the goal of dynamically adjusting the weights of user behavior characteristics based on user activity and market strength in real time, both for immediate and lifetime value. The ultimate result was that the server aggregated the low CTR weights of highly active users, generating a global model that pushed "member-only benefits" rather than "limited-time discounts" to these users, prioritizing long-term retention. This resulted in a 20% decrease in short-term click-through rates, but a 35% increase in 90-day repurchase rates. The cumulative LTV value in the market increased by 18%, extending the customer lifecycle. This achieved a balance between immediate and lifetime value characteristics, improving information push efficiency.

[0027] Optionally, as an optional implementation manner, obtaining the output objective function Q of the target DON model includes: Calculate the objective function Q: ; in, is the action value function for selecting action a in the current state s, It is the dynamic weight of the immediate characteristics of user behavior and the characteristics of user lifetime value. MLP is a neural network used to analyze the immediate characteristics of user behavior. MLP(S) is the output parameter of the MLP neural network processing the current state S. LSTM is a neural network used to analyze the characteristics of user lifetime value. LSTM(CLV) is the output parameter of the LSTM neural network processing the current state S.

[0028] Optionally, in this embodiment, the local DQN is based on Generate candidate action sets , calculate Q value, dynamic weight Adaptively adjusted by user activity, actions can be selected through, but are not limited to, the ε-greedy strategy: high-activity users: ε = 0.1 (focused on exploitation); low-activity users: ε = 0.5 (focused on exploring new strategies).

[0029] As a specific embodiment, the Q value is calculated for highly active users with ε=0.1: First, obtain the user profile, that is, obtain the user ID as A, obtain the user's activity level, that is, obtain the user who logs in every day, spends an average of 5,000 yuan per month, and has a historical preference for "maternal and child products." Generate a candidate action set based on the action instructions in the current state: Candidate coupon types: {maternal and child discount coupons, daily necessities discount coupons, clothing coupons, electronic product coupons}. Calculate the Q value to determine if the user is a highly active user, and then dynamically weight the user's immediate behavior characteristics and user life cycle value characteristics. The value is 0.9. At this point, the output of the MLP neural network, which analyzes the real-time characteristics of user behavior, dominates. MLP input: The user is currently browsing the stroller page → The prediction for "Mother and Baby Coupon" is the highest Q value. The LSTM input, historical CLV, shows that the mother and baby category has a high repurchase rate → The long-term value of "Mother and Baby Coupon" is strengthened.

[0030] The final Q value ranking is: maternal and child discount coupons > daily necessities coupons > clothing coupons > electronic product coupons.

[0031] Targeted information push action selection: 90% probability of selecting the action with the highest Q value: maternity and baby discount coupons, 10% random try other coupons such as electronics coupons. The final result accurately matches the user's known preferences, improving short-term conversion rates and avoiding strategy rigidity with limited exploration. This improves information push efficiency and user experience.

[0032] As a specific embodiment, the Q value is calculated for highly active users with ε=0.5 who focus on exploring new strategies: First, obtain a user profile, specifically user ID B, and determine the user's activity level. This means a new user has only purchased low-priced daily necessities once, resulting in a low predicted CLV. A candidate action set is generated based on the action instructions in the current state: candidate coupon types: {daily necessities coupons, beauty coupons, sports equipment coupons, book coupons}. The Q-value is calculated, and if the user is considered low-activity, the dynamic weight of the user's immediate behavior features and the user's lifetime value features is 0.2. In this case, the LSTM neural network used to analyze the user's lifetime value features has the dominant weight. MLP input: The user recently searched for "yoga mats" → the Q-value for "sports equipment coupons" is predicted to be higher; LSTM input: The CLV prediction model finds that "beauty users" have higher long-term value → the weight for "beauty coupons" is increased.

[0033] Final Q value ranking: beauty coupons > sports equipment coupons > daily necessities coupons > book coupons.

[0034] Targeted information push action selection: 50% probability is to select the action with the highest Q value, namely beauty coupons, and 50% probability is to randomly try to push other coupons, such as sports equipment coupons and book coupons. The final results are used to proactively explore potential user interests in areas such as beauty and sports, quickly accumulate behavioral data, and optimize long-term CLV (customer lifetime value). This improves information push efficiency and user experience.

[0035] Optionally, as an optional implementation, before inputting the multimodal data source into the target hierarchical spatiotemporal model, constructing the target hierarchical spatiotemporal model includes: S1, constructs a short-term behavior layer by processing second-level operation sequences through a bidirectional gated recurrent neural network; S2, constructs a long-term cycle layer by analyzing historical purchase intervals through fast Fourier transform and neural network; S3, fuses device and environment features through cross-modal attention and dynamically assigns weights.

[0036] Optionally, in this embodiment, a lightweight model-targeted hierarchical spatiotemporal model is constructed and executed on the device side. Specifically, it includes a short-term behavior layer, a long-term periodic layer, and cross-modal attention. The short-term behavior layer uses a bidirectional gated recurrent neural network (BiGRU) to process second-level operation sequences, such as processing the features of consecutive clicks on products in the same category. The long-term periodic layer uses a fast Fourier transform and neural network (FFT+MLP) to analyze historical purchase intervals, such as extracting weekly and monthly purchase frequency cycle features. Cross-modal attention (MHA) integrates device and environmental features and dynamically assigns weights. For example, when memory is low, the weight of real-time features is reduced to reduce memory usage.

[0037] Optionally, in this embodiment, the layered spatiotemporal model features layered joint optimization, where the outputs of each layer interact through cross-layer attention or feature concatenation. For example, long-term periodic features serve as contextual input to the short-term behavior layer. The layered spatiotemporal model features end-to-end training: all layers are trained jointly, rather than independently trained and then concatenated. For example, the output of the short-term behavior layer and the fast Fourier transform features are jointly input to the attention layer.

[0038] It should be noted that the hierarchical spatiotemporal model features topology-aware hypernetworks and conditional fusion. It dynamically generates attention weight matrices based on device performance, such as memory usage, rather than statically adjusting them. For example, when GPU usage exceeds 80%, the update frequency of real-time behavioral features is reduced while the weight of long-term periodic features is increased. To account for the impact of device performance on information push, the weights and update frequency are adjusted in real time based on GPU usage data.

[0039] Optionally, in this embodiment, the hierarchical spatiotemporal model considers both temporal proximity, such as the timing of operation sequences, and spatial correlation, such as similar behaviors of geographically close users, in its attention mechanism. The model uses hierarchical spatiotemporal embedding to encode timestamps and spatial coordinates into a joint vector that is input into each layer of the model.

[0040] Optionally, as an optional implementation, obtaining a multimodal data source for user behavior analysis includes: S1, collecting multimodal data sources, where the multimodal data sources include user behavior data, device and environment data, and historical preference data; S2, when any modality data source is missing, the impact weight of the missing modality data source is quantified by the Shapley value; S3, prioritize the collection or compensation of missing modality data sources with higher weights.

[0041] Optionally, in this embodiment, the user end collects real-time behavior through the SDK to obtain a multimodal data source. Based on the collected real-time user behavior instructions and environmental information, the multimodal data source can be categorized into user behavior data, device and environment data, and historical preference data. To improve the comprehensiveness of data collection, if any modal data source is missing, the impact weight of the missing modal data source is quantified using the Shapley value.

[0042] It's important to note that the Shapley value is derived from game theory and is used to fairly distribute the benefits of cooperation to participating modalities, contributing to the final model performance. In this example, the Shapley value quantifies the contribution of each modality to the model's prediction results. For example, device performance data contributes 20% to the push decision.

[0043] Optionally, in this embodiment, the influence weight of the actual modality on the task is quantified by the Shapley value, so that the actual modality with high weight can be prioritized for data collection, or the missing modality can be interpolated through the time feature fusion method to ensure the comprehensiveness of the data.

[0044] Optionally, in this embodiment, when the target model data on the local side is sent to the server, the server performs federated model aggregation. The target global model parameters are the model parameters after federated model aggregation, and the server then sends the target global model parameters to the client. The server uses a modal contribution quantization aggregation algorithm to dynamically adjust the client weights: Optionally, in this embodiment, Indicates the weight value of each client. According to the above formula, The client's local model parameters have a greater impact on the global model, so federated resource allocation is designed in the model data transmission from the server to the client, which prioritizes computing resources and allocates more bandwidth or training time to high-weight clients, thereby improving the utilization of data transmission.

[0045] Optionally, in this embodiment, after each round of training, global knowledge is transferred to the local model through mutual knowledge distillation (MKD) to prevent forgetting. For example, high-value user features need to be retained long-term. Mutual knowledge distillation is used to distill high-value features from the global model, such as long-term user preferences, into the local model, avoiding the loss of important information due to dynamic weight adjustments. Even if a client's weight decreases, its previously contributed knowledge is retained in the global knowledge base through MKD, improving data integrity and security.

[0046] Optionally, as an optional implementation method, the target DQN model is constructed by integrating the user behavior instantaneous features and the user lifetime value features, which then includes: S1, adds noise to the gradient of the target DQN model parameter data to ensure that a single parameter data cannot be reversely inferred; S2, encrypts the perturbed gradient using an encryption algorithm, determines the generated target ciphertext, and sends the target ciphertext to the server; S3, the server aggregates all target ciphertexts and updates the global model parameters after decryption; S4, the server sends the global model parameters to the client.

[0047] Optionally, in this embodiment, it can be understood as regularly uploading encrypted gradients To the parameter server, the global model is aggregated and sent down. Specifically, each device adds noise, such as Laplace noise, to the gradient of the target DQN model parameter data to ensure that a single parameter data cannot be reversely inferred, control the privacy budget, and set privacy parameters. To limit the risk of information leakage in each round of training, use encryption algorithms such as the additive homomorphic public key encryption algorithm (Paillier) to encrypt the gradient rather than the original data to generate the target ciphertext , the server directly sums the encrypted gradients to get the aggregated result The server decrypts the aggregated gradients, updates the global model, and distributes them to all devices. By adopting differentially private federated learning, user data is processed locally, and only gradients are encrypted for transmission, rather than uploading the raw data directly to the server. This overcomes the potential for data leakage and reduces end-side inference latency to 50ms through feature binning, improving transmission efficiency.

[0048] As an optional embodiment, in order to avoid pushing high-denomination coupons (abnormal behavior) to users with low user lifetime value (CLV), the steps are as follows: First, user A's device uses local data (purchase history, click behavior) to train a lightweight DQN model and generate gradients. , and then perform differential privacy processing: add Laplace noise to the gradient ,in is the random noise added to meet the differential privacy requirements, Indicates that the noise scale parameter is 0.5, which is related to data sensitivity and privacy budget. represents the privacy budget, which controls the privacy protection strength. The smaller it is, the stronger the privacy protection is, but the louder the noise is.

[0049] Then use the Paillier public key to encrypt the perturbed gradient E( Upload to the server; the server performs aggregation and anomaly detection, aggregating all encrypted gradients , and update the global model after decryption.

[0050] Optionally, this embodiment implements asynchronous updates and offline support. Local caching can be, but is not limited to, understanding the device caching decision logs during network outages, such as user B's push history, which is encrypted and uploaded after the network is restored. Delayed synchronization can be, but is not limited to, understanding the server aggregating delayed updates from offline devices according to a time window every six hours to ensure eventual consistency. This embodiment achieves the technical effect of maintaining user-level privacy in weak network environments without losing key behavioral data.

[0051] Optionally, as an optional implementation method, the server sends the global model parameters to the client, and then further includes: S1, calculates the federated target value and determines the target attribution report, where the target attribution report is used to explain push decisions in real time; S2, determining the target interception value based on the global model parameters sent by the server; S3: When the federal target value of the client is less than the target interception value, the target information push action is intercepted.

[0052] Optionally, the federated DeepSHAP target value is calculated, a target attribution report is determined, and a target interception value is determined based on the global model parameters. If the client's federated target value is less than the target interception value, an interception is performed. That is, if the client's DeepSHAP target value is detected to be less than the target interception value, the rule engine is triggered to intercept abnormal push notifications, such as high-value coupons being pushed to low-value users. The target attribution report is used to explain push notification decisions in real time. The federated DeepSHAP real-time attribution formula is as follows: Local calculation of feature contribution ; in, is the feature sensitivity, Local feature activation strength, the formula calculates the feature contribution by multiplying two parts The sensitivity of the function to the feature and the strength of local feature activation. The global value function in federated learning evaluates the expected benefits of taking action a in state s, such as model prediction accuracy, collaboration efficiency, etc.

[0053] As an optional embodiment, a high-risk example: the action is to push a 500 yuan high-priced commodity voucher to user B whose historical purchasing power is ≤ 100 yuan; the feature contribution analysis is the user purchasing power contribution Abnormal features, where the global mean = 0.6, ; The interception rule is <0.6−2×0.15=0.3 , which means the push is blocked. This enables server-side aggregation of attribution results, triggering the rule engine to block high-risk actions.

[0054] Optionally, in this embodiment, each device periodically calculates the DeepSHAP target value and generates an attribution report, such as "This push is due to the user frequently browsing this category in the past three days", which is used to explain the reason for the push and intercept abnormal decisions based on the global mean and interception rules, such as pushing high-value coupons to low-value users.

[0055] Optionally, as an optional implementation method, training can be performed in stages during the model building phase, using simulated behavior data for pre-training during the cold start phase, and gradually introducing real interaction data during the mature phase to avoid distribution shift.

[0056] Optionally, as an optional implementation method, during the online federated learning process, whether the user behavior changes is monitored at all times, and interest drift can be detected through KL divergence in the local model. Specifically, the client monitors the KL divergence changes of the behavior distribution in real time.

[0057] ; in, The definition is Kullback-Leibler divergence (KL divergence), which is used to measure the difference between two probability distributions. and The degree of difference between The definition is the distribution of new behaviors currently monitored by the client, such as user clicks, purchases, and other interaction data, which is used to reflect the latest status of system / user behavior; The definition is the historical baseline distribution, the distribution of behavior during the training phase or the last fine-tuning, which is used as a "normal state" reference for comparison; Definition is a preset threshold used to determine whether the distribution change is significant. It needs to be determined through experiments or business needs.

[0058] As an optional embodiment, when the end-side inference delay, for example, the target is <50ms, and the federated aggregation convergence speed, for example, the target error within 5 rounds is <1%, performance testing is performed. Optionally, it can be understood as verifying whether the new solution of the interest drift detection model is better than the traditional collaborative filtering solution, reducing the risk of full online launch. For example, if the user feedback score of the old solution is higher than the new solution for three consecutive days, it will automatically roll back to the old solution. For example, if the performance monitoring triggers an alarm if the single-point delay is greater than 50ms, and the backup server node is switched after three consecutive rounds of communication failures.

[0059] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0060] According to another aspect of the embodiment of the present application, an information push device for implementing the above-mentioned information push method is also provided. Figure 2 As shown, the device includes: A fusion unit 202 is configured to respond to user behavior instructions, fuse target multi-source data, and construct a target hierarchical spatiotemporal model. The target hierarchical spatiotemporal model is divided into multiple layers by different time scales and modalities, and is configured to process a variety of different types of data. An adjustment unit 204 is configured to determine values ​​of a first indicator and a second indicator based on a target global model parameter, and dynamically adjust weights of the first indicator and the second indicator based on user activity and market activity, wherein the target global model parameter is the aggregated model parameter of the federated model, the first indicator represents the immediate conversion rate, and the second indicator represents the long-term user value; The determination unit 206 is configured to determine a target execution action according to a target dynamic scheduling algorithm, wherein the target execution action is at least one of the candidate execution actions.

[0061] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned information push method is also provided, such as Figure 3 As shown, the electronic device includes a memory 302 and a processor 304. The memory 302 stores a computer program, and the processor 304 is configured to execute the steps in any of the above method embodiments through the computer program.

[0062] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0063] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program: S1, responds to user behavior instructions, obtains a multimodal data source for user behavior analysis, inputs the multimodal data source into a target hierarchical spatiotemporal model, and outputs a user state vector. The target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, which is used to fuse multimodal data sources and process multiple different types of data; S2, integrates the user behavior real-time features and user lifetime value features to build a target DQN model, and inputs the user state vector into the target DQN model, where the user feature vector serves as the state space of the target DQN model and the information push parameters serve as the action space of the target DQN model; S3, based on user activity and market strength, adjusts the dynamic weights of user behavior instant features and user lifetime value features in real time. The dynamic weights of user instant features and user lifetime value features are used to determine the output of the target DQN model; S4: Obtain an output objective function of the target DQN model, and determine at least one target information push action from the candidate information push actions based on the objective function.

[0064] Alternatively, those skilled in the art will appreciate that Figure 3 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 3 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 3 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 3 Different configurations shown.

[0065] Among them, the memory 302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the information push method and device in the embodiments of the present application. The processor 304 executes various functional applications and data processing by running the software programs and modules stored in the memory 302, that is, to realize the above-mentioned information push method based on user behavior analysis. The memory 302 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 302 may further include a memory remotely located relative to the processor 304, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 302 can be used specifically, but not limited to, to store information such as candidate information push actions. As an example, such as Figure 3 As shown, the memory 302 may include, but is not limited to, the acquisition unit 202, fusion unit 204, adjustment unit 206, and determination unit 208 in the information push device. In addition, it may also include, but is not limited to, other module units in the information push device based on user behavior analysis, which will not be repeated in this example.

[0066] Optionally, the transmission device 306 is configured to receive or transmit data via a network. Specific examples of the aforementioned network may include wired networks and wireless networks. In one embodiment, the transmission device 306 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to enable communication with the Internet or a local area network. In one embodiment, the transmission device 306 is a radio frequency (RF) module configured to communicate with the Internet wirelessly.

[0067] In addition, the electronic device further includes: a display 308 for displaying information such as the target information push action; and a connection bus 310 for connecting various module components in the electronic device.

[0068] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when run.

[0069] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps: S1, responds to user behavior instructions, obtains a multimodal data source for user behavior analysis, inputs the multimodal data source into a target hierarchical spatiotemporal model, and outputs a user state vector. The target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, which is used to fuse multimodal data sources and process multiple different types of data; S2, integrates the user behavior real-time features and user lifetime value features to build a target DQN model, and inputs the user state vector into the target DQN model, where the user feature vector serves as the state space of the target DQN model and the information push parameters serve as the action space of the target DQN model; S3, based on user activity and market strength, adjusts the dynamic weights of user behavior instant features and user lifetime value features in real time. The dynamic weights of user instant features and user lifetime value features are used to determine the output of the target DQN model; S4: Obtain an output objective function of the target DQN model, and determine at least one target information push action from the candidate information push actions based on the objective function.

[0070] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0071] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0072] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.

[0073] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0074] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, there may be other division methods, such as combining or integrating multiple units or components into another system, or ignoring or not implementing some features. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of units or modules, and may be electrical or other forms.

[0075] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0076] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0077] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An information push method based on user behavior analysis, characterized in that: The method comprises: Responding to user behavior instructions, obtaining a multimodal data source for user behavior analysis, inputting the multimodal data source into a target hierarchical spatiotemporal model, and outputting a user state vector, wherein the target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, and is used to fuse multimodal data sources and process multiple different types of data; A target DQN model is constructed by fusing the user behavior instantaneous features and the user lifetime value features, and the user state vector is input into the target DQN model, wherein the user feature vector serves as the state space of the target DQN model and the information push parameters serve as the action space of the target DQN model; Adjust the dynamic weights of user behavior instant features and user lifetime value features in real time based on user activity and market strength, where the dynamic weights of user instant features and user lifetime value features are used to determine the output of the target DQN model; An output objective function of the target DQN model is obtained, and at least one target information push action is determined from candidate information push actions based on the objective function.

2. The information push method based on user behavior analysis according to claim 1, characterized in that: The dynamic weighting of user behavior instantaneous features and user lifetime value features is adjusted in real time based on user activity and market strength, including: Determining the user's activity level based on the user behavior information, wherein the user behavior information includes user behavior frequency, number of user deep interactions, and number of user core actions; Obtaining the user activity level and market strength; When the user activity level is greater than the target high activity threshold, a weight formula is triggered to adjust the dynamic weights of the user behavior instant features and the user life cycle value features; When the market strength is greater than the target high activity threshold, a weight formula is triggered to adjust the dynamic weights of the user behavior instant features and the user life cycle value features.

3. The information push method based on user behavior analysis according to claim 2, characterized in that: Obtaining the output objective function Q of the target DON model includes: Calculate the objective function Q: ; Among them, the is the action value function for selecting action a in the current state s, It is the dynamic weight of the user behavior instant feature and the user life cycle value feature. The MLP is a neural network used to analyze the user behavior instant feature. The MLP(S) is the output parameter of the MLP neural network processing the current state S. The LSTM is a neural network used to analyze the user life cycle value feature. The LSTM(CLV) is the output parameter of the LSTM neural network processing the current state S.

4. The information push method based on user behavior analysis according to claim 1, characterized in that: Before inputting the multimodal data source into the target hierarchical spatiotemporal model, it is characterized in that constructing the target hierarchical spatiotemporal model includes: A short-term behavior layer is constructed by processing second-level operation sequences through a bidirectional gated recurrent neural network; Constructing long-term cycle layers by analyzing historical purchase intervals using Fast Fourier Transform and neural networks; Device and environment features are fused through cross-modal attention and weights are dynamically assigned.

5. The information push method based on user behavior analysis according to claim 1, characterized in that: The obtaining of a multimodal data source for user behavior analysis includes: Collecting multimodal data sources, wherein the multimodal data sources include user behavior data, device and environment data, and historical preference data; In the case that any modality data source is missing, the influence weight of the missing modality data source is quantified by the Shapley value; The missing modality data sources with higher weights are collected or compensated first.

6. The information push method based on user behavior analysis according to any one of claims 1 to 5, characterized in that: The target DQN model is constructed by integrating the immediate characteristics of user behavior and the characteristics of user lifetime value, which includes: Adding noise to the gradient of the target DQN model parameter data to ensure that a single parameter data cannot be reversely inferred; Encrypt the perturbed gradient using an encryption algorithm, determine to generate a target ciphertext, and send the target ciphertext to a server; The server aggregates all the target ciphertexts and updates the global model parameters after decryption; The server sends the global model parameters to the client.

7. The information push method based on user behavior analysis according to claim 6, characterized in that: The server sends the global model parameters to the client, and then further includes: Calculate the federation target value and determine the target attribution report, wherein the target attribution report is used to explain the push decision in real time; Determining a target interception value based on the global model parameters sent by the server; When the federal target value of the client is less than the target interception value, the target information push action is intercepted.

8. An information push device based on user behavior analysis, characterized in that: include: an acquisition unit, configured to respond to user behavior instructions, acquire a multimodal data source for user behavior analysis, input the multimodal data source into a target hierarchical spatiotemporal model, and output a user state vector, wherein the target hierarchical spatiotemporal model has a multi-layer architecture divided by different time scales and modalities, and is configured to fuse multimodal data sources and process multiple different types of data; a fusion unit, configured to fuse user behavior instantaneous features and user lifetime value features to construct a target DQN model, and input the user state vector into the target DQN model, wherein the user feature vector serves as the state space of the target DQN model, and the information push parameters serve as the action space of the target DQN model; An adjustment unit, configured to adjust in real time the dynamic weights of user behavior instantaneous features and user lifetime value features based on user activity and market strength, wherein the dynamic weights of user instantaneous features and user lifetime value features are used to determine the output of a target DQN model; A determination unit is configured to obtain an output objective function of the target DQN model and determine at least one target information push action from candidate information push actions based on the objective function.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Weight adaptive feedback adjustment method and device for information pushing

    CN109284437A

  • Push strategy acquisition method and device, equipment and storage medium

    CN114756756A

  • Product pushing method and device, computer equipment and storage medium

    CN116797316A

  • Intranet service quality optimization method and system based on deep reinforcement learning

    CN119496716A

  • Activity matching method and system based on user behaviors

    CN120146959A

Cited By

  • Message display method and device, equipment and storage medium

    CN122053551A