A design method for dynamic incentive mechanism of Internet points based on reinforcement learning

Through context-adaptive reward algorithms and multiple optimization algorithms, combined with user behavior data and contextual tags, the Internet points incentive mechanism is dynamically adjusted, which solves the static and single problems of the traditional points incentive mechanism, realizes the personalization and real-time optimization of the reward strategy, and improves user participation and loyalty.

CN120258886BActive Publication Date: 2025-09-30YUEJI ENTERPRISE MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510291778.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-09-30
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing Internet points incentive mechanism lacks flexibility and targeting, fails to fully consider the diversity of user behavior, situational changes and personalized needs, resulting in difficulty in maintaining user engagement and loyalty in the long term, and fails to effectively utilize the time series characteristics and situational information of user behavior.

Method used

It adopts a context-adaptive reward algorithm, a gated recurrent unit model, a chaotic search strategy, and a wolf pack optimization algorithm, and combines user behavior data, emotional state, and context labels to dynamically adjust the reward strategy. It optimizes reward calculation through a fuzzy logic system and feedback loop to achieve personalized and real-time adjustment of rewards.

Benefits of technology

It improves the accuracy and flexibility of the reward strategy, enhances user participation and satisfaction, improves user activity and loyalty, ensures that the reward mechanism adapts to users' long-term needs and behavioral changes, and optimizes the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258886B_ABST
    Figure CN120258886B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for designing a dynamic incentive mechanism for internet points based on reinforcement learning, comprising the following steps: S1, constructing a user behavior dataset; S2, generating context labels using natural language processing technology; S3, calculating reward values ​​using a context-adaptive reward algorithm; S4, dynamically adjusting the reward strategy based on the reward value and user behavior feedback, combined with the user's task participation and feedback loop; S5, predicting the user's future behavior trends using a gated recurrent unit model and generating user behavior prediction results; S6, dynamically optimizing the reward strategy using a chaotic search strategy and a wolf pack optimization algorithm to obtain an optimized reward strategy; S7, evaluating the effectiveness of the rewards and changes in user interests, and obtaining a final global reward strategy based on the A3C algorithm, thereby achieving real-time dynamic adjustment of the internet points incentive mechanism. The present invention utilizes a context-adaptive reward algorithm and optimization techniques, among others, to achieve dynamic adjustment of the internet points incentive strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of incentive mechanism design, and in particular to a method for designing an Internet points dynamic incentive mechanism based on reinforcement learning. Background Art

[0002] In the modern internet landscape, point-based incentive mechanisms have become widely used across various online platforms, particularly in e-commerce, social networking, and mobile applications, as an effective means of engaging users, increasing loyalty, and guiding behavior. Point-based rewards are often closely tied to user behavior. Users earn points by completing tasks and participating in activities, and these points can be redeemed for cash, coupons, or other rewards. However, traditional point-based incentive mechanisms often rely on fixed rules and fail to fully account for the diversity of user behavior, changing contexts, and personalized user needs. As user behavior becomes increasingly complex, traditional point-based incentive mechanisms face numerous challenges, necessitating innovative solutions.

[0003] Existing point-based incentive mechanisms often rely on simple rules and algorithms, such as awarding rewards based on user activity frequency, task completion, or simple historical data. While these methods can achieve a certain degree of motivation, they lack flexibility and targeting, making it difficult to motivate users over the long term. More importantly, most existing methods ignore users' individual needs and contextual factors, relying solely on a single reward mechanism and failing to fully leverage user behavioral feedback to dynamically adjust reward strategies. This makes it difficult to maintain long-term user engagement and loyalty.

[0004] Furthermore, existing technologies often employ static algorithmic models when processing user behavior data, ignoring the time series nature of user behavior. User behavior is time-dependent and can change significantly over time. Traditional reward strategies fail to fully account for this dynamic nature, potentially weakening their effectiveness over time. For example, some platforms may use simple rules to award a fixed number of points each month. While this approach can stimulate user activity in the short term, over time, users may experience a lack of novelty or insufficient motivation, leading to reduced engagement.

[0005] Against this backdrop, the shortcomings of traditional point-based incentive mechanisms are becoming increasingly apparent. Simple rules are no longer sufficient, especially in environments with large user bases, personalized needs, and changing contexts. With the development of artificial intelligence (AI), particularly reinforcement learning, natural language processing, and optimization algorithms, new approaches to incentive mechanism design have become a hot topic. These technologies enable dynamic modeling of user behavior, personalized adjustments to reward mechanisms, and contextual awareness, providing internet platforms with more intelligent and efficient incentive mechanism designs.

[0006] Reinforcement learning, a technology capable of self-adjusting based on user feedback, has garnered widespread attention in recent years. Through reinforcement learning, the system can dynamically adjust reward strategies based on historical user behavior and feedback, achieving more precise user incentives. For example, by calculating reward values ​​based on user behavioral characteristics and contextual information, and employing a context-adaptive reward algorithm, the reward mechanism can be adaptively adjusted to meet user needs in different contexts. The introduction of fuzzy logic and optimization algorithms can further enhance the flexibility and accuracy of reward calculation. However, existing reinforcement learning applications are mostly concentrated in specific fields, and there is still a lack of in-depth exploration and implementation of its practical application in internet point-based incentive mechanisms.

[0007] Furthermore, as users' emotional states and behavioral complexity increase, contextual information becomes increasingly important in point-based incentive mechanisms. Existing technologies generally fail to consider the impact of context on reward effectiveness, nor do they effectively utilize contextual tags to personalize user reward needs. By employing natural language processing technology and contextual tag generation methods, the system's understanding of user behavior can be further improved, enabling the generation of reward strategies more tailored to user needs and avoiding the inadequacy of single, fixed rules across all users.

[0008] However, even with this, existing approaches to designing incentive mechanisms based on reinforcement learning still have some limitations. For example, while reinforcement learning can dynamically adjust based on user feedback, managing the diversity of different user needs and behaviors while ensuring global optimization remains a challenge. Existing technologies often use a single optimization algorithm and fail to fully integrate advanced optimization methods such as chaotic search strategies and wolf pack optimization algorithms. These algorithms typically focus on balancing global search and local optimization, but in practical applications, dynamically adjusting reward strategies in multi-objective optimization problems remains a challenge.

[0009] At the same time, user behavioral feedback often presents complexities in practical applications. User feedback encompasses not only reward acceptance but also task engagement, point usage, emotional state, and other multi-dimensional information. Existing technologies typically focus solely on basic user behavioral data, such as activity frequency and task completion, while ignoring the impact of factors like emotional state and evolving interests on reward mechanisms. This makes it difficult for existing incentive mechanisms to make timely and precise adjustments to meet users' individual needs and changing contexts.

[0010] Therefore, how to provide a method for designing a dynamic incentive mechanism for Internet points based on reinforcement learning is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0011] One purpose of the present invention is to propose a method for designing a dynamic incentive mechanism for Internet points based on reinforcement learning. The present invention makes full use of the context-adaptive reward algorithm, the gated recurrent unit model, the chaotic search strategy, and the wolf pack optimization algorithm, and describes in detail the design method for a dynamic incentive strategy for Internet points based on user behavior prediction. By combining the user's behavioral data, emotional state, and contextual labels, the reward strategy is dynamically adjusted to ensure the long-term stability and adaptability of the strategy. This invention has significant advantages in improving user engagement, enhancing the accuracy of the reward strategy, and optimizing the user experience. It can respond to user needs in real time and further improve the user activity and loyalty of the Internet platform.

[0012] A method for designing a dynamic incentive mechanism for Internet points based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0013] S1. Collect user behavior data through sensors and mobile devices and pre-process it to build a user behavior dataset;

[0014] S2. Use natural language processing technology to obtain the user's emotional state and context information and generate context labels;

[0015] S3. Calculate the reward value using a context-adaptive reward algorithm, dynamically adjust the reward intensity and form based on the user behavior dataset and context labels, and introduce a fuzzy logic system into the reward generation;

[0016] S4. Adjust the type and frequency of rewards based on the generated reward values ​​and user behavior feedback, dynamically adjusting the reward strategy based on the user's task participation and feedback loop, which is optimized based on the user's reward acceptance and behavior changes;

[0017] S5. Based on the adjusted reward strategy, use the gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, engagement, and points of interest, and generate user behavior prediction results;

[0018] S6. Based on the user behavior prediction results, the reward strategy is dynamically optimized by combining the chaotic search strategy and the wolf pack optimization algorithm. The optimization process simulates the wolf pack's foraging behavior to find the optimal strategy in the multi-objective optimization space and obtain the optimized reward strategy.

[0019] S7. Based on the optimized reward strategy, user feedback is collected, including user acceptance of rewards, task participation, and point usage. The effectiveness of rewards and changes in user interests are evaluated, and a final global reward strategy is derived based on the A3C algorithm. The A3C algorithm combines exploration and exploitation to adapt the global reward strategy to the needs of different user scenarios, enabling real-time dynamic adjustment of the Internet points incentive mechanism.

[0020] Optionally, the sensors of S1 include GPS, accelerometer and gyroscope.

[0021] Optionally, the S3 specifically includes:

[0022] S31. Constructing user behavior characteristics based on the user behavior dataset, wherein the behavior characteristics include the user's activity frequency, task participation, and historical behavior;

[0023] S32. Combine the context labels and use the context-adaptive reward algorithm to calculate the reward value:

[0024]

[0025] Among them, R represents the reward value, α, β and δ represent the adjustment parameters, and w i Represents behavioral characteristics x i The weight of x i represents the i-th behavioral feature, n represents the total number of user behavioral features, θ j Represents situational feature y j The weight of y j represents the jth context feature, m represents the total number of user context features, φ k represents the feedback weight, b k represents the evaluation index of feedback, l represents the number of feedback items;

[0026] S33. After the reward value is calculated, a fuzzy logic system is introduced to further adjust the reward value. The fuzzy logic system performs fuzzification processing on the reward value through a membership function to further adjust the intensity of the reward to adapt to the user's long-term behavior pattern and immediate situation;

[0027] S34. Based on the adjustment result of the fuzzy logic system, an adaptive optimization algorithm is used to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the user's behavior changes and feedback information during the context adjustment process. The optimized reward value meets different user behavior patterns and task requirements;

[0028] S35. Dynamically adjust the intensity and form of rewards based on user behavior datasets and contextual tags. The adjustment process takes into account the user's task completion, behavior frequency, and contextual changes to ensure that rewards adapt to different behavior patterns.

[0029] S36. Generate a reward strategy based on the dynamically adjusted reward value, and update the reward strategy in real time to dynamically adjust the Internet points incentive mechanism.

[0030] Optionally, the S4 specifically includes:

[0031] S41. Determine the reward adjustment factor based on the generated reward value and user behavior feedback, combined with the user's task participation and context label:

[0032]

[0033] Among them, λ t Reward adjustment factor, w i Represents behavioral characteristics x i The weight of Δx i (t) represents the change of the behavior feature at time t, n represents the total number of user behavior features, θ j Represents situational feature y j The weight of Δy j (t) represents the change of contextual features at time t, m represents the total number of user contextual features, ΔB t represents the comprehensive feedback amount of reward acceptance and behavior change, α1, β1 and γ1 represent the adjustment parameters;

[0034] S42, based on reward adjustment factor λ t Dynamically adjust the intensity and type of rewards:

[0035]

[0036] Among them, R′ t Represents the adjusted reward value, R t represents the initial reward value, δ1 represents the adjustment coefficient, φ k represents the feedback weight, ΔB k (t) represents the change in the kth feedback item, and l represents the number of feedback items;

[0037] S43. Based on the adjusted reward value, the type and frequency of rewards are dynamically adjusted in combination with the user's acceptance of the reward and the task participation. The adjustment process is optimized based on user behavior feedback:

[0038] ΔR t =γ2·(A t ΔX t )+γ3·(P t ΔT t );

[0039] Where, ΔR t Indicates the change in the adjusted reward value, A t represents the user’s reward acceptance, P t represents the user's task participation, ΔX t Indicates the change in user behavior, ΔT t represents the change in task completion, γ2 and γ3 represent adjustment parameters;

[0040] S44. By collecting user feedback on rewards in real time, the type and frequency of rewards are adjusted using a feedback loop. The feedback loop includes user acceptance of rewards, task participation, and behavior change information. The behavior change information is updated in real time using user behavior datasets and context labels.

[0041] Optionally, the S5 specifically includes:

[0042] S51. Based on the user's behavior feature vector, the behavior data is processed using a gated recurrent unit model. The state update formula of the gated recurrent unit model is:

[0043] h t =(1-z t )·h t-1 +z t tanh(W h ·X t +b h );

[0044] Among them, h t represents the hidden state at time t, z t represents the update gate at time t, h t-1 represents the hidden state at time t-1, tanh represents the hyperbolic tangent function, W h represents the weight matrix, b h represents the bias term, X t represents the behavioral feature vector;

[0045] S52. Train the gated recurrent unit model to obtain a dynamic hidden state of the user's behavior, and map the dynamic hidden state into a prediction result of future behavior using a fully connected layer, including the user's future activity frequency, engagement, and points of interest;

[0046] S53. By combining the prediction results with the user's historical behavior data, the weight coefficient of the future behavior prediction model is further adjusted to generate a multi-step prediction result:

[0047]

[0048] in, represents the predicted k-th step behavior trend, ζ represents the weighting coefficient, represents the behavior prediction result at time t, Indicates the k-1th step behavior trend of the forecast, and optimizes the forecast accuracy by adjusting the balance between the current and previous forecasts;

[0049] S54. Calculate the comprehensive trend of the user's future behavior by weighted average method based on the multi-step prediction results to obtain the final user's behavior prediction result.

[0050] Optionally, the S6 specifically includes:

[0051] S61. Based on the user behavior prediction results, a multi-objective optimization problem is constructed. The objectives include maximizing the user reward value, improving user engagement, and maintaining the long-term stability of the reward strategy:

[0052]

[0053] Among them, L(w) represents the objective function, N represents the total number of reward targets, μ i , λ1 and λ2 represent weight coefficients, w represents the parameters of the reward strategy, Reward i (w) represents the value of the i-th reward target, Engagement(w) represents the user's engagement, and Stability(w) represents the stability of the reward strategy;

[0054] S62, initialize the wolf pack optimization algorithm population, generate an initial solution through a chaotic search strategy, and initialize the population using a chaotic mapping, wherein the chaotic mapping process is:

[0055] x t+1 =ξ·x t ·(1-x t )with x0∈[0,1];

[0056] Among them, x t+1 represents the individual wolf position at the t+1th iteration, x t represents the individual wolf position at the tth iteration, ξ represents the control parameter of the chaotic mapping, and x0 represents the initial value;

[0057] S63. After using chaos initialization, the fitness of each wolf individual is evaluated based on the objective function;

[0058] S64. Based on the fitness evaluation results, the position is updated by simulating the wolf pack's foraging behavior, and the update formula with chaotic perturbation is used to enhance the global search capability:

[0059] w new =w old +κ1·(p best -w old )+κ2·(g best -w old )+η·F(x t );

[0060] Among them, w newrepresents the updated reward strategy, w old represents the current reward strategy, κ1 represents the influence of the individual optimal position on the update, κ2 represents the influence of the global optimal position on the update, and p best represents the best solution of the current wolf individual in the search space, g best represents the optimal solution among all wolf individuals, η represents the chaotic perturbation coefficient, F(x t ) represents the chaotic perturbation function, which perturbs the position of individual wolves based on the current state of the chaotic map;

[0061] S65. Based on the updated individual positions, strategy optimization is achieved by combining the exploration and development phases. The exploration phase relies on chaotic perturbations to conduct extensive searches, while the development phase refines the reward strategy through local searches to obtain the optimized reward strategy.

[0062] Optionally, the S7 specifically includes:

[0063] S71. Collect user feedback information based on the optimized reward strategy, including user acceptance of rewards, task participation, point usage, reward consumption rate, and changes in user interests, to form a user feedback dataset.

[0064] S72: Construct a reward effect evaluation function based on the user feedback data set. The reward effect evaluation function is modeled in combination with the user feedback data set to define the behavior state s. t is the user's state at time step t, including the user's activity frequency and task participation information, and is based on the behavior state s t Evaluate user acceptance of rewards:

[0065]

[0066] Among them, RA(s t ) represents the reward effect evaluation function, ρ i Indicates the weight of the user feedback data item, Feedback i Represents various indicators of user feedback, τ represents the influence weight of user task participation on reward acceptance, Engagement t represents the user’s engagement at time step t;

[0067] S73. Calculate and optimize the user behavior strategy using the parallel agent in the A3C algorithm, where the policy network and value network in the A3C algorithm are used to predict the user's expected reward and actual reward under different states, respectively.

[0068] S74. Combine the exploration and utilization strategies to dynamically adjust the global reward strategy. The exploration part randomly selects actions to perform strategy exploration and obtain the final global reward strategy:

[0069]

[0070] Among them, θ newexplore represents the parameters of the updated exploration policy network, θ explore represents the parameters of the exploration strategy network before updating, θ newexploit represents the updated parameters of the utilization strategy network, θ exploit represents the parameters of the utilization strategy network before updating, ∈ represents the balance coefficient between exploration and utilization, represents the gradient of the loss function in the exploration phase, Represents the gradient of the loss function in the utilization phase.

[0071] The beneficial effects of the present invention are:

[0072] First, the present invention utilizes a context-adaptive reward algorithm that dynamically calculates reward values ​​based on user behavioral data and contextual tags, and adjusts the intensity and form of rewards. This approach significantly improves the static nature of traditional point-based incentive mechanisms, enabling them to adapt in real time to users' immediate needs and behavioral changes, rather than relying on fixed rules or patterns. By introducing contextual tags and personalized analysis, the system can better understand users' behavioral motivations in specific situations, thereby improving the accuracy and effectiveness of reward strategies and avoiding the decline in incentive effectiveness that has occurred in past reward mechanisms due to their overly simplistic and universal nature.

[0073] Secondly, this invention further adjusts reward values ​​by introducing a fuzzy logic system, allowing the reward calculation process to not only consider the user's current behavior but also incorporate fitness adjustments based on historical behavior feedback, thereby enhancing the flexibility and adaptability of the reward system. By processing complex behavioral data, the fuzzy logic system can better cope with the uncertainty of user behavior and feedback, ensuring that rewards are issued in a manner that better meets the user's actual needs and expectations, thereby enhancing user participation and satisfaction. This innovation enables the points incentive mechanism to not only adapt to current behavior but also form long-term adaptive adjustments to user needs based on historical data and long-term behavioral patterns, avoiding the rapid decline of short-term incentive effects.

[0074] In addition, the present invention uses a variety of advanced optimization algorithms, such as chaos search strategy and wolf pack optimization algorithm, to dynamically optimize the reward strategy. These optimization algorithms can find the optimal strategy in a complex multi-objective decision space. The optimization process can better balance the multi-dimensional objectives of the reward strategy, ensuring that the reward mechanism can not only stimulate active user participation but also meet the platform's needs in many aspects. By combining global optimization and local search, the optimization algorithm can improve the accuracy of the reward strategy, avoiding the limitations of narrow optimization space and local optimality in traditional methods, and providing the platform with more efficient and sophisticated incentive management capabilities.

[0075] In addition, the present invention combines user behavior prediction technology and uses a gated recurrent unit (GRU) model to predict future user behavior trends, thereby planning and adjusting reward strategies in advance. By accurately predicting users' future engagement, points of interest, and activity frequency, the platform can adjust the reward strategy in advance based on these prediction results, making the reward mechanism more forward-looking and better able to motivate users to participate in various activities on the platform. Unlike the reward adjustment based on past data in the prior art, the present invention can foresee user behavior changes in advance through real-time data and behavior prediction, so that the platform's reward strategy always remains efficient and targeted. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0077] Figure 1 This is an overall flow chart of the design method of an Internet points dynamic incentive mechanism based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION

[0078] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0079] refer to Figure 1 , a design method for a dynamic incentive mechanism for Internet points based on reinforcement learning, including the following steps:

[0080] S1. Collect user behavior data through sensors and mobile devices and pre-process it to build a user behavior dataset;

[0081] S2. Use natural language processing technology to obtain the user's emotional state and context information and generate context labels;

[0082] S3. Calculate the reward value using a context-adaptive reward algorithm, dynamically adjust the reward intensity and form based on the user behavior dataset and context labels, and introduce a fuzzy logic system into the reward generation;

[0083] S4. Adjust the type and frequency of rewards based on the generated reward values ​​and user behavior feedback, dynamically adjusting the reward strategy based on the user's task participation and feedback loop, which is optimized based on the user's reward acceptance and behavior changes;

[0084] S5. Based on the adjusted reward strategy, use the gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, engagement, and points of interest, and generate user behavior prediction results;

[0085] S6. Based on the user behavior prediction results, the reward strategy is dynamically optimized by combining the chaotic search strategy and the wolf pack optimization algorithm. The optimization process simulates the wolf pack's foraging behavior to find the optimal strategy in the multi-objective optimization space and obtain the optimized reward strategy.

[0086] S7. Based on the optimized reward strategy, user feedback is collected, including user acceptance of rewards, task participation, and point usage. The effectiveness of rewards and changes in user interests are evaluated, and a final global reward strategy is derived based on the A3C algorithm. The A3C algorithm combines exploration and exploitation to adapt the global reward strategy to the needs of different user scenarios, enabling real-time dynamic adjustment of the Internet points incentive mechanism.

[0087] In this embodiment, the sensors of S1 include GPS, accelerometer and gyroscope.

[0088] In this embodiment, S3 specifically includes:

[0089] S31. Constructing user behavior characteristics based on the user behavior dataset, wherein the behavior characteristics include the user's activity frequency, task participation, and historical behavior;

[0090] S32. Combine the context labels and use the context-adaptive reward algorithm to calculate the reward value:

[0091]

[0092] Among them, R represents the reward value, α, β and δ represent the adjustment parameters, and w i Represents behavioral characteristics x i The weight of x i represents the i-th behavioral feature, n represents the total number of user behavioral features, θ j Represents situational feature y j The weight of y jrepresents the jth context feature, m represents the total number of user context features, φ k represents the feedback weight, b k represents the evaluation index of feedback, l represents the number of feedback items;

[0093] S33. After the reward value is calculated, a fuzzy logic system is introduced to further adjust the reward value. The fuzzy logic system performs fuzzification processing on the reward value through a membership function to further adjust the intensity of the reward to adapt to the user's long-term behavior pattern and immediate situation;

[0094] S34. Based on the adjustment result of the fuzzy logic system, an adaptive optimization algorithm is used to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the user's behavior changes and feedback information during the context adjustment process. The optimized reward value meets different user behavior patterns and task requirements;

[0095] S35. Dynamically adjust the intensity and form of rewards based on user behavior datasets and contextual tags. The adjustment process takes into account the user's task completion, behavior frequency, and contextual changes to ensure that rewards adapt to different behavior patterns.

[0096] S36. Generate a reward strategy based on the dynamically adjusted reward value, and update the reward strategy in real time to dynamically adjust the Internet points incentive mechanism.

[0097] In this embodiment, the S4 specifically includes:

[0098] S41. Determine the reward adjustment factor based on the generated reward value and user behavior feedback, combined with the user's task participation and context label:

[0099]

[0100] Among them, λ t Reward adjustment factor, w i Represents behavioral characteristics x i The weight of Δx i (t) represents the change of the behavior feature at time t, n represents the total number of user behavior features, θ j Represents situational feature y j The weight of Δy j (t) represents the change of contextual features at time t, m represents the total number of user contextual features, ΔB t represents the comprehensive feedback amount of reward acceptance and behavior change, α1, β1 and γ1 represent the adjustment parameters;

[0101] S42, based on reward adjustment factor λ t Dynamically adjust the intensity and type of rewards:

[0102]

[0103] Among them, R′ t Represents the adjusted reward value, R t represents the initial reward value, δ1 represents the adjustment coefficient, φ k represents the feedback weight, ΔB k (t) represents the change in the kth feedback item, and l represents the number of feedback items;

[0104] S43. Based on the adjusted reward value, the type and frequency of rewards are dynamically adjusted in combination with the user's acceptance of the reward and the task participation. The adjustment process is optimized based on user behavior feedback:

[0105] ΔR t =γ2·(A t ΔX t )+γ3·(P t ΔT t );

[0106] Where, ΔR t Indicates the change in the adjusted reward value, A t represents the user’s reward acceptance, P t represents the user's task participation, ΔX t Indicates the change in user behavior, ΔT t represents the change in task completion, γ2 and γ3 represent adjustment parameters;

[0107] S44. By collecting user feedback on rewards in real time, the type and frequency of rewards are adjusted using a feedback loop. The feedback loop includes user acceptance of rewards, task participation, and behavior change information. The behavior change information is updated in real time using user behavior datasets and context labels.

[0108] In this embodiment, the S5 specifically includes:

[0109] S51. Based on the user's behavior feature vector, the behavior data is processed using a gated recurrent unit model. The state update formula of the gated recurrent unit model is:

[0110] h t =(1-z t )·h t-1 +z t tanh(W h ·X t +b h );

[0111] Among them, h t represents the hidden state at time t, zt represents the update gate at time t, h t-1 represents the hidden state at time t-1, tanh represents the hyperbolic tangent function, W h represents the weight matrix, b h represents the bias term, X t represents the behavioral feature vector;

[0112] S52. Train the gated recurrent unit model to obtain a dynamic hidden state of the user's behavior, and map the dynamic hidden state into a prediction result of future behavior using a fully connected layer, including the user's future activity frequency, engagement, and points of interest;

[0113] S53. By combining the prediction results with the user's historical behavior data, the weight coefficient of the future behavior prediction model is further adjusted to generate a multi-step prediction result:

[0114]

[0115] in, represents the predicted k-th step behavior trend, ζ represents the weighting coefficient, represents the behavior prediction result at time t, Indicates the k-1th step behavior trend of the forecast, and optimizes the forecast accuracy by adjusting the balance between the current and previous forecasts;

[0116] S54. Calculate the comprehensive trend of the user's future behavior by weighted average method based on the multi-step prediction results to obtain the final user's behavior prediction result.

[0117] In this embodiment, S6 specifically includes:

[0118] S61. Based on the user behavior prediction results, a multi-objective optimization problem is constructed. The objectives include maximizing the user reward value, improving user engagement, and maintaining the long-term stability of the reward strategy:

[0119]

[0120] Among them, L(w) represents the objective function, N represents the total number of reward targets, μ i , λ1 and λ2 represent weight coefficients, w represents the parameters of the reward strategy, Reward i (w) represents the value of the i-th reward target, Engagement(w) represents the user's engagement, and Stability(w) represents the stability of the reward strategy;

[0121] S62, initialize the wolf pack optimization algorithm population, generate an initial solution through a chaotic search strategy, and initialize the population using a chaotic mapping, wherein the chaotic mapping process is:

[0122] x t+1 =ξ·x t ·(1-x t )with x0∈[0,1];

[0123] Among them, x t+1 represents the individual wolf position at the t+1th iteration, x t represents the individual wolf position at the tth iteration, ξ represents the control parameter of the chaotic mapping, and x0 represents the initial value;

[0124] S63. After using chaos initialization, the fitness of each wolf individual is evaluated based on the objective function;

[0125] S64. Based on the fitness evaluation results, the position is updated by simulating the wolf pack's foraging behavior, and the update formula with chaotic perturbation is used to enhance the global search capability:

[0126] w new =w old +κ1·(p best -w old )+κ2π(g best -w old )+ηπF(x t );

[0127] Among them, w new represents the updated reward strategy, w old represents the current reward strategy, κ1 represents the influence of the individual optimal position on the update, κ2 represents the influence of the global optimal position on the update, and p best represents the best solution of the current wolf individual in the search space, g best represents the optimal solution among all wolf individuals, η represents the chaotic perturbation coefficient, F(x t ) represents the chaotic perturbation function, which perturbs the position of individual wolves based on the current state of the chaotic map;

[0128] S65. Based on the updated individual positions, strategy optimization is achieved by combining the exploration and development phases. The exploration phase relies on chaotic perturbations to conduct extensive searches, while the development phase refines the reward strategy through local searches to obtain the optimized reward strategy.

[0129] In this embodiment, the S7 specifically includes:

[0130] S71. Collect user feedback information based on the optimized reward strategy, including user acceptance of rewards, task participation, point usage, reward consumption rate, and changes in user interests, to form a user feedback dataset.

[0131] S72: Construct a reward effect evaluation function based on the user feedback data set. The reward effect evaluation function is modeled in combination with the user feedback data set to define the behavior state s. t is the user's state at time step t, including the user's activity frequency and task participation information, and is based on the behavior state s t Evaluate user acceptance of rewards:

[0132]

[0133] Among them, RA*s t (represents the reward effect evaluation function, ρ i Indicates the weight of the user feedback data item, Feedback i Represents various indicators of user feedback, τ represents the influence weight of user task participation on reward acceptance, Engagement t represents the user’s engagement at time step t;

[0134] S73. Calculate and optimize the user behavior strategy using the parallel agent in the A3C algorithm, where the policy network and value network in the A3C algorithm are used to predict the user's expected reward and actual reward under different states, respectively.

[0135] S74. Combine the exploration and utilization strategies to dynamically adjust the global reward strategy. The exploration part randomly selects actions to perform strategy exploration and obtain the final global reward strategy:

[0136]

[0137] Among them, θ newexplore represents the parameters of the updated exploration policy network, θ explore represents the parameters of the exploration strategy network before updating, θ newexploit represents the updated parameters of the utilization strategy network, θ exploit represents the parameters of the utilization strategy network before updating, ∈ represents the balance coefficient between exploration and utilization, represents the gradient of the loss function in the exploration phase, Represents the gradient of the loss function in the utilization phase.

[0138] Example 1:

[0139] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the user points incentive system of a certain e-commerce platform. The number of daily active users of this e-commerce platform reaches 5 million, with a huge user base and rich user behavior data. The user behavior data of the platform includes browsing history, shopping history, search behavior, sharing behavior, comment behavior, etc. These data not only reflect the frequency of users' activities on the platform, but also reflect the users' interest in different products, activities and promotional strategies. However, the platform's existing points incentive mechanism is too simple, the reward method is fixed, and it cannot be effectively adjusted according to the user's personalized needs and situational changes, resulting in a gradual decline in user participation and the platform's activity is also affected. Therefore, the platform urgently needs a solution that can adjust the reward strategy in real time and improve user activity and loyalty.

[0140] Specifically, the platform first collects user behavioral data and constructs a behavioral feature vector for each user, which includes multi-dimensional information such as the user's activity frequency, task participation, historical behavior, and situational factors. On this basis, the platform uses a context-adaptive reward algorithm, combined with situational labels, to calculate the reward value for each user. This reward value not only takes into account the user's current behavior, but also further adjusts the reward value through a fuzzy logic system to adapt it to the long-term behavior patterns and immediate situations of different users. In addition, the platform also improves users' task completion and participation by dynamically adjusting the reward strategy. For example, when a user has not participated in a task for a long time, the form of reward will be more attractive to the user to re-participate; and when the user actively participates, the intensity of the reward will gradually increase to ensure the diversity and adaptability of the reward strategy.

[0141] During implementation, the platform also leverages a gated recurrent unit (GRU) model to predict future user behavior and dynamically optimizes the reward strategy by combining a chaotic search strategy with a wolf pack optimization algorithm. This optimization process simulates the foraging behavior of a wolf pack to ensure the reward strategy remains highly effective across diverse user behavior patterns. Finally, the platform uses the A3C algorithm to adjust the reward strategy in real time, evaluating user acceptance of rewards, task engagement, and point usage to further optimize the overall reward strategy.

[0142] To verify the effectiveness of the present invention, the platform collected data before and after implementation for comparison. Before implementation, the number of daily active users on the platform was 5 million. After implementation, the number of daily active users increased to 6 million, with an increase of 20.10%. In terms of task completion, it was 50.25% before implementation and increased to 75.38% after implementation, with an increase of 50.25%. In terms of user retention rate, the platform's next-day retention rate was 30.10% before implementation, and increased to 45.45% after implementation. In addition, the number of tasks users participated in per month increased from an average of 5.12 times to 8.13 times, an increase of 58.88%. These data show that the present invention has significant effects in improving user activity, task completion, user engagement and loyalty.

[0143] Table 1 Comparison of user behaviors before and after the implementation of the Internet points incentive mechanism based on reinforcement learning

[0144] index Before implementation After implementation Increase Daily active users 5 million 6.005 million 20.10% Task completion 50.25% 75.38% 50.25% User retention rate 30.10% 45.45% 15.15% Average number of task participations 5.12 times / month 8.13 times / month 58.88% User Reward Acceptance 70.20% 85.15% 21.30%

[0145] As Table 1 clearly demonstrates, the implementation of this invention effectively improves user activity, task completion, engagement, and loyalty on the platform. In particular, by dynamically adjusting reward strategies, the platform is able to provide each user with personalized and flexible rewards, ensuring a consistently efficient reward mechanism and thus enhancing user engagement and loyalty to the platform.

[0146] In summary, the reinforcement learning-based design method for a dynamic internet points-based incentive mechanism can dynamically adjust reward strategies based on user behavior and contextual information, significantly improving user activity, task completion, and engagement on the platform. Furthermore, by continuously optimizing reward strategies, the present invention ensures the long-term adaptability of the points-based incentive mechanism, driving sustained user growth and loyalty improvements for the platform.

[0147] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for designing a dynamic incentive mechanism for Internet points based on reinforcement learning, characterized in that: The steps include: S1. Collect user behavior data through sensors and mobile devices and pre-process it to build a user behavior dataset; S2. Use natural language processing technology to obtain the user's emotional state and context information and generate context labels; S3. Calculate the reward value using a context-adaptive reward algorithm, dynamically adjust the reward intensity and form based on the user behavior dataset and context labels, and introduce the reward value into the fuzzy logic system; S4. Adjust the type and frequency of rewards based on the generated reward values ​​and user behavior feedback, dynamically adjusting the reward strategy based on the user's task participation and feedback loop, which is optimized based on the user's reward acceptance and behavior changes; S5. Based on the adjusted reward strategy, use the gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, engagement, and points of interest, and generate user behavior prediction results; S6. Based on the user behavior prediction results, the reward strategy is dynamically optimized by combining the chaotic search strategy and the wolf pack optimization algorithm. The optimization process simulates the wolf pack's foraging behavior to find the optimal strategy in the multi-objective optimization space and obtain the optimized reward strategy. S7. Based on the optimized reward strategy, user feedback is collected, including user acceptance of rewards, task participation, and point usage. The effectiveness of rewards and changes in user interest are evaluated, and the final global reward strategy is derived based on the A3C algorithm. The A3C algorithm combines exploration and exploitation to adapt the global reward strategy to the needs of different user scenarios, enabling real-time dynamic adjustment of the Internet points incentive mechanism. The S4 specifically includes: S41. Determine the reward adjustment factor based on the generated reward value and user behavior feedback, combined with the user's task participation and context label: ; in, represents the reward adjustment factor, Indicates behavioral characteristics The weight of Indicates behavioral characteristics over time The amount of change, Represents the total number of user behavior features, Representing situational characteristics The weight of Represents situational features in time The amount of change, represents the total number of user context features, represents the comprehensive feedback amount of reward acceptance and behavior change, 、 and represents the adjustment parameter; S42. Based on reward adjustment factor Dynamically adjust the intensity and type of rewards: ; in, Indicates the adjusted reward value, represents the initial reward value, represents the adjustment coefficient, represents the feedback weight, Indicates the The change in feedback, Indicates the number of feedback items; S43. Based on the adjusted reward value, combined with the user's acceptance of the reward and task participation, the type and frequency of rewards are dynamically adjusted. The adjustment process is optimized based on user behavior feedback: ; in, Indicates the change in the adjusted reward value. Indicates the user's reward acceptance, Indicates the user's task engagement, Indicates the amount of change in user behavior. Indicates the change in task completion status, and represents the adjustment parameter; S44. By collecting user feedback on rewards in real time, the type and frequency of rewards are adjusted using a feedback loop. The feedback loop includes user acceptance of rewards, task participation, and behavioral change information. The behavioral change information is updated in real time using the user behavior dataset and contextual tags. The S7 specifically includes: S71. Collect user feedback information based on the optimized reward strategy, including user acceptance of rewards, task participation, point usage, reward consumption rate, and changes in user interests, to form a user feedback dataset. S72: Construct a reward effect evaluation function based on the user feedback data set. The reward effect evaluation function is modeled in combination with the user feedback data set to define the behavior state. For the time step The user's status at the time of use, including the user's activity frequency and task participation information, and based on the behavior status Evaluate user acceptance of rewards: ; in, represents the reward effect evaluation function, represents the weight of the user feedback data item, Indicates various indicators of user feedback, represents the influence weight of user task participation on reward acceptance, Indicates that the user is at time step participation; S73. Utilize the parallel agent in the A3C algorithm to calculate and optimize the user behavior strategy, wherein the strategy network and value network in the A3C algorithm are used to predict the user's expected reward and actual reward under different states, respectively; S74. Combine the exploration and utilization strategies to dynamically adjust the global reward strategy. The exploration part randomly selects actions to perform strategy exploration and obtain the final global reward strategy: ; ; in, represents the parameters of the updated exploration strategy network, represents the parameters of the exploration strategy network before updating, represents the parameters of the updated utilization strategy network, represents the parameters of the utilization strategy network before updating, represents the balance coefficient between exploration and exploitation, represents the gradient of the loss function in the exploration phase, Represents the gradient of the loss function in the utilization stage.

2. The method for designing an Internet points dynamic incentive mechanism based on reinforcement learning according to claim 1 is characterized in that: The S1's sensors include GPS, accelerometer, and gyroscope.

3. The method for designing an Internet points dynamic incentive mechanism based on reinforcement learning according to claim 1 is characterized in that: The S3 specifically includes: S31. Constructing user behavior characteristics based on the user behavior dataset, wherein the behavior characteristics include the user's activity frequency, task participation, and historical behavior; S32. Combine the context labels and use the context-adaptive reward algorithm to calculate the reward value: ; in, Indicates the reward value. 、 and represents the adjustment parameter, Indicates behavioral characteristics The weight of Indicates the behavioral characteristics, Represents the total number of user behavior features, Representing situational characteristics The weight of Indicates the situational characteristics, represents the total number of user context features, represents the feedback weight, The evaluation index of feedback, Indicates the number of feedback items; S33. After the reward value is calculated, a fuzzy logic system is introduced to further adjust the reward value. The fuzzy logic system performs fuzzification processing on the reward value through a membership function to further adjust the intensity of the reward to adapt to the user's long-term behavior pattern and immediate situation; S34. Based on the adjustment result of the fuzzy logic system, an adaptive optimization algorithm is used to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the user's behavior changes and feedback information during the context adjustment process. The optimized reward value meets different user behavior patterns and task requirements; S35. Dynamically adjust the intensity and form of rewards based on user behavior datasets and contextual tags. The adjustment process takes into account the user's task completion, behavior frequency, and contextual changes to ensure that rewards adapt to different behavior patterns. S36. Generate a reward strategy based on the dynamically adjusted reward value, and update the reward strategy in real time to dynamically adjust the Internet points incentive mechanism.

4. The method for designing a dynamic incentive mechanism for Internet points based on reinforcement learning according to claim 1 is characterized in that: The S5 specifically includes: S51. Based on the user's behavior feature vector, the behavior data is processed using a gated recurrent unit model. The state update formula of the gated recurrent unit model is: ; in, Indicates time The hidden state of the moment, Indicates time The door of renewal at all times, Indicates time The hidden state of the moment, represents the hyperbolic tangent function, represents the weight matrix, represents the bias term, represents the behavioral feature vector; S52. Train the gated recurrent unit model to obtain a dynamic hidden state of the user's behavior, and map the dynamic hidden state into a prediction result of future behavior using a fully connected layer, including the user's future activity frequency, engagement, and points of interest; S53. By combining the prediction results with the user's historical behavior data, the weight coefficient of the future behavior prediction model is further adjusted to generate a multi-step prediction result: ; in, The predicted Walking is the trend. represents the weighting coefficient, Indicates time The behavior prediction results at each moment, The predicted Follow behavioral trends and optimize forecast accuracy by adjusting the balance between current and previous forecasts; S54. Calculate the comprehensive trend of the user's future behavior by weighted average method based on the multi-step prediction results to obtain the final user's behavior prediction result.

5. The method for designing an Internet points dynamic incentive mechanism based on reinforcement learning according to claim 1 is characterized in that: The S6 specifically includes: S61. Based on the user behavior prediction results, a multi-objective optimization problem is constructed. The objectives include maximizing the user reward value, improving user engagement, and maintaining the long-term stability of the reward strategy: ; in, represents the objective function, Indicates the total number of reward targets, 、 and represents the weight coefficient, Parameters representing the reward strategy, Indicates the The value of the reward target, Indicates the user's engagement. Indicates the stability of the reward strategy; S62, initialize the wolf pack optimization algorithm population, generate an initial solution through a chaotic search strategy, and initialize the population using a chaotic mapping, wherein the chaotic mapping process is: ; in, Indicates the The individual wolf position at the iteration, Indicates the The individual wolf position at the iteration, represents the control parameter of the chaotic map, Indicates the initial value; S63. After using chaos initialization, the fitness of each wolf individual is evaluated based on the objective function; S64. Based on the fitness evaluation results, the position is updated by simulating the wolf pack's foraging behavior, and the update formula with chaotic perturbation is used to enhance the global search capability: ; in, represents the updated reward strategy, Represents the current reward strategy, Indicates the degree of influence of the individual optimal position on the update, Indicates the influence of the global optimal position on the update, represents the best solution of the current wolf individual in the search space, represents the optimal solution among all wolf individuals, represents the chaotic perturbation coefficient, represents the chaotic perturbation function, which perturbs the position of individual wolves based on the current state of the chaotic map; S65. Based on the updated individual positions, strategy optimization is achieved by combining the exploration and development phases. The exploration phase relies on chaotic perturbations to conduct extensive searches, while the development phase refines the reward strategy through local searches to obtain the optimized reward strategy.