Internet integral dynamic incentive mechanism design method based on reinforcement learning

By combining situational adaptive reward algorithms and multiple optimization algorithms, the Internet points incentive strategy is dynamically adjusted, which solves the problem of flexibility and insufficient personalization of the traditional points incentive mechanism, and improves user participation and loyalty.

CN120258886AActive Publication Date: 2025-07-04YUEJI ENTERPRISE MANAGEMENT CO LTD

Patent Information

Application Number
CN202510291778.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing Internet points incentive mechanism lacks flexibility and personalization, fails to fully consider the diversity of user behavior and situational changes, which makes it difficult to maintain user participation and loyalty for a long time, and fails to effectively use user behavior feedback for dynamic adjustments.

Method used

The situational adaptive reward algorithm, gated loop unit model, chaotic search strategy and wolf pack optimization algorithm are adopted, and the reward strategy is dynamically adjusted, and reward calculation is optimized through the fuzzy logic system and feedback loop to achieve real-time response to user needs.

Benefits of technology

It improves the accuracy and adaptability of the reward strategy, enhances user participation and satisfaction, improves user activity and loyalty, and ensures the long-term adaptability and efficiency of the points incentive mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258886A_ABST
    Figure CN120258886A_ABST
Patent Text Reader

Abstract

The invention discloses an internet integral dynamic incentive mechanism design method based on reinforcement learning. The method comprises the following steps: S1, constructing a user behavior data set; s2, generating a situation label by adopting a natural language processing technology; s3, calculating a reward value by using a context-adaptive reward algorithm; s4, dynamically adjusting the reward strategy through the reward value and the user behavior feedback in combination with the task participation condition of the user and a feedback loop; s5, predicting a future behavior trend of the user by using the gating cycle unit model, and generating a behavior prediction result of the user; s6, dynamically optimizing the reward strategy in combination with a chaos search strategy and a wolf pack optimization algorithm to obtain an optimized reward strategy; and S7, evaluating effectiveness of rewards and user interest changes, obtaining a final global reward strategy based on an A3C algorithm, and realizing real-time dynamic adjustment of an internet integral incentive mechanism. According to the method, a situation self-adaptive reward algorithm, an optimization technology and the like are utilized, and the dynamic adjustment of an internet integral incentive strategy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of incentive mechanism design, and particularly to a method for designing an Internet integral dynamic incentive mechanism based on reinforcement learning. Background Art

[0002] In the modern Internet environment, integral incentive mechanisms have been widely applied to various online platforms, especially in the fields of e-commerce, social networks, and mobile applications, as effective means for user participation, loyalty improvement, and behavior guidance. Integral rewards are usually closely related to user behaviors. Users obtain points by completing tasks, participating in activities, etc., and these points can be exchanged for cash, coupons, or other rewards. However, traditional integral incentive mechanisms are often distributed based on fixed rules, without fully considering the diversity of user behaviors, changes in situations, and user personalized needs. With the increasing complexity of user behaviors, traditional integral incentive mechanisms are facing many challenges and urgently need innovative solutions.

[0003] Existing integral incentive mechanisms usually rely on simple rules and algorithms, such as distributing rewards according to the frequency of user participation in activities, task completion, or simple historical data. Although these methods can achieve a certain incentive effect, due to the lack of flexibility and pertinence, it is often difficult to mobilize the long-term enthusiasm of users. More critically, most existing methods ignore the personalized needs and situational factors of users, relying only on a single reward mechanism and failing to fully utilize user behavior feedback to dynamically adjust the reward strategy, resulting in the difficulty of maintaining user participation and loyalty in the long term.

[0004] In addition, when dealing with user behavior data, existing technologies often adopt static algorithm models, ignoring the time series characteristics of user behaviors. User behaviors are time-dependent and may change significantly over time. Traditional reward strategies do not fully consider this dynamic characteristic, resulting in a possible weakening of their effects over time. For example, some platforms may set a fixed number of integral rewards to be distributed monthly through simple rules. Although this method can stimulate user activity in the short term, over time, users may feel a lack of freshness or insufficient incentives, thus reducing their participation.

[0005] Under this background, the defects of traditional integral incentive mechanisms are gradually emerging, especially in an environment that needs to face a large number of users, personalized needs, and changing situations, where simple rules are difficult to meet the requirements. With the development of artificial intelligence technologies, especially the introduction of reinforcement learning, natural language processing, and optimization algorithms, new methods for designing incentive mechanisms have begun to become a research hotspot. These technologies provide the possibility for dynamic modeling of user behaviors, personalized adjustment of reward mechanisms, situation awareness, etc., and thus provide a more intelligent and efficient incentive mechanism design solution for Internet platforms.

[0006] As a technology that can self - adjust based on user feedback, reinforcement learning has received extensive attention in recent years. Through reinforcement learning, the system can dynamically adjust the reward strategy according to the user's historical behavior and feedback information to achieve more precise user motivation. For example, by calculating the reward value based on the user's behavior characteristics and context information and adopting a context - adaptive reward algorithm, the reward mechanism can be adaptively adjusted according to the user's needs in different contexts. By introducing fuzzy logic and optimization algorithms, the flexibility and accuracy of reward calculation can be further improved. However, most of the existing reinforcement learning applications are concentrated in individual fields, and there is still a lack of in - depth exploration and implementation in the actual application of the Internet integral incentive mechanism.

[0007] In addition, with the increase in the complexity of user emotional states and behaviors, the role of context information in the integral incentive mechanism becomes increasingly important. Most of the existing technologies fail to consider the impact of context on the reward effect and also fail to effectively use context tags to personalize the adjustment of users' reward needs. By adopting natural language processing technology and context tag generation methods, the system's understanding of user behavior can be further improved, thereby generating a reward strategy that better meets the user's needs and avoiding the lack of universality of a single fixed rule for all users.

[0008] However, even so, there are still some limitations in the existing design methods of incentive mechanisms based on reinforcement learning. For example, although reinforcement learning can dynamically adjust according to user feedback, how to handle the diversity of different users' needs and behaviors while ensuring global optimization is still a difficult point. Most of the existing technologies use a single optimization algorithm and fail to fully combine advanced optimization methods such as chaotic search strategies and wolf pack optimization algorithms. These algorithms usually focus on the balance between global search and local optimization, but in actual applications, how to achieve the dynamic adjustment of the reward strategy in multi - objective optimization problems is still a challenge.

[0009] At the same time, the user's behavior feedback in actual applications is often complex. The user's feedback includes not only the acceptance of rewards but also task participation, integral usage, emotional state, and other multi - dimensional information. Existing technologies usually only focus on the user's basic behavior data, such as activity frequency and task completion rate, and ignore the impact of factors such as emotional state and interest changes on the reward mechanism. This results in the difficulty of the existing incentive mechanism to make timely and accurate adjustments when facing users' personalized needs and changing contexts.

[0010] Therefore, how to provide a design method for an Internet integral dynamic incentive mechanism based on reinforcement learning is an urgent problem for those skilled in the art. Summary of the Invention

[0011] An object of the present invention is to propose a design method for an Internet integral dynamic incentive mechanism based on reinforcement learning. The present invention makes full use of a situation-adaptive reward algorithm, a gated recurrent unit model, a chaotic search strategy, and a wolf pack optimization algorithm, and details a design method for an Internet integral dynamic incentive strategy based on user behavior prediction. By combining the user's behavior data, emotional state, and situation labels, the reward strategy is dynamically adjusted to ensure the long-term stability and adaptability of the strategy. The invention has significant advantages in improving user participation, enhancing the accuracy of the reward strategy, and optimizing the user experience, can respond to user needs in real time, and further improve the user activity and loyalty of the Internet platform.

[0012] A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0013] S1. Collect the user's behavior data through sensors and mobile devices and perform preprocessing to construct a user behavior data set;

[0014] S2. Adopt natural language processing technology to obtain the user's emotional state and the situation information where the user is located, and generate situation labels;

[0015] S3. Use the situation-adaptive reward algorithm to calculate the reward value, dynamically adjust the reward intensity and form according to the user behavior data set and the situation labels, and introduce a fuzzy logic system in the reward generation;

[0016] S4. Adjust the type and frequency of the reward through the generated reward value and the user behavior feedback, combine the user's task participation situation and the feedback loop, and dynamically adjust the reward strategy, and the feedback loop is optimized based on the user's reward acceptance degree and behavior changes;

[0017] S5. Based on the adjusted reward strategy, use the gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, participation degree, and interest points, and generate the user's behavior prediction result;

[0018] S6. According to the user's behavior prediction result, combine the chaotic search strategy and the wolf pack optimization algorithm to dynamically optimize the reward strategy. The optimization process searches for the optimal strategy in the multi-objective optimization space by simulating the foraging behavior of the wolf pack to obtain the optimized reward strategy;

[0019] S7. According to the optimized reward strategy, collect the user's feedback information, including the user's acceptance degree of the reward, task participation degree, and integral usage situation, evaluate the effectiveness of the reward and the change of the user's interest, and obtain the final global reward strategy based on the A3C algorithm. The A3C algorithm combines the exploration and exploitation methods to make the global reward strategy adapt to the needs in different user situations, and realizes the real-time dynamic adjustment of the Internet integral incentive mechanism.

[0020] Optionally, the sensor of S1 includes GPS, an accelerometer, and a gyroscope.

[0021] Optionally, S3 specifically includes:

[0022] S31. Construct the behavior characteristics of the user based on the user behavior dataset, where the behavior characteristics include the user's activity frequency, task participation, and historical behavior;

[0023] S32. Combine the context tags and use the context adaptive reward algorithm to calculate the reward value:

[0024]

[0025] where R represents the reward value, α, β, and δ represent adjustment parameters, w i represents the weight of the behavior characteristic x i , x i represents the i-th behavior characteristic, n represents the total number of user behavior characteristics, θ j represents the weight of the context characteristic y j , y j represents the j-th context characteristic, m represents the total number of user context characteristics, φ k represents the feedback weight, b k represents the evaluation index of the feedback, and l represents the number of feedback items;

[0026] S33. After calculating the reward value, introduce a fuzzy logic system to further adjust the reward value. The fuzzy logic system performs fuzzy processing on the reward value through a membership function to further adjust the intensity of the reward for adapting to the user's long-term behavior pattern and immediate context;

[0027] S34. Based on the adjustment result of the fuzzy logic system, use an adaptive optimization algorithm to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the feedback information during the user's behavior change and context adjustment process, and the optimized reward value meets different user behavior patterns and task requirements;

[0028] S35. Dynamically adjust the reward intensity and form according to the user behavior dataset and context tags. The adjustment process combines the user's task completion degree, behavior frequency, and context change situation to ensure that the reward adapts to different behavior patterns;

[0029] S36. Generate a reward strategy based on the dynamically adjusted reward value and update the reward strategy in real time for dynamically adjusting the Internet integral incentive mechanism.

[0030] Optionally, S4 specifically includes:

[0031] S41. Determine a reward adjustment factor based on the generated reward value and user behavior feedback, in combination with the user's task participation and context tags:

[0032]

[0033] where λ t represents the reward adjustment factor, w i represents the weight of the behavior feature x i Δx i (t) represents the change in the behavior feature at time t, n represents the total number of user behavior features, θ j represents the weight of the context feature y j Δy j (t) represents the change in the context feature at time t, m represents the total number of user context features, ΔB t represents the comprehensive feedback amount of reward acceptance and behavior change, and α1, β1, and γ1 represent adjustment parameters;

[0034] S42. Dynamically adjust the intensity and type of the reward based on the reward adjustment factor λ t :

[0035]

[0036] where R′ t represents the adjusted reward value, R t represents the initial reward value, δ1 represents the adjustment coefficient, φ k represents the feedback weight, ΔB k (t) represents the change in the k-th feedback, and l represents the number of feedback items;

[0037] S43. Dynamically adjust the type and frequency of the reward according to the adjusted reward value, in combination with the user's acceptance of the reward and task participation, and optimize the adjustment process based on user behavior feedback:

[0038] ΔR t = γ2·(A t ·ΔX t ) + γ3·(P t ·ΔT t );

[0039] where ΔR t represents the change in the adjusted reward value, A t represents the user's reward acceptance, P t represents the user's task participation, ΔX t represents the change in the user's behavior, ΔT t represents the change in the task completion situation, and γ2 and γ3 represent adjustment parameters;

[0040] S44. By collecting the feedback information of users on rewards in real time, the type and frequency of rewards are adjusted using a feedback loop, and the feedback loop includes the acceptance of rewards by users, task participation, and behavior change information, and the behavior change information is updated in real time through the user behavior data set and the context label.

[0041] Optionally, the S5 specifically includes:

[0042] S51. Based on the behavior feature vector of the user, the gated recurrent unit model is used to process the behavior data, and the state update formula of the gated recurrent unit model is:

[0043] h t =(1 - z t )·h t-1 +z t ·tanh(W h ·X t +b h );

[0044] Among them, h t represents the hidden state at time t, z t represents the update gate at time t, h t-1 represents the hidden state at time t - 1, tanh represents the hyperbolic tangent function, W h represents the weight matrix, b h represents the bias term, and X t represents the behavior feature vector;

[0045] S52. By training the gated recurrent unit model, the dynamic hidden state of the user behavior is obtained, and based on the dynamic hidden state, it is mapped to the prediction result of the future behavior by using the fully connected layer, including the future activity frequency, participation degree, and interest points of the user;

[0046] S53. By combining the prediction result and the user historical behavior data, the weight coefficient of the future behavior prediction model is further adjusted to generate a multi-step prediction result:

[0047]

[0048] Among them, represents the k-step behavior trend of the prediction, ζ represents the weighting coefficient, represents the behavior prediction result at time t, represents the (k - 1)-step behavior trend of the prediction, and the prediction accuracy is optimized by adjusting the balance between the current and previous predictions;

[0049] S54. Calculate the comprehensive trend of the user's future behavior through weighted average based on the multi-step prediction results to obtain the behavior prediction result of the final user.

[0050] Optionally, the S6 specifically includes:

[0051] S61. Construct a multi-objective optimization problem based on the behavior prediction result of the user. The objectives include maximizing the user reward value, improving the user engagement, and the long-term stability of the reward strategy:

[0052]

[0053] Among them, L(w) represents the objective function, N represents the total number of reward objectives, μ i , λ1 and λ2 represent the weight coefficients, w represents the parameters of the reward strategy, Reward i (w) represents the value of the i-th reward objective, Engagement(w) represents the user engagement, and Stability(w) represents the stability of the reward strategy;

[0054] S62. Initialize the population of the wolf pack optimization algorithm, generate the initial solution through the chaotic search strategy, and initialize the population using the chaotic mapping. The chaotic mapping process is:

[0055] x t+1 = ξ·x t ·(1 - x t ) with x0 ∈ [0,1];

[0056] Among them, x t+1 represents the position of the wolf individual at the (t + 1)-th iteration, x t represents the position of the wolf individual at the t-th iteration, ξ represents the control parameter of the chaotic mapping, and x0 represents the initial value;

[0057] S63. After using chaotic initialization, evaluate the fitness of each wolf individual based on the objective function;

[0058] S64. Based on the fitness evaluation results, use the simulation of the foraging behavior of the wolf pack to update the position, and adopt an update formula with chaotic perturbation to enhance the global search ability:

[0059] w new = w old + κ1·(p best - w old ) + κ2·(g best - w old ) + η·F(x t );

[0060] Among them, w newRepresents the updated reward strategy, w old Represents the current reward strategy, κ1 represents the influence degree of the individual optimal position on the update, κ2 represents the influence degree of the global optimal position on the update, p best Represents the best solution of the current wolf individual in the search space, g best Represents the optimal solution among all wolf individuals, η represents the chaotic perturbation coefficient, F(x t ) represents the chaotic perturbation function, and perturbs the position of the wolf individual based on the current state of the chaotic map;

[0061] S65. According to the updated individual positions, achieve policy optimization by combining the exploration and exploitation phases. In the exploration phase, rely on chaotic perturbation for extensive search. In the exploitation phase, refine the reward policy through local search to obtain the optimized reward policy.

[0062] Optionally, the S7 specifically includes:

[0063] S71. According to the optimized reward policy, collect the feedback information of users. The feedback information includes the acceptance degree of rewards by users, task participation degree, integral usage situation, reward consumption rate, and changes in user interests, and form a user feedback data set;

[0064] S72. According to the user feedback data set, construct a reward effect evaluation function. The reward effect evaluation function is modeled by combining the user feedback data set, and define the behavior state s t as the state of the user at time step t, including the activity frequency and task participation degree information of the user, and evaluate the acceptance degree of the user for the reward according to the behavior state s t :

[0065]

[0066] where, RA(s t ) represents the reward effect evaluation function, ρ i represents the weight of the user feedback data item, Feedback i represents each index of the user feedback, τ represents the influence weight of the user task participation degree on the reward acceptance degree, Engagement t represents the participation degree of the user at time step t;

[0067] S73. Use the parallel agents in the A3C algorithm to calculate and optimize the strategy of user behavior. In the A3C algorithm, the policy network and the value network are respectively used to predict the reward expectation and the actual reward of the user in different states;

[0068] S74. Combine the exploration and exploitation strategies to dynamically adjust the global reward policy. The exploration part conducts policy exploration by randomly selecting actions to obtain the final global reward policy:

[0069]

[0070] where θ newexplore represents the parameters of the updated exploration policy network, and θ explore represents the parameters of the exploration policy network before update, and θ newexploit represents the parameters of the updated exploitation policy network, and θ exploit represents the parameters of the exploitation policy network before update, and ∈ represents the balance coefficient between exploration and exploitation, represents the gradient of the loss function in the exploration phase, represents the gradient of the loss function in the exploitation phase.

[0071] The beneficial effects of the present invention are as follows:

[0072] First, the present invention adopts a context - adaptive reward algorithm, which can dynamically calculate reward values according to the user's behavior data and context labels, and adjust the intensity and form of the rewards. This method significantly improves the static nature of the traditional integral incentive mechanism, enabling the reward mechanism to adapt in real - time according to the user's immediate needs and behavior changes, rather than relying on fixed rules or patterns. By introducing context labels and personalized analysis, the system can better understand the user's behavior motivation in a specific context, thereby improving the accuracy and effectiveness of the reward strategy, and avoiding the decline in incentive effects caused by the over - simplicity and universality of the past reward mechanism.

[0073] Second, the present invention further adjusts the reward values by introducing a fuzzy logic system, making the reward calculation process not only consider the user's current behavior, but also incorporate the fitness adjustment of historical behavior feedback, enhancing the flexibility and adaptability of the reward system. The fuzzy logic system can better handle the uncertainty of user behavior and feedback by processing complex behavior data, ensuring that the rewards are distributed more in line with the user's actual needs and expectations, thereby enhancing the user's sense of participation and satisfaction. This innovation enables the integral incentive mechanism not only to adapt to the current behavior, but also to form a long - term adaptive adjustment to the user's needs based on historical data and long - term behavior patterns, avoiding the rapid decline of short - term incentive effects.

[0074] In addition, the present invention adopts a variety of advanced optimization algorithms, such as chaotic search strategy and wolf pack optimization algorithm, to dynamically optimize the reward strategy. These optimization algorithms can find the optimal strategy in the complex multi-objective decision space, and the optimization process can better balance the multi-dimensional objectives of the reward strategy, ensuring that the reward mechanism can not only stimulate users to actively participate, but also meet the platform's various needs. Through the combination of global optimization and local search, the optimization algorithm can improve the accuracy of the reward strategy, avoiding the limitations of narrow optimization space and local optimum in traditional methods, and providing the platform with more efficient and refined incentive management capabilities.

[0075] Furthermore, the present invention combines user behavior prediction technology, uses a gated recurrent unit (GRU) model to predict the future behavior trends of users, and thus plans and adjusts the reward strategy in advance. By accurately predicting users' future participation, interests, and activity frequencies, the platform can adjust the reward strategy in advance according to these prediction results, making the reward mechanism more forward-looking and able to better motivate users to participate in various activities of the platform. Different from the reward adjustment based on past data in the prior art, the present invention can anticipate users' behavior changes in advance through real-time data and behavior prediction, so that the platform's reward strategy always remains efficient and targeted. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0077] Figure 1 is the overall flowchart of a method for designing an Internet integral dynamic incentive mechanism based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.

[0079] Refer to Figure 1 , a method for designing an Internet integral dynamic incentive mechanism based on reinforcement learning, includes the following steps:

[0080] S1. Collect the behavior data of users through sensors and mobile devices and perform preprocessing to construct a user behavior data set;

[0081] S2. Adopt natural language processing technology to obtain the emotional state and situational information of users and generate situational labels;

[0082] S3. Calculate the reward value using the context - adaptive reward algorithm, dynamically adjust the reward intensity and form according to the user behavior dataset and context labels, and introduce a fuzzy logic system in the reward generation;

[0083] S4. Adjust the type and frequency of rewards through the generated reward value and user behavior feedback, and dynamically adjust the reward strategy in combination with the user's task participation and feedback loop. The feedback loop is optimized based on the user's reward acceptance and behavior changes;

[0084] S5. Based on the adjusted reward strategy, use the gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, participation, and points of interest, and generate the user's behavior prediction results;

[0085] S6. According to the user's behavior prediction results, dynamically optimize the reward strategy by combining the chaotic search strategy and the wolf pack optimization algorithm. The optimization process searches for the optimal strategy in the multi - objective optimization space by simulating the foraging behavior of wolf packs to obtain the optimized reward strategy;

[0086] S7. According to the optimized reward strategy, collect the user's feedback information, including the user's acceptance of rewards, task participation, and integral usage, evaluate the effectiveness of rewards and changes in user interest, and obtain the final global reward strategy based on the A3C algorithm. The A3C algorithm adapts the global reward strategy to the needs in different user contexts by combining exploration and exploitation, realizing the real - time dynamic adjustment of the Internet integral incentive mechanism.

[0087] In this embodiment, the sensors in S1 include GPS, accelerometers, and gyroscopes.

[0088] In this embodiment, S3 specifically includes:

[0089] S31. Construct the user's behavior characteristics according to the user behavior dataset. The behavior characteristics include the user's activity frequency, task participation, and historical behavior;

[0090] S32. Combine the context labels and use the context - adaptive reward algorithm to calculate the reward value:

[0091]

[0092] Among them, R represents the reward value, α, β, and δ represent adjustment parameters, w i represents the weight of the behavior characteristic x i of, x i represents the i - th behavior characteristic, n represents the total number of user behavior characteristics, θ j represents the weight of the context characteristic y j of, y jrepresents the j-th situational feature, m represents the total number of user situational features, φ k represents the feedback weight, b k represents the evaluation index of the feedback, l represents the number of feedback items;

[0093] S33. After calculating the reward value, introduce a fuzzy logic system to further adjust the reward value. The fuzzy logic system performs fuzzy processing on the reward value through a membership function to further adjust the intensity of the reward, so as to adapt to the long-term behavior pattern and immediate situation of the user;

[0094] S34. Based on the adjustment result of the fuzzy logic system, use an adaptive optimization algorithm to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the feedback information in the process of the user's behavior change and situation adjustment. The optimized reward value meets the different user behavior patterns and task requirements;

[0095] S35. According to the user behavior data set and situation labels, dynamically adjust the reward intensity and form. The adjustment process combines the user's task completion degree, behavior frequency and situation change to ensure that the reward adapts to different behavior patterns;

[0096] S36. According to the dynamically adjusted reward value, generate a reward strategy and update the reward strategy in real time to dynamically adjust the Internet integral incentive mechanism.

[0097] In this embodiment, the specific content of S4 includes:

[0098] S41. According to the generated reward value and user behavior feedback, combined with the user's task participation situation and situation labels, determine the reward adjustment factor:

[0099]

[0100] where λ t represents the reward adjustment factor, w i represents the weight of the behavior feature x i Δx i (t) represents the change amount of the behavior feature at time t, n represents the total number of user behavior features, θ j represents the weight of the situational feature y j Δy j (t) represents the change amount of the situational feature at time t, m represents the total number of user situational features, ΔB t represents the comprehensive feedback amount of reward acceptance and behavior change, and α1, β1 and γ1 represent adjustment parameters;

[0101] S42. Based on the reward adjustment factor λ t dynamically adjust the intensity and type of the reward:

[0102]

[0103] Among them, R' t represents the adjusted reward value, R t represents the initial reward value, δ1 represents the adjustment coefficient, and φ k represents the feedback weight, and ΔB k (t) represents the change amount of the k-th feedback, and l represents the number of feedback items;

[0104] S43. According to the adjusted reward value, in combination with the user's acceptance of the reward and task participation, dynamically adjust the type and frequency of the reward, and optimize the adjustment process based on the user behavior feedback:

[0105] ΔR t =γ2·(A t ·ΔX t ) + γ3·(P t ·ΔT t );

[0106] Among them, ΔR t represents the change amount of the adjusted reward value, A t represents the user's acceptance of the reward, P t represents the user's task participation, ΔX t represents the change amount of the user's behavior, ΔT t represents the change amount of the task completion situation, and γ2 and γ3 represent the adjustment parameters;

[0107] S44. By collecting the user's feedback information on the reward in real time, adjust the type and frequency of the reward using a feedback loop, and the feedback loop includes the user's acceptance of the reward, task participation situation, and behavior change information, and the behavior change information is updated in real time through the user behavior data set and context labels.

[0108] In this embodiment, the specific steps of S5 are as follows:

[0109] S51. Based on the user's behavior feature vector, use the gated recurrent unit model to process the behavior data, and the state update formula of the gated recurrent unit model is:

[0110] h t =(1 - z t )·h t-1 +z t ·tanh(W h ·X t +b h );

[0111] Among them, h t represents the hidden state at time t, and zt denotes the update gate at time t, h t-1 denotes the hidden state at time t-1, tanh denotes the hyperbolic tangent function, W h denotes the weight matrix, b h denotes the bias term, X t denotes the behavioral feature vector;

[0112] S52. By training the gated recurrent unit model, obtain the dynamic hidden state of the user behavior, and based on the dynamic hidden state, use the fully connected layer to map it to the prediction result of the future behavior, including the future activity frequency, participation degree and interest points of the user;

[0113] S53. Through the combination of the prediction result and the user's historical behavior data, further adjust the weight coefficient of the future behavior prediction model to generate a multi-step prediction result:

[0114]

[0115] where, denotes the k-step behavior trend of the prediction, ζ denotes the weighting coefficient, denotes the behavior prediction result at time t, denotes the (k-1)-step behavior trend of the prediction, and optimize the prediction accuracy by adjusting the balance between the current and previous predictions;

[0116] S54. Calculate the comprehensive trend of the user's future behavior by the weighted average method according to the multi-step prediction result to obtain the final behavior prediction result of the user.

[0117] In this embodiment, the S6 specifically includes:

[0118] S61. According to the behavior prediction result of the user, construct a multi-objective optimization problem, and the objectives include the maximization of the user reward value, the improvement of the user participation degree, and the long-term stability of the reward strategy:

[0119]

[0120] where, L(w) represents the objective function, N represents the total number of reward objectives, μ i , λ1 and λ2 represent the weight coefficients, w represents the parameters of the reward strategy, Reward i (w) represents the value of the i-th reward objective, Engagement(w) represents the user participation degree, and Stability(w) represents the stability of the reward strategy;

[0121] S62. Initialize the population of the wolf pack optimization algorithm, generate the initial solution through the chaotic search strategy, and initialize the population by using the chaotic mapping, where the chaotic mapping process is:

[0122] x t+1 = ξ·x t ·(1 - x t ) with x0 ∈ [0, 1];

[0123] Wherein, x t+1 represents the position of the wolf individual at the (t + 1)-th iteration, x t represents the position of the wolf individual at the t-th iteration, ξ represents the control parameter of the chaotic map, and x0 represents the initial value;

[0124] S63. After initializing with chaos, evaluate the fitness of each wolf individual based on the objective function;

[0125] S64. Based on the fitness evaluation results, use the simulation of the wolf pack foraging behavior to update the position, and adopt an update formula with chaotic perturbation to enhance the global search ability:

[0126] w new = w old + κ1·(p best - w old ) + κ2π(g best - w old ) + ηπF(x t );

[0127] Wherein, w new represents the updated reward strategy, w old represents the current reward strategy, κ1 represents the influence degree of the individual optimal position on the update, κ2 represents the influence degree of the global optimal position on the update, p best represents the best solution of the current wolf individual in the search space, g best represents the optimal solution among all wolf individuals, η represents the chaotic perturbation coefficient, and F(x t ) represents the chaotic perturbation function, which perturbs the position of the wolf individual based on the current state of the chaotic map;

[0128] S65. According to the updated individual positions, achieve policy optimization by combining the exploration and exploitation phases. In the exploration phase, rely on chaotic perturbation for extensive search, and in the exploitation phase, refine the reward strategy through local search to obtain the optimized reward strategy.

[0129] In this embodiment, the specific steps of S7 include:

[0130] S71. According to the optimized reward strategy, collect the feedback information of the users. The feedback information includes the users' acceptance of the rewards, task participation, integral usage, reward consumption rate, and changes in user interests, and form a user feedback data set;

[0131] S72. Construct a reward effect evaluation function based on the user feedback dataset. The reward effect evaluation function is modeled in combination with the user feedback dataset, and the behavioral state s is defined. t As the state of the user at time step t, it includes the user's activity frequency and task participation information, and evaluates the user's acceptance of the reward according to the behavioral state s. t Evaluate the user's acceptance of the reward:

[0132]

[0133] Among them, RA*s t (represents the reward effect evaluation function, ρ i represents the weight of the user feedback data item, Feedback i represents the various indicators of the user feedback, τ represents the influence weight of the user task participation on the reward acceptance, and Engagement t represents the user's participation at time step t;

[0134] S73. Use the parallel agents in the A3C algorithm to calculate and optimize the strategy of the user behavior. In the A3C algorithm, the policy network and the value network are used to predict the reward expectation and the actual reward of the user in different states respectively.

[0135] S74. Combine the exploration and exploitation strategies to dynamically adjust the global reward strategy. The exploration part explores the strategy by randomly selecting actions to obtain the final global reward strategy:

[0136]

[0137] Among them, θ newexplore represents the parameters of the updated exploration policy network, θ explore represents the parameters of the exploration policy network before update, θ newexploit represents the parameters of the updated exploitation policy network, θ exploit represents the parameters of the exploitation policy network before update, ∈ represents the balance coefficient of exploration and exploitation, represents the loss function gradient in the exploration stage, represents the loss function gradient in the exploitation stage.

[0138] Example 1:

[0139] To verify the feasibility of the present invention in implementation, the present invention is applied to the user point incentive system of an e-commerce platform. The daily active user number of this e-commerce platform reaches 5 million, with a huge user base and rich user behavior data. The user behavior data of the platform includes browsing records, shopping records, search behaviors, sharing behaviors, comment behaviors, etc. These data not only reflect the activity frequency of users within the platform, but also can reflect users' interests in different products, activities, and promotion strategies. However, the existing point incentive mechanism of the platform is too single, and the reward method is fixed, and it cannot be effectively adjusted according to the personalized needs and situational changes of users, resulting in a gradual decline in user participation and the activity of the platform being affected. Therefore, the platform urgently needs a solution that can adjust the reward strategy in real time, improve user activity and loyalty.

[0140] Specifically, the platform first collects the behavior data of users, constructs the behavior feature vector of each user, which contains multi-dimensional information such as the activity frequency of users, task participation, historical behaviors, and situational factors. On this basis, the platform uses the context-adaptive reward algorithm, combines context tags, and calculates the reward value for each user. This reward value not only considers the current behavior of users, but also further adjusts the reward value through a fuzzy logic system to make it adapt to the long-term behavior patterns and immediate situations of different users. In addition, the platform also improves the task completion rate and participation rate of users by dynamically adjusting the reward strategy. For example, when a user has not participated in tasks for a long time, the reward form will be more attractive to encourage the user to participate again; while when a user actively participates, the reward intensity will gradually increase to ensure the diversity and adaptability of the reward strategy.

[0141] During the implementation process, the platform also uses the gated recurrent unit (GRU) model to predict the future behaviors of users, and combines the chaotic search strategy and the wolf pack optimization algorithm to dynamically optimize the reward strategy. The optimization process simulates the foraging behavior of wolf packs to ensure that the reward strategy always remains efficient under different user behavior patterns. Finally, the platform uses the A3C algorithm to adjust the reward strategy in real time, evaluates the acceptance of rewards by users, task participation, and point usage, so as to further optimize the global reward strategy.

[0142] To verify the effectiveness of the present invention, the platform collected data before and after implementation for comparison. Before implementation, the daily active user count of the platform was 5 million. After implementation, the daily active user count increased to 6.005 million, with a growth rate of 20.10%. In terms of task completion rate, it was 50.25% before implementation and increased to 75.38% after implementation, with an increase rate of 50.25%. In terms of user retention rate, the next-day retention rate of the platform was 30.10% before implementation and increased to 45.45% after implementation. In addition, the average number of task participations per user per month increased from 5.12 times to 8.13 times, with a growth rate of 58.88%. These data indicate that the present invention has significant effects in enhancing user activity, task completion rate, user participation, and loyalty.

[0143] Table 1 Comparison Table of User Behaviors before and after the Implementation of the Internet Point Incentive Mechanism Based on Reinforcement Learning

[0144] Indicator Before implementation After implementation Growth rate Daily active user count 5 million 6.005 million 20.10% Task completion rate 50.25% 75.38% 50.25% User retention rate 30.10% 45.45% 15.15% Average number of task participations 5.12 times per month 8.13 times per month 58.88% User reward acceptance rate 70.20% 85.15% 21.30%

[0145] As can be clearly seen from Table 1 above, the implementation of the present invention has effectively improved the user activity, task completion rate, participation, and loyalty of the platform. In particular, by dynamically adjusting the reward strategy, the platform can provide personalized and flexible rewards for each user, ensuring that the reward mechanism always remains efficient, thereby enhancing the stickiness and loyalty of users to the platform.

[0146] In summary, the design method of the Internet point dynamic incentive mechanism based on reinforcement learning of the present invention can dynamically adjust the reward strategy according to the user's behavior and context information, significantly improving the user activity, task completion rate, and user participation of the platform. At the same time, by continuously optimizing the reward strategy, the present invention can ensure the long-term adaptability of the point incentive mechanism, bringing continuous user growth and loyalty improvement to the platform.

[0147] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning, characterized in that It includes the following steps: S1. Collect the user's behavior data through sensors and mobile devices and perform preprocessing to construct a user behavior dataset; S2. Adopt natural language processing technology to obtain the user's emotional state and situational information, and generate situational labels; S3. Use a situation-adaptive reward algorithm to calculate the reward value, dynamically adjust the reward intensity and form according to the user behavior dataset and situational labels, and introduce a fuzzy logic system in the reward generation; S4. Adjust the type and frequency of rewards through the generated reward value and user behavior feedback, combine the user's task participation and feedback loop, and dynamically adjust the reward strategy. The feedback loop is optimized based on the user's reward acceptance and behavior changes; S5. Based on the adjusted reward strategy, use a gated recurrent unit model to predict the user's future behavior trends, including the user's future activity frequency, participation, and interest points, and generate the user's behavior prediction results; S6. According to the user's behavior prediction results, combine the chaotic search strategy and the wolf pack optimization algorithm to dynamically optimize the reward strategy. The optimization process searches for the optimal strategy in the multi-objective optimization space by simulating the foraging behavior of the wolf pack to obtain the optimized reward strategy; S7. According to the optimized reward strategy, collect the user's feedback information, including the user's acceptance of the reward, task participation, and integral usage, evaluate the effectiveness of the reward and the change of user interest, and obtain the final global reward strategy based on the A3C algorithm. The A3C algorithm adapts the global reward strategy to the needs in different user situations by combining exploration and exploitation, and realizes the real-time dynamic adjustment of the Internet integral incentive mechanism.

2. The design method of an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that The sensors in S1 include GPS, accelerometers, and gyroscopes.

3. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that, The specific content of S3 includes: S31. Construct the user's behavior characteristics according to the user behavior dataset. The behavior characteristics include the user's activity frequency, task participation, and historical behavior; S32. Combine the situational labels and use a situation-adaptive reward algorithm to calculate the reward value; where, R represents the reward value, α, β, and δ represent adjustment parameters, w i represents the weight of the behavioral feature x i and x i represents the i-th behavioral feature, n represents the total number of user behavioral features, θ j represents the weight of the situational feature y j and y j represents the j-th situational feature, m represents the total number of user situational features, φ k represents the feedback weight, b k represents the evaluation index of the feedback, and l represents the number of feedback items; S33. After calculating the reward value, introduce a fuzzy logic system to further adjust the reward value. The fuzzy logic system performs fuzzy processing on the reward value through a membership function to further adjust the reward intensity for adapting to the user's long-term behavior pattern and immediate situation; S34. Based on the adjustment result of the fuzzy logic system, use an adaptive optimization algorithm to optimize the reward value. The adaptive optimization algorithm gradually adjusts the reward value according to the user's behavior changes and feedback information during the situation adjustment. The optimized reward value meets different user behavior patterns and task requirements; S35. Dynamically adjust the reward intensity and form according to the user behavior dataset and situational labels. The adjustment process combines the user's task completion degree, behavior frequency, and situation change to ensure that the reward adapts to different behavior patterns; S36. Generate a reward strategy according to the dynamically adjusted reward value and update the reward strategy in real time for dynamically adjusting the Internet integral incentive mechanism.

4. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that, The specific content of S4 includes: S41. Determine a reward adjustment factor based on the generated reward value and user behavior feedback, in combination with the user's task participation and situation tags: Among them, λ t represents the reward adjustment factor, w i represents the weight of the behavior feature x i , Δx i (t) represents the change amount of the behavior feature at time t, n represents the total number of user behavior features, θ j represents the weight of the situation feature y j , Δy j (t) represents the change amount of the situation feature at time t, m represents the total number of user situation features, ΔB t represents the comprehensive feedback amount of the reward acceptance and the behavior change, and α1, β1, and γ1 represent the adjustment parameters; S42. Based on the reward adjustment factor λ t Dynamically adjust the intensity and type of rewards: Among them, R' t represents the adjusted reward value, R t represents the initial reward value, δ1 represents the adjustment coefficient, φ k represents the feedback weight, ΔB k (t) represents the change amount of the k-th feedback, and l represents the number of feedback items; S43. Dynamically adjust the type and frequency of rewards based on the adjusted reward value, in combination with the user's acceptance of the reward and task participation, and optimize the adjustment process based on user behavior feedback: ΔR t = γ2π(A t ·ΔX t ) + γ3·(P t ·ΔT t ); Among them, ΔR t represents the change in the adjusted reward value, A t represents the user's reward acceptance, P t represents the user's task participation, ΔX t represents the change in the user's behavior, ΔT t represents the change in the task completion situation, and γ2 and γ3 represent adjustment parameters; S44. Adjust the type and frequency of rewards through a feedback loop by collecting the user's feedback information on the rewards in real time. The feedback loop includes the user's acceptance of the reward, task participation, and behavior change information, and the behavior change information is updated in real time through the user behavior dataset and situation tags.

5. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that The specific steps of S5 are as follows: S51. Process the behavior data using a gated recurrent unit model based on the user's behavior feature vector. The state update formula of the gated recurrent unit model is: h t = (1 - z t )·h t-1 + z t ·tanh(W h ·X t + b h ); Among them, h t represents the hidden state at time t, z t represents the update gate at time t, h t-1 represents the hidden state at time t-1, tanh represents the hyperbolic tangent function, W h represents the weight matrix, b h represents the bias term, X t represents the behavioral feature vector; S52. Train the gated recurrent unit model to obtain the dynamic hidden state of the user's behavior, and based on the dynamic hidden state, map it to the prediction results of future behaviors using a fully connected layer, including the user's future activity frequency, participation, and interest points; S53. Further adjust the weight coefficients of the future behavior prediction model by combining the prediction results and the user's historical behavior data to generate multi-step prediction results: Among them, represents the predicted behavior trend at the k-th step, ζ represents the weighting coefficient, represents the behavior prediction result at time t, represents the predicted behavior trend at the (k - 1)-th step, and optimizes the prediction accuracy by adjusting the balance between the current and previous predictions; S54. Calculate the comprehensive trend of the user's future behavior by the method of weighted average according to the multi-step prediction results to obtain the final behavior prediction result of the user.

6. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that, The specific steps of S6 are as follows: S61. Construct a multi-objective optimization problem according to the user's behavior prediction results, with the objectives including maximizing the user's reward value, improving the user's participation, and the long-term stability of the reward strategy: Among them, L(w) represents the objective function, N represents the total number of reward objectives, μ i , λ1 and λ2 represent the weight coefficients, w represents the parameters of the reward strategy, Reward i (w) represents the value of the i-th reward objective, Engagement(w) represents the user engagement, and Stability(w) represents the stability of the reward strategy; S62. Initialize the population of the wolf pack optimization algorithm, generate the initial solution through a chaotic search strategy, and initialize the population using a chaotic map. The chaotic map process is as follows: x t+1 = ξ·x t ·(1 - x t ) with x0 ∈ [0, 1]; Among them, x t+1 represents the position of the wolf individual at the (t + 1)-th iteration, and x t represents the position of the wolf individual at the t-th iteration, ξ represents the control parameter of the chaotic map, and x0 represents the initial value; S63. After using chaotic initialization, evaluate the fitness of each wolf individual based on the objective function; S64. Based on the fitness evaluation results, update the position by simulating the foraging behavior of the wolf pack, and use an update formula with chaotic perturbation to enhance the global search ability: w new = w old + κ1·(p best - w old ) + κ2·(g best - w old ) + η·F(x t ); Among them, w new represents the updated reward strategy, w old represents the current reward strategy, κ1 represents the influence degree of the individual optimal position on the update, κ2 represents the influence degree of the global optimal position on the update, p best represents the best solution of the current wolf individual in the search space, g best represents the optimal solution among all wolf individuals, η represents the chaotic perturbation coefficient, F(x t ) represents the chaotic perturbation function, which perturbs the position of the wolf individual based on the current state of the chaotic map; S65. According to the updated individual positions, achieve strategy optimization by combining the exploration and exploitation phases. In the exploration phase, rely on chaotic perturbation for extensive search, and in the exploitation phase, refine the reward strategy through local search to obtain the optimized reward strategy.

7. A design method for an Internet integral dynamic incentive mechanism based on reinforcement learning according to claim 1, characterized in that The specific steps of S7 are as follows: S71. Collect the user's feedback information according to the optimized reward strategy. The feedback information includes the user's acceptance of the reward, task participation, integral usage, reward consumption rate, and user interest changes, and form a user feedback dataset; S72. Construct a reward effect evaluation function based on the user feedback dataset. The reward effect evaluation function is modeled by combining the user feedback dataset, and define the behavior state s t as the state of the user at time step t, including the user's activity frequency and task participation information, and evaluate the user's acceptance of the reward according to the behavior state s t : Among them, RA(s t ) represents the reward effect evaluation function, ρ i represents the weight of the user feedback data item, Feedback i represents each index of the user feedback, τ represents the influence weight of the user task participation degree on the reward acceptance degree, Engagement t represents the participation degree of the user at time step t; S73. Use the parallel agents in the A3C algorithm to calculate and optimize the strategy of the user's behavior. In the A3C algorithm, the policy network and value network are used to predict the user's reward expectation and actual reward in different states respectively; S74. Combine the exploration and exploitation strategies to dynamically adjust the global reward strategy. The exploration part conducts strategy exploration by randomly selecting actions to obtain the final global reward strategy: Among them, θ newexplore represents the parameters of the updated exploration policy network, θ explore represents the parameters of the exploration policy network before update, θ newexploit represents the parameters of the updated exploitation policy network, θ exploit represents the parameters of the exploitation policy network before update, ∈ represents the balance coefficient between exploration and exploitation, represents the gradient of the loss function in the exploration phase, represents the gradient of the loss function in the exploitation phase.

Citation Information

Patent Citations

  • Deep reinforcement learning-based incomplete information game method, device, system and storage medium

    CN110399920A

  • Internet of Things acquisition platform computing resource scheduling method based on reinforcement learning

    CN117687791A

  • End-to-end automatic driving control system and device based on human preference reinforcement learning

    CN119018181A

  • Data security dynamic protection method based on artificial intelligence

    CN119382949A

  • Digital factory operation virtual simulation teaching method and system

    CN119396096A

Cited By

  • Personalized reward distribution method and system based on deep reinforcement learning

    CN120634636A

  • Deep reinforcement learning based personalized reward allocation method and system

    CN120634636B

  • Reinforcement learning course intervention method oriented to learning motivation decline

    CN121010486A