Information processing device, information processing method, and information processing program

The information processing device enhances task completion rates in interventions by optimizing negative incentives based on individual sensitivity to initial and provisionally acquired incentives, addressing the limitations of uniform incentive approaches.

WO2025163774A1PCT designated stage Publication Date: 2025-08-07NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002946
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing interventions using negative incentives to encourage behaviors such as health and learning behaviors have not maximized task completion rates due to a lack of consideration for individual sensitivity to incentives, which is influenced by the initial, provisionally acquired, and negative incentives.

Method used

An information processing device optimizes an incentive policy by constructing a behavioral model that incorporates the user's sensitivity to incentives, including initial and provisionally acquired incentives, to determine the amount of negative incentives, using a mathematical model and machine learning techniques like Deep Q Network to maximize motivation and task completion rates.

Benefits of technology

The optimized incentive policy increases task completion rates by personalizing the negative incentive amounts based on individual sensitivity, surpassing previous studies that used uniform incentives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002946_07082025_PF_FP_ABST
    Figure JP2024002946_07082025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to one embodiment of the present invention optimizes an incentive policy used to determine the amount of negative incentive to be presented when imposing a task on a user, in an intervention using the negative incentive. The incentive policy is a function that takes the user's behavior history as input, and outputs an amount of negative incentive. An optimization unit optimizes the incentive policy by using a behavior model that incorporates the negative incentive and at least one of a temporarily granted incentive that is an incentive granted temporarily and an initially granted incentive that is an incentive granted at the initial stage of the intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and information processing program

[0001] The present invention relates to a technique for promoting user behavior through intervention using incentives.

[0002] Although humans have behaviors that should be continued daily, such as health behaviors and learning behaviors, it can be difficult to continue these behaviors voluntarily. Therefore, interventions using financial incentives are being carried out to encourage such behaviors to continue.

[0003] Non-Patent Document 1 conducts an experimental program that uses monetary incentives to encourage healthy behavior, and experiments were conducted using the following three incentive program designs. The tasks described below are the performance of specified actions. - Program using positive incentives If the task is completed, a fixed amount of incentive is given each time. If the task is not completed, no incentive is given. - Program using lottery-based incentives If the task is completed, the amount of incentive is determined by drawing lots and given. If the task is not completed, no incentive is given. - Program using negative incentives If the task is completed, the previously given incentive is maintained. If the task is not completed, a fixed amount of incentive is confiscated each time.

[0004] As a result, it was found that the program using negative incentives was the most effective in encouraging healthy behavior. Non-Patent Document 2 also claims that the program using negative incentives increased task completion rates.

[0005] A program using negative incentives has the following three characteristics: (1) An incentive is first granted when the program starts. Hereinafter, the incentive granted when the program starts will be referred to as the initially granted incentive. (2) When a task is assigned within the program, a provisionally granted incentive and a negative incentive are presented. Hereinafter, the provisionally granted incentive will be referred to as the provisionally acquired incentive, and the negative incentive will simply be referred to as the negative incentive. The provisionally acquired incentive is the initially granted incentive that has been maintained or reduced. The negative incentive is an incentive that will be forfeited if the task is not completed. (3) If the task is completed, the provisionally acquired incentive is maintained, but if the task is not completed, the negative incentive is forfeited from the provisionally acquired incentive.

[0006] Mitesh S. Patel et al., "Framing Financial Incentives to Increase Physical Activity Among Overweight and Obese Adults", Annals of Internal Medicine, Vol. 164, issue 6, p.385-394, March 15, 2016.Neel P. Chokshi et al., "Loss‐Framed Financial Incentives and Personalized Goal-Setting to Increase Physical Activity Among Ischemic Heart Disease Patients Using Wearable Devices: The ACTIVE REWARD Randomized Trial", Journal of the American Heart Association, Vol. 7, No. 12, June 9, 2018. Woohyeok Choi & Uichin Lee, "Loss-Framed Adaptive Microcontingency Management for Preventing Prolonged Sedentariness: Development and Feasibility Study", JMIR mHealth and uHealth, Vol 11, Jan 27, 2023.

[0007] Both Non-Patent Documents 1 and 2 conclude that programs using negative incentives increased task completion rates, so programs using negative incentives can be said to be a powerful intervention strategy among incentive-based intervention methods.

[0008] Sensitivity to incentives should differ between individuals, but previous studies such as those disclosed in Non-Patent Documents 1 and 2 have only involved a uniform intervention in which a fixed amount of incentive is presented to all subjects each time. In contrast, Non-Patent Document 3 linked the motivation to complete a task with the amount of negative incentive and the circumstances for performing the task (weather, work situation, etc.) so that the amount of negative incentive presented to each individual would vary, but this did not lead to an increase in the task completion rate. In other words, it is still unclear how to provide negative incentives that maximize the task completion rate while taking into account each individual's sensitivity to incentives.

[0009] Furthermore, in order to consider an individual's sensitivity to incentives, we must not only focus on the amount of negative incentives presented. As mentioned above, in a program using negative incentives, not only the negative incentives but also the initial incentives and provisionally acquired incentives are presented. In other words, it is thought that motivation to accomplish a task is determined by comparing the amount of negative incentives with the amount of the initial incentives and provisionally acquired incentives.

[0010] Therefore, when determining how to provide negative incentives that maximize task completion rates, taking into account each individual's sensitivity to incentives, it is necessary to consider the initial incentives and provisional incentives that are compared with the negative incentives.

[0011] The present invention aims to provide a technique that can increase task completion rates in interventions using negative incentives compared to previous studies.

[0012] An information processing device according to one aspect of the present invention optimizes an incentive policy used to determine the amount of a negative incentive to present to a user when assigning a task to the user in an intervention using a negative incentive. The incentive policy is a function that takes the user's behavioral history as input and outputs the amount of the negative incentive, and an optimization unit optimizes the incentive policy using a behavioral model that incorporates the negative incentive and at least one of a provisionally acquired incentive, which is an incentive that is provisionally granted, and an initially granted incentive, which is an incentive that is granted at the beginning of the intervention.

[0013] According to the present invention, a technique can be provided that can increase the task completion rate in interventions using negative incentives compared to previous studies.

[0014] Fig. 1 is a block diagram showing the functional configuration of an information processing apparatus according to an embodiment, Fig. 2 is a block diagram showing the hardware configuration of a computer that can realize the information processing apparatus according to an embodiment, and Fig. 3 is a flowchart showing an information processing method according to an embodiment.

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0016] In an embodiment, a mathematical model (hereinafter referred to as a behavioral model) is prepared for each user, which inputs the negative incentive, the provisionally acquired incentive, and / or the initially granted incentive, and outputs motivation for completing a task. An incentive policy for each user is optimized based on the behavioral model of each user. Here, the incentive policy is defined as a function that inputs past behavioral history and outputs the amount of negative incentive to be presented next time. The optimized incentive policy can output an amount of negative incentive that increases the user's motivation to complete a task.

[0017] In constructing the behavioral model, a comparison target for a negative incentive is introduced to take into account each user's sensitivity to incentives. This comparison target is at least one of the provisionally acquired incentive and the initial incentive. Various amounts of the negative incentive can be considered, with the provisionally acquired incentive as the upper limit. However, it is more desirable to consider not only the absolute amount of the negative incentive but also its magnitude relationship with the provisionally acquired incentive or the initial incentive. By optimizing the incentive policy under a behavioral model that also takes into account the absolute amount of the provisionally acquired incentive or the initial incentive, it is possible to realize an intervention using a negative incentive that maximizes average motivation during the intervention period.

[0018] 1 is a schematic diagram illustrating an example of the functional configuration of an information processing device 10 according to an embodiment of the present invention. As shown in FIG. 1, the information processing device 10 includes an estimation unit 11, an optimization unit 12, and an output unit 13.

[0019] The estimation unit 11 receives experimental data 14 and a motivation function 15 corresponding to the behavioral model. The experimental data 14 and motivation function 15 are prepared in advance. The experimental data 14 is generated for each user, and the motivation function 15 is common to all users. Here, we will explain the case where there is one user. If there are multiple users, the series of processes described below will be performed for each user.

[0020] The experimental data 14 is obtained by implementing an experimental program in which a user is intervened with a negative incentive and made to accomplish a task. The task indicates an action to be performed, such as exercising for a predetermined period of time (for example, 30 minutes of exercise) or studying for a predetermined period of time. The experimental program is carried out over T days, and a task is assigned to the user every day. The user is given a negative incentive p on the first day of the experimental program period. 0 Then, when each task is assigned, the provisional incentive value p t-1 and the value of the negative incentive x tFor example, on the Tth day, "When the task on the tth day is completed, the provisional incentive will be p t-1 If it fails, p t-1 From x t The message "x negative incentive will be confiscated." is displayed. t is a value randomly selected from the set X, where 0≦x t ≦p t-1 and the provisional acquisition incentive p t-1 satisfies the following formula: Here, t' denotes the time step t when the user fails to complete the task. Failure in completing the task means that the task cannot be completed, and success in completing the task means that the task can be completed. The initial incentive p 0 is the provisional acquisition incentive on the first day.

[0021] The experimental data 14 obtained from the above experimental program is the initial incentive p 0 , the negative incentive x presented at each time step 1 , x 2 , …, x T , the provisional incentive p presented at each time step 1 , p 2 , ..., p T-1 , and information y indicating whether the task is accomplished at each time step. 1 , y 2 , ..., y T Contains y t = 1 indicates that the user successfully completed the task on the tth day, and y t = 0 indicates that the user failed to complete the task on day t.

[0022] The motivation function 15 is a function that receives at least one of the negative incentive, the provisionally acquired incentive, and the initially granted incentive as input, and outputs the motivation for accomplishing the task.

[0023] The probability of success in completing the task at each time step t is u tand the situation is defined as one-hot encoded vector C = {c 1 , ..., c n Here, n is the number of situations to be considered, such as the weather and whether or not there is work.

[0024] The motivation function (behavioral model) is created taking into consideration the following three points: (1) Motivation increases as negative incentives increase, (2) Motivation changes depending on the situation, and (3) Motivation increases sharply when negative incentives and provisionally acquired incentives exceed a certain threshold.

[0025] The above items (1) and (2) have been considered in conventional research, and the above item (3) is a new item that is taken into consideration in the present invention.

[0026] Items (1) and (3) can be expressed by using a sigmoid function as the motivation function, and item (2) can be expressed by adding a term representing the situation to the incentive term.

[0027] In addition, the following factors are thought to be involved in motivation for task completion: (a) Absolute amount of negative incentive (b) Absolute amount of provisionally acquired incentive (c) Absolute amount of initially granted incentive (d) Ratio of negative incentive to provisionally acquired incentive (e) Ratio of negative incentive to initially granted incentive (f) Ratio of provisionally acquired incentive to initially granted incentive

[0028] In conventional research, only factor (a) is incorporated into the motivation function. In this embodiment, one of factors (b) to (f) is incorporated into the motivation function to introduce a comparison target for negative incentives. Preferably, factor (d) or factor (e) is incorporated into the motivation function. In this case, it is permitted to incorporate any of the other factors (b), (c), and (f) into the motivation function. Here, two specific motivation functions are described as examples. Motivation function incorporating factors (b) and (d) ・Motivation function incorporating factors (b) and (e) Here, β 0 is the response coefficient to motivation of negative incentives, γ is the response coefficient to motivation of provisionally acquired incentives, and α th is the threshold for negative incentives and provisionally acquired incentives or the threshold for negative incentives and initially granted incentives, and B = {β 1 , …, β n} is the response coefficient for task success depending on the situation. In other words, the parameter that represents the sensitivity of individuals to incentives is β 0 , γ, α th There are four types:

[0029] The advantage of defining the motivation function in this way is that by introducing a comparison target for negative incentives and freely combining the factors involved in the behavioral model, it is possible to construct a personalized behavioral model that expresses each individual's sensitivity to incentives.

[0030] The estimation unit 11 estimates the amount representing the user's sensitivity to incentives based on the experimental data 14 and the motivation function 15. Specifically, the estimation unit 11 estimates the amount representing the user's sensitivity to incentives based on the experimental data 14. 0 , γ, α th , B is calculated.

[0031] The experimental data 14 includes, as observed values, the negative incentive presented at each time step, the provisionally acquired incentive presented at each time step, and information indicating whether the task was successfully accomplished at each time step. In summary, the experimental data 14 is 1 , p 0 , y 1 ), ..., (x T , p T-1 , y T At each time step, the probability model of task completion can be written as a binomial distribution:

[0032] This is P(yt |u t ) = P(y t |u t (s)), where s is a user-specific parameter β 0 , γ, α th , B. Then, the likelihood L(s) can be written as follows:

[0033] Therefore, the parameters can be found using maximum likelihood estimation and expressed as follows:

[0034] The estimation unit 11 calculates the parameter set s * is output as the parameter estimation result 16. The following is an example of the parameter estimation result 16 obtained by the estimation unit 11. In this example, five types of situations are assumed.

[0035] A user behavior model (a behavior model customized for the user) is obtained by applying the parameter estimation result 16 obtained by the estimation unit 11 to the motivation function 15. The user behavior model is a motivation function that reflects the user's sensitivity to incentives (specifically, negative incentives and at least one of the provisionally acquired incentives and the initially granted incentives).

[0036] The optimization unit 12 receives as input the motivation function 15 and the parameter estimation result 16 obtained by the estimation unit 11. The optimization unit 12 optimizes the amount of negative incentive for the user based on the motivation function 15 and the parameter estimation result 16. Specifically, the optimization unit 12 optimizes the incentive policy based on the motivation function 15 and the parameter estimation result 16, i.e., calculates the optimal incentive policy.

[0037] The incentive policy is the success or failure of the task completion on the (t-1)th day at the current time step t. t-1 , the provisional incentive p presented on the tth day t-1 , and the initial grant incentive p 0The amount of negative incentive to be presented on the tth day is x t The optimal incentive policy may be a function f that outputs: Here, E[·] represents the expected value.

[0038] Success or failure of task completion t follows the following Markov decision process (hereinafter referred to as MDP): State on day t: ・Possible incentive set on day t: The state V on day (t+1) is conditioned on the state on day t and the incentive. t+1 The probability of generating: ・Reward on day t: y t

[0039] However, negative incentive x t The possible values ​​of are N discrete values ​​{a 1 , a 2 , ..., a N}, provisional acquisition incentive p t-1 It is known that in an MDP, a policy that maximizes the expected value of the sum of rewards can be obtained by solving the Bellman optimal equation. In this embodiment, an incentive policy f that satisfies the above formula (3) is * is obtained by solving the Bellman optimal equation. There are several methods for solving the Bellman optimal equation. One example is the Deep Q Network, which uses a neural network. A method for solving the Bellman optimal equation using the Deep Q Network is disclosed, for example, in Volodymyr Mnih et al., "Playing Atari with Deep Reinforcement Learning", arXiv, 2013.

[0040] When using Deep Q Network, the incentive policy f * is the action value function Q(V t , x t) is given as follows:

[0041] Therefore, what is output in this example is a set of model parameters for the neural network as an incentive policy.

[0042] The output unit 13 outputs the optimized incentive policy 17. For example, the output unit 13 transmits the optimized incentive policy to an external terminal device. When assigning a task, the terminal device determines the amount of negative incentive using the optimized incentive policy and displays the determined negative incentive on a display device. Specifically, the terminal device outputs the amount of negative incentive x on the tth day. t In order to determine the success or failure of the task at the current time step t, the (t-1)th day, y t-1 , the provisional incentive p to be presented on the (t-1)th day t-1 , and the initial grant incentive p 0 is input to the incentive policy, and the negative incentive amount output from the incentive policy is the negative incentive amount on day t x t In another example, the output unit 13 may determine the amount of the negative incentive using the optimized incentive policy and may display the determined negative incentive on the display device.

[0043] 2 schematically illustrates an example of the hardware configuration of a computer 20 that can realize the information processing device 10. As illustrated in Fig. 2, the computer 20 includes a processor 21, a RAM (random access memory) 22, a storage device 23, and an input / output interface 24. The processor 21 exchanges signals with the RAM 22, the storage device 23, and the input / output interface 24.

[0044] The processor 21 includes a general-purpose processor such as a CPU (central processing unit) or a GPU (graphics processing unit). The RAM 22 is a volatile memory and is used as a working area for the processor 21. The storage device 23 is a non-volatile memory such as an HDD (hard disk drive) or an SSD (solid state drive). The storage device 203 stores programs such as an information processing program and various data. When executed by the processor 21, the information processing program causes the processor 21 to perform a series of processes described in the embodiments. In other words, the processor 21 is configured to function as an estimation unit 11, an optimization unit 12, and an output unit 13.

[0045] The input / output interface 24 is an interface for communicating with external devices. The input / output interface 24 is used to connect input devices such as a keyboard and a mouse, output devices such as a liquid crystal display device, and a communication network that may include the Internet to the computer 20. The processor 21 communicates with the input devices, output devices, and devices on the communication network via the input / output interface 24.

[0046] A program such as an information processing program may be provided to the computer 20 in a state where it is stored on a computer-readable recording medium. In this case, the computer 20 is equipped with a drive that reads data from the recording medium and acquires the program from the recording medium. Examples of recording media include magnetic disks, optical disks (CD-ROM, CD-R, DVD-ROM, DVD-R, etc.), magneto-optical disks (MO, etc.), and semiconductor memories. The program may also be distributed via a communications network. Specifically, the program may be stored on a server on the communications network, and the computer 20 may download the program from the server.

[0047] The processor 21 may include a dedicated processor such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit) instead of or in addition to a general-purpose processor.

[0048] 3 is a diagram illustrating an example of the procedure of an information processing method according to an embodiment of the present invention. The information processing method illustrated in FIG. 3 is executed by the information processing device 10 illustrated in FIG.

[0049] In step S31 of FIG. 3 , the estimation unit 11 estimates a quantity representing the user's sensitivity to incentives based on experimental data 14 obtained from an experimental program that was previously implemented on the user and involves intervention using a negative incentive, and a motivation function 15 that includes, as a parameter, a quantity representing sensitivity to incentives. The motivation function 15 is configured to receive as input the negative incentive, the provisionally acquired incentive, or at least one of the initially granted incentive, and to output motivation for completing a task. The motivation function 15 can be the function shown in Equation (1) or Equation (2) above. For example, the estimation unit 11 calculates the parameters included in the motivation function 15 from the experimental data 14 using maximum likelihood estimation.

[0050] In step S32, the optimization unit 12 optimizes the incentive policy based on the motivation function 15 and the estimation result (parameter estimation result) of the quantity representing sensitivity to incentives obtained in step S31. The incentive policy is configured to receive as input time step t, success or failure of task completion at time step (t-1), the provisionally acquired incentive presented at time step t, and the initially granted incentive, and to output the amount of negative incentive to be presented at time step t. The optimization unit 12 applies the parameter estimation result obtained in step S31 to the motivation function 15 to obtain a user behavior model, and optimizes the incentive policy using the user behavior model.

[0051] In step S33, the output unit 13 outputs the optimized incentive policy. For example, the output unit 13 transmits the optimized incentive policy to an external terminal device.

[0052] As described above, the information processing device 10 optimizes an incentive policy used to determine the amount of negative incentive to present to a user when assigning a task in an intervention that uses negative incentives to encourage the user to perform a task. The information processing device 10 estimates a quantity representing a user's sensitivity to negative incentives based on experimental data obtained from an experimental program that previously implemented an intervention using negative incentives on the user and a behavioral model (motivation function) that includes, as a parameter, a quantity representing sensitivity to incentives. The information processing device 10 then optimizes the incentive policy based on the behavioral model and the parameter estimation results. Specifically, the incentive policy is optimized to output the amount of negative incentive that maximizes average motivation during the intervention period.

[0053] The above configuration makes it possible to calculate the amount of negative incentive taking into account each individual's sensitivity to incentives, thereby increasing the task completion rate compared to previous studies.

[0054] In one example, the behavioral model may be a function that inputs the negative incentive and the provisionally acquired incentive, as shown in the above formula (1), and outputs motivation for completing the task. In this case, it is possible to calculate the amount of the negative incentive taking into account sensitivity to the negative incentive and the provisionally acquired incentive. In another example, the behavioral model may be a function that inputs the negative incentive and the initially granted incentive, as shown in the above formula (2), and outputs motivation for completing the task. In this case, it is possible to calculate the amount of the negative incentive taking into account sensitivity to the negative incentive and the initially granted incentive.

[0055] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected components from the disclosed components. For example, if the problem can be solved and the effects can be obtained even if some components are removed from all the components shown in the embodiments, the configuration from which these components are removed can be extracted as an invention.

[0056] REFERENCE SIGNS LIST 10: Information processing device 11: Estimation unit 12: Optimization unit 13: Output unit 20: Computer 21: Processor 22: RAM 23: Storage device 24: Input / output interface

Claims

1. An information processing device comprising: an optimization unit that optimizes an incentive policy used to determine the amount of negative incentive to present when assigning a task to a user in an intervention using a negative incentive, wherein the incentive policy is a function that takes the user's behavioral history as input and outputs the amount of negative incentive, and the optimization unit optimizes the incentive policy using a behavioral model that incorporates the negative incentive and at least one of a provisionally acquired incentive, which is an incentive that is temporarily granted, and an initially granted incentive, which is an incentive that is granted at the beginning of the intervention.

2. The information processing device according to claim 1, wherein the behavioral model is configured to receive at least one of the negative incentive, the provisionally acquired incentive, and the initially granted incentive as input, and to output motivation for accomplishing a task.

3. An information processing device as described in claim 2, further comprising an estimation unit that estimates the quantity representing sensitivity to incentives from the behavioral model including a quantity representing sensitivity to incentives as a parameter and experimental data obtained from an experimental program that performs the intervention and that has been implemented in advance on the user, and wherein the optimization unit optimizes the incentive policy based on the behavioral model and the estimation results obtained by the estimation unit.

4. The information processing device of claim 3, wherein the user's behavioral history includes a provisionally acquired incentive presented at a first time step, success or failure of task completion at a second time step prior to the first time step, and an initial granted incentive.

5. An information processing method executed by a computer, comprising optimizing an incentive policy used to determine the amount of negative incentive to present when assigning a task to a user in an intervention using negative incentives, wherein the incentive policy is a function that takes the user's behavioral history as input and outputs the amount of negative incentive, and optimizing the incentive policy includes optimizing the incentive policy using a behavioral model that incorporates the negative incentive and at least one of a provisionally acquired incentive, which is an incentive that is temporarily granted, and an initially granted incentive, which is an incentive that is granted at the beginning of the intervention.

6. An information processing program for causing a computer to execute a means for optimizing an incentive policy used to determine the amount of negative incentive to be presented when a task is assigned to a user in an intervention using a negative incentive, wherein the incentive policy is a function that takes the user's behavioral history as input and outputs the amount of negative incentive, and the means for optimizing the incentive policy comprises optimizing the incentive policy using a behavioral model that incorporates the negative incentive and at least one of a provisionally acquired incentive, which is an incentive that is temporarily granted, and an initially granted incentive, which is an incentive that is granted at the beginning of the intervention.

Citation Information

Patent Citations

  • Information processing device, incentive measure calculation method, and program

    WO2023042382A1

  • Information processing device, information processing method, and information processing program

    WO2023242941A1