Information processing device, information processing method, and information processing program

The information processing device optimizes incentive measures based on user-specific behavioral models to address individual responses, enhancing the effectiveness and cost-efficiency of incentive-based interventions.

JP7758188B2Active Publication Date: 2025-10-22NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024527943
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-10-22
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

Conventional incentive methods do not account for individual differences in response to incentives, leading to inefficiencies in achieving target behaviors, and the effect of incentives fluctuates based on a person's internal state, making it difficult to administer effective incentives.

Method used

An information processing device that acquires behavioral history data, estimates user-specific parameter values for a behavioral model incorporating self-efficacy and self-restoration effects, and calculates optimal incentive measures to sustain target behaviors.

Benefits of technology

Identifies cost-effective incentive measures for each individual, enabling businesses to support target behaviors at a lower cost, thereby increasing profits or reducing service fees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758188000019
    Figure 0007758188000019
  • Figure 0007758188000020
    Figure 0007758188000020
  • Figure 0007758188000021
    Figure 0007758188000021
Patent Text Reader

Abstract

An information processing device according to one embodiment of the present invention comprises: an acquisition unit for acquiring behavior history data for each user and a condition for optimizing an incentive measure for each user; a parameter estimation unit for estimating, on the basis of the behavior history data, a parameter value of a behavior model for each user and having, as an internal variable, success stock indicating the accumulated psychological amount of successful experiences in the past; an optimization unit for calculating an optimum incentive measure for each user on the basis of the estimated parameter value and the condition; and an output unit for outputting the optimum incentive measure.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] In order to achieve a certain target behavior, it is possible to provide an incentive and use that incentive to achieve the target behavior.

[0003] Non-Patent Document 1 describes the achievement of target behavior or the formation of target habits through incentives. For example, Non-Patent Document 1 discloses that, with the aim of forming an exercise habit, the formation of a person's exercise habit is promoted by providing incentives (monetary) according to the amount of exercise. Furthermore, Non-Patent Document 2 discloses that the effect of incentives differs depending on the method of providing the incentive. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Finkelstein, Eric. A., et al., "A Randomized Study of Financial Incentives to Increase Physical Activity among Sedentary Older Adults", Preventive medicine, 47(2), pp.182-187, 2008. [Non-patent document 2] Bachireddy Chethan, et al., "Effect of Different Financial Incentive Structures on Promoting Physical Activity Among Adults: A Randomized Clinical Trial", JAMA Network Open, 2(8), pp.1-13, 2019. Summary of the Invention [Problem to be solved by the invention]

[0005] The magnitude of the effect of incentives in achieving a certain target behavior varies from person to person, even if the amount of incentive is the same. However, conventional techniques do not take into account individual differences in response to incentives. As a result, incentives may not be used effectively for each person. Furthermore, conventional techniques assume that the amount of incentive given each time (daily, weekly, etc.) is either constant, monotonically decreasing, or monotonically increasing, but it is thought that the effect of incentives also fluctuates depending on the person's internal state, which fluctuates from day to day. Therefore, it may be difficult to administer effective incentives using simple incentive giving methods.

[0006] For managers who implement incentive-based interventions, incentives (such as cash or coupons) are directly linked to costs, so it is desirable to achieve high cost-effectiveness, that is, to achieve large effects with fewer incentives.

[0007] The object of the present invention is to address the above-mentioned circumstances, and to provide a technology that can identify the most cost-effective incentive measures for each individual to sustain a target behavior. [Means for solving the problem]

[0008] In order to solve the above problems, one aspect of the present invention is an information processing device comprising: an acquisition unit that acquires behavioral history data for each user and conditions for optimizing incentive measures; a parameter estimation unit that estimates parameter values ​​of a behavioral model for each user based on the behavioral history data, the behavioral model having as an internal variable a success stock that represents the psychological accumulation of past successful experiences; an optimization unit that calculates an optimal incentive measure for each user based on the estimated parameter values ​​and the conditions; and an output unit that outputs the optimal incentive measure. [Effects of the Invention]

[0009] According to one aspect of the present invention, it is possible to identify the most cost-effective incentive measures for each individual to encourage the continuation of target behavior. By using cost-effective incentive measures, businesses can support the achievement of each user's target behavior at a lower cost. This allows businesses to increase profits or lower service fees. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing apparatus according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing the software configuration of the information processing device according to the first embodiment in relation to the hardware configuration shown in FIG. [Figure 3] FIG. 3 is a flowchart showing an example of a parameter estimation operation of the information processing device. [Figure 4] FIG. 4 is a flowchart showing an example of the operation of the information processing device to calculate the optimal incentive policy. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the following, elements that are the same as or similar to elements that have already been described will be designated by the same or similar reference numerals, and duplicated descriptions will basically be omitted.

[0012] First, social cognitive theory research has reported that high self-efficacy increases the probability of achieving a goal. Here, self-efficacy refers to a person's recognition that they have the ability to achieve a goal. In other words, self-efficacy refers to a state in which one believes they can achieve a goal. It has also been reported that past experience of achieving a goal increases self-efficacy. In other words, achieving a goal (for example, achieving 10,000 steps a day) induces further achievement of the goal through self-efficacy. Therefore, the more goals one achieves, the higher one's self-efficacy becomes.

[0013] On the other hand, for people who have a personal threshold for the frequency of achieving a target behavior, achieving the target behavior does not necessarily induce further achievement of the target behavior and may cause a temporary decline in motivation for the target behavior. For example, if a person's goal is to consistently walk 10,000 steps per day, but their threshold is to walk 30,000 steps per week, they are likely to reduce their daily step count in the latter half of the week after achieving close to 30,000 steps in the middle of the week. Conversely, if they are less than 10,000 steps in the middle of the week, they are likely to actively try to increase their daily step count in the latter half of the week.

[0014] In other words, a personal baseline for the frequency of achieving a target behavior has the effect of bringing a person's behavior closer to that baseline. This effect will be referred to as the self-restoring effect hereafter. For example, if a person achieves close to the baseline in the first half of a given period, due to the self-restoring effect, they will not actively try to achieve the target behavior in the second half. Conversely, if a person achieves only a value far from the baseline in the first half of a given period, they will actively try to achieve the target behavior in the second half.

[0015] In this invention, the above-mentioned problem is solved by constructing a mathematical model (hereinafter referred to as a behavioral model) in which incentives are input and the degree of achievement of target behavior is output, while simultaneously taking into consideration self-efficacy and self-restoration effects, and determining the incentive granting method based on the behavioral model.

[0016] [Embodiment] (composition) FIG. 1 is a block diagram showing an example of a hardware configuration of an information processing device 1 according to the first embodiment. The information processing device 1 is realized by a computer such as a PC (Personal Computer). The information processing device 1 includes a control unit 11, an input / output interface 12, and a storage unit 13. The control unit 11, the input / output interface 12, and the storage unit 13 are connected to each other via a bus so as to be able to communicate with each other.

[0017] The control unit 11 controls the information processing device 1. The control unit 11 includes a hardware processor such as a central processing unit (CPU).

[0018] The input / output interface 12 is an interface that enables transmission and reception of information between the input device 2 and the output device 3. The input / output interface 12 may include a wired or wireless communication interface. That is, the information processing device 1, the input device 2, and the output device 3 may transmit and receive information via a network such as a LAN or the Internet.

[0019] The storage unit 13 is a storage medium. The storage unit 13 is configured by combining a nonvolatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), a nonvolatile memory such as a read only memory (ROM), and a volatile memory such as a random access memory (RAM). The storage unit 13 has a storage area including a program storage area and a data storage area. The program storage area stores an operating system (OS), middleware, and application programs required to execute various processes.

[0020] The input device 2 includes, for example, a keyboard, a pointing device, etc., which are used by the owner of the information processing device 1 (for example, an assignor, an administrator, or a supervisor) to input instructions to the information processing device 1. The input device 2 may also include a reader for reading data to be stored in the storage unit 13 from a memory medium such as a USB memory, or a disk device for reading such data from a disk medium. Furthermore, the input device 2 may include an image scanner.

[0021] The output device 3 includes a display that displays output data to be presented to the owner from the information processing device 1, a printer that prints the output data, etc. The output device 3 may also include a writer that writes data to be input to another information processing device 1 such as a PC or a smartphone onto a memory medium such as a USB memory, and a disk device that writes such data onto a disk medium.

[0022] FIG. 2 is a block diagram showing the software configuration of the information processing device 1 according to the first embodiment in relation to the hardware configuration shown in FIG. The storage unit 13 includes an acquired data storage unit 131, a parameter storage unit 132, and an optimized incentive policy storage unit 133.

[0023] The acquired data storage unit 131 stores various data acquired by the later-described acquisition unit 111 of the control unit 11. The data stored in the acquired data storage unit 131 may be acquired by externally importing behavior history data, conditions, etc. via the input device 2, or may include data generated by the control unit 11. The behavior history data and conditions will be described later.

[0024] The parameter storage unit 132 stores the parameter values ​​of the behavior model estimated by the parameter estimation unit 112. The behavior model and the parameter values ​​of the behavior model will be described later.

[0025] The optimized incentive policy storage unit 133 stores the optimal incentive policy calculated by the optimization unit 113, which will be described later. The optimal incentive policy will be described later.

[0026] The control unit 11 includes an acquisition unit 111, a parameter estimation unit 112, an optimization unit 113, and an output control unit 114. These functional units are realized by the hardware processor executing an application program stored in the storage unit 13.

[0027] The acquisition unit 111 acquires necessary data and stores it in the acquired data storage unit 131. The acquisition unit 111 includes a behavior history data acquisition unit 1111 and a condition acquisition unit 1112.

[0028] The behavior history data acquisition unit 1111 acquires behavior history data for each user from the input device 2 via the input / output interface 12, and stores the acquired behavior history data in the acquired data storage unit 131. The behavior history data acquisition unit 1111 may acquire behavior history data for one user separately, or may acquire behavior histories of multiple users at once in a form that allows them to be distinguished from one another. The behavior history data acquisition unit 1111 may also output a signal indicating that behavior history data has been acquired to the parameter estimation unit 112. The acquired behavior history data will be described later.

[0029] The condition acquisition unit 1112 acquires conditions for each user from the input device 2 via the input / output interface 12, and stores the acquired conditions in the acquired data storage unit 131. The condition acquisition unit 1112 may also acquire conditions for one user separately, or may acquire conditions for multiple users at once in a form that allows them to be distinguished from one another. The condition acquisition unit 1112 may also output a signal indicating that the conditions have been acquired to the optimization unit 113. The acquired conditions will be described later.

[0030] The parameter estimation unit 112 estimates, for each user, parameter values ​​of a mathematical model (behavioral model) that inputs an incentive amount and outputs the degree of achievement of a target behavior, based on the behavior history data stored in the acquired data storage unit 131. Furthermore, the parameter estimation unit 112 stores the estimated parameter values ​​in the parameter storage unit 132. The incentive amount, target behavior, and behavioral model will be described later.

[0031] The optimization unit 113 calculates an optimal incentive policy based on the parameter values ​​estimated by the parameter estimation unit 112 and the conditions stored in the acquired data storage unit 131. The optimization unit 113 calculates this optimal incentive policy for each user. The optimization unit 113 also stores the calculated optimal incentive policy in the optimized incentive policy storage unit 133. Details of the optimal incentive policy will be described later.

[0032] After parameter values ​​for a given user are estimated based on the behavior history data of the given user, the output control unit 114 outputs the optimal incentive policy stored in the optimized incentive policy storage unit 133 to the output device 3 via the input / output interface 12 in response to acquiring conditions from the input device 2. Furthermore, after the optimal incentive policy for a given user is calculated based on parameter values ​​and conditions for the given user, the output control unit 114 may output the optimal incentive policy for the given user stored in the optimized incentive policy storage unit 133 to the output device 3 via the input / output interface 12 in response to an operation by the user of the information processing device 1.

[0033] (operation) FIG. 3 is a flowchart showing an example of the parameter estimation operation of the information processing device 1. The control unit 11 of the information processing device 1 reads out and executes a program stored in the storage unit 13, thereby realizing the operation of this flowchart.

[0034] The operation may be started at any timing, for example, automatically at regular time intervals, or may be started in response to an operation by the owner of the information processing device.

[0035] In step ST101, the behavior history data acquisition unit 1111 acquires behavior history data from the input device 2 via the input / output interface 12. For example, a user may input behavior history data to the input device 2. Alternatively, the behavior history data acquisition unit 1111 may acquire behavior history data stored in an external server or the like via the input / output interface 12. Then, the behavior history data acquisition unit 1111 stores the acquired behavior history data in the acquired data storage unit 131. Furthermore, the behavior history data acquisition unit 1111 may output a signal indicating that the behavior history data has been acquired to the parameter estimation unit 112. Alternatively, the behavior history data acquisition unit 1111 may output the behavior history data to the parameter estimation unit 112.

[0036] Here, the behavior history data includes various information for each user at each observation time. For example, the behavior history data includes a user ID (hereinafter, referred to as u), the total number of users (hereinafter, referred to as U), the length of the period of the target behavior of user u (hereinafter, referred to as T u ), and the sequence of observed values ​​of the target behavior of user u at each observation time (hereafter,

[0037]

number

[0038] ), and the series of incentive amounts presented to user u at each observation time (hereinafter,

[0039]

number

[0040] ), the explanatory variable sequence at each observation time of user u (hereinafter,

[0041]

number

[0042] Here, the observed value of the target behavior {y u t} is a numerical value that evaluates the success or failure of the target behavior, and takes the value 0 (failure) or 1 (success). Furthermore, the explanatory variable {e u t} is information such as the day of the week, weather, etc., which may affect the user's target behavior other than the incentive. u t} may be, for example, money or points, etc. Furthermore, the behavior history data may be, for example, data obtained by acquiring the above information for each user using a behavior observation device or the like including a sensor, etc.

[0043] In step ST102, the parameter estimation unit 112 estimates parameter values. When receiving a signal indicating that behavior history data has been acquired from the behavior history data acquisition unit 1111, the parameter estimation unit 112 acquires the behavior history data stored in the acquired data storage unit 131. Alternatively, when behavior history data is received directly from the behavior history data acquisition unit 1111, the parameter estimation unit 112 may use the received behavior history data. Then, the parameter estimation unit 112 estimates, for each user u, parameter values ​​of a behavior model that takes the amount of incentive included in the behavior history data as input and the degree of achievement of the target behavior as output.

[0044] The behavioral model uses the success stock (hereinafter, x u t The stock of success is the psychological accumulation of past successful experiences, which decays over time and follows the following equation:

[0045]

number

[0046] Here, β u represents the forgetting rate. The forgetting rate is a value that indicates, for example, how much of something that has been memorized can be remembered over time. Formula (1) is a formula that indicates that the success stock at the next observation time is larger if the interval between the current observation time and the next observation time is close, and also takes into account if the target behavior has been achieved (success). The internal variable (hereinafter referred to as m u t If we call the success rate (denoted as "success stock") "motivation," motivation can be expressed as follows, as it is determined by the stock of success, the amount of incentive offered, and explanatory variables:

[0047]

number

[0048] Here, h(a u t |θ u h ) is a function that represents the sensitivity of user u to the amount of incentive, and the parameter value θ u h Also, g(e u t |θ u e ) is a function that represents the influence of user u on explanatory variables, and the parameter value θ u e Furthermore, k(x u t |θ u e ) is a function that represents the influence of user u on the success stock, and the parameter value θ u e And self-efficacy and self-restoration effects are k(x u t |θ u x ) is implemented in the behavioral model via k(x u t |θ u x) is a monotonically increasing function, the higher the frequency of past success, the higher the motivation, and the behavioral model will be a model that reflects self-efficacy. Also, if it is a function that changes from increasing to decreasing at a certain success stock value, the behavioral model will be a model that reflects the self-restoration effect. Alternatively, if it is a function that changes from decreasing to increasing at a certain success stock value, the behavioral model will be a model that reflects the self-restoration effect. The influence of self-efficacy and self-restoration effect, which differ depending on the user, can be determined by the parameter value θ u x It is expressed as:

[0049] Here, based on the motivation, the observed value y of the target behavior for each user at time t is u t is the following binomial distribution P(y u t ) are assumed to be generated probabilistically from

[0050]

number

[0051] where σ(·|θ u σ ) is a non-negative function that satisfies the following conditions, and the parameter value θ u σ It has.

[0052]

number

[0053] The behavioral model defined above is based on the following user-specific parameter values ​​(hereafter referred to as θ u (denoted as

[0054]

number

[0055] The parameter values ​​are estimated by the parameter estimation unit 112 based on the maximum likelihood estimation method shown in the following equation.

[0056]

number

[0057] That is, the parameter estimation unit 112 estimates the parameter value θ of the behavior model for each user based on the behavior history data. u Estimate.

[0058] In step ST103, the parameter estimation unit 112 stores the estimated parameter values ​​in the parameter storage unit 132.

[0059] FIG. 4 is a flowchart showing an example of the operation of the information processing device 1 to calculate the optimal incentive policy. The control unit 11 of the information processing device 1 reads out and executes a program stored in the storage unit 13, thereby realizing the operation of this flowchart.

[0060] The operation may be started at any timing, for example, automatically at regular time intervals, or may be started in response to an operation by the owner of the information processing device.

[0061] In step ST201, the condition acquisition unit 1112 acquires conditions from the input device 2 via the input / output interface 12. For example, the user may input the conditions to the input device 2. Alternatively, the behavior history data acquisition unit 1111 may acquire conditions stored in an external server or the like via the input / output interface 12. Then, the condition acquisition unit 1112 stores the acquired conditions in the acquired data storage unit 131. Furthermore, the condition acquisition unit 1112 may output a signal indicating that the conditions have been acquired to the optimization unit 113. Alternatively, the condition acquisition unit 1112 may output the conditions to the optimization unit 113.

[0062] The condition is the length of the target period (hereafter, Ξ u), the total budget used for incentives in the target period (hereinafter referred to as B), the series of explanatory variables in the target period (hereinafter referred to as

[0063]

number

[0064] The objective function (hereinafter referred to as Z) is used to evaluate the optimality of the incentive policy. Here, the incentive policy that maximizes the expected value of the objective function is defined as the optimal incentive policy. The objective function Z is, for example, the total number of successful target actions during the target period.

[0065]

number

[0066] , the weighted sum of the total number of successes and the total amount of incentives paid

[0067]

number

[0068] etc. Here, c is a weight. It goes without saying that the objective function Z is not limited to the above example.

[0069] In step ST202, the optimization unit 113 acquires the parameter values ​​stored in the parameter storage unit 132. Upon receiving the signal indicating that the conditions have been acquired, the optimization unit 113 acquires the parameter values ​​stored in the parameter storage unit 132. Furthermore, the optimization unit 113 acquires the conditions stored in the acquired data storage unit 131. Furthermore, when the conditions are received directly from the condition acquisition unit 1112, the optimization unit 113 may use the received conditions.

[0070] In step ST203, the optimization unit 113 calculates an optimal incentive policy. The optimization unit 113 calculates an optimal incentive policy for each user u∈{1, 2, . . . , U} based on reinforcement learning theory. Here, the incentive policy is calculated based on the following formula: u t , the remaining available budget of the total budget at time t (hereinafter, b u t ), and the explanatory variable e at time t u t The amount of incentive a presented at time t is u t A function f that outputs u and is expressed by the following formula:

[0071]

number

[0072] Furthermore, the optimal incentive policy is the policy that maximizes the expected value of the objective function Z as described above, and is expressed by the following equation.

[0073]

number

[0074] Here, E[·] represents the expected value. Under the behavioral model described in step ST102 with reference to FIG. 3, the state V u t of

[0075]

number

[0076] Then, state V u t follows the following Markov decision process (hereafter referred to as MDP): where, the state V at time t u thas the success stock, remaining budget, explanatory variables, and observed values ​​of behavior as functions. At time t, the incentive amount a u t The observed value of the target behavior when presented with y u t is generated stochastically according to equation (3), where the incentive amount a u t The possible values ​​of are the remaining budget b u t Given the following: · Observed value of target behavior y u t After creation, the state transition from time t to time (t+1) is executed with probability 1:

[0077]

number

[0078] In MDP, the policy that maximizes the expected value of the objective function Z can be obtained by, for example, solving the Bellman optimal equation. For example, the incentive policy f that satisfies Equation (8) is * can also be obtained by solving the Bellman optimal equation. Here, a method for solving the Bellman optimal equation may be, for example, a Deep Q Network using a neural network. This Deep Q Network using a neural network is described in, for example, the non-patent document "Volodymyr Mnih et al., "Playing Atari with Deep Reinforcement Learning", arXiv, 2013."

[0079] Optimized incentive policy f u* For example, when solving the Bellman optimal equation using a Deep Q Network, the action value function approximated by the neural network is

[0080]

number

[0081] Using

[0082]

number

[0083] The optimization unit 113 stores the calculated optimal incentive policy in the optimized incentive policy storage unit 133. The optimization unit 113 may also output a signal to the output control unit 114 indicating that the optimal incentive policy has been stored in the optimized incentive policy storage unit 133. Alternatively, the optimization unit 113 may directly output the optimal incentive policy to the output control unit 114.

[0084] In step ST204, the output control unit 114 outputs the optimal incentive policy. When receiving a signal from the optimization unit 113 indicating that the optimal incentive policy has been stored in the optimized incentive policy storage unit 133, the output control unit 114 outputs the optimal incentive policy f u* from the optimization incentive policy storage unit 133. Alternatively, the optimization unit 113 acquires the optimal incentive policy f u* If the output control unit 114 directly receives the optimal incentive policy f , the output control unit 114 may use the received optimal incentive policy f . Then, the output control unit 114 transmits the optimal incentive policy f to the output device 3 via the input / output interface 12. u* Here, the optimal incentive policy f is output to the output device 3 as shown in equation (10). u* are the parameter values ​​of the neural network model.

[0085] In this way, by inputting the behavior history data and the conditions into the input device 2, the user can select the optimal incentive policy f u* can be obtained from the output device 3.

[0086] (Action and effect) According to the embodiment, it is possible to identify the most cost-effective incentive measures for achieving target behavior for each individual. Furthermore, by using cost-effective incentive measures, businesses can support the achievement of each user's target behavior at a lower cost. This allows businesses to increase profits or lower service fees.

[0087] [Other embodiments] It should be noted that the present invention is not limited to the above-described embodiment. For example, in the present invention, an example of solving the Bellman optimal equation using a Deep Q Network has been shown, but the present invention is not limited to this. For example, the Bellman optimal equation may be solved by approximation using a multilayer perceptron. In other words, a general method can be applied to solve the Bellman optimal equation.

[0088] The techniques described in the above embodiments can be stored as a program (software means) that can be executed by a computer on a storage medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), and can also be distributed by transmitting the program via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only executable programs but also tables and data structures) that the computer executes. The computer that implements this device loads the program stored on the storage medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by controlling the operation of the software means. The term "storage medium" as used herein is not limited to storage media for distribution, but also includes storage media such as magnetic disks and semiconductor memories installed inside the computer or in devices connected via a network.

[0089] In short, this invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in combination as appropriate as possible, and in such cases, the combined effects can be obtained. Furthermore, the above-described embodiments include inventions at various stages, and various inventions can be extracted by appropriately combining the disclosed multiple constituent elements. [Explanation of symbols]

[0090] 1...Information processing device 11...Control unit 111…Acquisition Department 1111...Behavioral history data acquisition unit 1112...Condition acquisition section 112...Parameter estimation unit 113...Optimization section 114...Output control unit 12...Input / output interface 13...Storage section 131...Acquired data storage unit 132...parameter storage unit 133...Optimization incentive policy memory unit 2...Input device 3...Output device

Claims

1. an acquisition unit that acquires behavior history data for each user and conditions for optimizing incentive measures; a parameter estimation unit that estimates parameter values ​​of a behavioral model for each user, the parameter value having a success stock representing a psychological accumulation of past successful experiences as an internal variable based on the behavioral history data; an optimization unit that calculates an optimal incentive policy for each user based on the estimated parameter values ​​and the conditions; an output unit that outputs the optimal incentive policy; An information processing device comprising:

2. the behavior history data includes a series of incentive amounts at each observation time for each user, The information processing apparatus according to claim 1 , wherein the parameter estimation unit estimates parameter values ​​of a behavioral model for each user, the parameter estimation unit receiving the series of incentive amounts as an input and receiving a degree of achievement of a target behavior for each user as an output.

3. the behavior history data further includes an observation value of a target behavior that evaluates the success or failure of the target behavior at each observation time for each user, and an explanatory variable that is information that influences the target behavior at each observation time for each user, 3. The information processing device according to claim 2, wherein the behavioral model for each user further includes, as the internal variable, a motivation that determines the success or failure of the target behavior, and the motivation is determined by a function that represents the influence on the success stock for each user, a function that represents the sensitivity to the incentive amount for each user, and a function that represents the influence on the explanatory variable for each user.

4. 4. The information processing device according to claim 3, wherein the function representing the influence on the successful stock for each user is either a monotonically increasing function, a function that increases up to a predetermined value and then decreases after the predetermined value, or a function that decreases up to a predetermined value and then increases after the predetermined value.

5. the behavioral model for each user is probabilistically generated from a binomial distribution in which the behavior of each user at each observation time is greater than 0 and less than 1 and is expressed by a non-negative function having the motivation as an internal variable, and the parameter estimation unit estimates parameter values ​​of the behavioral model for each user based on a maximum likelihood estimation method; 5. The information processing device according to claim 3, wherein the conditions include a length of a target period, a total budget to be used for incentives in the target period, a series of the explanatory variables in the target period, and an objective function for evaluating optimality of an incentive policy, wherein the incentive policy is a function that takes as input a time, the successful stock at the time, the remaining budget out of the total budget usable for the incentive policy, and the explanatory variables, and outputs an incentive amount to be presented at the time, and the optimal incentive policy is an incentive policy that maximizes an expected value of the objective function.

6. 6. The information processing device according to claim 5, wherein the state at the time is the success stock, the remaining budget, the explanatory variables, and the observed value of the behavior, the observed value of the target behavior when the incentive amount is presented at the time is generated stochastically according to the binomial distribution, the possible values ​​of the incentive amount are equal to or less than the remaining budget, and the transition from the time to the next time has a probability of 1, in a Markov decision process in which the optimization unit calculates the optimal incentive policy by solving a Bellman optimal equation.

7. An information processing method executed by an information processing device having a processor, The processor acquires behavior history data for each user; obtaining a condition under which the processor optimizes an incentive policy; The processor estimates parameter values ​​of a behavioral model for each user, the behavioral model having an internal variable representing a psychological accumulation of past successful experiences, based on the behavioral history data; calculating an optimal incentive policy for each user based on the estimated parameter values ​​and the conditions; the processor outputting the optimal incentive policy; An information processing method comprising:

8. Obtaining behavioral history data for each user and conditions for optimizing incentive measures; estimating parameter values ​​of a behavioral model for each user, the behavioral model having a success stock representing a psychological accumulation of past successful experiences as an internal variable, based on the behavioral history data; calculating an optimal incentive policy for each user based on the estimated parameter values ​​and the conditions; outputting the optimal incentive policy; An information processing program having instructions to be executed by a processor provided in an information processing device.

Citation Information

Patent Citations

  • Information analysis apparatus, information analysis method, and program

    JP2019046172A

  • Information processing device, information processing method, and program

    JP2021064337A

  • Behavior modification system, program, and behavior modification method

    JP2022013990A

  • Information processing device, information processing method, and information processing program

    JP2022030321A

  • JPP7456485B