Behavior change assistance device, behavior change assistance system, behavior change assistance method, and program
The system addresses privacy and burden issues in behavior change support by using a pre-trained preference estimator and agent learning to selectively present interventions, enhancing efficiency and reducing unnecessary personal data acquisition.
Patent Information
- Application Number
- PCT/JP2024/016898
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-02
- Publication Date
- 2025-11-06
AI Technical Summary
Conventional behavior change support systems acquire a wide range of personal characteristics from users, leading to an undesirable burden and privacy risks.
A behavior change support system that includes a preference estimator pre-trained to estimate intervention preferences as a probability distribution from user answers, an agent learning unit to reinforce learning for suitable behaviors, and an output unit to present tailored interventions, while minimizing unnecessary personal characteristic acquisition.
Reduces user burden and privacy risks by efficiently selecting and presenting interventions based on necessary personal characteristics, using multitask reinforcement learning to formulate efficient question and intervention selection.
Smart Images

Figure JP2024016898_06112025_PF_FP_ABST
Abstract
Description
Behavioral change support device, behavioral change support system, behavioral change support method, and program
[0001] The present invention relates to a behavior change support device, a behavior change support system, a behavior change support method, and a program.
[0002] There are behavior change support systems that support users in changing their behavior. For example, studies are being conducted to use a system to support interviewers in understanding individuals and providing appropriate motivation tailored to each individual during health guidance interviews for preventing lifestyle-related diseases (see, for example, Non-Patent Document 1).
[0003] Furthermore, a technology is known in which, when a subject performs a certain target behavior, information that influences the subject's behavior, such as environmental data and schedules, is classified into motivating factors and inhibiting factors, and a behavior change support message is sent based on the classification results (see, for example, Patent Document 1).
[0004] Tae Sato, Kaori Fujimura, Reiko Ariga, Yasuo Ishigure, Asami Miyajima, Tamae Ogata, Akina Mine, Emiko Kikuchi, Yasuhiro Nishizaki, "Analysis of Motivational Dialogue Processes in Health Guidance Interviews for the Prevention of Lifestyle-Related Diseases," Transactions of the Human Interface Society, 1344-7262, Human Interface Society, 2023-08-25,25,3,189-202.
[0005] Japanese Patent Application Laid-Open No. 2021-86282
[0006] In a behavior change support system, selecting an appropriate intervention message for each user who is the target of the intervention is important for realizing the user's behavior change. However, the personal characteristic information required to select an appropriate intervention message for a user is diverse, ranging from physical information such as medical history and test results to social information such as home environment and workplace, and psychological information such as personality traits.
[0007] In the conventional technology, since all of a wide range of personal characteristics are acquired from the user, there are problems that this is undesirable from the viewpoint of the burden on the user and privacy.
[0008] The embodiments of the present invention have been made in consideration of the above-mentioned problems, and suppress the acquisition of unnecessary personal characteristics in a behavior change support system, thereby reducing the burden on the user and the risk to privacy.
[0009] In order to solve the above problems, a behavior change support device according to an embodiment of the present invention includes a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions using learning data collected in advance; a reception unit that receives inputs of answers to questions for a user and evaluation values for interventions for the user; an agent learning unit that reinforces learning an agent to predict behavior suitable for the user from the answers to questions for the user using the probability distribution estimated by the preference estimator from the answers to questions for the user; and an output unit that presents questions suitable for the user or interventions suitable for the user to the user based on the behavior predicted using the agent from the answers to questions for the user.
[0010] According to an embodiment of the present invention, in a behavior change support system, it is possible to suppress the acquisition of unnecessary personal characteristics, thereby reducing the burden on the user and the risk to privacy.
[0011] It is a diagram showing an example of the configuration of a behavior change support system according to the present embodiment. It is a flowchart showing an example of a pre-learning process according to the present embodiment. It is a flowchart showing an example of a behavior change support process according to the present embodiment. It is a diagram showing an example of the hardware configuration of a computer.
[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.
[0013] The behavior change support system according to this embodiment interactively asks a user, who is a target of support, questions about personal characteristics, accepts input of answers, and selects and presents an intervention appropriate for the user based on the accepted information. Here, personal characteristics include, for example, biological characteristics such as health checkup results, social characteristics such as home environment, schedule, and work environment, psychological characteristics such as personality and time preference, and lifestyle habits. Furthermore, interventions include, for example, motivational messages or video content.
[0014] The question to the user may be, for example, a speech or a text message, but the intervention to the user is not limited to a message and may be any content.
[0015] In this embodiment, we consider a subject who has been determined to need support due to risk items found in the results of a health check, or a subject who has voluntarily agreed to receive motivational intervention to support behavioral change.
[0016] In the health field, the subject's personal characteristics may include items closely related to lifestyle-related diseases, such as the subject's health checkup results (weight, waist circumference, BMI, systolic blood pressure, triglycerides, fasting blood glucose, etc.) and lifestyle habits (exercise, daily activities, diet, sleep, smoking, drinking, etc.). The subject's personal characteristics may also include biological characteristics (family history, etc.), psychological characteristics (personality traits, time preference, etc.), and social characteristics (family environment such as family structure, workplace position, workplace environment such as overtime hours, etc.), which are considered important for understanding patient characteristics in medical care. In the learning field, the subject's personal characteristics may include test scores, school curriculum, etc., as well as whether or not the subject participates in club activities, whether or not they use cram schools, and their dreams for the future.
[0017] The behavior change support system according to this embodiment efficiently selects and asks questions about the information necessary to select an intervention that will motivate each subject based on such a wide range of personal characteristics. Furthermore, the behavior change support system has a first feature of reducing the input burden on the subject and reducing privacy risks caused by obtaining more information than necessary by selecting and presenting an intervention only from the items asked about.
[0018] Furthermore, the behavior change support system according to this embodiment formulates efficient question and intervention selection as a multitask reinforcement learning problem. This allows the selection of questions and interventions to be formulated as a series of decision-making problems, including task estimation based on questions and interventions based on task estimation values. This is a second feature of the system.
[0019] Furthermore, the behavior change support system according to this embodiment introduces a preference estimator as a structure for solving multitask reinforcement learning problems. This preference estimator is pre-trained using pre-training data that compiles evaluation scores for all interventions and answers to questions about some or all personal characteristics for multiple individuals. A third feature of the behavior change support system is that the preference estimator predicts intervention preferences as a probability distribution based on the immediately preceding question and its answer.
[0020] Furthermore, the behavior change support system according to this embodiment uses the output probability distribution of the preference estimator to train the agent in the agent learning unit. During this process, the output probability distribution of the preference estimator is compared before and after the questioning behavior, and the amount of entropy reduction in the probability distribution is fed back to the agent as an auxiliary reward. A fourth feature of the behavior change support system is that this auxiliary reward allows the agent to feed back the quality of its task estimation behavior in multitask reinforcement learning to the learning process.
[0021] This embodiment can be commonly applied to various situations requiring behavioral modification, such as in the fields of health and learning, but here, a behavioral modification support system in the field of health will be described.
[0022] The variables used in this embodiment are as follows: S: state space s∈S: state A: action space A q : Questioning space A i :Intervention action space A=A q ∪A q ・a∈A: Action ・R: Reward function ・r: Reward calculated by the reward function (auxiliary reward r MT (excluding) ・T: transition function ・r MT: Auxiliary reward calculated using a preference estimator Ω: Task distribution in multitask reinforcement learning ω~p(Ω): Task variable in multitask reinforcement learning ω^: Estimated value of task variable calculated using a preference estimator Note that the symbol "ω^" is originally the letter "ω" with a hat symbol (or circumflex) added above it, but in the main text of the specification, it is written as "ω^" because a hat symbol cannot be added to a string of characters.
[0023] In this embodiment, a dialogue with subjects with various characteristics is formulated as a multitask reinforcement learning problem. A dialogue with each subject is treated as a Markov decision process (MDP), and a task variable representing the subject's characteristics is included in the MDP variables. The MDP for each subject is defined in the form of the following equation (1):
[0024] <System Configuration> Fig. 1 is a diagram showing an example of the configuration of a behavior change support system according to this embodiment. In the example of Fig. 1, the behavior change support system 1 includes a behavior change support device 100 and an external device 10 that can communicate with the behavior change support device 100.
[0025] (Behavior Change Support Device) The behavior change support device 100 is, for example, an information processing device having a computer configuration, or a system including multiple computers. The behavior change support device 100 realizes, for example, each functional configuration as shown in Fig. 1 by executing a predetermined program on the computer provided in the behavior change support device 100. In the example of Fig. 1, the behavior change support device 100 includes a reception unit 101, a pre-learning data DB (Database) 102, an estimator pre-learning unit 103, a preference estimator 104, an agent learning unit 105, an agent 106, a prediction unit 107, an output unit 108, a question master 109, and an intervention master 110.
[0026] The reception unit 101 executes a reception process for receiving input of answers to questions posed to the user and evaluation values for the user's intervention. For example, if the agent's most recent action was a question, the reception unit 101 receives the answer to the question from the user as an option via an external device 10 such as a smartphone used by the user. Furthermore, if the agent's most recent action was an intervention, the reception unit 101 receives an evaluation value for the intervention (for example, a subjective evaluation motivation score, an action record such as the number of steps, or an improvement in a test value).
[0027] The pre-learning data DB 102 is a database that stores pre-learning data collected in advance, for example, through a questionnaire survey or the like.
[0028] The estimator pre-learning unit 103 uses the learning data stored in the pre-learning data DB 102 to perform an estimator pre-learning process in which the preference estimator 104 pre-learns a discrete probability distribution of the proportion of interventions that are most highly rated in the learning data for each pair of question and answer.
[0029] The preference estimator 104 is an estimator that has been trained in advance to estimate intervention preferences as a probability distribution from answers to questions posed to the user.
[0030] The agent learning unit 105 executes an agent learning process that uses a probability distribution estimated by the preference estimator 104 from the answers to questions posed to the user to reinforce learning the agent 106 to predict actions suitable for the user from the answers to questions posed to the user. The reinforcement learning algorithm that reinforces learning the agent 106 may be any algorithm that can output actions discretely, but here, DQN (Deep Q-Network), a basic deep reinforcement learning algorithm, is used. Note that components commonly used in DQN, such as a replay memory, are included in the agent learning unit 105.
[0031] If the immediately preceding agent action was a question, the agent learning unit 105 uses the preference estimator E pre-trained by the estimator pre-training unit 103 to calculate the auxiliary reward r for the immediately preceding question using the following equations (2) and (3): t MTCalculate.
[0032] However, ω t ^ is a discrete probability distribution. t The symbol "^" is, as mentioned above, the character "ω" t It represents a symbol with a hat symbol (or circumflex) added above it.
[0033] This supplementary reward r t MT The reward given to the agent is the sum of the negative reward (penalty) r related to the turn consumption for the question and the previous action a t , previous and current state s t-1 , s t , reward r t MT +r is stored in the replay memory. After that, the agent learning unit 105 randomly selects a set (s, s', a, r t MT +r) is extracted in batches of a predefined batch size, and a neural network (agent 106) is updated, which takes s as input and outputs vector Q(s, a). The objective function of this update is expressed by the following equation (4).
[0034] In this way, in this embodiment, in addition to the general reinforcement learning reward r, the auxiliary reward r t MT One of its features is that it contains
[0035] On the other hand, if the immediately preceding agent action was an intervention, the agent learning unit 105 stores the immediately preceding action a t , the previous state s t-1 , and stores the evaluation value r for the intervention. Thereafter, the agent learning unit 105 updates the neural network (agent 106) in the same manner as when the immediately preceding behavior was a question behavior.
[0036] The agent 106 is a neural network that predicts appropriate actions for the user based on the user's answers.
[0037] The prediction unit 107 receives the answer to the previous question from the reception unit 101, and executes a prediction process to output the next topic ID using the agent 106, which is a model learned by the agent learning unit 105. First, the prediction unit 107 calculates the state s t The state s is a vector whose length is equal to the number of predefined questions. The I-th element s i The answer (option number > 1) to the question with IDi in the question master is stored in s. However, if the i-th question has not been asked to the current subject, i = 0 is stored.
[0038] If there is no prior knowledge, the initial state is s 0 is a 0 vector, and each time a question is asked, an answer value is filled into the element. t is input to the agent 106, which is a neural network trained by the agent learning unit 105, to obtain the vector Q(s, a). After that, the prediction unit 107 obtains the element number that maximizes Q(s, a) and outputs it as the optimal action ID.
[0039] The output unit 108 executes an output process to present a question or an intervention suited to the user to the user based on the behavior predicted by the agent 106 from the user's answer. For example, the output unit 108 converts the behavior (behavior ID) output from the prediction unit 107 into a specific utterance or the like, and outputs the utterance to the external device 10. First, the output unit 108 determines whether the input behavior ID is a question behavior or an intervention behavior. Note that each behavior ID is predefined as either a question or an intervention. If the behavior (behavior ID) is a question, the output unit 108 acquires content (e.g., a question sentence) corresponding to the behavior ID from the question master 109 and outputs it to the external device 10, etc. If the behavior is an intervention, the output unit 108 acquires content (e.g., an intervention message, motivational video content, etc.) corresponding to the behavior ID from the intervention master 110 and outputs (presents) it to the external device 10, etc.
[0040] (External Device) The external device 10 is an information processing device such as a smartphone, tablet terminal, or PC (Personal Computer) used by a user who is a target of support. The external device 10 displays question messages, intervention messages, etc. presented by the behavior change support device 100, for example, by executing a predetermined program corresponding to the behavior change support system 1. The external device 10 also transmits the user's answers to the question messages, the user's evaluation values for the intervention messages, etc. to the behavior change support device 100. Note that in this embodiment, the external device 10 may have any configuration.
[0041] The system configuration of the behavior change support system 1 shown in Fig. 1 is an example. For example, the functional components of the behavior change support device 100 in Fig. 1 may be distributed across multiple devices. For example, the reception unit 101, the output unit 108, the question master 109, and / or the intervention master 110 may be provided outside the behavior change support device 100.
[0042] Furthermore, the behavior change support device 100 may have the functions (such as a user interface) of the external device 10. In short, each functional configuration of the behavior change support device 100 in FIG. 1 may be possessed by any one of one or more devices included in the behavior change support system 1.
[0043] <Processing Flow> Next, the processing flow of the behavior change support method according to this embodiment will be described.
[0044] (Pre-learning process) Fig. 2 is a flowchart showing an example of pre-learning process according to this embodiment. This process shows an example of pre-learning process that the behavior change support system 1 executes before executing the behavior change support process described in Fig. 3.
[0045] In step S201, a person in charge of the pre-learning process or the like stores pre-learning data collected in advance by a questionnaire survey or the like in the pre-learning data DB 102. In the questionnaire survey or the like, the same questions as those output by the prediction unit 107 and the output unit 108 or the like are asked, and for each question of each respondent, a set of "question number, answer number to that question, and the respondent's evaluation value for all interventions" is stored as one piece of data in the pre-learning data DB 102. Note that the evaluation value for the interventions includes, for example, subjective motivation score, behavioral performance such as the number of steps, and improvement in test values.
[0046] In step S202, the estimator pre-training unit 103 pre-trains the preference estimator 104. For example, the estimator pre-training unit 103 uses the training data stored in the pre-training data DB 102 to cause the preference estimator 104 to pre-train, for each question-answer pair, a discrete probability distribution of the proportion of interventions that are most highly rated in the training data.
[0047] The preference estimator 104 can be realized in a number of ways. For example, the preference estimator 104 can be trained by a method of obtaining a posterior distribution by Bayesian estimation or the like, or by the following enumeration format: 1) A set of intervention rewards (R I 1 , R I 2 , ...R I N ) is one-hot vectorized (R Io 2) Transition T=(s t-1 , s t , a t-1 ) for each R Io Take the sum of R IT where s is the state, a is the action, and t is the timestamp (number of question turns). 3) R IT is normalized so that the sum of all elements is 1, and is set as p(ω^). This allows it to be treated as a discrete probability distribution. As mentioned above, the symbol "ω^" represents the letter "ω" with a hat symbol (or circumflex) added above it. 4) A dictionary that returns p(ω^) using each transition T as a key is obtained as a preference estimator E.
[0048] The process of FIG. 2 may be executed by a pre-learning device or the like external to the behavior change support device 100 .
[0049] (Behavior Change Support Processing) Figure 3 is a flowchart showing an example of processing of the behavior change support system according to this embodiment. This processing shows an example of behavior change support processing executed, for example, when the behavior change support device 100 described in Figure 1 is conversing with a user who is a target of support via the external device 10. It is assumed that, at the start of the processing in Figure 3, the behavior change support system 1 has already executed the pre-learning processing described in Figure 2.
[0050] In step S301, the prediction unit 107 initializes a state vector. For example, the prediction unit 107 initializes the state s to the initial state s 0 Initialize to.
[0051] In step S302, the prediction unit 107 predicts an optimal behavior ID. For example, the prediction unit 107 inputs the initialized state vector to the agent 106, which is a trained neural network, and acquires the optimal behavior ID predicted by the agent 106.
[0052] In step S303, the output unit 108 converts the behavior ID acquired by the prediction unit 107 into an output. For example, the output unit 108 determines whether the behavior ID acquired by the prediction unit 107 is a behavior ID corresponding to a question or a behavior ID corresponding to an intervention. If the behavior ID is a question, the output unit 108 acquires content corresponding to the behavior ID (e.g., a question sentence) from the question master 109. On the other hand, if the behavior ID is a behavior ID corresponding to an intervention, the output unit 108 acquires content corresponding to the behavior ID (e.g., an intervention message, motivational video content, etc.) from the intervention master 110.
[0053] In step S304, the output unit 108 outputs the acquired question sentence or intervention content to the external device 10.
[0054] In step S305, the behavior change support device 100 branches the process depending on whether the output behavior (question or intervention) is an intervention. If the output behavior is an intervention, the behavior change support device 100 transitions the process to step S310. On the other hand, if the output behavior is not an intervention (if it is a question), the behavior change support device 100 transitions the process to step S306.
[0055] When the process proceeds to step S306, the receiving unit 101 receives an answer to the question. For example, the question output by the output unit 108 is displayed on the display screen of the external device 10. The external device 10 also transmits an answer selected by the user to the question to the behavior change support device 100. The receiving unit 101 receives the answer to the question transmitted by the external device 10.
[0056] In step S307, the prediction unit 107 updates the state vector. For example, the prediction unit 107 receives an answer to the question from the reception unit 101 and calculates the state s t As mentioned above, the state s is a vector whose length is equal to the number of predefined questions. The a-th element s i The answer (option number > 1) to the question with IDi in the question master is stored in the field.
[0057] In step S308, the preference estimator 104 estimates intervention preferences as a probability distribution from the answers to the questions posed to the user.
[0058] In step S309, the agent learning unit 105 learns the agent 106 using the probability distribution estimated by the preference estimator 104. For example, the agent learning unit 105 calculates an auxiliary reward for the question using the probability distribution estimated by the preference estimator 104, includes the calculated auxiliary reward in the reward, and updates the behavioral policy of the agent 106 using a reinforcement learning algorithm.
[0059] Furthermore, the agent learning unit 105 returns the process to step S302 after learning the agent 106. As a result, in step S302, the prediction unit 107 predicts the optimal behavior ID for the question received in step S306 using the agent 106 learned in steps S307 to S309, and executes the processes from step S303 onwards again.
[0060] On the other hand, when the process proceeds from step S305 to S310, the receiving unit 101 receives an evaluation value for the intervention. For example, the intervention content output by the output unit 108 is displayed (or output) by the external device 10. Furthermore, the external device 10 transmits an evaluation value (e.g., a subjective evaluation value, a behavioral performance indicator such as the number of steps) to the behavior change support device 100 based on the user's reaction to the intervention. The receiving unit 101 receives the evaluation value for the intervention transmitted by the external device 10.
[0061] In step S311, the agent learning unit 105 learns the agent 106. For example, the agent learning unit 105 uses the evaluation value for the intervention input from the receiving unit 101 to update the agent's behavior policy by a reinforcement learning algorithm.
[0062] 2 , the behavior change support system 1 can understand the characteristics of the subject in motivating behavior change and formulate the process of intervening in accordance with the characteristics as a multi-task reinforcement learning problem. Furthermore, by calculating the auxiliary reward using the pre-trained preference estimator 104, the behavior change support system 1 can realize the effect of allowing the agent 106 to solve the multi-task reinforcement learning problem, ask efficient questions, and select the optimal intervention. Therefore, according to this embodiment, the behavior change support system 1 can suppress the acquisition of unnecessary personal characteristics, thereby reducing the burden on the user and privacy risks.
[0063] Next, a specific example will be described.
[0064] First Embodiment In a first embodiment, a specific example of processing will be described in which the behavior change support system 1 presents a walking promotion message in a healthcare application.
[0065] In accordance with this embodiment, the elements of the Markov decision process representing the dialogue with each user (subject) are detailed as follows.
[0066] State space S: All combinations of one-dimensional vectors of length 6 that store the answer to each question (2 values: 1: yes, 2: no) or 0: not asked as each element. This state space S is 6 = 729 different states.
[0067] Behavior space A: Assume that behavior IDs are assigned to six types of questions and four types of intervention messages. The six types of questions are, for example, "Do you think you are highly cooperative?", "Do you think you have a high tendency toward neuroticism?", "Did you put off doing your homework as a child?", "Do you prefer to lose weight through exercise rather than diet?", "Do you currently place importance on reaching your ideal weight?", and "Do you often laze around at home?". The four types of intervention messages are, for example, "Walking for about 30 minutes will burn about 100 kcal," "Watching TV casually will burn only about 26 kcal in 30 minutes, but exercising for 30 minutes will burn about 100 kcal," "Exercising large muscles such as the thigh muscles increases energy metabolism," "You look like you're in good shape. If you keep up this good physical condition, you'll have fun. Why don't you try exercising for about 30 minutes?"
[0068] Reward function R ω : A reward of -1 for the questioning behavior and a subjective motivation score from the user. Here, the subjective motivation score is an evaluation value of the user's motivation in response to the intervention message presented to the user, and is expressed on a six-point scale, for example, from 1: not motivated at all to 6: very motivated.
[0069] Transition function T ω ω: is a function that updates the state after receiving answers to questions. Task ω: is the state vector when each user (subject) answers all questions. The initial entropy when no questions are asked is the entropy of a uniform distribution where all interventions have the same probability.
[0070] It is assumed that the preference estimator 104 has been pre-trained using a questionnaire in which all questions included in the behavior space A are presented in random order and evaluations of all intervention messages are requested at the end.
[0071] (Example of specific processing) As an example, suppose we are interacting with a user (subject) whose scores, if all questions are answered, are "Agreeableness: 1, Neuroticism: 2, Puts off homework: 1, Prefers exercise to diet: 2, Places importance on achieving ideal weight: 1, Often lazes around at home: 1."
[0072] In this case, the prediction unit 107 predicts the state vector s 0 is initialized as a 0 vector and input to the agent 106 to obtain the optimal behavior ID. Here, it is assumed that the prediction unit 107 has obtained a behavior ID corresponding to the question, for example, "Did you put off doing your homework when you were a child?"
[0073] As a result, the output unit 108 outputs the question "Did you put off doing your homework when you were a child?" to the external device 10, and the receiving unit 101 receives the answer "2" to the question from the external device 10.
[0074] The agent learning unit 105 calculates the auxiliary reward using the preference estimator 104 trained by the estimator pre-training unit 103. At this time, since information increases due to the question, entropy decreases, and the auxiliary reward becomes a positive value. The agent learning unit 105 uses this auxiliary reward and the reward of -1 for the questioning behavior as rewards to be fed back to the agent 106.
[0075] Assume that the behavior change support system 1 repeatedly asks questions, such as "Do you think you are highly cooperative?", "Did you put off your homework as a child?", and "Do you often slack off at home?". Furthermore, it is assumed that the behavior change support system 1 then asks the question "Do you think you are highly neurotic?" and the preference estimator 104 calculates the supplementary reward. It is also assumed that the entropy did not change before and after the question "Do you think you are highly neurotic?", i.e., the supplementary reward was 0. In this case, the agent only receives a reward of -1 for the questioning behavior, and the question "Do you think you are highly neurotic?" is less likely to be selected for the next user (subject) and thereafter. In this way, for example, the behavior change support device 100 improves the efficiency of questioning.
[0076] The behavioral change support device 100 repeats the decision-making process, and when the agent 106 selects an intervention action (for example, "walking for about 30 minutes will burn about 100 kcal"), it presents the selected message to the external device 10 and ends the dialogue with the user.
[0077] Second Embodiment In a second embodiment, a specific example of processing will be described in which the behavior change support system 1 presents daily learning materials in a video learning system for students taking exams.
[0078] In accordance with this embodiment, the elements of the Markov decision process representing the dialogue with each user (subject) are detailed as follows.
[0079] State space S: All combinations of one-dimensional vectors of length 6 that store the answers to each question (an integer value of 1 or greater, with different values set depending on the question) or 0 (not asked) as each element, and a one-dimensional vector of length 7 that combines the six questions and the prior information on the desired school.
[0080] Behavior space A: Behavior IDs are assigned to six types of question items and four types of video content. Here, the six types of question items are, for example, "Do you think you are highly extroverted?", "Do you think you have a high tendency toward neuroticism?", "Are you busy with club activities?", "Do you find school classes enjoyable?", "Do you think it is important to get into the school of your choice?", "Do you often laze around at home?", etc.
[0081] Reward function R ω : A reward of -1 for the questioning behavior and a subjective motivation score from the user. Here, the subjective motivation score is an evaluation value of the user's motivation for the video content presented to the user, and is expressed on a six-point scale, for example, from 1: not motivated at all to 6: very motivated.
[0082] Transition function T ω ω: A function that updates the state after receiving answers to questions. Task ω: The state vector when each user (subject) answers all questions. The initial entropy, when no questions are asked, is the entropy change from a 0 vector when information about the desired school is learned.
[0083] It is assumed that the preference estimator 104 has been pre-trained using a questionnaire in which all questions included in the behavior space A are presented in random order and evaluations of all intervention messages are requested at the end.
[0084] (Example of specific processing) As an example, let us say we are interacting with a user (subject) whose answer to all questions would be "extroversion: 1 (low), neuroticism: 2 (high), are you busy with club activities: 3 (there are tournaments on weekends too), do you find school classes fun: 1 (yes), do you think it is important to get into the school of your choice: 1 (yes), do you often laze around at home: 1 (yes)."
[0085] In this case, the prediction unit 107 predicts the state vector s 0 The only dimension that contains the information about the desired school is the information about the desired school, and the other elements are initialized to 0. The agent 106 is then given a state vector s 0 In this example, it is assumed that the action ID corresponding to the question "Are you busy with club activities?" is acquired.
[0086] As a result, the output unit 108 outputs the question "Are you busy with club activities?" to the external device 10, and the receiving unit 101 receives the answer to the question "3 (there are tournaments etc. on weekends too)" from the external device 10.
[0087] The agent learning unit 105 calculates the auxiliary reward using the preference estimator 104. At this time, since the question increases information and decreases entropy, the auxiliary reward becomes a positive value. The agent learning unit 105 uses this auxiliary reward and the reward of -1 for the questioning behavior as rewards to be fed back to the agent 106.
[0088] The behavioral change support system 1 repeatedly asks questions such as, "Do you often laze around at home?", "Do you think it is important to get into your desired school?", and "Do you find school classes enjoyable?". The agent 106 then selects an intervention action that will allow the agent to effectively utilize the time they can laze around at home (for example, a video that allows them to grasp the main points and prepare in a short amount of time without any study materials) because they have little time due to club activities, etc., but are highly motivated to get into their desired school. In this case, the output unit 108 presents this video to the user, and the dialogue with the user ends.
[0089] In this way, the behavioral change support system 1 according to this embodiment can be customized to be applied to various fields, such as health, learning, medical care, nursing care, and business.
[0090] <Hardware Configuration> The behavior change support device 100 according to this embodiment has, for example, the hardware configuration of a computer 400 as shown in Fig. 4. Alternatively, the behavior change support device 100 is realized by a plurality of computers 400.
[0091] 4 is a diagram showing an example of the hardware configuration of a computer. In the example of Fig. 4, a computer 400 includes a processor 401, a memory 402, a storage device 403, a communication device 404, an input device 405, an output device 406, a bus B, etc.
[0092] The processor 401 is an arithmetic unit such as a CPU (Central Processing Unit) that executes a predetermined program to realize, for example, each functional configuration of the behavior change support device 100 described in Fig. 1. The memory 402 is a storage medium readable by the computer 400 and includes, for example, a RAM (Random Access Memory) and a ROM (Read Only Memory). The storage device 403 is a computer-readable storage medium and may include, for example, a HDD (Hard Disk Drive), an SSD (Solid State Drive), various optical disks, and a magneto-optical disk.
[0093] The communication device 404 includes one or more pieces of hardware (communication devices) for communicating with other devices via a wireless or wired network. The input device 405 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 406 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 405 and the output device 406 may be integrated into one device (e.g., an input / output device such as a touch panel display).
[0094] The bus B is commonly connected to the above components and transmits, for example, address signals, data signals, and various control signals. The processor 401 is not limited to a CPU, and may be, for example, a DSP (Digital Signal Processor), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array).
[0095] (Supplementary Note) The behavior change support device 100 in this embodiment is not limited to being realized by a dedicated device, but may also be realized by a general-purpose computer. In this case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed. Note that the term "computer system" here includes hardware such as an OS and peripheral devices.
[0096] Furthermore, "computer-readable recording media" includes various storage devices such as portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as hard disks built into computer systems. Furthermore, "computer-readable recording media" may also include devices that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and devices that store programs for a certain period of time, such as volatile memory within computer systems that serve as servers or clients in such cases.
[0097] The above program may be for realizing some of the above functions, or may be capable of realizing the above functions in combination with a program already recorded in a computer system, or may be realized using hardware such as a PLD or FPGA. Furthermore, the computer that executes the above program is not limited to a physical machine, and may be, for example, a virtual machine on a cloud.
[0098] <Effects of the embodiment> According to the present embodiment, the behavior change support system 1 can suppress the acquisition of unnecessary personal characteristics, thereby reducing the burden on the user and the risk to privacy.
[0099] For example, the behavior change support device 100 according to this embodiment formulates questions to grasp personal characteristics and selection of interventions for users (subjects) with diverse personal characteristics as a multitask reinforcement learning problem, and introduces supplemental rewards using a pre-trained preference estimator 104. This enables the behavior change support device 100 to select interventions according to the personal characteristics while reducing the burden on the user.
[0100] Summary of Embodiments This specification discloses at least the following behavior change support device, behavior change support system, behavior change support method, and program: (Item 1) A behavior change support device comprising: a preference estimator that is pre-trained to estimate intervention preferences as a probability distribution from answers to questions using learning data collected in advance; a reception unit that receives input of answers to questions for a user and evaluation values for interventions for the user; an agent learning unit that reinforces learning an agent for predicting behavior suitable for the user from the answers to the questions for the user using the probability distribution estimated by the preference estimator from the answers to the questions for the user; and an output unit that presents questions suitable for the user or interventions suitable for the user to the user based on the behavior predicted using the agent from the answers to the questions for the user. (2) The behavior change support device according to paragraph 1, wherein, when the reception unit receives an answer to a question for the user, the agent learning unit calculates an auxiliary reward for the question using the probability distribution estimated by the preference estimator, and includes the auxiliary reward in the reward to perform reinforcement learning on the agent. (3) The behavior change support device according to paragraph 1 or 2, wherein, when the reception unit receives an evaluation value for an intervention with the user, the agent learning unit performs reinforcement learning on the agent using the evaluation value. (Section 4) A behavior change support system comprising: a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions using learning data collected in advance; a reception unit that receives input of answers to questions for a user and evaluation values for interventions for the user; an agent learning unit that reinforces learning an agent to predict behavior appropriate for the user from the answers to the questions for the user using the probability distribution estimated by the preference estimator from the answers to the questions for the user; and an output unit that presents questions appropriate for the user or interventions appropriate for the user to the user based on the behavior predicted using the agent from the answers to the questions for the user.(5) The behavior change support system according to claim 4, further comprising: a database for storing the training data collected in advance; and a pre-training unit for causing the preference estimator to pre-train, for each pair of question and answer, a discrete probability distribution of the proportion of the intervention that is most highly rated in the training data, using the training data stored in the database. (6) The behavior change support system according to claim 4 or 5, wherein the pre-collected training data includes evaluation values for all interventions and answers to questions regarding some or all personal characteristics, collected from a plurality of users. (Clause 7) A behavior change support method in a behavior change support system having a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions based on learning data collected in advance, wherein a computer performs the following processes: accepting input of answers to questions for a user and evaluation values for interventions for the user, using the probability distribution estimated by the preference estimator from the answers to the questions for the user, to reinforce learning an agent that predicts behavior suitable for the user from the answers to the questions for the user, and presenting to the user questions suitable for the user or interventions suitable for the user based on the behavior predicted using the agent from the answers to the questions for the user. (Clause 8) A program, or a storage medium storing a program, that causes a computer to execute the behavior change support method described in Clause 7.
[0101] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0102] REFERENCE SIGNS LIST 1 Behavior change support system 10 External device 100 Behavior change support device 101 Reception unit 102 Pre-learning data DB 103 Estimator pre-learning unit 104 Preference estimator 105 Agent learning unit 106 Agent 107 Prediction unit 108 Output unit 400 Computer
Claims
1. A behavior change support device comprising: a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions using learning data collected in advance; a reception unit that receives input of answers to questions for a user and evaluation values for interventions for the user; an agent learning unit that reinforces learning an agent to predict behavior appropriate for the user from the answers to questions for the user using the probability distribution estimated by the preference estimator from the answers to questions for the user; and an output unit that presents questions appropriate for the user or interventions appropriate for the user to the user based on the behavior predicted using the agent from the answers to questions for the user.
2. The behavior change support device of claim 1, wherein when the reception unit receives an answer to a question for the user, the agent learning unit calculates an auxiliary reward for the question using the probability distribution estimated by the preference estimator, includes the auxiliary reward in the reward, and reinforces learning the agent.
3. The behavior change support device according to claim 1 or 2, wherein when the reception unit receives an evaluation value for the intervention with the user, the agent learning unit uses the evaluation value to reinforce learning the agent.
4. A behavior change support system comprising: a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions using learning data collected in advance; a reception unit that receives input of answers to questions for a user and evaluation values for interventions for the user; an agent learning unit that reinforces learning an agent to predict behavior appropriate for the user from the answers to questions for the user using the probability distribution estimated by the preference estimator from the answers to questions for the user; and an output unit that presents questions appropriate for the user or interventions appropriate for the user to the user based on the behavior predicted using the agent from the answers to questions for the user.
5. A behavior change support system as described in claim 4, comprising: a database for storing the learning data collected in advance; and a pre-learning unit for causing the preference estimator to pre-learn, for each question and answer pair, a discrete probability distribution of the proportion of the intervention most highly rated in the learning data, using the learning data stored in the database.
6. A behavior change support system as described in claim 4 or 5, wherein the pre-collected learning data includes evaluation values for all interventions and answers to questions regarding some or all personal characteristics collected from multiple users.
7. A behavior change support method in a behavior change support system having a preference estimator that has been pre-trained to estimate intervention preferences as a probability distribution from answers to questions based on learning data collected in advance, wherein a computer performs the following processes: a process of accepting input of answers to questions for a user and evaluation values for interventions for the user; a process of reinforcement learning an agent that predicts behavior appropriate for the user from the answers to the questions for the user using the probability distribution estimated by the preference estimator from the answers to the questions for the user; and a process of presenting questions appropriate for the user or interventions appropriate for the user to the user based on the behavior predicted using the agent from the answers to the questions for the user.
8. A program for causing a computer to execute the behavioral change support method according to claim 7.
Citation Information
Patent Citations
Action support message generator, method, and program
JP2021086282A
Intervention control program, device, system and method for determining intervention policy from estimated intervention result
JP2024048666A
Recommendation device
WO2022163203A1
Cited By
Information processing system, information processing method, and program
JP7851059B1