Message generation device, method, and program
Patent Information
- Application Number
- PCT/JP2023/044272
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-06-19
Smart Images

Figure JP2023044272_19062025_PF_FP_ABST
Abstract
Description
Message generating device, method and program
[0001] FIELD Embodiments of the present invention relate to a message generating device, method, and program.
[0002] Application programs (sometimes referred to as applications or apps) that encourage healthy behaviors such as exercise are becoming increasingly popular. These applications display behavioral change messages to users via pop-up notifications and other means, encouraging them to take better actions. For example, a message might say, "Try to walk 8,000 steps today."
[0003] However, in order for behavior change messages to actually change a user's behavior, there are several issues, such as the following (A1) to (A3). (A1) Content of the message Regarding the content of the message, the following issues (A1-1) can be mentioned. (A1-1) The content needs to be tailored to the user's condition, habits, or surrounding environment. For example, on a very hot day, it is desirable to send a message encouraging indoor muscle training rather than a message encouraging a walk. Also, it is meaningless to send a message encouraging someone who habitually aims to walk 10,000 steps a day to aim to walk 8,000 steps a day.
[0004] (A2) Message text Regarding the message text, the following issues (A2-1) and (A2-2) can be raised.
[0005] (A2-1) The message text needs to take into consideration the characteristics or preferences of the user.
[0006] User characteristics include, for example, present bias or loss aversion. For example, a message to a user with a strong loss aversion tendency is preferably a statement that conveys the loss of not taking action, such as, "If you don't eat vegetables, you're more likely to develop a vitamin deficiency and develop a disease like *****." For example, a message is more effective if it includes lines from a character the user likes, rather than a mechanical message. (A2-2) If the same or similar message is sent multiple times to the same user, the message will lose its effectiveness due to boredom or other reasons.
[0007] (A3) Message Timing The following issues (A3-1) to (A3-4) can be raised regarding message timing. (A3-1) The appropriate message differs depending on the timing. (A3-2) The appropriate timing differs depending on the content of the message. For example, it is meaningless to send a message saying "Taking caffeine before exercise will increase the fat burning effect" to a user after the user has finished exercising.
[0008] (A3-3) If a message is sent to a user at a time when the user cannot check the message, the user cannot check the message immediately. Also, when the user checks the message, the relevant action may not be taken in time. (A3-4) If messages are sent to a user frequently, the user will feel stressed.
[0009] Considering the above-mentioned issues, personalized behavior change messages (sometimes referred to as "personalized behavior change messages") are needed that are tailored to the user's condition, habits, characteristics, preferences, or surrounding environment. However, manually creating messages for each user is not practical. Furthermore, while automatically generating messages, it is necessary to adjust the timing of message presentation to the user. For example, Non-Patent Document 1 discloses a method for optimizing whether to send an intervention message at a specific timing.
[0010] Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical Activity<https: / / dl.acm.org / doi / abs / 10.1145 / <3381007>
[0011] In Non-Patent Document 1, a method is disclosed that encourages more exercise while preventing excessive message sending. However, in this method, the above-mentioned personalization is not applied to the messages.
[0012] However, personalization is implemented at the timing when the message is sent. On the other hand, for the actions recommended to the user, the text of the message is single and no personalization is implemented.
[0013] This invention has been made in view of the above circumstances, and an object thereof is to provide a message generation device, method, and program that can appropriately generate a message that promotes a change in the user's behavior.
[0014] A message generation device according to an aspect of the present invention includes an estimation unit that estimates a feature amount of a message that promotes a change in the user's subsequent behavior based on the user's behavior state from the past to the present, the user's behavior goal from the past to the present, the degree of achievement of the user's behavior goal with respect to the user's behavior from the past to the present, and the characteristics of the message that promoted a change in the user's past behavior, and an output unit that outputs a message that promotes a change in the user's subsequent behavior based on the feature amount estimated by the estimation unit.
[0015] A message generation method according to one aspect of the present invention is a method performed by a message generation device, and includes an estimation unit of the message generation device estimating features of a message that will encourage a change in the user's future behavior based on the user's behavioral state from the past to the present, the user's behavioral goal from the past to the present, the degree to which the user's behavioral goal has been achieved through the user's behavior from the past to the present, and features of messages that encouraged the user to change their past behavior; and an output unit of the message generation device that outputs a message that will encourage a change in the user's future behavior based on the features estimated by the estimation unit.
[0016] According to the present invention, it is possible to appropriately generate a message that encourages a user to change their behavior.
[0017] FIG. 1 is a diagram illustrating an application example of a message generation device according to a first embodiment of the present invention. FIG. 2 is a diagram illustrating an application example of a policy formulation unit of a message generation device according to a first embodiment of the present invention. FIG. 3 is a flowchart illustrating an example of a procedure of processing operations by the message generation device according to the first embodiment of the present invention. FIG. 4 is a flowchart illustrating an example of processing related to learning model parameters of the text conversion unit of the message generation device according to the first embodiment of the present invention. FIG. 5 is a flowchart illustrating an example of processing related to offline reinforcement learning of model parameters of the policy formulation unit of the message generation device according to the first embodiment of the present invention. FIG. 6 is a flowchart illustrating an example of processing related to online reinforcement learning of model parameters of the policy formulation unit of the message generation device according to the first embodiment of the present invention. FIG. 7 is a diagram illustrating an application example of a message generation device according to a second embodiment of the present invention. FIG. 8 is a flowchart illustrating an example of processing related to learning model parameters of the text generation unit of the message generation device according to the second embodiment of the present invention. FIG. 9 is a block diagram illustrating an example of a hardware configuration of a message generation device according to an embodiment of the present invention.
[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. (First Embodiment) First, the first embodiment will be described. In this embodiment, a behavioral goal of a user is set.
[0019] This behavioral goal is set by the user himself / herself or a doctor, etc. Examples of behavioral goals include "exercise to the extent of walking 10,000 steps or more per day" and "limit calorie intake to **."
[0020] In this embodiment, sensing related to a user is performed using a mobile device such as a smartphone, or a non-mobile device such as a sensor device installed inside a car or indoors. For example, the user's state, numerical values related to a behavioral goal, and other behaviors are observed in real time. In this embodiment, appropriate personalized behavior change messages are sent to the user at appropriate times to encourage the user to achieve their behavioral goal.
[0021] In this embodiment, the following values (B1) to (B8) can be used: (B1) t: time step The time step is the timing of an intervention candidate. This timing can be, for example, every hour, every minute, or any predetermined time.
[0022] (B2) G t : Behavioral Goal A behavioral goal is an action that the user is encouraged to take within a predetermined real time from the current time, or until a predetermined time. The above-mentioned real time may be, for example, 30 minutes. For example, a behavioral goal may be the amount of exercise to be performed within 30 minutes from the current time, or the calorie intake of the next meal. Behavioral goals are similar to rewards, which will be described later, but the two differ in that behavioral goals relate to target values, while rewards are measured values used in reinforcement learning.
[0023] In this embodiment, the action goal (G t+i ) can be set outside the system. In this embodiment, the behavioral goal can be changed depending on the degree of achievement of the behavioral goal at each timing.
[0024] (B3) M: Behavioral Change Message The behavioral change message M is a specific message that encourages the user to change their behavior. In addition, this behavioral change message M may be omitted and not sent to the user.
[0025] (B4) S t = {I t , Z t , U t}: user state at time t The meanings of I, Z, and U of this state are as follows (B4-1) to (B4-3). (B4-1) I: Value indicating whether intervention is possible or not State I indicates that intervention is not possible, for example, when the user is driving or sleeping. State I is determined from data obtained from a device such as a smartphone.
[0026] (B4-2) Z: Context State Z is, for example, the state of the user or the surroundings from the past to the present, such as a location, a value of sensor data, weather, or temperature.
[0027] (B4-3) U: User's characteristics, personality, or preferences. State U is, for example, the strength of a current bias or a tendency toward loss aversion. State U may also be something that does not change over time. State U can be obtained, for example, from answers to questions posed to the user.
[0028] (B5) A t = {F t}: Time of intervention F is a feature of the behavior change message sent to the user.
[0029] (B6) R t: Time reward This reward is, for example, a value calculated from the measured value of the behavioral goal during a predetermined real time from the current time or until a predetermined time. Examples of reward calculations are as follows (B6-1) to (B6-4). (B6-1) If the larger the increase value, the better the reward, the actual value is used. (B6-2) If the larger the decrease value, the better the reward, the value multiplied by a minus is used. (B6-3) When reducing the error with the target, the value multiplied by a minus is used. (B6-4) If it is set so that the better the result, the larger the value, an arbitrary value is used.
[0030] A reward for a behavioral goal related to the amount of exercise is, for example, the amount of exercise performed in 30 minutes from the current time. A reward for a behavioral goal related to calorie intake is, for example, a negative value of the calorie intake of the next meal.
[0031] (B7) A target reward, which is the sum of rewards from time t onwards, or an estimated value thereof. The value of (B7) is expressed, for example, by the following (1).
[0032]
[0033] The period for counting rewards may be until the end of data collection or until a predetermined timing.
[0034] (B8) F: Feature of behavior change message Any multiple contents can be used simultaneously as the content of the feature. The content of the feature is, for example, (B8-1) and (B8-2) below. (B8-1) Sentence style The style of the sentence is, for example, the sentence length, the amount of emoticons, or the proportion of function words. (B8-2) Gain frame and loss frame A gain frame is a sentence that focuses on the positive aspects or benefits. A loss frame is a sentence that focuses on the negative aspects or losses. The content of the feature includes a feature corresponding to whether or not there is a message. When there is no message, the content of the feature indicates that no intervention was made to the user.
[0035] 1 is a diagram showing an application example of a message generation device according to a first embodiment of the present invention. As shown in Fig. 1, the message generation device 100 according to the first embodiment of the present invention includes an observation unit 10, a policy formulation unit 20, a template DB (database) 30, a text conversion unit 40, a feature calculation unit 50, and a notification unit 60.
[0036] The policy formulation unit 20 has a policy formulation model. The policy formulation model is any online reinforcement learning model such as Online Decision Transformer. In this embodiment, by providing a learning phase and an operation phase for each user, it is also possible to use any offline reinforcement learning model as the policy formulation model. In this embodiment, the policy formulation model is updated to a shared model for everyone by offline reinforcement learning. Also, in this embodiment, online reinforcement learning is performed while the policy formulation model is in operation, thereby updating the model to a model specialized for each user.
[0037] In this embodiment, the user's state S t By adding information that identifies individuals, such as a user ID, the policy formulation model can be used as a shared model even after online reinforcement learning.
[0038] The policy formulation model takes as input a state, an action, and a reward. A state is a time t, a user's state S t , and behavioral goal G t The behavior includes the features of the message, which are expressed as (2) below.
[0039]
[0040] Rewards are based on the degree of achievement of the behavioral goal. t The policy formulation model includes the feature F t Output.
[0041] The template DB 30 stores example messages calling for actions, i.e., messages M calling for actions corresponding to various behavioral goals G. The feature amounts of each sentence of the messages stored in the template DB 30 may be uniform or may vary.
[0042] The text conversion unit 40 has, as a text conversion model, any text style conversion model implemented by any model such as BART (Bidirectional Auto-Regressive Transformer) or a text generation model such as GPT-3. The text conversion model is implemented by using a feature F of a message. t , and a message template. The message template is expressed as (3) below:
[0043]
[0044] The text conversion model outputs a personalized message, which is expressed as (4) below.
[0045]
[0046] The feature calculation unit 50 has a feature calculation model. The feature calculation model is an arbitrary model that extracts a feature expressed as in the following (5) from the message output by the text conversion unit 40.
[0047]
[0048] Feature computation models depend on the features used and include morphological analyzers such as mecab, or transformer-based classification and regression models.
[0049] 2 is a diagram showing an application example of the policy formulation unit of the message generation device according to the first embodiment of the present invention. Here, an example in which an Online Decision Transformer is applied as a policy formulation model will be described.
[0050] 2, the policy formulation unit 20 includes a feature calculation unit 21, a reply buffer 22, and an update unit 23. The feature calculation unit 21 is realized using, for example, a neural network model. The feature calculation unit 21 includes embedding units, including a behavioral goal embedding unit 21a, a state embedding unit 21b, a feature embedding unit 21c, and a target-reward embedding unit 21d. The feature calculation unit 21 also includes an estimation unit 21e.
[0051] The embedder receives time steps and data and projects them into a latent space. The various embedders can be implemented using different models depending on the type of input. The time steps are used for positional encoding. When the input to the various embedders does not include a sequence, the model of the various embedders can be constructed using a fully connected layer. When the input to the various embedders includes a variable-length sequence, the model of the various embedders can be implemented using a Transformer or a Recurrent Neural Network (RNN), etc.
[0052] Regarding the state embedding unit 21b, since the state S includes a plurality of components, a model provided for each element may be applied, or a model for processing each element collectively may be applied.
[0053] The estimation unit 21e is realized by, for example, the following causal transformer. The causal transformer is disclosed, for example, on the following website: (Website) <https: / / proceedings.mlr.press / v162 / melnychuk22a.html> The estimation unit 21e inputs data expressed as in the following (6) and calculates a feature quantity F of the message. t+1 Estimate.
[0054]
[0055] In (6) above, k is a hyperparameter indicating the number of steps to consider.
[0056] The update unit 23 uses a gradient method to update model parameters applied to the behavioral goal embedding unit 21 a, the state embedding unit 21 b, the feature embedding unit 21 c, the target reward embedding unit 21 d, and the estimation unit 21 e, which are the various units in the feature calculation unit 21. The reply buffer 22 is a database that stores data used for updating by the update unit 23.
[0057] Next, a processing operation when the time step is t will be described. Fig. 3 is a flowchart showing an example of a procedure of a processing operation by the message generating device according to the first embodiment of the present invention. First, the observation unit 10 determines the user's state S t The policy formulation unit 20 acquires the behavioral goal G t , the feature calculated by the feature calculation unit 50, and the state S obtained in S11 t Based on this, the feature value F of the message t is estimated and output to the text conversion unit 40 (S12).
[0058] Next, if the user can be prompted to send a message (Yes in S13), the feature quantity F t If the feature indicating that there is user intervention is included in the template DB 30 (Yes in S14), the text conversion unit 40 selects the behavioral goal G t A message expressed by the following (7) corresponding to the above is extracted (S15).
[0059]
[0060] The text conversion unit 40 inputs this extracted message, generates a message expressed by the following (8), and outputs this to the feature calculation unit 50 and the notification unit 60 (S16).
[0061]
[0062] The notification unit 60 notifies the user of the message output from the text conversion unit 40 (S17).
[0063] Furthermore, the feature calculation unit 50 calculates the feature expressed by the following (9) from the message output from the text conversion unit 40 and outputs it to the policy formulation unit 20.
[0064]
[0065] When the determination in S13 or S14 is "No" and no intervention is made into the feature quantity of the message (S19), or after S18, the observation unit 10 checks the user's behavior and calculates the reward R t is calculated and output to the policy formulation unit 20 (S20).
[0066] Then, the update unit 23 of the policy formulation unit 20 calculates the reward R t Based on the feature quantities from the feature quantity calculation unit 50, the parameters of the models of each unit in the policy formulation unit 20 are updated (S21).
[0067] Next, the processing related to parameter learning of the model of the text conversion unit 40 will be described as follows (C1) to (C3). Figure 4 is a flowchart showing an example of processing related to parameter learning of the model of the text conversion unit of the message generation device according to the first embodiment of the present invention. The following processing is an example, and any text conversion model and method can be used, with the required dataset and algorithm differing depending on the type of method.
[0068] The model of the text conversion unit 40 is trained using a training data set D M , model f t , and the feature calculation model f F The learning dataset is a set of sentences that encourage behavioral change, as shown in (10) below.
[0069]
[0070] Model F t It is desirable that the feature calculation model f is a pre-trained model such as a transformer. Fis an arbitrary model that calculates features from messages. This model may be a machine learning-based model, a rule-based model, or a combination of multiple models. However, the feature calculation model f F When is not differentiable, a differentiable feature estimator created by supervised learning is used.
[0071] (C1) Feature calculation model f F is the training dataset D M Each sentence M i Regarding the feature F i (C2) The text conversion unit 40 calculates the model f as follows (C2-1) and (C2-2): t (C2-1) The text conversion unit 40 converts the sentence M i Sentence M with noise added to ´i The sentences to which noise has been added are, for example, sentences to which a mask process has been applied.
[0072] (C2-2) The text conversion unit 40 converts the model f t to [<F i >, M ´i In (C2-2), the text conversion unit 40 inputs <F i > and M ´i The above is <F i >Feature F i is a vectorized value, and the vectorization method is arbitrary. In (C2-2), the text conversion unit 40 converts the message M i Model f so that t Learn the parameters of
[0073] (C3) To learn the text conversion model, the text conversion unit 40 uses the model f as shown in (C3-1) below. t (C3-1) The text conversion unit 40 learns [<F ´i >, M i ]. F ´i Is F iIn (C3-1), the text conversion unit 40 converts the output sentence into a feature calculation model f F (or the feature estimator shown in (11) below) to obtain the feature shown in (12) below.
[0074]
[0075] In (C3-1), the text converter 40 calculates the model f so as to minimize the error shown in the following (13). t Learn the parameters of
[0076]
[0077] This error may be, for example, a mean square error (MSE) or a cross-entropy error.
[0078] Next, offline reinforcement learning of the model of the policy formulation unit 20 will be described as follows (D1) to (D3). Fig. 5 is a flowchart showing an example of processing related to offline reinforcement learning of the parameters of the model of the policy formulation unit of the message generation device according to the first embodiment of the present invention. The following processing is an example, and any offline reinforcement learning method can be used, with the required dataset and algorithm differing depending on the type of method.
[0079] The offline reinforcement learning of the model parameters of the policy formulation unit 20 uses a learning dataset D T and feature calculation model f F The training data set D T is expressed as follows in (14):
[0080]
[0081] Training dataset D T The following (15) in the above is generated by manual creation and text generated by the text conversion unit 40.
[0082]
[0083] In this embodiment, a plurality of subjects are actually collected and a training dataset D T The data for each of the following is collected. T can be created by the same procedure as the procedure for updating the policy formulation model shown in FIG. t 3 in that the message M is sent randomly instead of determining the message M and preparing it. The above (15) also includes the case where there is no message (no intervention).
[0084] (D1) Feature calculation model f F obtains the feature quantities expressed by the following (16) corresponding to each of the above (15) (S41).
[0085]
[0086] (D2) The policy formulation unit 20 inputs the following (17) into the policy formulation model and obtains an estimated value of the output expressed by the following (18) (S42).
[0087]
[0088] (D3) The policy formulation unit 20 calculates a loss function and updates the parameters of the policy formulation model using stochastic gradient descent and Lagrange multipliers to minimize the loss function (S43). The loss function is the negative log-likelihood expressed by the following (19), with a lower limit on the entropy magnitude of the model output.
[0089]
[0090] Next, online reinforcement learning of the model of the policy formulation unit 20 will be described as follows (E1) to (E8). FIG. 6 is a flowchart showing an example of processing related to online reinforcement learning of the parameters of the model of the policy formulation unit of the message generation device according to the first embodiment of the present invention. The following processing is an example, and any online reinforcement learning method can be used. The required dataset and algorithm differ depending on the type of method.
[0091] For online reinforcement learning of the model parameters of the policy formulation unit 20, an offline data set D T and the feature calculation model f F Offline set D is used. T is expressed as in (14) above. As a preliminary step, the policy formulation unit 20 extracts a plurality of data for learning from the offline data set and stores them in the reply buffer 22.
[0092] (E1) As a process at time t, the policy formulation unit 20 receives the target reward expressed by the following (20) and the state expressed by the following (21) (S51).
[0093]
[0094] The target reward expressed by (22) below is the value expressed by (23) below.
[0095]
[0096] The target reward expressed by the following (24) is a value expressed by the following (25), and is determined inductively.
[0097]
[0098] (E2) The policy formulation unit 20 obtains latent expressions from each embedding unit for the value expressed by the following (26) (S52).
[0099]
[0100] (E3) The policy formulation unit 20 inputs the latent expression obtained in (E2) to the estimation unit 21e and calculates the feature F t is obtained (S53).
[0101] (E4) The policy formulation unit 20 calculates the feature of the message transmitted from the feature calculation unit 50, expressed by the following (27), and the reward R t and updates the contents stored in the reply buffer 22 in accordance with the received result (S54).
[0102]
[0103] If the capacity of the reply buffer 22 is insufficient, the policy formulation unit 20 deletes the oldest data stored in the reply buffer 22 and writes new data to the reply buffer 22 .
[0104] (E5) The policy formulation unit 20 randomly extracts a plurality of data represented by the above (26) from the reply buffer 22 (S55). (E6) The policy formulation unit 20 inputs this extracted data into the policy formulation model and obtains an estimated value of the output represented by the following (28) (S56).
[0105]
[0106] (E7) The policy formulation unit 20 calculates a loss function similar to the loss function calculated in offline reinforcement learning, and updates the parameters of the policy formulation model by stochastic gradient descent to minimize this loss function (S57).
[0107] (E8) If the predetermined number of repetitions for the update has not been reached (No in S58), the policy formulation unit 20 returns to S55. If the predetermined number of repetitions for the update has been reached (Yes in S58), the policy formulation unit 20 ends the process.
[0108] Second Embodiment Next, a second embodiment will be described. Detailed descriptions of parts of the second embodiment that overlap with those of the first embodiment will be omitted. FIG. 7 is a diagram showing an application example of a message generation device according to a second embodiment of the present invention. As shown in FIG. 7, a message generation device 100a according to the second embodiment of the present invention includes an observation unit 10a, a policy formulation unit 20a, a text generation unit 40a, a feature calculation unit 50a, and a notification unit 60a.
[0109] The second embodiment includes a text generation unit 40a instead of the text conversion unit 40 in the first embodiment, and does not include a template DB. The functions of the observation unit 10a, policy formulation unit 20a, feature calculation unit 50a, and notification unit 60a are similar to the functions of the observation unit 10, policy formulation unit 20, feature calculation unit 50, and notification unit 60 in the first embodiment.
[0110] Furthermore, the text conversion unit 40 described in the first embodiment converts the message feature Ft and a message template are input, but the text generating unit 40a does not input the above template, but instead generates the message feature F t In addition, Action Goal G t Enter more.
[0111] Next, the process for learning the parameters of the model of the text generator will be described as follows (F1) to (F5). Fig. 8 is a flowchart showing an example of the process for learning the parameters of the model of the text generator of the message generator according to the second embodiment of the present invention.
[0112] The following process is an example, and any text generation model and method can be used. The required data set and algorithm differ depending on the type of method. Also, a general-purpose large-scale language model may be used as a text generation model without training. For training the model of this text generation unit 40a, model f t , feature calculation model f F and the behavioral goal estimation model f G is used.
[0113] Model F t is a pre-trained general-purpose large-scale language model. F As in the first embodiment, σ is an arbitrary model that calculates features from messages. This model may be a machine learning-based model or a rule-based model, or may be a combination of multiple models.
[0114] Behavioral goal estimation model f G is an arbitrary model that computes action goals from messages. This model can be a machine learning-based model or a rule-based model.
[0115] (F1) The text generator 40a randomly generates a feature amount F and a behavioral goal G (S51). (F2) The text generator 40a receives the feature amount F and the behavioral goal G and generates a plurality of messages (S52).
[0116] The content of the input data for generation in (F2) is, for example, "Please generate a message recommending a 30-minute walk. However, please limit the length of the sentence to 50 characters and use 3 to 5 emoticons." In the above example, the behavioral goal G is "30-minute walk," and the feature F is: "limit the length of the sentence to 50 characters" and "use 3 to 5 emoticons." (F3) The text generation unit 40a calculates the feature F of each message generated in (F2) using the feature calculation model f F The behavioral goal G of each message generated in (F2) is estimated by the behavioral goal estimation model f G The estimated feature amount is expressed by the following (29), and the estimated action goal is expressed by the following (29).
[0117]
[0118] (F4) The text generator 40a calculates the error between the feature value F of each message generated in (F2) and the estimated value in (F3), and calculates the error between the behavioral goal G of each message generated in (F2) and the estimated value in (F3).
[0119] (F5) The text generator 40a learns the parameters of the model of the text generator 40a, for example, according to Instruct GPT, so as to reduce the error calculated in (F4). Instruct GPT is disclosed, for example, on the following website: (Website) <https: / / arxiv.org / abs / 2203.02155> For example, the text generator 40a learns the parameters of the model of the text generator 40a, for example, according to Instruct GPT, so as to reduce the error calculated in (F4). θ is replaced by the value obtained by multiplying the error calculated in (F4) by a negative number.
[0120] In each of the above embodiments, it is possible to generate a behavior change message that is tailored to each individual user and notify the user appropriately. When simply combining existing methods, for example, using ChatGPT as a text conversion model and updating the parameters with Online Decision Transformer, it is not possible to generate the behavior change message described above due to the following issues.
[0121] The first challenge concerns training data. Personalizing a large-scale language model requires a large amount of training data, i.e., pairs of text messages and rewards indicating whether the user's behavior actually changed after sending the messages, which is difficult to prepare. For example, if you intervene with a user once a day, it takes 100 days to collect 100 training data sets. In addition, there is a limit to the number of times you can increase the number of interventions per day.
[0122] The second challenge is related to the message feature space. Messages have various characteristics, making it difficult to identify which aspects of behavior change messages are suitable for individuals and which are not. Therefore, a large amount of training data is required to identify how to improve these characteristics.
[0123] On the other hand, in this embodiment, the above-mentioned behavior change message can be generated by separating the policy formulation model and the model that takes personalization into account as a feature from the text conversion model or text generation model.
[0124] Furthermore, in this embodiment, the policy formulation model is based on online reinforcement learning, and therefore it is possible to notify messages that also correspond to changes in the user's preferences or habits.
[0125] 9 is a block diagram showing an example of the hardware configuration of a message generation device 100 according to an embodiment of the present invention. In the example shown in FIG. 9, the message generation device 100 according to the above embodiment is configured, for example, by a server computer or a personal computer, and has a hardware processor 111A such as a CPU. A program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to this hardware processor 111A via a bus 115. The same applies to the message generation device 100a.
[0126] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information to and from a communication network NW. As the wireless interface, for example, an interface that adopts a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.
[0127] An input device 300 and an output device 400 attached to the message generation device 100 and used by a user or the like are connected to the input / output interface 113. The input / output interface 113 receives operation data input by a user or the like via the input device 300, such as a keyboard, a touch panel, a touchpad, or a mouse, and outputs output data to an output device 400, which may include a display device using a liquid crystal or an organic electroluminescence (EL) display, or an audio output device, for display. The input device 300 and the output device 400 may be devices built into the message generation device 100, or may be input devices and output devices of other information terminals that can communicate with the message generation device 100 via a network NW.
[0128] The program memory 111B is a non-transitory tangible storage medium that is a combination of a non-volatile memory that can be written to and read from at any time, such as a hard disk drive (HDD) or a solid state drive (SSD), and a non-volatile memory such as a read only memory (ROM), and stores programs necessary to execute various control processes, etc., according to one embodiment.
[0129] The data memory 112 is a tangible storage medium that is, for example, a combination of the above-mentioned nonvolatile memory and a volatile memory such as RAM (Random Access Memory), and is used to store various data acquired and created during various processing steps.
[0130] The message generation device 100 according to one embodiment of the present invention may be configured as a data processing system or information processing device having a software-based processing function unit. The storage system used as a work memory or the like by the message generation device 100 may be configured using the data memory 112 shown in FIG. 9. However, these configured storage areas are not essential components within the message generation device 100, and may be areas provided in a storage system such as an external storage medium such as a USB (Universal Serial Bus) memory, or a database server located in the cloud.
[0131] The processing function unit can be realized by having the hardware processor 111A read and execute a program stored in the program memory 111B, but the processing function unit may also be realized in various other forms, including an integrated circuit such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).
[0132] The methods described in each embodiment may be stored as a program (software means) that can be executed by a computer on a recording medium such as a magnetic disk (e.g., a floppy disk, a hard disk, etc.), an optical disk (e.g., a CD-ROM, a DVD, an MO, etc.), or a semiconductor memory (e.g., a ROM, a RAM, a flash memory, etc.), or may be transmitted and distributed via a communication medium. The program stored on the medium also includes a configuration program that configures the software means (including not only executable programs but also tables and data structures) that the computer executes. The computer that realizes this device reads the program stored on the recording medium and, in some cases, configures the software means using the configuration program, and executes the above-described processing by having the operation controlled by this software means. The term "recording medium" as used herein is not limited to a storage medium for distribution, but also includes a storage medium such as a magnetic disk or semiconductor memory installed inside the computer or in a device connected via a network.
[0133] The present invention is not limited to the above-described embodiments, and various modifications can be made in the implementation stage without departing from the spirit of the invention. Furthermore, the embodiments may be implemented in appropriate combinations, in which case the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combining selected elements from the disclosed elements. For example, if the problem can be solved and the desired effect can be obtained even if some elements are deleted from all elements shown in the embodiments, the configuration from which these elements are deleted can be extracted as an invention.
[0134] DESCRIPTION OF SYMBOLS 100, 100a... Message generation device 10, 10a... Observation unit 20, 20a... Policy formulation unit 21... Feature calculation unit 21a... Action goal embedding unit 21b... State embedding unit 21c... Feature embedding unit 21d... Target reward embedding unit 21e... Estimation unit 22... Reply buffer 23... Update unit 30... Template DB 40... Text conversion unit 40a... Text generation unit 50, 50a... Feature calculation unit 60, 60a... Notification unit
Claims
1. An estimation unit that estimates a feature amount of a message that promotes a change in the user's subsequent behavior based on the user's past to current behavior state, the user's past to current behavior goal, the degree of achievement of the user's behavior goal by the user's past to current behavior, and the characteristics of a message that promoted a change in the user's past behavior; and an output unit that outputs a message that promotes a change in the user's subsequent behavior based on the feature amount estimated by the estimation unit. A message generation device comprising:
2. The output unit converts, based on the feature amount estimated by the estimation unit, a template of a message that promotes a change in the user's behavior according to the user's behavior goal into a message that promotes a change in the user's subsequent behavior and outputs the message. The message generation device according to claim 1.
3. The output unit generates and outputs a message that promotes a change in the user's subsequent behavior based on the feature amount estimated by the estimation unit and the user's behavior goal. The message generation device according to claim 1.
4. A feature calculation unit that calculates the features of the message output by the output unit; a notification unit that notifies the user of the message output by the output unit; a second output unit that outputs information indicating the degree of achievement of the user's behavior goal by the user's behavior in response to the notification by the notification unit; and an update unit that updates the parameters of the model used for the estimation by the estimation unit based on the information output by the second output unit and the features calculated by the feature calculation unit. The message generation device according to claim 1, further comprising:
5. The features calculated by the feature calculation unit include at least one of a message indicating a profit or loss incurred by the user according to the user's behavior along the message that promotes a change in the behavior, and the number of characters of a predetermined type included in the message. The message generation device according to claim 4.
6. A method performed by a message generation device, comprising: estimating, by an estimation unit of the message generation device, a feature amount of a message that prompts a change in the user's future behavior based on the state of the user's actions from the past to the present, the goals of the user's actions from the past to the present, the degree of achievement of the user's action goals by the user's actions from the past to the present, and the characteristics of a message that prompted a change in the user's past actions; and outputting, by an output unit of the message generation device, a message that prompts a change in the user's future behavior based on the feature amount estimated by the estimation unit. A message generation method comprising an output unit.
7. A message generation processing program that causes a processor to function as each unit of the message generation device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Device, method and program for supporting lifestyle habit improvement
JP2011128851A
Behavior prediction
JP2017188089A
Information presentation device, learning device, information presentation method, learning method, information presentation program, and learning program
JP7380691B2