Information processing device, information processing method, program

The described system addresses the issue of inappropriate retraining in machine learning models by incorporating feedback-based correction and retraining mechanisms, resulting in improved user suggestions.

JP2026081505APending Publication Date: 2026-05-19NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2024-11-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing machine learning models lack a clear method for modifying labels during retraining, leading to inappropriate suggestions to users, as seen in Patent Documents 1 and 2, where incorrect labels can lead to ineffective retraining.

Method used

An information processing device and method that includes an estimation unit to obtain estimated values from a machine learning model, an acquisition unit to gather feedback information, a correction unit to modify these estimates based on feedback, and a learning unit to retrain the model using corrected information.

Benefits of technology

Enables the generation of a machine learning model capable of making appropriate suggestions to users by refining estimates based on user responses, ensuring accurate retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026081505000001_ABST
    Figure 2026081505000001_ABST
Patent Text Reader

Abstract

The machine learning model is unable to provide appropriate suggestions to the user. [Solution] The information processing device of the present disclosure includes: an estimation unit that acquires estimation information including estimated values ​​for each user action output from a machine learning model by inputting user state information to the machine learning model; an acquisition unit that acquires feedback information representing the user's response to recommendation information that recommends user actions based on the estimation information; a modification unit that generates modified estimation information by modifying the estimated values ​​for each action in the estimation information based on the feedback information; and a learning unit that retrains the machine learning model based on the state information and the modified estimation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Patent Document 1 discloses an information processing apparatus that outputs a recommended image to a subject using a model learned using teacher data related to communication. In this apparatus, the situation of the subject obtained after generating output information is acquired as feedback information, and the learning model is relearned by correcting the label. Also, Patent Document 2 discloses a system that learns a model based on customer attributes and answers to a plurality of questions and proposes beer beverages. This system relearns the learning model using feedback from customers regarding the proposed beverages.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

[0005] However, while Patent Document 1 modifies the labels used for retraining based on feedback information, it does not show a specific method for modifying the labels, making it unclear whether proper retraining is performed. Furthermore, in Patent Document 2, if the proposed beer yields a negative result, other beers are used as the correct answer for training, but these modified other beers are not necessarily appropriate, again making it unclear whether proper retraining is performed. As a result, the retraining methods disclosed in Patent Documents 1 and 2 have the problem of not being able to generate a machine learning model that can make appropriate suggestions to the user.

[0006] Therefore, one of the purposes of this disclosure is to solve the aforementioned problem of being unable to generate machine learning models that can make appropriate suggestions to users. [Means for solving the problem]

[0007] An information processing device, which is one form of this disclosure, An estimation unit that obtains estimation information, including estimated values ​​for each user action output from a machine learning model, by inputting user state information into the machine learning model, An acquisition unit that acquires feedback information representing the user's response to recommendation information that recommends the user's actions based on the estimated information, A correction unit generates corrected estimation information by correcting the estimated values ​​for each action in the estimation information based on the feedback information, A learning unit that retrains the machine learning model based on the state information and the corrected estimation information, Equipped with, This is the structure it takes. Furthermore, the information processing method, which is one form of this disclosure, Information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. This is the structure it takes. Furthermore, one form of this disclosure is a program, In an information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. To execute the process This is the structure it takes. [Effects of the Invention]

[0008] This disclosure, when configured as described above, enables the generation of a machine learning model that can make appropriate suggestions to users. [Brief explanation of the drawing]

[0009] [Figure 1] This block diagram shows an example of the configuration of the estimation system related to this disclosure. [Figure 2] This figure shows an example of how the estimation system described in this disclosure processes data. [Figure 3] This figure shows an example of how the estimation system described in this disclosure processes data. [Figure 4] It is a flowchart showing an example of the processing operation of the estimation system according to the present disclosure. [Figure 5] It is a diagram showing an example of information used in the processing of the estimation system according to the present disclosure. [Figure 6] It is a diagram showing an example of information used in the processing of the estimation system according to the present disclosure. [Figure 7] It is a diagram showing an example of information used in the processing of the estimation system according to the present disclosure. [Figure 8] It is a flowchart showing an example of the processing operation of the estimation system according to the present disclosure. [Figure 9] It is a diagram showing an example of information used in the processing of the estimation system according to the present disclosure. [Figure 10] It is a diagram showing an example of information used in the processing of the estimation system according to the present disclosure. [Figure 11] It is a flowchart showing an example of the processing operation of the estimation system according to the present disclosure. [Figure 12] It is a diagram showing an example of the state of processing of the estimation system according to the present disclosure. [Figure 13] It is a block diagram showing an example of the hardware configuration of the information processing apparatus according to the present disclosure. [Figure 14] It is a block diagram showing an example of the configuration of the information processing apparatus according to the present disclosure.

Mode for Carrying Out the Invention

[0010] <First Embodiment> The first embodiment of the present disclosure will be described with reference to the drawings. Note that the drawings may be relevant to any of the embodiments.

[0011] [Configuration] The estimation system 10 of this disclosure generates an estimation model using machine learning with training data, outputs an estimated behavior value by inputting input data to the generated estimation model, and recommends a behavior based on the estimated behavior value. At this time, it can make appropriate recommendations by obtaining feedback information from the subject regarding the recommended behavior and retraining the model with the corrected estimated behavior value as the target value according to the feedback information.

[0012] For example, estimation system 10 is used to recommend products on an e-commerce (EC) site. The model generated by estimation system 10 takes data generated from customer browsing history of product pages and purchase history as input data, and outputs behavioral estimates representing the likelihood of the customer purchasing each product next as output data, and recommends products to the customer based on these behavioral estimates. At this time, estimation system 10 can acquire feedback information such as whether the customer's reaction to the recommended product is positive or negative, and retrain the model to reflect this feedback. For example, whether a customer has a positive or negative reaction to a recommended product can be determined by whether or not they have viewed the page related to the recommended product or by purchasing the product.

[0013] However, the estimation system 10 in this disclosure is not limited to being used to recommend products as described above. The estimation system 10 may be used to recommend any action by a person, not limited to the purchase of products. In other words, the estimation model generated and retrained by the estimation system 10 in this disclosure may be an estimation model that recommends any action by a person.

[0014] The estimation system 10 consists of one or more information processing devices equipped with an arithmetic unit and a memory device. As shown in Figure 1, the estimation system 10 includes an acquisition unit 11, a storage unit 12, a data generation unit 13, a first generation unit 14, an estimation unit 15, an output unit 16, and a regeneration unit 17. The functions of the acquisition unit 11, data generation unit 13, first generation unit 14, estimation unit 15, output unit 16, and regeneration unit 17 can be realized by the arithmetic unit executing a program for realizing each function stored in the memory device. The storage unit 12 is composed of a memory device.

[0015] With the above configuration, the estimation system 10 performs machine learning on an estimation model that estimates the next action to be taken based on the history of the subject's actions and the state at the time those actions were taken, as described below, and uses the estimation model to recommend actions according to the subject's state. At this time, the estimation system 10 learns the estimation model from the history of the subject's actions and the state at the time those actions were taken, using various past cases that have been labeled as success or failure cases.

[0016] In the following, "target person (user)" refers to a person who receives action recommendations from the estimation system 10, and may also refer to the state and action history information used by the estimation model for learning. Furthermore, "action" includes not only active actions taken by the target person themselves, but also passive actions influenced by external factors. Additionally, "state" includes not only the state of the target person themselves, but also the state of the environment in which the target person acts.

[0017] The acquisition unit 11 acquires behavioral data of the subject that has not yet been used to train the estimation model and supplies it to the storage unit 12. Behavioral data includes, for example, the browsing history and purchase history of product introduction pages by the customer, as shown in Figure 5, and will be explained in detail in the operation description later. The acquisition unit 11 also acquires state data at the time the subject performed the action and supplies it to the storage unit 12. State data includes, for example, the cumulative number of views and cumulative number of purchases for each product by the customer, as shown in Figure 7, and will be explained in detail in the operation description later. Alternatively, the acquisition unit 11 may not acquire state data, and instead generate state data using the data generation unit 13 as described later.

[0018] The acquisition unit 11 acquires positive / negative information labels indicating whether a subject's actions in any given state contributed to a success or failure, using positive examples for success cases and negative examples for failure cases, and supplies them to the storage unit 12. The positive / negative information labels are data that uses flags to indicate good customers as success cases and poor customers as failure cases, as shown in Figure 6, for example. Details will be explained later in the operation description. Note that the positive / negative information labels may be determined and assigned to each subject, rather than being determined and assigned to each subject's actions. Also, the acquisition unit 11 may not acquire the positive / negative information labels from an external source, but rather generate them using the data generation unit 13, which will be described later.

[0019] The memory unit 12 stores the subject's behavioral data. The memory unit 12 also stores data extracted from the behavioral data and data generated based on the behavioral data. Furthermore, the memory unit 12 stores the subject's state data at the time the behavior was performed. The memory unit 12 also stores data extracted from the state data and data generated based on the state data. Finally, the memory unit 12 stores positive and negative information labels for the behavior in relation to the state.

[0020] The data generation unit 13 extracts further data from the subject's behavioral data. The data generation unit 13 also generates state data from the subject's behavioral data. Furthermore, the data generation unit 13 extracts further data from the subject's state data. The data generation unit 13 generates positive / negative information labels for actions corresponding to states from the subject's behavioral data and state data. Specific examples of the behavioral data, state data, positive / negative information labels, etc., processed by the data generation unit 13 will be explained later in the operation description.

[0021] The data generation unit 13 determines whether the subject contributed to a success or failure based on the history of behavioral data and state data, and generates positive or negative information labels. Whether a case is a success or a failure can be determined, for example, based on criteria defined using a predetermined KPI (Key Performance Indicator). Note that whether an action in relation to a state is a success or a failure may not be determined immediately after the action is performed, but rather from a long-term perspective based on the results of performing the action multiple times afterward.

[0022] The first generation unit 14 generates an estimation model (machine learning model) that estimates the next action to be performed, based on the subject's actions, the history of their state at the time the actions were performed, and positive / negative information labels, using machine learning. The first generation unit 14 supplies the generated estimation model to the estimation unit 15, which stores the estimation model. When the subject's state data is input to the estimation model generated by the first generation unit 14, it outputs an action estimate, which is an estimated value regarding the next action to be performed. The action estimate is data with a numerical value for each candidate action, as shown in Figure 10, for example, and a higher numerical value indicates a higher expectation that the action will be performed next.

[0023] The first generation unit 14 may generate an estimation model using, for example, machine learning techniques such as the ACIL (Adversarial Cooperative Imitation Learning) method and the SAiL (Skill Acquisition Learning) method shown in Non-Patent Document 1.

[0024] The first generation unit 14 can generate an estimation model using the SAiL method. Figure 2 schematically shows the generation operation of the estimation model in the SAiL method. In the SAiL method, the behavior mimetic units (behavior mimetic units A, B, etc.) estimate the next action that the subject will take. For example, when past case A is input, the behavior mimetic unit outputs estimated behavior case B. The arrow shown below past case A on the left indicates the input of past case A to the first generation unit 14. The arrow below estimated behavior case B indicates the output of estimated behavior case B from the first generation unit 14. The circles in past case A and estimated behavior case B indicate the subject's state. The triangles in past case A and estimated behavior case B indicate the subject's action. The arrow between past case A indicates that they are the same case.

[0025] When estimating the next action, the first generation unit 14 compares the input with the estimation result in the action strategy selector and selects the optimal action mimic based on the estimation accuracy. The arrow between past example A and estimated action example B in Figure 2 shows the comparison between past example A, which is the input, and estimated action example B, which is the estimation result. The arrows to the left of the comparison arrow, pointing to the action strategy selector and action mimic, indicate that the comparison result is input to the action strategy selector and action mimic. Based on the comparison result between the input and the estimation result, the first generation unit 14 generates an estimation model by simultaneously training the action strategy selector and the action mimic.

[0026] Figure 3 schematically illustrates the operation of optimizing the behavior mimic. The first generation unit 14 can generate a behavior mimic that avoids the actions of failures that it does not want to mimic and accurately mimics the actions of successes by using the ACIL method. In the success case classifier, which is part of the action policy selector, ACIL compares cases containing the actions estimated by the behavior mimic with the actions of past successes. Also, in the failure case classifier, which is part of the action policy selector, it compares cases containing the actions estimated by the behavior mimic with the actions of past failures. The success case classifier performs the operation of distinguishing (or classifying) past successes from cases containing the actions estimated by the behavior mimic. Therefore, the behavior mimic and the success case classifier learn in a competitive manner, with the behavior mimic trying to approximate past successes and the success case classifier trying to distinguish them. Adversarial behavior is a process in which a behavior mimic attempts to estimate behaviors that differ little from successful examples, while a success classifier attempts to further identify those small differences. The goal is to learn to minimize the difference between the input data (successful examples) and the examples containing the estimated behavior.

[0027] On the other hand, the failure case classifier distinguishes (or classifies) past success cases from failure cases. Therefore, the behavior imitator and the failure case classifier learn in cooperation, with the behavior imitator trying to move away from past failure cases and the failure case classifier trying to distinguish them. Cooperation means that the behavior imitator tries to estimate actions that have a small difference from success cases and a large difference from failure cases, while the failure case classifier distinguishes between success cases and failure cases, so that the learning process progresses so that the difference between the input data of failure cases and cases containing the estimated actions becomes large. In this way, by performing machine learning using both adversarial and cooperative approaches, it becomes possible to obtain an estimation model that can make highly accurate estimations without making fatal mistakes.

[0028] The estimation unit 15 uses the estimation model generated by the first generation unit 14 to estimate the most likely next action the subject will take, based on the subject's state data (state information). At this time, the estimation unit 15 stores the action estimate (estimated information) output from the estimation model, which is input to the subject's state data, in the storage unit 12. The action estimate is an estimated value that represents the level of expectation for each action to be performed, as shown in Figure 12, for example.

[0029] The output unit 16 outputs recommendation information that recommends an action to the subject based on the action estimate value, which is the estimation result of the estimation unit 15. The recommendation information is, for example, information corresponding to the single action with the highest numerical value among the action estimate values. At this time, the output unit 16 may also output as recommendation information an action to be taken by the subject in order to encourage the subject to take the action estimated by the estimation model in the estimation unit 15. The storage unit 12 then stores the recommendation information output by the output unit 16.

[0030] Next, as mentioned above, we will explain the functions of each component after estimating the target person's behavior and outputting recommendation information to the target person.

[0031] The acquisition unit 11 acquires feedback information, which is the response of multiple subjects to the recommendation results output using the estimation model, and supplies it to the storage unit 12. The feedback information may include, for example, whether or not there was a response to the recommended behavior and an evaluation of the recommendation. For example, if a subject gives a positive response to the recommended behavior, the feedback information is represented by a positive label, and if a subject gives a negative response to the recommended behavior, it is represented by a negative label.

[0032] The regeneration unit 17 (correction unit, learning unit) regenerates the estimation model by relearning using machine learning, using the action estimates output from the estimation model, the feedback information from multiple subjects regarding the outputted recommendation information, the state data corresponding to the subjects' actions, and the estimation model stored in the estimation unit 15, all of which are stored in the memory unit 12. At this time, the regeneration unit 17 takes the state data as input data and relearns the estimation model so that the action estimates become the output data, but it relearns the action estimates based on the feedback information.

[0033] When the regeneration unit 17 regenerates the estimation model by retraining using machine learning, it may generate a new estimation model by fine-tuning the already generated estimation model. In other words, it may regenerate the estimation model by performing retraining while fixing some of the parameters of the original estimation model.

[0034] For example, if the subject responds positively to the recommended action (a specific action), the regeneration unit 17 modifies the behavior estimates for the recommended action by adding some or all of the behavior estimates for other actions to the behavior estimates for the recommended action, and then retrains the estimation model using the modified values ​​(modified estimation information) as the target values. Also, as an example, if the subject responds negatively to the recommended action, the regeneration unit 17 modifies the behavior estimates for other actions by distributing some or all of the behavior estimates for the recommended action to the behavior estimates for other actions, and then retrains the estimation model using the modified values ​​(modified estimation information) as the target values.

[0035] The regeneration unit 17 may regenerate the estimation model by relearning using machine learning methods such as the ACIL method and the SAiL method, using the state data used as input when the recommendation for which feedback was received, and the target value obtained by correcting the action estimate. When using the ACIL method and the SAiL method, the state data corresponding to the target value obtained by correcting the action estimate can be considered as a state-action sequence that represents a successful case, and an estimation model that estimates actions similar to the successful case can be regenerated. The regeneration unit 17 supplies the regenerated estimation model to the estimation unit 15, and when the estimation unit 15 receives the estimation model, it newly stores it internally.

[0036] Subsequently, the estimation unit 15 uses the estimation model regenerated by the regeneration unit 17 to estimate the most likely next action the subject will take.

[0037] [Operation] Next, we will describe an example of how the estimation system 10 operates. Here, as an example, we will describe the operation of recommending products on an e-commerce site to a target person.

[0038] First, we will explain the generation of estimation models using machine learning. Figure 4 is a diagram showing the flow of the estimation model generation process.

[0039] The acquisition unit 11 acquires behavioral data and positive / negative information labels for multiple subjects (step S1). Once the acquisition unit 11 has acquired behavioral data for multiple subjects, it supplies it to the storage unit 12. As an example, as shown in Figure 5, it acquires behavioral data that includes the browsing history and purchase history of product introduction pages by the target customer. The acquisition unit 11 also acquires positive / negative information labels assigned to each subject, indicating whether they contributed to a success or failure case, and supplies them to the storage unit 12. As an example, as shown in Figure 6, a flag indicates a good customer (1) as a success case and a bad customer (-1) as a failure case. Whether a customer is a good customer or a bad customer can be determined based on, for example, the total amount of past purchases. The storage unit 12 stores the supplied behavioral history data and positive / negative information labels.

[0040] The data generation unit 13 reads the behavioral data stored in the storage unit 12 and extracts the behaviors to be estimated from the training data. The data generation unit 13 also generates state data that includes the state of the subject at the time the behavior to be estimated was performed (step S2). As an example, as shown in Figure 7, it generates state data that includes the cumulative number of views and cumulative number of purchases for each product by the customer. The data generation unit 13 stores the generated state data in the storage unit 12.

[0041] The first generation unit 14 reads the state data, behavior data, and positive / negative information labels stored in the memory unit 12 and generates an estimation model that performs estimations about the next action to be performed using machine learning with the ACIL method and the SAiL method (step S3). As an example, using state data consisting of cumulative view counts and cumulative purchase counts shown in Figure 7 as input data, and using good customers as success cases and non-good customers as failure cases, machine learning is performed using the SAiL method and the ACIL method to generate an estimation model that predicts purchasing behavior that will transform customers into good customers.

[0042] Once an estimation model is generated, the first generation unit 14 supplies the generated estimation model to the estimation unit 15 (step S4). Upon receiving the supplied estimation model, the estimation unit 15 stores the estimation model internally.

[0043] Next, we will explain the process of performing action recommendations using the estimation model. Figure 8 is a diagram showing the flow of the process of performing action recommendations using the estimation model.

[0044] The acquisition unit 11 acquires behavioral data about the estimated subject (step S11). Once the behavioral history data is acquired, the acquisition unit 11 supplies the acquired behavioral data to the storage unit 12, and the storage unit 12 stores the supplied behavioral data.

[0045] The data generation unit 13 reads the behavior data stored in the storage unit 12 and generates state data at the time of behavior estimation, as shown in Figure 9 (step S12). The data generation unit 13 stores the generated state data in the storage unit 12.

[0046] The estimation unit 15 reads the state data stored in the memory unit 12 at the time of performing the behavior estimation, and inputs it into the estimation model generated by the first generation unit 14, thereby outputting behavior estimates regarding the next action the subject will perform, as shown in Figure 10 (step S13). Specifically, the behavior estimates consist of estimates for each action and represent the expected level of each action. The estimation unit 15 sends the behavior estimates output by the estimation model to the output unit 16 and then stores them in the memory unit 12.

[0047] The output unit 16 receives the behavior estimate from the estimation unit 15 and outputs recommendation information for the next action the subject will take (step S14). For example, the output unit 16 can output recommendation information for the action with the highest value among the behavior estimates. As an example, if product A has the highest value among the behavior estimates, the output unit 16 can output recommendation information that encourages the customer to view the product A page by displaying the product A page to the customer.

[0048] Next, we will explain the regeneration of the estimation model using machine learning. Figure 11 is a diagram showing the flow of the estimation model regeneration process.

[0049] The acquisition unit 11 acquires feedback information from multiple subjects regarding the recommendation information output from the output unit 16 and stores it in the storage unit 12 (step S21). For example, the acquisition unit 11 can acquire feedback information for the outputted action recommendation as positive if the subject actually performed the action, and negative if the subject did not, and can store it with labels attached to the recommended state and action. As an example, if a subject was recommended product A but did not show any response such as viewing the product page, a negative label is attached to the action estimate value that formed the basis of the recommendation and the state data column used as input.

[0050] The regeneration unit 17 reads from the storage unit 12 the state data input to the estimation model for performing behavior estimation, the behavior estimates which are the output results of the estimation model, and feedback information from multiple subjects regarding the output recommendation results, as training data to be used to regenerate the estimation model. The regeneration unit 17 uses the behavior estimates that have been assigned negative labels from the training data used for relearning as the target value for relearning by distributing some or all of the behavior estimates related to the recommended behavior to the behavior estimates of other behaviors and adding them together (step S22). As an example, as shown in Figure 12, if the output unit 16 recommends product A as an action to the recommended subject based on the estimation results of the estimation unit 15, but the recommended subject does not view the page for product A, the target value is corrected by distributing the estimated value of product A from the estimated purchase behavior estimated by the estimation model to the other behavior estimates. Note that the distribution of the estimated value of product A may be performed based on weights calculated according to the magnitude of each behavior estimate of products other than product A. For example, the larger the estimated behavior for products other than product A, the larger the estimated value of that product may be allocated.

[0051] The method of generating target values ​​(corrected estimation information) by modifying the behavior estimates by the regeneration unit 17 described above is just one example, and may be carried out by methods such as those described below. For example, the regeneration unit 17 may modify the target value by adding some or all of the behavior estimates other than the recommended behavior estimates to the recommended behavior estimates, for behavior estimates with positive labels among the training data used for retraining. For example, if the output unit 16 recommends product A as an action for the target person based on the estimation results of the estimation unit 15, and the target person views the product A page, the regeneration unit 17 modifies the target value by adding some or all of the behavior estimates other than product A among the purchase behavior estimates estimated by the estimation model to the behavior estimates for product A. At this time, the regeneration unit 17 may calculate a value to add from the behavior estimates other than product A to the behavior estimates for product A based on weights calculated according to the magnitude of the behavior estimates other than product A, and transfer the behavior estimates from products other than product A to product A. For example, the larger the magnitude of the behavior estimates for products other than product A, the larger the value of the estimates transferred to product A.

[0052] As described above, the regeneration unit 17 corrects the behavior estimates by decreasing the estimated value of a predetermined behavior and increasing the estimated values ​​of other behaviors based on feedback information from the subject regarding the recommendation information. More specifically, it increases the estimated values ​​of other behaviors based on the value obtained by decreasing the estimated value of a predetermined behavior, by transferring the value obtained by decreasing the estimated value of a predetermined behavior to the estimated value of other behaviors. This makes it possible to correct the behavior estimates to reflect the subject's response to the recommendation information.

[0053] The regeneration unit 17 reads the estimation model stored in the estimation unit 15 and uses the state data used as input when making a recommendation and the target value obtained by correcting the behavior estimate to retrain (fine-tune) the read estimation model using machine learning with the ACIL method and the SAiL method (step S23). For example, if product A was recommended but the customer had a negative reaction, retraining is performed using the corrected behavior estimate. At this time, the state data used as input when recommending product A and the obtained target value are considered as a sequence of states and behaviors that represent a successful case, and the ACIL method and the SAiL method are used to regenerate an estimation model that estimates behaviors similar to the successful case.

[0054] The regeneration unit 17 supplies the regenerated estimated model to the estimation unit 15, and upon receiving the estimated model, the estimation unit 15 newly stores it internally (step S24).

[0055] As described above, according to the estimation system 10 in this disclosure, a machine learning model capable of making appropriate suggestions to users can be generated by retraining the estimation model using behavioral estimates that reflect the target person's response to recommendation information.

[0056] Here, we will describe an example of the application of the estimation system disclosed in this disclosure. One example of application is its use in health support through an application. The target person is someone who is interested in health and receives exercise advice from, for example, a medical professional such as a doctor. Based on data such as the target person's age, weight, and diet, state data is generated, and an appropriate exercise is recommended to the target person from among candidate exercise menus such as running or walking. If the target person does not perform the recommended exercise, this is considered a negative response, and the behavior estimate is modified to recommend other exercises, and the estimation model is retrained. Then, by using the retrained estimation model, it is possible to recommend more appropriate exercise to the target person. Through such use, it can also be applied to support decision-making by medical professionals such as doctors.

[0057] <Second Embodiment> Next, a second embodiment of this disclosure will be described with reference to the drawings. This embodiment shows an outline of the estimation system and the like described in the embodiments described above. Note that the drawings may be relevant to any of the embodiments.

[0058] First, the hardware configuration of the information processing device 100 in this disclosure will be described. The information processing device 100 is composed of a general information processing device, and as an example, it is equipped with the following hardware configuration as shown in Figure 13. ·CPU(Central Processing Unit)101(Arithmetic unit) ROM (Read Only Memory) 102 (Storage Device) • RAM (Random Access Memory) 103 (Storage Device) • Program group 104 loaded into RAM 103 • Storage device 105 for storing the program group 104 • Drive device 106 for reading and writing to external storage medium 110 of the information processing device. • Communication interface 107 connecting to a communication network 111 outside the information processing device. • Input / output interface 108 for data input and output. • Bus 109 connecting each component

[0059] Figure 13 shows an example of the hardware configuration of the information processing device 100, and the hardware configuration of the information processing device is not limited to the case described above. For example, the information processing device may consist of only a part of the configuration described above, such as not having a drive device 106. In addition, the information processing device may use a GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating point number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof instead of the CPU described above.

[0060] The information processing device 100 can be equipped with the estimation unit 121, acquisition unit 122, modification unit 123, and learning unit 124 shown in Figure 15 by having the CPU 101 acquire the program group 104 and execute it. The program group 104 is, for example, stored in advance in a storage device 105 or ROM 102, and the CPU 101 loads it into RAM 103 and executes it as needed. The program group 104 may also be supplied to the CPU 101 via a communication network 111, or it may be stored in advance in a storage medium 110, and the drive device 106 reads the program and supplies it to the CPU 101. However, the estimation unit 121, acquisition unit 122, modification unit 123, and learning unit 124 described above may be constructed with dedicated electronic circuits to realize such means.

[0061] The estimation unit 121 inputs user state information into a machine learning model and obtains estimation information, including estimated values ​​for each user action output from the machine learning model. The acquisition unit 122 obtains feedback information representing the user's response to recommendation information that recommends user actions based on the estimation information. The modification unit 123 generates modified estimation information by modifying the estimated values ​​for each action in the estimation information based on the feedback information. The learning unit 124 retrains the machine learning model based on the state information and the modified estimation information.

[0062] This disclosure, configured as described above, retrains the estimation model using action-specific estimates that have been modified to reflect user responses to recommendation information. This makes it possible to generate a machine learning model that can make appropriate suggestions to users.

[0063] Furthermore, at least one of the functions of the estimation unit 121, acquisition unit 122, modification unit 123, and learning unit 124 described above may be executed on an information processing device installed and connected at any location on the network, that is, it may be executed using so-called cloud computing.

[0064] Furthermore, the aforementioned programs can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). Programs may also be supplied to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can be supplied to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.

[0065] Although the present disclosure has been described above with reference to embodiments, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure are possible, as can be understood by those skilled in the art within the scope of the present disclosure. Furthermore, each of the embodiments described above can be combined with other embodiments as appropriate.

[0066] <Note> Some or all of the above embodiments may also be described as follows. The general configuration of the information processing apparatus, information processing method, and program in this disclosure is described below. However, this disclosure is not limited to the configurations described below. Furthermore, some or all of the configurations and functions described in Appendices 2 to 8.1, which are dependent on Appendice 1 below, may also be dependent on other Appendices 9 and 10 in the same way as in Appendices 2 to 8.1. Moreover, not limited to Appendices 1, 9, and 10, some or all of the configurations and functions described as appendices may also be dependent on similar hardware, software, various recording means for recording software, or systems, without departing from the embodiments described above. (Note 1) An estimation unit that obtains estimation information, including estimated values ​​for each user action output from a machine learning model, by inputting user state information into the machine learning model, An acquisition unit that acquires feedback information representing the user's response to recommendation information that recommends the user's actions based on the estimated information, A correction unit generates corrected estimation information by correcting the estimated values ​​for each action in the estimation information based on the feedback information, A learning unit that retrains the machine learning model based on the state information and the corrected estimation information, Equipped with an information processing device. (Note 2) The information processing device described in Appendix 1, The modification unit generates the modified estimation information by decreasing the estimated value of a predetermined action and increasing the estimated value of other actions in the estimation information based on the feedback information. Information processing device. (Note 3) The information processing device described in Appendix 2, The modification unit generates the modified estimation information by increasing the estimated values ​​of other actions based on a value obtained by decreasing the estimated value of the predetermined action. Information processing device. (Note 4) The information processing device described in Appendix 1, The modification unit generates the modified estimation information by increasing the estimated value of one specific action based on the feedback information and decreasing the estimated values ​​of several other actions that are different from the specific action. Information processing device. (Note 5) The information processing device described in Appendix 4, The modification unit generates the modified estimation information by increasing the estimated value of a specific action based on a value obtained by decreasing the estimated values ​​of a plurality of other actions in the estimation information. Information processing device. (Note 5.1) The information processing device described in Appendix 5, The modification unit generates the modified estimation information by adding the sum of the values ​​obtained by reducing the estimated values ​​of a plurality of other actions to the estimated value of the specific action, thereby increasing it. Information processing device. (Appendix 5.2) The information processing device described in Appendix 5.1, The modification unit reduces the estimated values ​​of multiple other actions in the estimation information according to the weights based on those estimated values, adds the sum of the reduced values ​​to the estimated value of the specific action to increase it, and generates the modified estimation information. Information processing device. (Note 6) The information processing device described in Appendix 1, The modification unit generates the modified estimation information by decreasing the estimated value of one specific action based on the feedback information and increasing the estimated values ​​of several other actions that are different from the specific action. Information processing device. (Note 7) The information processing device described in Appendix 6, The modification unit generates the modified estimation information by increasing the estimated values ​​of a plurality of other actions based on a value obtained by decreasing the estimated value of the specific action in the estimation information. Information processing device. (Note 8) The information processing device described in Appendix 7, The modification unit generates the modified estimation information by distributing the value obtained by reducing the estimated value of the specific action to the estimated values ​​of a plurality of other actions, thereby increasing the estimated values. Information processing device. (Note 8.1) The information processing device described in Appendix 8, The modification unit generates the modified estimation information by reducing the estimated value of the specific action in the estimation information, distributing the reduced value to the estimated values ​​of a plurality of specific actions according to the weights based on the estimated values, thereby increasing the estimated values. Information processing device. (Note 9) Information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. Information processing methods. (Note 10) In an information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. A program that executes a process. [Explanation of Symbols]

[0067] 10 Estimation System 11 Acquisition Department 12 Storage section 13 Data Generation Unit 14 First generation part 15 Estimation part 16 Output section 17 Regeneration part 100 Information Processing Devices 101 CPU 102 ROM 103 RAM 104 Program Groups 105 Storage device 106 Drive unit 107 Communication Interface 108 Input / Output Interfaces 109 Bus 110 Storage medium 111 Communication Network 121 Estimation Department 122 Acquisition Department 123 Correction section 124 Learning Department

Claims

1. An estimation unit that obtains estimation information, including estimated values ​​for each user action output from a machine learning model, by inputting user state information into the machine learning model, An acquisition unit that acquires feedback information representing the user's response to recommendation information that recommends the user's actions based on the estimated information, A correction unit generates corrected estimation information by correcting the estimated values ​​for each action in the estimation information based on the feedback information, A learning unit that retrains the machine learning model based on the state information and the corrected estimation information, Equipped with an information processing device.

2. An information processing apparatus according to claim 1, The modification unit generates the modified estimation information by decreasing the estimated value of a predetermined action and increasing the estimated value of other actions in the estimation information based on the feedback information. Information processing device.

3. An information processing apparatus according to claim 2, The modification unit generates the modified estimation information by increasing the estimated values ​​of the other actions based on the value obtained by decreasing the estimated value of the predetermined action. Information processing device.

4. An information processing apparatus according to claim 1, The modification unit generates the modified estimation information by increasing the estimated value of one specific action based on the feedback information and decreasing the estimated values ​​of several other actions that are different from the specific action. Information processing device.

5. An information processing apparatus according to claim 4, The modification unit generates the modified estimation information by increasing the estimated value of a specific action based on a value obtained by decreasing the estimated values ​​of a plurality of other actions in the estimation information. Information processing device.

6. An information processing apparatus according to claim 1, The modification unit generates the modified estimation information by decreasing the estimated value of one specific action based on the feedback information and increasing the estimated values ​​of several other actions that are different from the specific action. Information processing device.

7. An information processing apparatus according to claim 6, The modification unit generates the modified estimation information by increasing the estimated values ​​of a plurality of other actions based on a value obtained by decreasing the estimated value of the specific action in the estimation information. Information processing device.

8. An information processing apparatus according to claim 7, The modification unit generates the modified estimation information by distributing the value obtained by reducing the estimated value of the specific action to the estimated values ​​of a plurality of other actions, thereby increasing the estimated values. Information processing device.

9. Information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. Information processing methods.

10. In an information processing device, By inputting user state information into a machine learning model, estimated information including estimated values ​​for each user action output from the machine learning model is obtained. Based on the estimated information, feedback information is obtained that represents the user's response to recommendation information that recommends the user's actions. Based on the feedback information, modified estimation information is generated by correcting the estimated values ​​for each action in the estimation information. Based on the state information and the corrected estimation information, the machine learning model is retrained. A program that executes a process.