Policy evaluation apparatus, policy evaluation method, and policy evaluation program

JP2024088222A5Pending Publication Date: 2025-11-07RAKUTEN GROUP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022203285
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Offline evaluations of new policies are subject to bias due to changes in requirements during the implementation of the first policy, leading to inaccurate performance assessments.

Method used

A policy evaluation device and method that segments historical data based on unchanged requirements, generates a learning model for each category using a machine learning algorithm, and performs offline evaluation using estimated amounts to minimize bias.

Benefits of technology

Enables accurate offline evaluation of new policies by addressing bias, optimizing advertising strategies, and maximizing cumulative rewards without the risks and costs associated with online experimentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a policy evaluation apparatus, a policy evaluation method, and a policy evaluation program capable of coping with bias in offline evaluation.SOLUTION: A policy evaluation apparatus 30 includes one or more processors 32 and one or more memories 34. The memory 34 stores history data 38 including a plurality of data sets recorded when a first policy is executed. The processor 32 is configured to: generate a plurality of segment data by segmenting the plurality of data sets based on the fact that at least one requirement for the first policy is unchanged; generate a learning model for a second policy for each segment by training a machine learning algorithm on each segment data; and evaluate the second policy offline using an estimator approximated from each segment data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a policy evaluation device, a policy evaluation method, and a policy evaluation program. [Background technology]

[0002] For example, when determining the content of an advertisement to be displayed in one advertisement space, such as a banner advertisement on a web page, a machine learning model may be used to predict the effectiveness of the advertisement. For example, Patent Document 1 discloses a learning model that outputs a prediction result of the advertisement effectiveness when an image that is an advertisement design proposal is input.

[0003] When actually placing an advertisement, after deciding on an image for the advertisement, it is necessary to decide on a policy, such as when and to which users the image should be displayed. In order to evaluate this policy, for example in the case of online advertisements, online experiments may be conducted in which the advertisement is temporarily implemented. However, not only is implementing an advertisement costly, but if the policy is not appropriate, there are also risks such as negative user impressions and reduced sales.

[0004] In response to this, Non-Patent Document 1 discloses offline evaluation, which is a method for evaluating low-risk measures. Such evaluation of measures can be used in various fields, such as various types of marketing or communication with customers, not limited to advertising. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent Publication No. 2022-032204 [Non-patent literature]

[0006] [Non-Patent Document 1] Yuta Saito, 3 others, “Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation”, [online], October 26, 2021 (v5), arXivLabs, [searched on December 16, 2020], Internet<URL:https: / / arxiv.org / abs / 2008.07146> Summary of the Invention [Problem to be solved by the invention]

[0007] In the offline evaluation described above, the performance of a new measure that differs from past measures is evaluated using data from past measures, but such evaluations may be subject to bias.

[0008] The present disclosure aims to provide a policy evaluation device, a policy evaluation method, and a policy evaluation program that can address bias in offline evaluation. [Means for solving the problem]

[0009] A policy evaluation device according to one embodiment of the present disclosure includes one or more processors and one or more memories, and the memories store historical data including multiple data sets that are records of when a first policy was implemented. The processor is configured to perform the following: generate multiple partitioned data by partitioning the multiple data sets based on the fact that at least one requirement related to the first policy is constant; generate a learning model related to a second policy for each partition by training a machine learning algorithm using each partitioned data; and evaluate the second policy offline using an approximate estimator from each partitioned data.

[0010] A policy evaluation method according to one embodiment of the present disclosure includes having one or more computers acquire historical data including multiple data sets that are records of when a first policy was implemented, generate multiple partitioned data by partitioning the multiple data sets based on the fact that at least one requirement related to the first policy is constant, generate a learning model related to a second policy for each partition by training a machine learning algorithm using each partitioned data, and evaluate the second policy offline using an approximate estimator from each partitioned data.

[0011] A policy evaluation program according to one embodiment of the present disclosure is a program for causing one or more computers to execute the following steps: acquire historical data including multiple data sets that are records of when a first policy was implemented; generate multiple partitioned data by partitioning the multiple data sets based on the fact that at least one requirement related to the first policy is constant; generate a learning model related to a second policy for each partition by training a machine learning algorithm using each partitioned data; and evaluate the second policy offline using an approximate estimator from each partitioned data. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram showing the configuration of a system including a policy evaluation device according to an embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of a website on which advertisements are displayed. [Diagram 3] FIG. 3 is a flowchart showing a policy evaluation method according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] An example of a policy evaluation device and a policy evaluation method according to the present disclosure will be described below with reference to the drawings. [System Overview] 1, a policy evaluation system 11 according to the present disclosure includes a web server 20 and a policy evaluation device 30. The web server 20 and the policy evaluation device 30 communicate with each other via a network 14. Each of the web server 20 and the policy evaluation device 30 is an example of a computer.

[0014] The web server 20 communicates with one or more terminals 13 via a network 14. The terminals 13 are, for example, information processing devices such as smartphones, personal computers, and tablets. The terminals 13 are user terminals operated by users. The web server 20 may provide a website 40 for providing or recommending products or services. The website 40 may have a search window (e.g., a text box, not shown) for searching for products or services. The web server 20 provides various information or processing results to each terminal 13 in response to a request from one or more terminals 13.

[0015] The network 14 includes, for example, the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), a provider terminal, a wireless communication network, a wireless base station, a dedicated line, etc. It is not necessary for all combinations of the devices shown in Fig. 1 to be able to communicate with each other, and the network 14 may include a local network in part.

[0016] The web server 20 includes a processor 22, a memory 24, and a communication device 23. The communication device 23 enables communication with other devices, such as the terminal 13 and a policy evaluation device 30, via the network 14. The memory 24 stores a display program 25 for displaying advertisements, advertisement data 27, and history data 28. The memory 24 may further store customer data 29 that may be the target of advertisements.

[0017] Customer data 29 includes, for example, registration information of users who use website 40. Registration information includes, for example, but is not limited to, each user's name, age, gender, address, account, email address, and payment information. Examples of payment information are credit card numbers, debit card numbers, or withdrawal account numbers. Customer data 29 may also include each user's purchase history at website 40 and other stores. Purchase history includes, for example, the items of merchandise purchased, the date and time of purchase, and the purchase amount.

[0018] The history data 28 may include a browsing history, purchase history, or usage history of each user on the website 40. The history data 28 includes an access history (e.g., a click history) of each user to links (e.g., links to advertisements or product screens) included in the website 40. The history data 28 may include at least a part of the user's registration information as user attribute information by linking with the customer data 29. The user attribute information included in the history data 28 is, for example, each user's age, sex, browsing history, usage history, and access history.

[0019] The policy evaluation device 30 includes one or more processors 32, one or more memories 34, and a communication device 33. The policy evaluation device 30 is a computer such as a server. The communication device 33 enables communication with other devices, such as the web server 20, via the network 14. The memory 34 stores a classification data generation program 35, a learning program 36, and an evaluation program 37. The memory 34 is further configured to store history data 38 and one or more learning models 39. The processor 32 of the policy evaluation device 30 executes the learning program 36 stored in the memory 34 to generate the learning model 39.

[0020] The web server 20 transmits the history data 28 stored in the memory 24 to the policy evaluation device 30 in response to a request from the policy evaluation device 30 or in accordance with a program for sending data. The policy evaluation device 30 stores the received history data 28 in the memory 34 as history data 38. The history data 28 may be stored as history data 38 in the memory 34 of the policy evaluation device 30 via another portable memory.

[0021] The processors 22 and 32 include, for example, arithmetic units such as a CPU, a GPU, and a TPU. The processors 22 and 32 are processing circuits configured to execute various software processes. The processing circuits may include dedicated hardware circuits (e.g., ASICs, etc.) that process at least a part of the software processes. That is, the software processes may be executed by processing circuitry that includes at least one of one or more software processing circuits and one or more dedicated hardware circuits.

[0022] The memories 24 and 34 are computer-readable media. The memories 24 and 34 include non-transitory storage media such as a random access memory (RAM), a hard disk drive (HDD), a flash memory, a read only memory (ROM), etc. The processors 22 and 32 execute a series of instructions included in a program stored in the memories 24 and 34 in response to a signal provided thereto or in response to a predetermined condition being satisfied.

[0023] [Advertising System] 2, the website 40 presents information for providing, for example, various products or services. An example of the website 40 is a shopping site that sells multiple products. Examples of the products or services provided by the website 40 include, but are not limited to, travel plans, accommodations, tickets, books, magazines, music, videos, movies, insurance, or securities.

[0024] The website 40 may include one or more advertisement frames 41 (41A, 41B, 43C, 43D). The website 40 may include a plurality of recommendation frames 42 that display recommended products. The advertisement frames 41 and the recommendation frames 42 are examples of display frames. One advertisement image (e.g., a banner advertisement) may be displayed in one advertisement frame 41A. Alternatively, a plurality of candidate images may be displayed sequentially in one advertisement frame 41B. When there are a plurality of advertisement frames 41A to 41D, the display method and contents may be different from each other. Also, there may be a plurality of advertisement frames 41B (or 41C) of the same type.

[0025] The advertisement images or product images displayed on the website 40 may be changed according to a specified rule. For example, the images may be changed every certain period of time, or may be changed according to the attributes of the user browsing the website. As an example, a plurality of candidate images (e.g., five) to be displayed in one banner advertisement space during a promotional event period (e.g., one month) may be prepared, and the images may be displayed in a specified order. Furthermore, some of the candidate images (e.g., discount coupons) may be replaced with other candidate images (other discount coupons) during the event period.

[0026] The advertisement data 27 includes the type of each advertisement space 41, one or more candidate images corresponding to each advertisement space 41, a display change history for each advertisement space 41, and advertisement measures for each period. The display change history includes, for each advertisement unit (e.g., for each event), one or more advertisement images to be used or candidates for use, the date and time when each advertisement image will start to be used, and data related to the advertisement space 41 in which each advertisement is displayed.

[0027] The history data 38 includes multiple data sets, and each data set includes multiple data acquired each time a display request for a website 40 (and an advertisement included in the website 40) is made from a user to the web server 20. For each advertisement space 41, the multiple data include data related to, for example, the attributes of the user who requested the display, the display time, the advertisement displayed at each time, and whether or not the advertisement was clicked. Furthermore, each data set may include data related to whether or not the clicked user purchased the product related to the advertisement.

[0028] [Policy formulation] A policy refers to a method or plan for recommending some content to a user or encouraging some action. For example, an advertising policy using banner ads aims to have users click on banner ads (advertising images with links) to encourage users to purchase the products related to the ads. Or, in email marketing, the aim is to have users who receive emails click on links included in the email body.

[0029] Click-through rate is an example of an indicator for achieving the goal of "making people buy the product displayed in the link they click on." Other indicators include conversion rate per click (conversion rate via click) or revenue per advertising cost.

[0030] The click rate for an advertising measure that has already been implemented can be calculated based on the history data 38. The click rate can vary depending on the advertising measure, i.e., the location and timing of displaying the advertisement, or the attributes of the target user. Therefore, when formulating a new measure, it is desirable to optimize the advertising content in order to further increase the advertising effect.

[0031] A machine learning model can be used to optimize the new strategy. The processor 32 is configured to optimize the strategy by executing the learning program 36, for example, using a multi-armed bandit algorithm, which is an example of reinforcement learning. The multi-armed bandit algorithm is a problem of sequentially searching for the best one among multiple candidates called arms. The learning program 36 uses Thompson Sampling, for example.

[0032] For example, by solving a multi-armed bandit problem with a feature (feature vector) x∈X and m possible actions a∈A={1,2,...,m}, an action a that maximizes reward Y in a certain period of time is searched for. Feature x is, for example, a user attribute such as the user's age, gender, and purchase history. Action a is, for example, a product that is the subject of advertising. Reward Y may be, for example, a click-through rate or sales when a certain product is advertised.

[0033] A measure π is defined as a mapping π:X→Δ(A) from feature x∈X to a probability distribution in action space A. π(a|x) is the probability of selecting action a (action selection probability) for data characterized by a vector x. π(a|x) can also be said to be a decision-making measure that governs which action should be taken in what situation. For example, π(a|x) can be a function that indicates which product should be advertised (which advertising image should be displayed) when x (a certain user attribute) is input.

[0034] [Evaluation of the measures] The processor 32 is configured to evaluate a new policy by executing the evaluation program 37. Evaluating a new policy without implementing it is called Off-Policy Evaluation (OPE). A policy to be evaluated is called an evaluation policy π, and a policy to be compared with the evaluation policy is called a behavior policy πb.

[0035] In this disclosure, the behavior measure may be referred to as the first measure, and the evaluation measure may be referred to as the second measure. The first measure is generally a past measure that has already been implemented, and in this example, the history data 38 is a record of when the first measure was implemented. The history data 38 includes multiple data sets. However, both the first measure and the second measure may be measures that have already been implemented.

[0036] The performance of a policy π is defined as the objective variable V(π) as follows:

[0037]

number

[0038] The performance of the second measure π can be evaluated using history data 38 when the first measure πb was implemented. For example, the history data 38 includes T data sets represented by the following Equation 2.

[0039]

number

[0040] Here, when a user selects a certain action a, the history data 38 contains only data on reward Yt based on the action a. For example, the result (whether or not there is a click) of recommending the purchase of the first product to the user out of a first product and a second product (action a = first product) is known, but the result of not recommending the first product is not known. In other words, if the first product was recommended in a past measure and reward Y was obtained as a result, and the second product is recommended in a new measure, data on the second product being recommended (action a = second product) is not included in the history data 38. Therefore, it is difficult to obtain true performance for a new measure π based on the history data 38.

[0041] The performance of the new measure π can be substituted for the true performance V(π) by the following Equation 3 as an approximate estimate based on historical data 38.

[0042]

number

[0043] Instead of Equation 3, the estimated amount can also be obtained by Equation 4 below. Equation 4 is a method called DR (Doubly Robust).

[0044]

number

[0045] As an example, a case will be described in which a second measure optimizing the user attributes targeted by the advertisement is evaluated when a banner advertisement is displayed in the advertisement space 41A. The banner advertisement includes multiple (e.g., five) candidate images, and these candidate images are displayed in the advertisement space 41A according to a rule stipulated by the measure. For example, the first measure (behavior measure πb) stipulates the ratio at which the five candidate images are displayed in the advertisement space 41A. In the first measure, the five candidate images are displayed at the same ratio, for example, in a fixed order or randomly. In this case, the click rate calculated based on the history data 38 indicates the performance of the first measure.

[0046] The second measure (evaluation measure π) in this example specifies which candidate image is to be displayed based on user attributes when the same five candidate images as the first measure are displayed one by one in the same advertising space 41A. Therefore, in the second measure, the ratio at which the five candidate images are displayed may vary depending on the user. In this example, an estimated amount (e.g., click rate) when the second measure is implemented is calculated. When the click rate is calculated as the estimated amount, it becomes possible to estimate the expected profit when a new measure is applied based on the conversion rate per click and the revenue per conversion.

[0047] Here, the estimator calculated by Equation 3 or Equation 4 is premised on the premise that the advertisement content is unchanged during the implementation of the first measure. However, in reality, one or more requirements related to the advertisement content may change during the implementation of the first measure. For example, during the period in which the history data 38 is recorded, some of the five candidate images may have been replaced with other candidate images once or multiple times. The requirement that changes in this case is the lineup of candidate images used in the advertisement.

[0048] In addition, the estimated quantity calculated by Equation 3 or Equation 4 is biased due to the fact that the advertisement content (lineup) is constant. For example, when an image with the lowest click rate among five candidate images is changed to an image with a higher advertising effect, the click rates of the changed image and the remaining four images may vary before and after the change.

[0049] To address the above bias, in the present disclosure, multiple data sets included in the history data 38 are segmented based on the fact that at least one requirement (e.g., a lineup of multiple candidate images) related to the first measure is invariant. Then, using the segmented data generated by the segmentation, an estimator is calculated for each segment.

[0050] For example, the processor 32 executes the divided data generation program 35 to divide n data sets included in the history data 38 into K data sets to create the K data sets of divided data. At this time, the processor 32 is used as a divided data generation unit configured to generate a plurality of divided data sets based on a division in which at least one requirement related to the first measure is unchanged. Furthermore, the processor 32 executes the divided data generation program 35 or another division program (not shown) to divide each of the divided data sets into training data and evaluation data.

[0051] Each data set includes, for example, data on the time when the advertisement was displayed (timestamp), feature quantity x (user attribute), corresponding action a (displayed image at each time), and result of the action r (whether the image was clicked or not). For example, in an advertising event carried out over four weeks, if one of five candidate images is changed every week, four segmented data for one week in which the five candidate images remain unchanged are created.

[0052] For partitioned data, the kth partition is defined as follows:

[0053]

number

[0054]

number

[0055]

number

[0056] In this example, the historical data 38 includes multiple data sets acquired during the implementation of the advertising campaign. Each data set includes multiple data acquired at each display request of the website 40 from the user. The website 40 includes one or more advertising frames 41 (or one or more recommendation frames 42), and the first and second measures relate to displaying advertising images (or product images) in each advertising frame 41 (or recommendation frame 42). In addition, the feature amount x is a user attribute, the action a is multiple candidate images, and the result r of the action is whether or not the user clicked on the displayed advertising image (or product image). The index value indicating the result of the first measure is the click rate. In addition, it is assumed that the multiple candidate images are invariant in each category.

[0057] First, in step S11, the processor 32 acquires a machine learning algorithm (for example, a multi-armed bandit algorithm) for formulating the second measure. Specifically, the processor 32 acquires a learning program 36 and stores it in the memory 34.

[0058] In step S12, the processor 32 acquires the history data 38. Specifically, the processor 32 receives the history data 28 from the web server 20 and stores it in the memory 34 as the history data 38. Step S12 may be performed before step S11 or may be performed simultaneously with step S11.

[0059] In step S13, the processor 32 executes the divided data generation program 35 to generate a plurality of divided data from a plurality of data sets included in the history data 38, based on the fact that at least one requirement related to the first measure is unchanged. Furthermore, the processor 32 executes the divided data generation program 35 or another division program to divide each divided data into training data and evaluation data.

[0060] In step S14, the processor 32 performs a simulation using the partitioned data generated in step S13 in the algorithm obtained in step S11, and optimizes the second measure. Specifically, the processor 32 executes the learning program 36 using training data obtained by dividing the partitioned data, and trains the machine learning algorithm with each partitioned data, thereby creating a learning model 39 optimized for each partition. At this time, the processor 32 is used as a learning model generation unit configured to generate a learning model related to the second measure for each partition. The processor 32 also stores the generated learning model 39 in the memory 34.

[0061] In step S14, the processor 32 executes the evaluation program 37 to calculate an approximate offline evaluation (OPE) estimate from the section data. Specifically, the processor 32 calculates D=D in the above formula 3 or formula 4. kand calculates an estimator for the second measure using the learning model generated for each section and the evaluation data obtained by dividing the section data. In this case, in the above formula 3 or 4, πb is calculated as the frequency distribution of the data set of the entire history data 38. At this time, the processor 32 is used as an evaluation unit configured to evaluate the second measure offline using an estimator approximated from the section data.

[0062] [Effects of the present disclosure] In offline evaluation, by using past history data 38, it is possible to evaluate a new measure (second measure) while avoiding the risks and increased costs associated with online evaluation. However, if one or more requirements (e.g., images for advertising) related to the first measure are changed while the first measure is being implemented, the history data 38 may include data fluctuations caused by the change in the requirements. Therefore, if the history data 38 of the first measure is used for offline evaluation of the second measure different from the first measure, the estimator for evaluation will include a bias.

[0063] In this regard, in the present disclosure, a plurality of data sets included in the history data 38 are divided into sections in which one or more requirements related to the first measure are invariant, thereby generating a plurality of section data. Then, a learning model 39 related to the second measure is generated for each section, and the second measure is evaluated offline using an approximate estimator from the section data. This addresses bias in the offline evaluation.

[0064] [Effects of this disclosure] According to the present disclosure, the following effects can be achieved. (1) After generating a learning model for the second measure for each category, the second measure is evaluated offline using an approximate estimator from the category data, thereby addressing bias in the offline evaluation.

[0065] (2) By generating a learning model using a reinforcement learning algorithm such as a multi-armed bandit algorithm, a policy for maximizing cumulative rewards can be developed. (3) When calculating the estimator, the index value indicating the result of the first measure is weighted using the ratio of the behavioral selection probability of the first measure to that of the second measure, so that the second measure can be evaluated using the historical data 38 when the first measure was implemented.

[0066] (4) By generating segmentation data based on the invariance of multiple candidate images that may be displayed in each ad space 41, bias in offline evaluation can be addressed. (5) By using the history data 38 including the user attributes for offline evaluation of the second measure, it is possible to formulate a second measure that can display advertisements according to the user attributes. This is expected to increase the click-through rate of advertisements.

[0067] This embodiment can be modified as follows: This embodiment and the following modifications can be combined with each other to the extent that there is no technical contradiction. [Change Example 1] The first and second measures may include displaying an image in any display frame, not limited to the advertisement frame 41 or recommendation frame 42 of the website 40. For example, the display frame may be a display frame in a digital signage that transmits information by installing a video display device such as a display or projector in a station, store, facility, or office. When transmitting information by a video display device, the feature amount x may be an attribute of the location where the information is installed or an attribute of the method of transmitting the information. Moreover, the action a may be each candidate image, and the result r of the action may be an inquiry about the display content or a sales promotion effect.

[0068] Alternatively, the first and second measures may include displaying one of multiple search result candidates in a display frame for displaying search results on a search site. In this case, the feature amount x may be an attribute of a user who performed a search, the behavior a may be each search result, and the result r of the behavior may be whether or not the user clicked on the displayed search result.

[0069] [Change Example 2] The first and second measures are not limited to images, and may include selecting one or more targets from a plurality of target candidates. As an example, the first and second measures may include determining a product, service, user, or user attribute to be the target of advertising or marketing. For example, the target candidate may be a user or delivery address to which a postal advertisement is sent, a target area for posting, or a mailing address. Alternatively, the target candidate may be a user, user attribute, or telephone number to be a candidate for telephone sales or telephone appointments to provide information regarding monitor recruitment, sales, or services. In addition, the target candidate may be a coupon or product sample to be distributed to users by various media (paper or electronic media).

[0070] In this case, the feature x may be an attribute of a target candidate, the action a may be multiple target candidates, and the result r of the action may be the presence or absence of a reaction or the quality of the reaction (i.e., whether or not a desired response was obtained).

[0071] According to a second modification, bias in offline evaluation can be addressed by generating segmentation data based on the invariance of multiple candidate subjects. [Change Example 3] The first and second measures may include determining a display order of a plurality of images to be displayed in sequence in one display frame. In this case, it is sufficient that the plurality of images are invariant in each section. In this case, the display frame may be, for example, a display frame for a carousel advertisement, a display frame for a banner advertisement, a display frame for a recommended product, or a notification frame for notifying update information. These display frames may be included in an image displayed by a video display device such as the first modification, or may be included in a website.

[0072] According to a third variant, bias in the offline evaluation can be addressed by generating segmentation data based on invariance across multiple images. [Change Example 4] The first and second measures may include selecting one or more items or services to be recommended to the user from a plurality of items or services. In this case, the prices of the items or services may be constant for each category, the lineup of the items or services may be constant for each category, or both the prices and the lineup may be constant. The services may be, for example, music, movies, videos, television programs, insurance products, financial products, etc.

[0073] In this case, the feature x may be a user attribute (e.g., age, gender, past viewing history, past purchasing history, etc.), the behavior a may be a type of product or service, and the result r of the behavior may be the presence or absence of a reaction to a recommendation, the presence or absence of a purchase, or the viewing time.

[0074] According to the fourth modification, bias in offline evaluation can be addressed by generating segment data based on at least one of price and product lineup remaining constant.

[0075] [Change Example 5] The website 40 may be configured to present a search screen for searching for multiple products or multiple services and a search result screen thereof. In this case, each data set included in the history data may include multiple data acquired each time a search request is made by a user on the search screen. The first and second measures include determining a display order (ranking) of multiple search results to be displayed on the search result screen, and displaying multiple search results on the search result screen in an order according to the search intent. The multiple search results may change, for example, due to a change in the search algorithm. In this case, the multiple search results may be unchanged in each category, the lineup of products or services to be searched may be unchanged in each category, or both the search results and the search target may be unchanged.

[0076] According to the fifth modified example, bias in offline evaluation can be addressed by generating classification data based on the invariance of at least one of the multiple search targets and the multiple search results.

[0077] [Change Example 6] The policy evaluation device 30 may be a policy evaluation system including a classification data generating device (e.g., a computer) that generates classification data, a learning device (e.g., a computer) that generates a learning model related to the second policy, and an evaluation device (e.g., a computer) that evaluates the second policy offline. This policy evaluation system may further include a web server 20. In other words, the policy evaluation device 30 is not limited to a device physically housed in one housing, but may be a virtual device including multiple computers for realizing multiple functional units such as a data acquisition unit that acquires history data, a classification data generating unit, a learning model generating unit, and an evaluation unit.

[0078] [Change Example 7] The machine learning algorithm for formulating the second measure is not limited to the multi-armed bandit algorithm. For example, any algorithm, such as an algorithm related to other reinforcement learning, can be used.

[0079] [Change Example 8] The index value indicating the result of the policy, the feature amount x, the action a, and the result of the action r are not limited to those exemplified in this disclosure and can be changed arbitrarily. For example, when attribute information of a user who accesses a display related to the policy cannot be obtained, the feature amount x may be an attribute of the product or service to be advertised, or can be changed to an attribute of the display frame (e.g., size or display position), or to an arbitrary feature amount.

[0080] [Change Example 9] The conditions for generating the classification data can be changed arbitrarily. For example, the starting point of classification may be environmental changes surrounding the user, such as changes in policies or tax rates, changes in social conditions or behavior, and changes in seasons or weather. A certain margin may be provided between classifications until the changes settle down. This is because user behavior may change with such environmental changes. In addition, when creating classification data, multiple data sets may be divided into classifications in which multiple requirements are constant. In particular, it is preferable to generate classification data so that one or more requirements that are expected to contain a lot of bias are constant. Conversely, if the impact of a change in a requirement is small, classification based on that requirement may not be performed. That is, "at least one requirement" in this disclosure means a requirement that may have a significant impact on offline evaluation.

[0081] Below, aspects that can be understood from the above embodiment and modified examples are listed. [1] A system having one or more processors and one or more memories; The memory stores history data including a plurality of data sets that are records of when the first measure was implemented; The processor, Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm with the respective divided data to generate a learning model related to the second measure for each divided data; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A policy evaluation device configured to execute the above.

[0082] [2] The algorithm is a multi-armed bandit algorithm, Each of the data sets includes data on features, actions corresponding to the features, and results of the actions. The policy evaluation device according to [1] above.

[0083] [3] The calculation of the estimator includes weighting an index value indicating a result of the first measure using a ratio of a behavior selection probability of the first measure to a behavior selection probability of the second measure. The policy evaluation device according to [2] above.

[0084] [4] The first measure and the second measure include displaying one of a plurality of candidate images in a display frame; the plurality of candidate images are invariant in each of the segments; A policy evaluation device according to any one of the above [1] to [3].

[0085] [5] each of the data sets includes a plurality of data acquired at each time a user requests a website to be displayed; the website includes one or more advertising spaces; The first measure and the second measure include displaying an advertisement image in each of the advertisement spaces; the feature amount is an attribute of the user, the action is each candidate image, and the result of the action is whether or not the user clicked on the displayed advertisement image; the plurality of candidate images are invariant in each of the segments; The policy evaluation device according to the above [2] or [3].

[0086] [6] the first measure and the second measure include selecting one or more targets from a plurality of target candidates; the plurality of candidate objects are invariant in each of the partitions; A policy evaluation device according to any one of the above [1] to [3].

[0087] [7] The first and second measures include determining a display order of a plurality of images to be displayed sequentially in one display frame; the plurality of images are invariant in each of the sections; A policy evaluation device according to any one of the above [1] to [3].

[0088] [8] The first measure and the second measure include selecting one or more products or services to recommend to the user from a plurality of products or a plurality of services; In each of the categories, at least one of the prices of the plurality of items or the plurality of services and the lineup of the plurality of items or the plurality of services is unchanged. A policy evaluation device according to any one of the above [1] to [3].

[0089] [9] Each of the data sets includes a plurality of data acquired upon each search request from a user on a search screen; The first measure and the second measure include determining a display order of a plurality of search results to be displayed on a search result screen; In each of the segments, at least one of the plurality of search targets and the plurality of search results is unchanged. A policy evaluation device according to any one of the above [1] to [3].

[0090]

[10] A memory in which historical data including a plurality of data sets that are records of when the first measure was implemented is stored; A classification data generation unit configured to generate a plurality of classification data by classifying the plurality of data sets based on at least one requirement related to the first measure being unchanged; a learning model generation unit configured to generate a learning model related to the second measure for each of the categories by training a machine learning algorithm with the category data; an evaluation unit configured to perform an offline evaluation of the second policy for each of the segments by using the segment data; A policy evaluation device comprising:

[0091]

[11] A system comprising one or more processors and one or more memories, The memory stores history data including a plurality of data sets that are records of when the first measure was implemented; The processor, Segmenting the plurality of data sets based on environmental changes during the period in which the first measure was implemented, thereby generating a plurality of segmented data; training a machine learning algorithm with the respective divided data to generate a learning model related to the second measure for each divided data; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A policy evaluation device configured to execute the above.

[0092]

[12] A policy evaluation device according to any one of [1] to

[11] above; A server having a memory in which history data including a plurality of data sets that are records of when the first measure was implemented is stored; A policy evaluation system that is equipped with the following:

[0093]

[13] One or more computers, Obtaining historical data including a plurality of data sets that are records of when the first measure was implemented; Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm with the respective divided data to generate a learning model related to the second measure for each divided data; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A method for evaluating policies, including carrying out the above.

[0094]

[14] One or more computers, Obtaining historical data including a plurality of data sets that are records of when the first measure was implemented; Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm with the respective divided data to generate a learning model related to the second measure for each divided data; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A policy evaluation program to implement the above. [Explanation of symbols]

[0095] 11...system, 13...terminal, 14...network, 20...web server, 22...processor, 23...communication device, 24...memory, 25...display program, 27...advertising data, 28...history data, 29...customer data, 30...policy evaluation device, 32...processor, 33...communication device, 34...memory, 35...categorised data generation program, 36...learning program, 37...evaluation program, 38...history data, 39...learning model, 40...website, 41 (41A-41D)...advertising space, 42...recommendation space.

Claims

1. One or more processors and one or more memories; The memory stores history data including a plurality of data sets that are records of when the first measure was implemented; The processor, Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm using the data for each of the segments to generate a learning model related to the second measure for each segment; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A policy evaluation device configured to execute the above.

2. The algorithm is a multi-armed bandit algorithm, Each of the data sets includes data on features, actions corresponding to the features, and results of the actions. The policy evaluation device according to claim 1 .

3. The calculation of the estimator includes weighting an index value indicating a result of the first measure using a ratio of a behavior selection probability of the first measure to a behavior selection probability of the second measure. The policy evaluation device according to claim 2 .

4. The first measure and the second measure include displaying one of a plurality of candidate images in a display frame; the plurality of candidate images are invariant in each of the segments; The policy evaluation device according to any one of claims 1 to 3.

5. Each of the data sets includes a plurality of data acquired each time a user requests to display a website; the website includes one or more advertising spaces; The first measure and the second measure include displaying an advertisement image in each of the advertisement spaces; the feature amount is an attribute of the user, the action is each candidate image, and the result of the action is whether or not the user clicked on the displayed advertisement image; the plurality of candidate images are invariant in each of the segments; The policy evaluation device according to claim 2 or 3.

6. The first measure and the second measure include selecting one or more targets from a plurality of target candidates; the plurality of candidate objects are invariant in each of the partitions; The policy evaluation device according to any one of claims 1 to 3.

7. The first measure and the second measure include determining a display order of a plurality of images to be sequentially displayed in one display frame; the plurality of images are invariant in each of the sections; The policy evaluation device according to any one of claims 1 to 3.

8. The first measure and the second measure include selecting one or more products or services to recommend to the user from a plurality of products or a plurality of services; In each of the categories, at least one of the prices of the plurality of items or the plurality of services and the lineup of the plurality of items or the plurality of services is unchanged. The policy evaluation device according to any one of claims 1 to 3.

9. Each of the data sets includes a plurality of data acquired each time a search request is made by a user on a search screen, The first measure and the second measure include determining a display order of a plurality of search results to be displayed on a search result screen, In each of the segments, at least one of the plurality of search targets and the plurality of search results is unchanged. The policy evaluation device according to any one of claims 1 to 3.

10. On one or more computers, Obtaining historical data including a plurality of data sets that are records of when the first measure was implemented; Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm using the data for each of the segments to generate a learning model related to the second measure for each segment; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A method for evaluating policies, including carrying out the following:

11. On one or more computers, Obtaining historical data including a plurality of data sets that are records of when the first measure was implemented; Segmenting the plurality of data sets based on at least one requirement related to the first measure being unchanged, thereby generating a plurality of segmented data; training a machine learning algorithm using the data for each of the segments to generate a learning model related to the second measure for each segment; evaluating the second measure offline using an approximate estimator from each of the partitioned data; A policy evaluation program to implement the above.