Information processing method, information processing device, and program

JP2023170924A5Active Publication Date: 2025-05-26FUJIFILM CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022083028
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-26
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Existing information recommendation technologies face challenges in domain generalization, as they struggle to maintain prediction accuracy when transferred from one facility to another due to domain shifts, and existing methods lack effective ways to select suitable models without historical user behavior data from the target domain.

Method used

An information processing method that evaluates the similarity between facilities using metadata and facility-related information to select a model suitable for the target domain, even without user behavior history data, by training multiple models on different facilities and evaluating their performance based on facility characteristics.

Benefits of technology

Enables robust information recommendation by selecting a model that performs well in the target domain, despite lacking user behavior history, thereby improving prediction accuracy and adaptability across different facilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an information processing method, an information processing device, and a program that can perform appropriate information recommendation at an introduction destination facility even if data of a user's action history on items at the introduction destination facility cannot be used to evaluate the performance of a model.SOLUTION: The information processing method is performed by one or more processors. A plurality of models trained using one or more of datasets including a user's action history on items collected at each of a plurality of mutually different first facilities are prepared. The information processing method includes acquiring, by the one or more processors, characteristics of a second facility different from the plurality of first facilities and each of the plurality of first facilities, assessing a similarity between the acquired characteristics of the second facility and the characteristics of the first facility for which the datasets used to train the models are collected, and selecting a model suitable for the second facility from among the plurality of models on the basis of the similarity.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing method, an information processing device, and a program, and in particular to an information recommendation technology that makes recommendations robust against domain shifts. [Background technology]

[0002] In systems that provide users with a variety of items, such as electronic commerce (EC) sites or document information management systems, it is difficult in terms of both time and cognitive ability for users to select the best item from the many items available. In the case of EC sites, items are the products sold on the site, while in the case of document information management systems, items are the document information stored in the system.

[0003] Research is being conducted into recommendation technology, which presents selection candidates from a large number of items to assist users in selecting items. Generally, when a recommendation system is introduced into a facility, the model is trained based on data collected at the facility. However, if the same recommendation system is introduced into a facility other than the facility where the data used for training was collected, the model's predictive accuracy can decline. The problem of machine learning models not functioning well in unknown facilities is called domain shift, and in recent years, research into domain generalization, which is research into improving robustness against domain shifts, has become active, particularly in the field of image recognition. However, there are still very few examples of research into domain generalization in recommendation technology.

[0004] Non-Patent Document 1 describes a method for selecting a model to be used for transfer learning, i.e., a pre-trained model for fine-tuning, from multiple models trained in several different languages ​​in interlingual transfer learning applied to cross-language translation. In Non-Patent Document 1, to perform interlingual transfer learning, the similarity between a target domain and a source domain is estimated based on several feature quantities. The feature quantities used to estimate the similarity include dataset size, word overlap, geographic distance, genetic distance, and phonological distance.

[0005] Non-Patent Document 2 describes a configuration in which, when predicting a user's item rating value, there is data from a target domain and multiple source domains, predictions based on the source domain data are weighted by the similarity between the source domain and the target domain and added together. The domain similarity in Non-Patent Document 2 is configured to learn from data so as to minimize the prediction error of the target domain.

[0006] Patent Document 1 describes a configuration in which the similarity between a plurality of pre-stored models and feature data of a patient is determined, and a model having features similar to the feature data of a target patient is searched for and used from among the plurality of models. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Patent No. 6782802 [Non-patent literature]

[0008] [Non-Patent Document 1] Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang,Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig “Choosing Transfer Languages ​​for Cross-Lingual Learning”(ACL 2019) [Non-patent document 2] “Transfer Learning for Collective Link Prediction in Multiple Heterogenous Domains” (ICML 2010) [Non-patent document 3] Ivan Cantador, Ignacio Fenandez-Tobias, Shlomo Bwrkovsky, Paolo Cremonesi, Chapter 27:"Cross-domain Recommender System" (2015 Springer) Summary of the Invention [Problem to be solved by the invention]

[0009] Non-Patent Document 1 is a study on translation technology, not on information recommendation technology. In the technology described in Non-Patent Document 1, if language-specific features are not used when estimating similarity, the similarity estimation performance decreases.

[0010] The technology described in Non-Patent Document 2 cannot learn domain similarity without historical data on user behavior or evaluations of the target domain.

[0011] Furthermore, both Non-Patent Document 1 and Non-Patent Document 2 are studies on domain adaptation, and do not aim at generalization to unknown domains.

[0012] The technology described in Patent Document 1 can be applied to models for each patient, but cannot be applied to models for each facility (domain) that are assumed in information recommendation. Furthermore, even if the features of patients are similar, the predictive performance of the model may not be sufficient due to differences in domains.

[0013] When the learning domain and the domain where the system is to be introduced are different, one possible way to achieve information recommendation that is robust to domain shifts is to train multiple models using datasets collected in advance at multiple different facilities, evaluate the performance of these multiple models using a dataset containing users' behavioral history regarding items collected at the facility where the system is to be introduced before introduction, and then select the most appropriate model from among the multiple models.

[0014] However, when evaluating a model before implementation, it is possible that a dataset cannot be prepared for the facility where the model will be implemented, or that a sufficient amount of data required for model evaluation is not available. In such cases, it is not possible to evaluate the performance of the candidate models prepared in advance, and it is not possible to select a model that is appropriate for the facility where the model will be implemented.

[0015] The present disclosure has been made in consideration of these circumstances, and aims to provide an information processing method, an information processing device, and a program that make it possible to recommend information using a model that is suitable for the facility where the method is being implemented, even if data on the user's behavioral history regarding items at the facility where the method is being implemented cannot be used to evaluate the performance of the model. [Means for solving the problem]

[0016] An information processing method according to a first aspect of the present disclosure is an information processing method executed by one or more processors, which includes preparing a plurality of models trained using one or more datasets containing user behavioral histories regarding items collected at each of a plurality of first facilities that are different from each other, and the one or more processors acquiring characteristics of each of the plurality of first facilities and a second facility that is different from the plurality of first facilities, evaluating the similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the datasets used to train each model were collected, and selecting a model suitable for the second facility from the plurality of models based on the similarity.

[0017] According to this aspect, each of the prepared models is a model trained using one or more datasets from multiple datasets collected at multiple first facilities. Instead of directly evaluating the performance of each model at the second facility, the one or more processors evaluate the similarity between the facilities using the characteristics of the first facility from which the datasets used to train each model were collected and the characteristics of the second facility. A model trained using a dataset collected at a first facility with characteristics similar to those of the second facility as the main dataset for training may also exhibit relatively high performance at the second facility. According to this aspect, even if data on the user's behavior history with respect to items at the second facility is not available, a model suitable for recommending information at the second facility can be selected based on the similarity of the facility's characteristics.

[0018] The facility includes the concept of a group including a plurality of users, such as a company, a hospital, a store, a government agency, an e-commerce site, etc. The plurality of first facilities and second facilities may be in different domains.

[0019] An information processing method according to a second aspect of the present disclosure may be configured such that, in the information processing method according to the first aspect, one or more processors extract statistical information of a dataset of metadata that is an explanatory variable used in training the model, and the characteristics include the statistical information.

[0020] An information processing method according to a third aspect of the present disclosure may be configured such that, in the information processing method according to the second aspect, the metadata includes at least one of a user attribute and an item attribute.

[0021] An information processing method according to a fourth aspect of the present disclosure may be configured in such a way that, in the information processing method according to any one of the first to third aspects, one or more processors acquire facility-related information other than metadata included in a dataset used to train the model, and the characteristics include the facility-related information.

[0022] An information processing method according to a fifth aspect of the present disclosure may be configured such that, in the information processing method according to the fourth aspect, the facility-related information is extracted by web crawling.

[0023] An information processing method according to a sixth aspect of the present disclosure may be the information processing method according to the fourth or fifth aspect, wherein one or more processors receive the facility-related information via a user interface.

[0024] An information processing method according to a seventh aspect of the present disclosure may be configured such that, in the information processing method according to any one of the first to sixth aspects, one or more processors obtain an evaluation value of predictive performance at a first facility where datasets used to train each of the multiple models were collected, and select a model suitable for a second facility from among the multiple models based on the similarity and the evaluation value of predictive performance.

[0025] An information processing method according to an eighth aspect of the present disclosure may be configured such that, in the information processing method according to any one of the first to seventh aspects, one or more processors acquire compatibility evaluation information indicating an evaluation of the suitability of the model for the second facility, separately from the similarity, and select a model suitable for the second facility from among multiple models based on the similarity and the compatibility evaluation information.

[0026] An information processing method according to a ninth aspect of the present disclosure is the information processing method according to the eighth aspect, wherein the compatibility assessment information includes results of a questionnaire given to users of the second facility.

[0027] An information processing method according to a tenth aspect of the present disclosure may be configured such that, in the information processing method according to any one of the first to ninth aspects, one or more processors use characteristics of a plurality of third facilities whose similarity with the characteristics of the first facility has been evaluated, and evaluate the similarity between the characteristics of the second facility and the characteristics of the first facility based on the similarity between the characteristics of the second facility and the characteristics of the plurality of third facilities.

[0028] An information processing method according to an eleventh aspect of the present disclosure may be configured in the information processing method according to the tenth aspect, wherein one or more processors store, in a storage device, characteristics of a plurality of third facilities and respective degrees of similarity between the characteristics of the first facility and the characteristics of the plurality of third facilities.

[0029] An information processing method according to a twelfth aspect of the present disclosure is an information processing method according to any one of the first to eleventh aspects, wherein the model may be a predictive model used in a recommendation system that recommends items to a user.

[0030] An information processing method according to a thirteenth aspect of the present disclosure may be configured in such a way that, in the information processing method according to any one of the first to twelfth aspects, one or more processors store a plurality of models in a storage device.

[0031] An information processing method according to a fourteenth aspect of the present disclosure may be configured such that, in the information processing method according to the thirteenth aspect, one or more processors store in a storage device the characteristics of the first facility from which the data sets used to train each model were collected, in association with the model.

[0032] An information processing device according to a fifteenth aspect of the present disclosure is an information processing device comprising one or more processors and one or more storage devices in which instructions to be executed by the one or more processors are stored, wherein a plurality of models trained using one or more datasets containing user behavioral histories regarding items collected at each of a plurality of first facilities that are different from each other are stored in the storage device, and the one or more processors acquire characteristics of a second facility that is different from the plurality of first facilities and each of the plurality of first facilities, evaluate the similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the datasets used to train each model were collected, and select a model suitable for the second facility from the plurality of models based on the similarity.

[0033] The information processing device of the fifteenth aspect may have a configuration including the same specific aspect as the information processing method according to any one of the second to fourteenth aspects described above.

[0034] A program according to a sixteenth aspect of the present disclosure enables a computer to perform the following functions: store multiple models trained using one or more datasets containing user item behavior histories collected at each of multiple different first facilities; acquire characteristics of a second facility different from the multiple first facilities and each of the multiple first facilities; evaluate the similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the datasets used to train each model were collected; and select a model suitable for the second facility from the multiple models based on the similarity.

[0035] The program of the sixteenth aspect may have a configuration including the same specific aspect as the information processing method according to any one of the second to fourteenth aspects described above. [Effects of the Invention]

[0036] According to the present disclosure, even if data on the user's item behavior history at a second facility different from the first facility where the dataset used to train the model was collected cannot be used to evaluate the model's performance, it is possible to select a model suitable for the second facility from among multiple models, thereby making it possible to recommend appropriate information at the second facility using the selected model. [Brief explanation of the drawings]

[0037] [Figure 1] Figure 1 is a conceptual diagram of a typical recommendation system. [Figure 2] Figure 2 is a conceptual diagram showing an example of supervised machine learning, which is widely used to build recommendation systems. [Figure 3] FIG. 3 is an explanatory diagram showing a typical implementation flow of a recommendation system. [Figure 4] FIG. 4 is an explanatory diagram of the introduction flow of the recommendation system when data on the facility where the system is to be introduced cannot be obtained. [Figure 5] FIG. 5 is an explanatory diagram of model learning using domain adaptation. [Figure 6] FIG. 6 is an explanatory diagram of the recommendation system implementation flow, which includes a step of evaluating the performance of the trained model. [Figure 7] FIG. 7 is an explanatory diagram showing examples of learning data and evaluation data used in machine learning. [Figure 8] FIG. 8 is a graph that schematically shows the difference in model performance depending on the data set. [Figure 9] FIG. 9 is an explanatory diagram showing an example of a recommendation system introduction flow when the learning domain and the introduction domain are different. [Figure 10] FIG. 10 is an explanatory diagram showing the problem that occurs when there is no user behavior history in the facility where the system is introduced. [Figure 11] FIG. 11 is an explanatory diagram showing an overview of the information processing method according to the first embodiment. [Figure 12]FIG. 12 is an explanatory diagram illustrating an example of processing executed by the information processing device according to the embodiment. [Figure 13] FIG. 13 is a block diagram schematically illustrating an example of a hardware configuration of an information processing device. [Figure 14] FIG. 14 is a functional block diagram showing the functional configuration of the information processing device. [Figure 15] FIG. 15 is a flowchart showing an example of the operation of the information processing device. [Figure 16] FIG. 16 is an explanatory diagram showing an example of extracting statistical information on user attributes as facility characteristics. [Figure 17] FIG. 17 is an explanatory diagram showing an example of extracting statistical information on item attributes as facility characteristics. [Figure 18] FIG. 18 is an explanatory diagram showing an example of extracting information about facilities by web crawling. [Figure 19] FIG. 19 is an explanatory diagram showing an overview of an information processing method according to the second embodiment. [Figure 20] FIG. 20 is an explanatory diagram that schematically shows the characteristics of each facility in a vector space that represents the characteristics of the facility. [Figure 21] FIG. 21 is an example of a directed acyclic graph (DAG) that represents the dependency between variables in a joint probability distribution P(X, Y). [Figure 22] FIG. 22 is a diagram showing a specific example of the probability expression of the conditional probability distribution P(Y|X). [Figure 23] Figure 23 is an explanatory diagram showing the relationship between an equation expressing the conditional probability of a user's behavior toward an item (Y=1) for a combination of the user's behavioral characteristics and the item's characteristics, and a DAG expressing the dependency between variables in the joint probability distribution P(X, Y). [Figure 24] Figure 24 is an explanatory diagram showing the relationship between a user's behavioral characteristics defined by a combination of user attribute 1 and user attribute 2, an item's behavioral characteristics defined by a combination of item attribute 1 and item attribute 2, and a DAG that represents the dependency between variables. DETAILED DESCRIPTION OF THE INVENTION

[0038] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings.

[0039] Overview of information recommendation technology First, we will provide an overview of information recommendation technology and provide specific examples of its challenges. Information recommendation technology is a technology for recommending (suggesting) items to users.

[0040] FIG. 1 is a conceptual diagram of a typical recommendation system 10. The recommendation system 10 receives user information and context information as input, and outputs information about items to recommend to the user based on the context. Context refers to various "situations," such as the day of the week, the time of day, or the weather. Items can be various objects, such as books, videos, or restaurants.

[0041] A recommendation system 10 typically recommends multiple items simultaneously. Figure 1 shows an example in which the recommendation system 10 recommends three items IT1, IT2, and IT3. The recommendation is generally considered successful if the user responds positively to the recommended items IT1, IT2, and IT3. A positive response could be, for example, a purchase, a viewing, or a visit. This type of recommendation technology is widely used, for example, on e-commerce sites and gourmet sites that introduce restaurants.

[0042] The recommendation system 10 is constructed using machine learning technology. Figure 2 is a conceptual diagram showing an example of supervised machine learning, which is widely used in constructing the recommendation system 10. In general, positive and negative examples are prepared based on the user's past behavioral history, and combinations of the user and context are input into the prediction model 12, which is then trained to reduce the prediction error. For example, viewed items viewed by the user are considered positive examples, and unviewed items not viewed by the user are considered negative examples. Machine learning is continued until the prediction error converges, and the target prediction performance is achieved.

[0043] The trained prediction model 12 is used to recommend items with a high predicted viewing probability for a combination of a user and a context. For example, when a combination of a user A and a context β is input to the trained prediction model 12, the prediction model 12 infers that there is a high probability that user A will view a document such as item IT3 under the conditions of context β, and recommends an item similar to item IT3 to user A. Note that, depending on the configuration of the recommendation system 10, items are often recommended to users without taking the context into consideration.

[0044] [Examples of data used in developing recommendation systems] A user's behavioral history is roughly equivalent to the "ground truth" in machine learning. Strictly speaking, it can be understood as a task setting to infer the next (unknown) behavior from the past behavioral history, but it is common to learn latent features based on the past behavioral history.

[0045] The user's behavior history may include, for example, a book purchase history, a video viewing history, or a restaurant visit history.

[0046] Furthermore, the main features include user attributes and item attributes. User attributes can include various elements such as gender, age, occupation, family structure, and residential area. Item attributes can include various elements such as book genre, price, video genre, length, restaurant genre, and location.

[0047] [Model construction and operation] Figure 3 is an explanatory diagram showing a typical implementation flow of a recommendation system. This diagram shows a typical flow when implementing a recommendation system in a facility. Implementing a recommendation system involves first constructing a model 14 that performs the desired recommendation task (Step 1), and then introducing and operating the constructed model 14 (Step 2). In the case of a machine learning model, "constructing" a model 14 includes training the model 14 using learning (training) data to create a predictive model (recommendation model) that meets a practical level of recommendation performance. "Operating" a model 14 means, for example, obtaining an output list of recommended items from the trained model 14 in response to an input combination of a user and a context.

[0048] Learning data is required to build the model 14. As shown in Figure 3, the model 14 of a recommendation system is generally trained based on data collected at the facility where it is installed. By training using data collected from the facility where it is installed, the model 14 learns the behavior of users at the facility where it is installed, and is able to accurately predict recommended items for users at the facility where it is installed.

[0049] However, for various reasons, data from the facility where the system is to be implemented may not be available. For example, in the case of a document information recommendation system for a company's in-house system or a hospital's in-house system, the company developing the recommendation model often does not have access to data from the facility where the system is to be implemented. When data from the facility where the system is to be implemented is not available, it is necessary to train the system using data collected at a different facility instead.

[0050] Figure 4 is an explanatory diagram of the implementation flow of the recommendation system when data from the facility where the system is being implemented is unavailable. If a model 14 trained using data collected at a facility other than the facility where the system is being implemented is operated at the facility where the system is being implemented, there is a problem that the prediction accuracy of the model 14 will decrease due to differences in user behavior between the facilities.

[0051] The problem of machine learning models not performing well at unknown facilities different from the facilities they were trained on can be broadly understood as a technical challenge of improving robustness against the problem of domain shift, where the source domain in which the model 14 was trained differs from the target domain to which the model 14 is applied. A problem setting related to domain generalization is domain adaptation, which is a learning method that uses data from both the source and target domains. The purpose of using data from a different domain even when data from the target domain exists is to compensate for the small amount of data in the target domain that is insufficient for learning.

[0052] Fig. 5 is an explanatory diagram of the case where model 14 is trained by domain adaptation. Although the amount of data collected at the facility where the model is introduced, which is the target domain, is relatively small compared to the amount of data collected at a different facility, by training using both sets of data, model 14 becomes able to predict with a certain degree of accuracy the behavior of users at the facility where the model is introduced.

[0053] [Domain Description] The difference in "facility" mentioned above is a type of difference in domain. In Non-Patent Document 3 (Ivan Cantador et al., Chapter 27: "Cross-domain Recommender System"), which is a document on research into domain adaptation in information recommendation, differences in domain are classified into the following four types:

[0054] [1] Item attribute level: For example, comedy movies and horror movies are separate domains.

[0055] [2] Item type level: For example, movies and TV series are separate domains.

[0056] [3] Item level: For example, movies and books are separate domains.

[0057] [4] System level: For example, movies in cinemas and movies broadcast on television are separate domains.

[0058] The differences in the "facilities" shown in Figure 5 etc. correspond to the system-level domain [4] among the four classifications mentioned above.

[0059] To formally define a domain, it is defined by the joint probability distribution P(X,Y) of the response variable Y and the explanatory variable X, and when Pd1(X,Y) ≠ Pd2(X,Y), d1 and d2 are different domains.

[0060] The joint probability distribution P(X,Y) can be expressed as the product of the distribution of the explanatory variables P(X) and the conditional probability distribution P(Y|X), or the product of the distribution of the objective variable P(Y) and the conditional probability distribution P(Y|X).

[0061] P(X,Y)=P(Y|X)P(X)=P(X|Y)P(Y) Therefore, a change in one or more of P(X), P(Y), P(Y|X), and P(X|Y) results in a different domain.

[0062] [Typical pattern of domain shift] [Covariate shift] When the distribution of explanatory variables P(X) differs, it is called a covariate shift. For example, when the distribution of user attributes differs between datasets, or more specifically, when the ratio of men to women differs, this corresponds to a covariate shift.

[0063] [Prior probability shift] When the distribution of the objective variable P(Y) differs, it is called a prior probability shift. For example, when the average view rate or average purchase rate differs between datasets, this corresponds to a prior probability shift.

[0064] [Concept shift] When the conditional probability distributions P(Y|X) and P(X|Y) differ, it is called a concept shift. For example, the probability that a company's R&D department will read a data analysis document is P(Y|X), but this differs between data sets, which is an example of a concept shift.

[0065] Research on domain adaptation or domain generalization can be divided into two types: one that assumes one of the above patterns as the main factor, and one that considers how to deal with changes in P(X,Y) without considering which pattern is the main factor. In the former case, many studies assume covariate shifts in particular.

[0066] [Why domain shifts affect your business] Prediction / classification models that perform prediction or classification tasks make inferences based on the relationship between explanatory variable X and target variable Y, so naturally, prediction / classification performance declines when P(Y|X) changes. Furthermore, when machine learning a prediction / classification model, the prediction / classification error within the training data is minimized. However, for example, when the frequency with which the explanatory variable X = X_1 is greater than the frequency with which X = X_2 is achieved, i.e., when P(X=X_1)>P(X=X_2), there is more data with X=X_1 than with X=X_2, and so reducing the error for X=X_1 is prioritized over reducing the error for X=X_2. Therefore, prediction / classification performance declines when P(X) changes between facilities.

[0067] Domain shift can be a problem not only for information recommendation but also for models of various tasks. For example, when a model for predicting the risk of employee resignation is trained using data from one company, domain shift can become a problem when it is used by another company.

[0068] In addition, for models that predict antibody production by cells, domain shift can become a problem when a model trained using data on one antibody is used for another antibody.Furthermore, for models that classify Voice of Customer (VOC), such as a model that classifies VOC into "product features," "support responses," and "other," domain shift can become a problem when a classification model trained using data on one product is used for another product.

[0069] [Evaluation of the model before its implementation] Before introducing a trained model 14 into an actual facility, the performance of the model 14 is often evaluated. Performance evaluation is necessary to determine whether or not to introduce the model, and for research and development of models or learning methods.

[0070] FIG. 6 is an explanatory diagram of a recommendation system implementation flow that includes a step of evaluating the performance of the trained model 14. In FIG. 6, a step of evaluating the performance of the model 14 is added as "Step 1.5" between Step 1 (the step of training the model 14) and Step 2 (the step of operating the model 14) described in FIG. 5. The rest of the configuration is the same as in FIG. 5. As shown in FIG. 6, in a typical recommendation system implementation flow, data collected at the facility where the system is to be implemented is often divided into training data and evaluation data. After the predictive performance of the model 14 is confirmed using the evaluation data, operation of the model 14 begins.

[0071] However, when constructing a domain generalization model 14, the training data and the evaluation data must be from different domains. Furthermore, in domain generalization, it is preferable to use data from multiple domains for the training data, and it is even more preferable to have as many domains as possible for training.

[0072] [About generalization] FIG. 7 is an explanatory diagram showing examples of training data and evaluation data used in machine learning. A data set obtained from a joint probability distribution Pd1(X,Y) of a certain domain d1 is divided into training data and evaluation data. The evaluation data of the same domain as the training data is called "first evaluation data" and is denoted as "evaluation data 1" in FIG. 7. In addition, a data set obtained from a joint probability distribution Pd2(X,Y) of a domain d2 different from domain d1 is prepared and used as evaluation data. The evaluation data of a domain different from the training data is called "second evaluation data" and is denoted as "evaluation data 2" in FIG. 7.

[0073] The model 14 is trained using the training data of the domain d1, and the performance of the trained model 14 is evaluated using the first evaluation data of the domain d1 and the second evaluation data of the domain d2.

[0074] Fig. 8 is a graph showing a schematic diagram of differences in model performance due to differences in data sets. If the performance of model 14 in the training data is performance A, the performance of model 14 in the first evaluation data is performance B, and the performance of model 14 in the second evaluation data is performance C, then as shown in Fig. 8, the relationship is usually performance A > performance B > performance C.

[0075] High generalization performance of Model 14 generally refers to high performance B or a small difference between performance A and B. In other words, the aim is to achieve high predictive performance even for untrained data without overfitting to the training data.

[0076] In the context of domain generalization in this specification, this refers to high performance C, or a small difference between performance B and performance C. In other words, the goal is to achieve consistently high performance even in domains different from the domain used for learning.

[0077] Even if it is not possible to use the behavioral history data of the facility where the system will be introduced during learning, if the data can be prepared before introduction, it is possible to train multiple models using data collected at a facility other than the facility where the system will be introduced, and then evaluate the performance of these multiple models before introduction using data collected at the facility where the system will be introduced, and based on the evaluation results, select the most appropriate model from among the multiple models and apply it to the facility where the system will be introduced. Figure 9 shows an example.

[0078] FIG. 9 is an explanatory diagram illustrating an example of a recommendation system implementation flow when the learning domain and the implementation domain are different. As shown in FIG. 9, multiple models may be trained using data collected at a facility other than the implementation facility. Here, an example is shown in which models M1, M2, and M3 are trained using datasets DS1, DS2, and DS3 collected at different facilities. For example, model M1 is trained using dataset DS1, model M2 is trained using dataset DS2, and model M3 is trained using dataset DS3. Note that the datasets used to train each of models M1, M2, and M3 may be a combination of multiple datasets collected at different facilities. For example, model M1 may be trained using a dataset that combines dataset DS1 and dataset DS2.

[0079] After training multiple models M1, M2, and M3 in this way, the performance of each model M1, M2, and M3 is evaluated using data Dtg collected at the facility where the models will be introduced. In Figure 9, the symbols "A," "B," and "C" shown below each model M1, M2, and M3 represent the evaluation results of each model. A rating of A indicates good predictive performance that meets the introduction criteria. A rating of B indicates performance that is inferior to A. A rating of C indicates performance that is even inferior to B and is not suitable for introduction.

[0080] For example, as shown in Figure 9, if the evaluation result of model M1 is "A," the evaluation result of model M2 is "B," and the evaluation result of model M3 is "C," model M1 will be selected as the optimal model for the facility where it will be introduced, and the recommendation system 10 that applies model M1 will be introduced.

[0081] [Explanation of the problem] In this embodiment, we consider cases where data on users' behavioral history regarding items at the facility where the system is being implemented cannot be obtained either during model learning or during evaluation before implementation, or where data is available but the amount of data is small and a sufficient amount of data cannot be obtained to evaluate the model.

[0082] Figure 10 is an explanatory diagram showing the issues that arise when there is no user behavior history at the facility where the system is installed. The data used for information recommendation is broadly divided into user behavior history and metadata such as user attributes and item attributes. The user behavior history is the objective variable, and the metadata such as user attributes and item attributes are explanatory variables.

[0083] Without data on user behavior history at the facility where the system is being implemented, it is impossible to evaluate the performance of the model. In other words, since the models used in information recommendation systems predict the objective variable based on explanatory variables, the predictive accuracy of the model cannot be evaluated without correct data for the objective variable. Without the ability to evaluate the performance of a model, it is difficult to select from multiple models the model that is most suitable for the facility where the system is being implemented. This issue also exists when there is a history of user behavior at the facility where the system is being implemented, but the amount of data is small and insufficient data is available for model evaluation.

[0084] In this embodiment, a means is provided that enables a facility to select a high-performance model from multiple models even when there is no behavioral history data at the facility where the system is installed, or when the amount of data required for model evaluation is insufficient. In the following description, the facility where the recommendation system 10 is installed is referred to as the "installation facility," and the facility that collects data for training a candidate model is referred to as the "learning facility." The installation facility corresponds to the target domain, and the learning facility corresponds to the learning domain.

[0085] [Outline of information processing method according to the first embodiment] 11 and 12 are explanatory diagrams showing an overview of the information processing method according to the first embodiment. Here, an example will be described in which a model M1 and a model M2 are prepared in advance as candidate models, but in reality, a larger number of models may be prepared.

[0086] Candidate model M1 is a predictive model trained using data collected at a learning facility FA1, which is different from the installation facility FAt. Similarly, another candidate model M2 is a predictive model trained using data collected at a learning facility FA2, which is different from the installation facility FAt and the learning facility FA1.

[0087] It is assumed that the user's behavior history at the facility FAt where the system is implemented is unavailable. "Unavailable" includes cases where no behavior history exists, where the data is inaccessible even if it exists, or where the amount of data is insufficient for model evaluation. On the other hand, it is assumed that a metadata dataset Dmt, such as user attributes and / or item attributes, exists for the facility FAt where the system is implemented, as shown in Figure 11.

[0088] In this case, the information processing device 100 according to this embodiment performs processing according to the following procedure (steps 1 to 3).

[0089] [Step 1] In step 1, the information processing device 100 extracts information indicating the characteristics of each of the learning facilities FA1 and FA2 and the introduction facility FAt. The information indicating the characteristics of the learning facilities FA1 and FA2 may be statistical values ​​or distributions extracted by statistical processing or the like from the data sets Dm1 and Dm2 of the metadata (explanatory variables) used in learning.

[0090] Facility characteristic information, such as statistical information extracted from the metadata of explanatory variables, is referred to as "metadata-derived facility characteristic information." When the explanatory variables are continuous values, the statistical information as metadata-derived facility characteristic information may be, for example, a statistical value such as the mean or standard deviation, or a combination of these. When the explanatory variables are discrete values, the statistical information as metadata-derived facility characteristic information may be, for example, a mode or a probability distribution, or a combination of these.

[0091] Furthermore, the information indicating the characteristics of the learning facilities FA1 and FA2 may be external information separate from the dataset used for learning. The external information may be, for example, information collected from the Internet by web crawling, or statistics or distributions extracted from the collected information. The external information may also be information input via a user interface based on available publicly available materials.

[0092] In other words, the information processing device 100 can obtain information about facilities from sources other than the data sets collected at each facility in two ways: the algorithm of the information processing device 100 or another system automatically collects and / or extracts the information by web crawling, etc., or an operator researches and / or inputs publicly known materials, etc.

[0093] Such external information outside the dataset is facility characteristic information that cannot be extracted from the metadata included in the dataset. Facility characteristic information that cannot be extracted from the metadata is called "facility-related information outside the metadata."

[0094] The facility characteristic information acquired by the information processing device 100 for each facility may include both facility characteristic information derived from metadata and facility-related information outside the metadata, or may include only one of these pieces of information. For the installation destination facility FAt shown in Fig. 11, a data set Dmt of explanatory variables (metadata) such as user attributes and item attributes is assumed to be prepared.

[0095] The information processing device 100 acquires, as facility characteristic information for the learning facility FA1, facility characteristic information ST1 derived from the metadata about the learning facility FA1 and facility-related information EI1 outside the metadata. Also, the information processing device 100 acquires, as facility characteristic information for the learning facility FA2, facility characteristic information ST2 derived from the metadata about the learning facility FA2 and facility-related information EI2 outside the metadata.

[0096] Similarly, the information processing device 100 acquires facility characteristic information STt derived from metadata such as statistical values ​​and distributions extracted from a metadata dataset Dmt collected from the facility FAt where the device is installed, and facility-related information EIt outside the metadata by web crawling, etc.

[0097] [Step 2] In step 2, the information processing device 100 evaluates the similarity between each of the learning facilities FA1, FA2 and the introduction facility FAt based on the acquired facility characteristic information of each facility. For example, the facility characteristic information of each facility is expressed as a multidimensional vector, and the similarity is evaluated by the Euclidean distance between the vectors in vector space.

[0098] [Step 3] In step 3, the information processing device 100 selects a model trained using data collected at a learning facility that has a high similarity to the introduction destination facility FAt, based on the similarity obtained in step 2. For example, if the evaluation result in step 2 shows that the similarity between the learning facility FA1 and the introduction destination facility FAt is low and the similarity between the learning facility FA2 and the introduction destination facility FAt is high, the information processing device 100 selects model M2 from the candidate models M1 and M2 as the model suitable for the introduction destination facility FAt.

[0099] In addition, when the information processing device 100 uses only facility characteristic information derived from metadata as facility characteristic information for each facility, it is not necessary to acquire facility-related information EI1, EI2, EIt outside the metadata in Fig. 11. In addition, when the information processing device 100 uses only facility-related information derived from metadata as facility characteristic information for each facility, it is not necessary to acquire facility characteristic information ST1, ST2, STt derived from metadata in Fig. 11.

[0100] [Relationship between the datasets and models for each learning facility] The basic idea behind the relationship between the datasets and models collected at each learning facility is to train one model from one dataset using only the dataset from that one learning facility, as shown in Figures 11 and 12. In this case, the dataset used to train the model becomes the main dataset, and the source of this dataset is the main learning facility.

[0101] However, it is also possible to train a model using a dataset that combines two or more datasets collected from multiple different learning facilities. For example, if a model is trained using a dataset that includes a dataset of 10,000 records of behavioral history collected at learning facility 1 and a dataset of 100 records of behavioral history collected at learning facility 2, the majority of the dataset used to train the model will be data from learning facility 1, and the proportion of data from learning facility 2 will be relatively small. In such a case, the dataset collected at learning facility 1 is considered to be the primary dataset, and learning facility 1 is considered to be the primary learning facility. A model trained under these conditions will have high predictive performance at learning facility 1, which is the primary learning facility. Therefore, for models trained using multiple datasets collected from multiple learning facilities, the adoption or rejection of the model may be determined based on the similarity between the characteristics of the facility where the model is introduced and the characteristics of the primary learning facility.

[0102] Furthermore, when there are multiple primary learning facilities for a single model, the adoption of the model may be determined based on a representative value, such as the average, maximum, or minimum, of the degree of similarity between the characteristics of the facility where the model is introduced and the characteristics of each primary learning facility. For example, if a model is trained using a dataset including a dataset of 5,000 records of behavioral history collected at learning facility 1 and a dataset of 5,000 records of behavioral history collected at learning facility 2, the proportion of data from learning facility 1 and the proportion of data from learning facility 2 in the entire dataset used to train the model will be equal, making it difficult to identify a single learning facility as the primary learning facility. In such a case, learning facility 1 and learning facility 2 may each be treated as the primary learning facility, and the similarity of facility characteristics for multiple combinations of learning facilities may be evaluated by, for example, calculating the average, maximum, or minimum of the similarity between the characteristics of the facility where the model is introduced and the characteristics of learning facility 1, and the similarity between the facility where the model is introduced and learning facility 2. The adoption of the model may be determined based on the evaluation results.

[0103] When a model is trained using a dataset that combines two or more datasets collected at multiple learning facilities, the dataset from the learning facility that has the largest proportion of data from each facility to the total number of data used for training may be the "primary dataset." On the other hand, a dataset from a learning facility whose proportion of data from each facility to the total number of data used for training is less than a certain threshold may be excluded from the "primary dataset." The threshold may be set appropriately within the scope of technical purposes as a criterion for determining whether a dataset's contribution to training can be considered relatively small; for example, it may be 10% or 5%. A dataset from a learning facility whose proportion of data from each facility to the total number of data used for training is equal to or greater than the threshold may be the "primary dataset."

[0104] In addition, when a model is trained using a dataset that combines two or more datasets from multiple datasets collected at multiple learning facilities, the similarity of the facility characteristics for a combination of multiple learning facilities can be evaluated by a method such as calculating a weighted average of the similarity between the characteristics of the facility where the model is introduced and the characteristics of each learning facility, using the proportion of the number of data from each learning facility to the total number of data used for training as a weight, and a decision can be made whether to adopt the model based on the evaluation results.

[0105] Overview of Information Processing Device FIG. 13 is a block diagram schematically illustrating an example of the hardware configuration of an information processing device 100 according to an embodiment. The information processing device 100 can be realized using computer hardware and software. The physical form of the information processing device 100 is not particularly limited, and it may be a server computer, a workstation, a personal computer, a tablet terminal, or the like. Here, an example in which the processing functions of the information processing device 100 are realized using one computer will be described, but the processing functions of the information processing device 100 may also be realized by a computer system configured using multiple computers.

[0106] The information processing device 100 includes a processor 102 , a non-transitory tangible computer-readable medium 104 , a communication interface 106 , an input / output interface 108 , and a bus 110 .

[0107] The processor 102 includes a CPU (Central Processing Unit). The processor 102 may also include a GPU (Graphics Processing Unit). The processor 102 is connected to a computer-readable medium 104, a communication interface 106, and an input / output interface 108 via a bus 110. The processor 102 reads various programs, data, etc. stored in the computer-readable medium 104, and executes various processes. The term "program" includes the concept of a program module and includes instructions equivalent to a program.

[0108] The computer-readable medium 104 is, for example, a storage device including a memory 112 serving as a primary storage device and a storage 114 serving as an auxiliary storage device. The storage 114 is configured using, for example, a hard disk drive (HDD), a solid state drive (SSD), an optical disk, a magneto-optical disk, or a semiconductor memory, or an appropriate combination of these. The storage 114 stores various programs, data, and the like.

[0109] The memory 112 is used as a working area for the processor 102, and as a storage unit that temporarily stores programs and various data read from the storage 114. When a program stored in the storage 114 is loaded into the memory 112 and the processor 102 executes the instructions of the program, the processor 102 functions as a means for performing various processes defined by the program.

[0110] The memory 112 stores various programs, such as a facility characteristic acquisition program 130, a similarity evaluation program 132, and a model selection program 134, which are executed by the processor 102, as well as various data.

[0111] The facility characteristic acquisition program 130 is a program that executes a process for acquiring information indicating the characteristics of the learning facility and the introducing facility. The facility characteristic acquisition program 130 may acquire information indicating the characteristics of the learning facility, for example, by statistically processing data included in a dataset collected at the learning facility. Furthermore, the facility characteristic acquisition program 130 may, for example, accept input of information indicating the characteristics of the facility via a user interface, or may include a web crawling program that automatically collects public information indicating the characteristics of the facility from the Internet.

[0112] The similarity evaluation program 132 is a program that executes a process to evaluate the similarity of the facility characteristics between the introducing facility and each learning facility based on the facility characteristic information of each facility. The model selection program 134 is a program that executes a process to select a model suitable for the introducing facility from multiple candidate models based on the similarity evaluation results.

[0113] The memory 112 includes a facility information storage unit 136 and a candidate model storage unit 138. The facility information storage unit 136 is a storage area for storing facility information including facility characteristic information of each facility acquired by the facility characteristic acquisition program 130. The facility information storage unit 136 may also include a storage area for storing metadata collected at the facility where the system is introduced.

[0114] The candidate model storage unit 138 is a storage area for storing a plurality of trained models that have been learned using the respective data sets of a plurality of learning facilities. The candidate model storage unit 138 may include a storage area for storing the data sets used to learn each model. The candidate model storage unit 138 may also include a storage area for storing facility characteristic information of each learning facility in association with (linked to) the model.

[0115] The communication interface 106 performs communication processing with an external device via a wired or wireless connection, and exchanges information with the external device. The information processing device 100 is connected to a communication line (not shown) via the communication interface 106. The communication line may be a local area network, a wide area network, or a combination of these. The communication interface 106 can function as a data acquisition unit that accepts input of various data such as data sets.

[0116] The information processing device 100 may include an input device 152 and a display device 154. The input device 152 and the display device 154 are connected to the bus 110 via the input / output interface 108. The input device 152 may be, for example, a keyboard, a mouse, a multi-touch panel, or another pointing device, or a voice input device, or an appropriate combination thereof. The display device 154 may be, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination thereof. Note that the input device 152 and the display device 154 may be integrally configured as a touch panel, or the information processing device 100, the input device 152, and the display device 154 may be integrally configured as a touch panel tablet terminal.

[0117] 14 is a functional block diagram showing the functional configuration of the information processing device 100. The information processing device 100 includes a data acquisition unit 220, a data storage unit 222, a facility characteristic acquisition unit 230, a similarity evaluation unit 240, and a model selection unit 244. The data acquisition unit 220 acquires various data such as metadata about the installation destination facility FAt. The data acquisition unit 220 may be configured to include a communication interface 106.

[0118] The data acquired via the data acquisition unit 220 is stored in the data storage unit 222. The data storage unit 222 includes an introduction destination metadata storage unit 224 and a candidate model storage unit 138. The introduction destination metadata storage unit 224 stores a metadata data set Dmt such as user attributes and / or item attributes of the introduction destination facility FAt.

[0119] The candidate model storage unit 138 stores multiple candidate models M1, M2,..., Mn. The candidate model storage unit 138 may also store datasets DS1, DS2,..., DSn used to train each model, in association with the model. Here, it is assumed that each of the datasets DS1, DS2,..., DSn has been collected from a different learning facility, and model Mk (k=1, 2,..., n) is a model trained using dataset DSk collected at learning facility k. Note that some learning facilities may have a contract stipulating that the dataset used to train the model be discarded after model training, and there may be cases where the dataset used to train the model is not stored.

[0120] The facility characteristic acquisition unit 230 acquires information indicating the characteristics of each of the multiple learning facilities k and the introduction facility FAt. The facility characteristic acquisition unit 230 includes a statistical information extraction unit 232 and a non-metadata facility information extraction unit 234. The statistical information extraction unit 232 performs statistical processing on the metadata included in the dataset DSk of each learning facility k, and extracts statistical information such as statistical values ​​and / or distributions.

[0121] The non-metadata facility information extraction unit 234 performs web crawling on the Internet to extract facility-related information outside the metadata about the target facility. The non-metadata facility information extraction unit 234 may also accept information input by an operator via a user interface and acquire facility-related information outside the metadata about the target facility.

[0122] The similarity evaluation unit 240 evaluates the similarity between each learning facility k and the introduction destination facility FAt based on the facility characteristic information of each learning facility k and the introduction destination facility FAt.

[0123] The model selection unit 244 selects a model suitable for the introduction facility FAt from among a plurality of models based on the similarity evaluated by the similarity evaluation unit 240.

[0124] [Flowchart example] Fig. 15 is a flowchart showing an example of the operation of the information processing device 100. It is assumed that a plurality of pre-trained models Mk (k = 1, 2, n) are prepared. When the flowchart of Fig. 15 starts, in step S111, the processor 102 acquires the characteristics of the training facility k from which the dataset DSk used in training each of the prepared models Mk has been collected.

[0125] In step S112, the processor 102 acquires the characteristics of the installation destination facility FAt. The order of the processes in steps S111 and S112 may be reversed.

[0126] In step S113, the processor 102 evaluates the degree of similarity between the introduction destination facility FAt and each learning facility k based on the characteristics of each facility acquired in steps S111 and S112.

[0127] In step S114, the processor 102 selects a model trained using a dataset collected at a facility highly similar to the installation destination facility FAt. The processor 102 may extract, as the optimal model, a model trained using the dataset of the facility with the highest similarity. Alternatively, if there are multiple facilities with similarities equal to or greater than a reference value (threshold), the processor 102 may extract, as models applicable to the installation destination, two or more models trained using datasets of facilities that satisfy the tolerance conditions for similarity. If multiple models applicable to the installation destination are extracted, the processor 102 may prioritize the models with the highest similarity in order of the similarity of the facility characteristics and present the application candidate models with the highest similarity, or may present the application candidate models together with the similarity evaluation results. Information on the one or more models selected by the processor 102 is output to the display device 154 or the like as the processing result of the model selection. After step S114, the processor 102 ends the flowchart of FIG. 14.

[0128] [Specific application examples] Here, an example of a recommendation system for a retail store will be described. The data used for learning includes behavioral history (purchase history) data for each of Store 1, Store 2, and Store 3, as well as user attributes (age) and item attributes (price). Each of Stores 1 to 3 is a learning facility and is an example of a "first facility" in this disclosure. The goal is to develop a recommendation system for Store 4, which is to be newly opened, but behavioral history data for Store 4 does not yet exist. On the other hand, the product lineup planned for sale at Store 4 has been decided, so a dataset of item attributes exists. In addition, store members are being recruited in advance of the opening, so a dataset of user attributes for Store 4 also exists. Store 4 is the facility where the system will be introduced and is an example of a "second facility" in this disclosure.

[0129] Under the above conditions, the information processing device 100 performs the following processing steps 1 to 4.

[0130] [Processing Step 1] Processor 102 extracts characteristics of each of stores 1 to 4 from the dataset of user attributes and item attributes. For example, the average ages of users extracted from the user attributes of each of stores 1 to 4 are 35, 45, 50, and 40 years old, respectively, starting from store 1 (see FIG. 16). Also, for example, the average prices of items extracted from the item attributes of each of stores 1 to 4 are 500 yen, 300 yen, 600 yen, and 400 yen, respectively, starting from store 1 (see FIG. 17). The characteristics of the stores extracted from the dataset may be statistical values, distributions, etc. extracted from metadata such as explanatory variables.

[0131] [Processing Step 2] Next, processor 102 extracts the characteristics of each of stores 1 to 4 not from the data set of each store, but from external information separate from the data set. Processor 102 acquires, as facility-related information outside the metadata, for example, the floor area of ​​each store and the average annual household income of the city, town, or village where each store is located (see FIG. 18). The floor area of ​​each of stores 1 to 4 is, for example, 1000 m2, in order from store 1. 2 , 1500m 2 , 500m 2 , 2000m 2 Furthermore, the average annual household income of the municipalities where each of Stores 1 to 4 is located is, for example, 6 million yen, 4 million yen, 7 million yen, and 5 million yen, respectively, starting from Store 1. Store characteristics extracted from external information separate from the dataset may include various data related to the characteristics of the stores themselves that are not included in the dataset. The floor area of ​​each store and the average annual household income of the municipalities where each store is located are each an example of "facility-related information other than metadata" in this disclosure.

[0132] [Processing Step 3] The characteristics of each store are expressed as a multidimensional vector from multiple types of numerical values ​​indicating the characteristics of each store obtained by processing step 1 and processing step 2. In the above example, processor 102 expresses the characteristics of each store as a four-dimensional vector of the average age of users, the average price of items, the floor area of ​​the store, and the average annual household income of the city or town where the store is located. Specifically, the characteristic vectors of Store 1 to Store 4 are (35,500,1000,600), (45,300,1500,400), (50,600,500,700), and (40,400,2000,500).

[0133] [Processing Step 4] Next, processor 102 calculates the similarity of the characteristics of each store. When evaluating the similarity of the characteristic vectors in the vector space representing the characteristics of the stores, in order to align the range of values ​​of each dimension of the vector, processor 102 calculates the mean value and standard deviation for each dimension, and standardizes the value by subtracting the mean value from the value of each dimension and dividing by the standard deviation.

[0134] Then, processor 102 uses the standardized facility characteristic vector of each store to calculate the Euclidean distance between the vectors of store 4 and each of stores 1 to 3. The Euclidean distance between vectors is an example of an index (evaluation value) for evaluating the similarity of facility characteristics.

[0135] For example, the Euclidean distance between the vectors of Store 4 and Store 1 is 2.05, the Euclidean distance between the vectors of Store 4 and Store 2 is 1.55, and the Euclidean distance between the vectors of Store 4 and Store 3 is 3.55. As a result, it is determined that Store 2 has the highest similarity.

[0136] [Processing Step 5] Based on the similarity evaluation results from processing step 4, processor 102 selects the model trained using the data set of store 2 with the highest similarity as the model suitable for store 4.

[0137] Due to the similarity of the store characteristics, it is expected that among stores 1 to 3, store 2 will have the most similar behavioral characteristics for users' items to store 4. Therefore, by introducing a model trained using the dataset of store 2 to store 4, it will be possible to achieve highly effective information recommendations even for store 4, which does not yet have a behavioral history.

[0138] [Example of extracting statistical information on user attributes] FIG. 16 is an explanatory diagram showing an example of extracting statistical information on user attributes as facility characteristics. Here, a specific example of a facility is a retail store such as a supermarket. As mentioned above, each of stores 1 to 3 is a learning facility, and store 4 is an installation facility. The same applies to FIGS. 17 and 18.

[0139] FIG. 16 shows an example of data on "age," which is one of the user attributes of users at each of stores (facilities) Store 1 to Store 4.

[0140] The processor 102 obtains the average age of users at each store from a data set on the user attributes of users at each store, such as that shown in Fig. 16. The average age is an example of statistical information, and is one piece of information indicating the characteristics of each store.

[0141] For example, the average age of users at Store 1 calculated from the user attribute data of Store 1 is 35. Similarly, the average age of users at Store 2 calculated from the user attribute data of Store 2 is 45, the average age of users at Store 3 is 50, and the average age of users at Store 4 is 40.

[0142] The processor 102 may further calculate a standard deviation for each store. The processor 102 may also calculate a histogram of the ages of users or a density distribution of the ages instead of or in addition to the average age.

[0143] [Example of extracting statistical information on item attributes] Fig. 17 is an explanatory diagram showing an example of extracting statistical information on item attributes as facility characteristics. Fig. 17 shows an example of data on "price," which is one of the item attributes of each store (facility) Store 1 to Store 4.

[0144] Processor 102 calculates the average price of items at each store from data on the prices of items (commodities in this case) sold at each store, as shown in Fig. 17. The average price is an example of statistical information, and is one piece of information that indicates the characteristics of each store.

[0145] For example, the average price of items at store 1 calculated from the item attribute data for store 1 is 500 yen. Similarly, the average price of items at store 2 calculated from the item attribute data for store 2 is 300 yen, the average price of items at store 3 is 600 yen, and the average price of items at store 4 is 400 yen.

[0146] The processor 102 may further calculate the standard deviation for each store. The processor 102 may also calculate a histogram of item prices or a price density distribution instead of or in addition to the average price.

[0147] [Example of extracting facility-related information outside of metadata] FIG. 18 is an explanatory diagram showing an example of extracting information about facilities by web crawling. FIG. 18 shows an example of extracting facility-related information other than metadata for each store (facility), Store 1 to Store 4, such as the floor area of ​​each store and the average annual household income of the city, town, or village where the store is located. The information processing device 100 crawls information on the Internet and acquires information about the floor area of ​​each store 1 to Store 4 and information about the average annual household income of the city, town, or village where the store is located. The store's address, size (scale), average annual household income of residents living near the store, etc. can be store characteristics related to user behavior at the store.

[0148] Note that the crawling is not limited to being performed by the information processing device 100, and an information processing device (not shown) other than the information processing device 100 may perform web crawling, and the information processing device 100 may acquire information extracted by the crawling.

[0149] While FIG. 18 shows an example of a retail store, the information to be extracted may vary depending on the type of facility being targeted. For example, if the facility is a medical facility such as a hospital, facility characteristic information outside the dataset may include the type of hospital, the size of the hospital, or the types of medical departments offered. The type of hospital may be, for example, a national hospital, a public hospital, a university hospital, a general hospital, or the like, based on the establishment of the hospital. The type of hospital may also be categorized by function, such as a designated function hospital, a regional medical support hospital, or others. The size of the hospital may be, for example, the number of hospital beds.

[0150] Second Embodiment Fig. 19 is an explanatory diagram showing an overview of an information processing method according to the second embodiment. In Fig. 19, elements common to Fig. 12 are given the same reference numerals, and duplicate explanations will be omitted. The prerequisites explained in Fig. 12 are also applicable to Fig. 19. In the second embodiment, step 0 is added before step 1 in Fig. 12, and steps 3 and 4 in Fig. 19 are included instead of step 3 in Fig. 12.

[0151] In the second embodiment, the information processing device 100 performs processing according to the following procedure (steps 0 to 4).

[0152] [Step 0] In step 0, the information processing device 100 or another machine learning device evaluates the performance of each model when it has learned the model. The machine learning device may be a computer system different from the information processing device 100. The evaluation data used to evaluate the predictive performance of each model may be data collected at the same facility as the dataset used for learning. The predictive performance of a model is quantified using an index (evaluation value) such as prediction accuracy. The information processing device 100 stores the evaluation value of the predictive performance of each model in association with the model. For example, the evaluation value indicating the predictive performance of model M1 is 0.5, and the evaluation value indicating the predictive performance of model M2 is 0.2.

[0153] [Step 1 and Step 2] The processing in steps 1 and 2 is the same as that in FIG.

[0154] [Step 3] In step 3, the information processing device 100 calculates a composite score based on the prediction performance of the model and the similarity between facilities. Here, an example is shown in which the composite score is calculated by taking the product of the evaluation value of the prediction performance and the similarity, but an average value may be used instead of the product.

[0155] If the similarity between learning facility FA1 and the facility where it is introduced FAt is 0.6, and the similarity between learning facility FA2 and the facility where it is introduced FAt is 0.8, the composite score based on the predictive performance of model M1 and the similarity to learning facility FA1 is calculated to be 0.3, and the composite score based on the predictive performance of model M2 and the similarity to learning facility FA2 is calculated to be 0.16.

[0156] [Step 4] In step 4, the information processing device 100 selects a model with a high composite score based on the composite score obtained in step 3. In the example of Fig. 19, the information processing device 100 selects model M1 with a high composite score from models M1 and M2 as a model suitable for the facility FAt where the model is to be introduced.

[0157] In this way, the model may be selected taking into consideration not only the similarity between facilities but also the predictive performance of each model at each learning facility.

[0158] [Modification] In the second embodiment, an example has been described in which a composite score taking into account the predictive performance of each model is used. However, instead of or in combination with the predictive performance of the model, a composite score taking into account some kind of compatibility evaluation value that evaluates the compatibility of the model at the facility FAt where the model is introduced may be used. The compatibility evaluation value may be, for example, the result of a questionnaire given to users at the facility FAt where the model is introduced. The compatibility evaluation value based on the result of the questionnaire is an example of "compatibility evaluation information" in the present disclosure.

[0159] [Data sets collected at the learning facility are used after model training. inability Example of what to do if Due to contractual provisions such as the destruction of confidential information, data sets collected at learning facilities and data such as characteristics extracted from those data sets are required to be destroyed after learning, and it is anticipated that such data cannot be retained for long.

[0160] In such cases, when evaluating the similarity between facilities, it becomes impossible to use the metadata contained in the dataset collected at the learning facility, or the statistical values ​​of the metadata, etc. An example of how to deal with such a case will be explained using Figure 20.

[0161] FIG. 20 is an explanatory diagram that schematically illustrates the characteristics of each facility in a vector space that represents the characteristics of the facility. The facility characteristic LD indicated by the dashed line in FIG. 20 is data indicating the characteristics of a learning facility that will be discarded after learning and become unavailable. A learning facility with the facility characteristic LD is referred to as a "non-retained learning facility." In this case, the information processing device 100 may use multiple facility characteristics Dum1, Dum2, and Dum3 whose similarity with the facility characteristic LD of the non-retained learning facility has been evaluated, and evaluate the similarity between the facility characteristic LD of the non-retained learning facility and the facility characteristic TG of the introduction destination facility based on the similarity between the facility characteristic TG of the introduction destination facility and each of the multiple facility characteristics Dum1, Dum2, and Dum3. Each of the multiple facility characteristics Dum1, Dum2, and Dum3 may be dummy data.

[0162] The information processing device 100 or another information processing device may generate these multiple facility characteristics Dum1, Dum2, and Dum3 based on the facility characteristic LD of the non-held learning facility. The information processing device 100 can store data on the multiple facility characteristics Dum1, Dum2, and Dum3 in association with the similarity to the facility characteristic LD, instead of the facility characteristic LD of the non-held learning facility. This makes it possible to evaluate the similarity to the facility characteristic TG of the facility where the facility is being introduced, even if the facility characteristic LD of the non-held learning facility is not available. Each of the multiple facility characteristics Dum1, Dum2, and Dum3 is an example of a "third facility characteristic" in this disclosure.

[0163] [Explanation of learning methods] Next, a model learning method will be described. Here, the case of matrix factorization, which is often used in information recommendation, will be described as an example. Note that, although the following description shows an example in which the information processing device 100 executes the learning process, the device that executes the learning process may be a computer system separate from the information processing device 100.

[0164] When there is a data set containing the behavioral histories of multiple users with respect to multiple items at a learning facility, processor 102 first learns the dependencies between variables based on this data. More specifically, processor 102 uses a model in which users and items are represented by vectors, and the sum of their dot products represents the behavior probability, and updates the parameters of the model to minimize the error in behavior prediction.

[0165] A vector representation of a user can be expressed, for example, by adding up the vector representations of each of the user's attributes. The same is true for the vector representation of an item. A model that has learned the dependencies between variables is equivalent to expressing the joint probability distribution P(X,Y) between the objective variable Y and each explanatory variable X in a given behavioral history dataset.

[0166] Fig. 21 is an example of a directed acyclic graph (DAG) that represents the dependency relationships between variables in a joint probability distribution P(X, Y). Fig. 21 shows an example in which four variables, user attribute 1, user attribute 2, item attribute 1, and item attribute 2, are used as explanatory variables X. The relationship between each of these explanatory variables X and the objective variable Y, which is the user's behavior toward an item, can be represented, for example, by a graph like the one shown in Fig. 21.

[0167] During learning, a vector representation of the joint probability distribution P(X,Y) is obtained based on the dependency relationships between variables, such as the DAG shown in Fig. 21. The graph in Fig. 21 shows that the target variable, a user's behavior with respect to an item, depends on the user's behavioral characteristics and the item's characteristics, and that the user's behavioral characteristics depend on user attribute 1 and user attribute 2, and that the item's characteristics depend on item attribute 1 and item attribute 2.

[0168] As shown in Figure 21, the combination of user attribute 1 and user attribute 2 defines the user's behavioral characteristics. Also, the combination of item attribute 1 and item attribute 2 defines the item's characteristics. The user's behavior toward an item is defined by the combination of the user's behavioral characteristics and the item's characteristics.

[0169] Generally, the relationship P(X, Y) = P(X) × P(Y|X) holds, and when the graph in Figure 21 is applied to this equation, it is expressed as follows. P(X) = P(user attribute 1, user attribute 2, item attribute 1, item attribute 2) P(Y|X) = P(user's behavior toward an item | user attribute 1, user attribute 2, item attribute 1, item attribute 2) P(X,Y) = P(user attribute 1, user attribute 2, item attribute 1, item attribute 2) × P(user's behavior toward the item | user attribute 1, user attribute 2, item attribute 1, item attribute 2)

[0170] Furthermore, the graph shown in FIG. 21 shows that it can be decomposed into elements as follows: P(Y|X) = P(user's behavior toward an item | user's behavioral characteristics, item's characteristics) × P(user's behavioral characteristics | user attribute 1, user attribute 2) × P(item's behavioral characteristics | item attribute 1, item attribute 2)

[0171] [Example of probability expression of conditional probability distribution P(Y|X)] For example, the probability that a user will view an item (Y=1) is expressed as a sigmoid function of the inner product of the user characteristic vector and the item characteristic vector. This type of expression is called matrix factorization. The reason for using the sigmoid function is that the value of the sigmoid function ranges from 0 to 1, and the value of the function directly corresponds to the probability. The model expression is not limited to the sigmoid function, and other functions may also be used.

[0172] A specific example of the probability expression of P(Y|X) is shown in Figure 22. Formula F22A shown in the upper part of Figure 22 is an example of a formula that expresses the user characteristic vector θu and the item characteristic vector φi as five-dimensional vectors using matrix decomposition, and expresses the sigmoid function σ(θu·φi) of their inner product (θu·φi) as the conditional probability P(Y=1|user, item).

[0173] u is an index value that distinguishes users. i is an index value that distinguishes items. Note that the dimension of the vector is not limited to five dimensions, and can be set to an appropriate number of dimensions as a hyperparameter of the model.

[0174] The user characteristic vector θu is expressed as the sum of the user's attribute vectors. For example, as shown in formula F22B in the middle of FIG. 22, the user characteristic vector θu is expressed as the sum of the user attribute 1 vector and the user attribute 2 vector. Furthermore, the item characteristic vector φi is expressed as the sum of the item's attribute vectors. For example, as shown in formula F22C in the bottom of FIG. 22, the item characteristic vector φi is expressed as the sum of the item attribute 1 vector and the item attribute 2 vector.

[0175] Figure 23 is an explanatory diagram showing the relationship between formula F22A, which expresses the conditional probability of a user's behavior toward an item (Y=1) for a combination of the user's behavioral characteristics and the item's characteristics, and a DAG, which expresses the dependency relationship between variables in the joint probability distribution P(X, Y). As shown in Figure 23, formula F22A expresses the conditional probability of the part enclosed by the dashed-line frame FR1 in the DAG shown in Figure 23.

[0176] Figure 24 is an explanatory diagram showing the relationship between a user's behavioral characteristics defined by a combination of user attribute 1 and user attribute 2, an item's characteristics defined by a combination of item attribute 1 and item attribute 2, and a DAG that expresses the dependency relationships between variables. As shown in Figure 24, formula F22B represents the relationship of the part surrounded by a frame FR2 indicated by a dashed line in the DAG shown in Figure 24. Furthermore, formula F22C represents the relationship of the part surrounded by a frame FR3 indicated by a dashed line in the DAG shown in Figure 24.

[0177] The value of each vector shown in FIG. 23 is determined by learning from data (learning data) included in the dataset of user behavior history for a given domain.

[0178] For example, the vector values ​​are updated using, for example, stochastic gradient descent (SGD) so that P(Y=1|user, item) is large for pairs of users and items who viewed the item, and P(Y=1|user, item) is small for pairs of users and items who did not view the item.

[0179] In the case of the joint probability distribution P(X, Y) shown in FIGS. 23 and 24, the parameters to be learned from the data are as follows: User characteristic vector: θu Item characteristic vector: φi User attribute 1 vector: Vk_u^1 User attribute 2 vector: Vk_u^2 Item attribute 1 vector: Vk_i^1 Item attribute 2 vector: Vk_i^2 However, these parameters satisfy the following relationship: θu=Vk_u^1+Vk_u^2 φi=Vk_i^1+Vk_i^2 k is an index value that distinguishes attributes. For example, if user attribute 1 has 10 types of department, user attribute 2 has 6 levels of age, item attribute 1 has 20 types of product type, and item attribute 2 has 5 types of product price, the number of attribute types is 10 + 6 + 20 + 5 = 41, so k can take on values ​​from 1 to 41. For example, k=1 corresponds to user attribute 1's sales department, and the index value of user attribute 1 for user u is expressed as k_u^1.

[0180] The values ​​of each vector, user attribute 1 vector Vk_u^1, user attribute 2 vector Vk_u^2, item attribute 1 vector Vk_i^1, and item attribute 2 vector Vk_i^2, are determined by learning from the learning data.

[0181] As a loss function for learning, for example, logloss shown in the following equation (1) is used.

[0182] L=-{Y ui log σ(θu·φi)+(1-Y ui ) log (1-σ(θu·φi))}(1) Y if user u views item i ui = 1, and the larger the predicted probability σ(θu φi), the smaller the loss L. Conversely, if user u does not view item i, Y ui =0, and the smaller σ(θu·φi) is, the smaller the loss L is.

[0183] The parameters of the vector representation are learned so that the above loss L is reduced. For example, when optimizing using stochastic gradient descent, one record is randomly selected from all training data (if it does not depend on the context, one ui pair is selected from all ui pairs), the partial derivative (gradient) of each parameter of the loss function is calculated for the selected record, and the parameters are changed in the direction that reduces the loss L in proportion to the magnitude of the gradient.

[0184] For example, the parameters of the user attribute 1 vector (Vk_u^1) are updated according to the following equation (2).

[0185]

number

[0186] Generally, among a large number of items, there are far more items with Y=0 than items with Y=1, so when saving behavioral history data as shown in Figure 19 as a table, only Y=1 is retained, and pairs of user u and item i that are not included in the behavioral history data are learned as Y=0. In other words, if only positive example data is saved, negative examples can be easily generated as those not included in the positive example data.

[0187] [Model representation] The means of expressing the joint probability distribution of explanatory variable X and target variable Y are not limited to matrix factorization. For example, logistic regression or Naive Bayes may be applied instead of matrix factorization. Any predictive model can also be used to express the joint probability distribution by calibrating it so that the output score is close to the probability P(Y|X). For example, SVM (Support Vector Machine), GDBT (Gradient Boosting Decision Tree), and neural network models of any architecture can also be used.

[0188] [Regarding programs that operate computers] A program that causes a computer to realize some or all of the processing functions of the information processing device 100 can be recorded on a computer-readable medium such as an optical disk, a magnetic disk, a semiconductor memory, or other tangible non-transitory information storage medium, and the program can be provided through this information storage medium.

[0189] In addition, instead of providing the program by storing it on such a tangible, non-transitory computer-readable medium, it is also possible to provide the program signal as a download service using a telecommunications line such as the Internet.

[0190] Furthermore, some or all of the processing functions of the information processing device 100 may be realized by cloud computing, and may also be provided as SaaS (Software as a Service).

[0191] [Hardware configuration of each processing unit] The hardware structure of the processing units in the information processing device 100 that perform various processes, such as the data acquisition unit 220, facility characteristic acquisition unit 230, statistical information extraction unit 232, non-metadata facility information extraction unit 234, similarity evaluation unit 240, and model selection unit 244, is, for example, various processors as shown below.

[0192] Various types of processors include CPUs, which are general-purpose processors that execute programs and function as various processing units, GPUs, programmable logic devices (PLDs), such as FPGAs (Field Programmable Gate Arrays), which are processors whose circuit configuration can be changed after manufacture, and dedicated electrical circuits, such as ASICs (Application Specific Integrated Circuits), which are processors with circuit configurations designed specifically to execute specific processes.

[0193] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types. For example, a single processing unit may be configured with multiple FPGAs, or a combination of a CPU and an FPGA, or a combination of a CPU and a GPU. Alternatively, multiple processing units may be configured with a single processor. A first example of multiple processing units configured with a single processor is a configuration in which one or more CPUs and software are combined to form a single processor, as typified by client or server computers, and this processor functions as multiple processing units. A second example is a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0194] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0195] Advantages of the embodiment According to each of the above-described embodiments, even if data on the user's item behavior history at a facility where the model is being introduced, which is different from the facility where the dataset used to train the model was collected, cannot be used to evaluate the model's performance, a model suitable for the facility where the model is being introduced can be selected from multiple models based on the similarity of the facility's characteristics.

[0196] According to each embodiment, when the domain (learning domain) of a facility or the like where data used for model learning is collected is different from the domain (destination domain) of a facility or the like where the model is introduced, it becomes possible to provide a recommended item list that is robust to domain shifts.

[0197] [Other application examples] In the above-described embodiment, the example of user purchasing behavior at a retail store has been described, but the scope of application of the present disclosure is not limited to this example. For example, the technology of the present disclosure can be applied to a model that predicts user behavior regarding various items, regardless of the purpose, such as viewing documents at a company, viewing medical images and various documents at a medical facility such as a hospital, or watching content such as videos on a content providing site.

[0198] 〔others〕 The present disclosure is not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the technical idea of ​​the present disclosure. [Explanation of symbols]

[0199] 10 Recommendation Systems 12 Predictive Models 14 models 100 Information processing device 102 processors 104 Computer-readable medium 106 Communication Interface 108 Input / Output Interface 110 Bus 112 memory 114 Storage 130 Facility Characteristics Acquisition Program 132 Similarity Evaluation Program 134 Model Selection Program 136 Facility Information Storage Department 138 Candidate Model Storage Unit 152 Input Device 154 Display device 220 Data Acquisition Unit 222 Data Storage Unit 224 Destination Metadata Storage 230 Facility Characteristics Acquisition Department 232 Statistical information extraction section 234 Metadata non-facility information extraction unit 240 Similarity Evaluation Unit 244 Model Selection Section DS1 dataset DS2 dataset DS3 dataset Dtg Data Dm1 dataset Dm2 dataset Dmt dataset Dum1 Facility Characteristics Dum2 facility characteristics Dum3 Facility Characteristics LD Facility Characteristics TG Facility Characteristics EI1 Facility Information EI2 Facility Information EIT Facility Information ST1 Facility Characteristics Information ST2 Facility Characteristics Information STt Facility Characteristics Information TG Facility Characteristics F22A type F22B formula F22C formula FR1 frame FR2 frame FR3 frame IT1 Item IT2 Items IT3 items M1 model M2 model Mn model S111 to S114: Steps of processing performed by the information processing device

Claims

1. 1. An information processing method executed by one or more processors, comprising: a plurality of models trained using one or more of datasets including behavioral histories of users with respect to items collected at a plurality of first facilities different from each other are prepared; the one or more processors: acquiring a characteristic of each of a second facility different from the plurality of first facilities and the plurality of first facilities; Evaluating the degree of similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the datasets used to train each of the models were collected; selecting a model suitable for the second facility from the plurality of models based on the similarity; An information processing method including:

2. the one or more processors: extracting statistical information of a dataset of metadata that are explanatory variables used to train the model; the characteristics include the statistical information; The information processing method according to claim 1 .

3. the metadata includes at least one of user attributes and item attributes; The information processing method according to claim 2 .

4. the one or more processors: obtaining facility-related information other than metadata included in the dataset used to train the model; the characteristics include the facility-related information; The information processing method according to claim 1 .

5. the facility-related information is extracted by web crawling; The information processing method according to claim 4.

6. the one or more processors: accepting the facility-related information via a user interface; 6. The information processing method according to claim 4 or 5.

7. the one or more processors: acquiring an evaluation value of prediction performance at the first facility where the dataset used in training each of the plurality of models was collected; selecting a model suitable for the second facility from the plurality of models based on the similarity and the evaluation value of the prediction performance; 5. The information processing method according to claim 1.

8. the one or more processors: Acquire compatibility evaluation information indicating an evaluation of compatibility of the model with the second facility separately from the similarity; selecting a model suitable for the second facility from among the plurality of models based on the similarity and the compatibility evaluation information; 5. The information processing method according to claim 1.

9. The compatibility assessment information includes the results of a questionnaire survey conducted on users of the second facility. The information processing method according to claim 8.

10. the one or more processors: using characteristics of a plurality of third facilities whose similarities with the characteristics of the first facility have been evaluated; evaluating a similarity between a characteristic of the second facility and a characteristic of the first facility based on a similarity between a characteristic of the second facility and a characteristic of the plurality of third facilities; The information processing method according to claim 1 .

11. the one or more processors: characteristics of the plurality of third facilities; storing in a storage device the similarities between the characteristics of the first facility and the characteristics of the plurality of third facilities, The information processing method according to claim 10.

12. The model is a predictive model used in a recommendation system that recommends items to users. The information processing method according to claim 1 .

13. the one or more processors: storing the plurality of models in a storage device; The information processing method according to claim 1 .

14. the one or more processors: and storing in the storage device, characteristics of the first facility from which the data sets used to train each of the models were collected, in association with the models. The information processing method according to claim 13.

15. one or more processors; and one or more storage devices in which instructions to be executed by the one or more processors are stored, a plurality of models trained using one or more of datasets including behavioral histories of users with respect to items collected at each of a plurality of first facilities different from one another are stored in the storage device; the one or more processors: acquiring characteristics of a second facility different from the plurality of first facilities and the plurality of first facilities; Evaluating the degree of similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the main datasets used to train each of the models were collected; selecting a model suitable for the second facility from the plurality of models based on the similarity; Information processing device.

16. On the computer, a function of storing a plurality of models trained using one or more of datasets including behavioral histories of users with respect to items collected at each of a plurality of first facilities different from each other; a function of acquiring a characteristic of a second facility different from the plurality of first facilities and a characteristic of each of the plurality of first facilities; a function of evaluating the degree of similarity between the acquired characteristics of the second facility and the characteristics of the first facility from which the datasets used to train each of the models were collected; a function of selecting a model suitable for the second facility from the plurality of models based on the similarity; A program to make this happen.