Information Processing Apparatus, Information Processing Method, and Program
The information processing apparatus and method enable the extraction of similar data across different data sets by converting feature amounts based on attribute relationships, addressing the limitations of existing technologies in data analysis by providing interpretability for the extraction process.
Patent Information
- Application Number
- JP2023527179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-06-07
AI Technical Summary
Existing technologies are unable to extract similar data across different data sets, and they lack the capability to interpret and explain the reasons behind the extraction, which is crucial for data analysis.
An information processing apparatus and method that includes a first data set with a first attribute expressed in natural language and a second data set with a second attribute expressed in natural language. The apparatus calculates a second feature amount for the second data set and converts it into a first feature amount based on the relationship between the first and second attributes, enabling the extraction of similar data across different data sets and providing interpretability for the extraction process.
The solution allows for the effective extraction of similar data across different data sets, providing interpretability and explainability for the extraction process, which is essential for data analysis and user understanding.
Smart Images

Figure 0007697509000013 
Figure 0007697509000014 
Figure 0007697509000015
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and the like that can be used for data analysis.
Background Art
[0002] In recent years, it has become possible to use data from a variety of data sources, and technologies for analyzing and using such data from a variety of data sources have also been developed. For example, Patent Document 1 below discloses a method for extracting a plurality of similar users having preferences similar to those of a user for whom information is to be provided, and enabling the user for whom information is to be provided to select other users having the same hobbies and preferences as himself / herself as additional similar users.
[0003] Specifically, in the information providing method of Patent Document 1, based on the evaluation values of users for restaurants and the like registered in a database, a plurality of similar users having preferences similar to those of the user are extracted, and information on restaurants and the like highly evaluated by the similar users is provided to the user. Further, in the information providing method of Patent Document 1, information on users who were not extracted as similar users is also provided, and the user is allowed to select a user to be added to the similar users.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] The technology of Patent Document 1 extracts similar users based on information registered in one database. Therefore, in this technology, when there are a plurality of data sets, it is impossible to extract data similar to the data included in a certain data set from other data sets, that is, to extract similar data across different data sets.
[0006] In recent years, in data analysis, interpretability of analysis results is often required. For example, when automatically extracting similar data, it is desirable to be able to interpret or explain the reason or basis for the extraction of the data.
[0007] An object of the present invention is to provide an information processing apparatus and the like that enable extraction of similar data across different data sets and also enable interpretation and explanation of the reason for the extraction.
Means for Solving the Problems
[0008] An information processing apparatus according to an aspect of the present invention includes: a first data set including at least one data and associated with a first attribute expressed in natural language; among a second data set including a plurality of data and associated with a second attribute expressed in natural language, feature amount calculation means for calculating a second feature amount regarding the second attribute for the data included in the second data set; and conversion means for converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0009] An information processing method according to one aspect of the present invention includes: at least one processor calculating a second feature amount regarding the second attribute for data included in the second data set out of a first data set including at least one data and associated with a first attribute expressed in natural language and a second data set including a plurality of data and associated with a second attribute expressed in natural language; and converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0010] A program according to one aspect of the present invention causes a computer to function as a feature amount calculation means for calculating a second feature amount regarding the second attribute for data included in the second data set out of a first data set including at least one data and associated with a first attribute expressed in natural language and a second data set including a plurality of data and associated with a second attribute expressed in natural language, and a conversion means for converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
Advantages of the Invention
[0011] According to one aspect of the present invention, it is possible to extract similar data across different data sets and to interpret and explain the reasons for the extraction.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0013] 〔Exemplary Embodiment 1〕 The first exemplary embodiment of the present invention will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of the exemplary embodiments described later.
[0014] (Configuration of Information Processing Apparatus) The configuration of the information processing apparatus 1 according to this exemplary embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1. As shown in the drawing, the information processing apparatus 1 includes a feature amount calculation unit (feature amount calculation means) 11 and a conversion unit (conversion means) 12.
[0015] The feature quantity calculation unit 11 calculates a second feature quantity regarding the second attribute for the data included in the second data set out of the first data set and the second data set. Note that the first data set includes a plurality of data, and a first attribute expressed in natural language is associated with the first data set. Also, the second data set includes at least one data, and a second attribute expressed in natural language is associated with the second data set.
[0016] The conversion unit 12 converts the second feature quantity calculated by the feature quantity calculation unit 11 into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0017] As described above, in the information processing apparatus 1 according to this exemplary embodiment, among the first data set including a plurality of data and associated with a first attribute expressed in natural language, and the second data set including at least one data and associated with a second attribute expressed in natural language, for the data included in the second data set, a feature quantity calculation unit 11 that calculates a second feature quantity regarding the second attribute, and a conversion unit 12 that converts the second feature quantity into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute are provided.
[0018] That is, according to the above configuration, for both the data included in the first data set and the data included in the second data set, with which different attributes, i.e., the first attribute and the second attribute, are associated, it becomes possible to calculate feature quantities using a common measure, i.e., the first attribute. And thereby, it also becomes possible to extract data having similar feature quantities to the data included in the second data set from among the plurality of data included in the first data set.
[0019] Also, since the feature quantity serving as the basis for the above extraction is a feature quantity regarding the first attribute associated with the first data set, it is possible to explain or interpret the reason for the extraction using the first attribute expressed in natural language.
[0020] As described above, according to the information processing apparatus 1 according to the present exemplary embodiment, it is possible to extract similar data across different data sets, namely, the first data set and the second data set, and it is possible to explain and interpret the reason for the extraction.
[0021] (Program) The functions of the above-described information processing apparatus 1 can also be realized by a program. The program according to the present exemplary embodiment causes a computer to function as a feature amount calculation unit that calculates a second feature amount regarding the second attribute for data included in the second data set among a first data set including a plurality of data and associated with a first attribute expressed in a natural language and a second data set including at least one data and associated with a second attribute expressed in a natural language, and a conversion unit that converts the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute. Therefore, according to the program according to the present exemplary embodiment, it is possible to extract similar data across different data sets and to interpret and explain the reason for the extraction.
[0022] (Flow of information processing method) The flow of the information processing method according to the present exemplary embodiment will be described with reference to FIG. 2. FIG. 2 is a flowchart showing the flow of the information processing method. Note that the execution subject of each step in this information processing method may be a processor included in the information processing apparatus 1, or may be a processor included in another apparatus, or may be processors provided in different apparatuses for each step.
[0023] In S11, at least one processor calculates a second feature quantity regarding the second attribute for data included in the second data set out of a first data set including a plurality of data and associated with a first attribute expressed in natural language and a second data set including at least one data and associated with a second attribute expressed in natural language.
[0024] In S12, at least one processor converts the second feature quantity into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0025] As described above, in the information processing method according to this exemplary embodiment, at least one processor includes calculating a second feature quantity regarding the second attribute for data included in the second data set out of a first data set including a plurality of data and associated with a first attribute expressed in natural language and a second data set including at least one data and associated with a second attribute expressed in natural language, and converting the second feature quantity into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute. Thus, according to the information processing method according to this exemplary embodiment, it is possible to extract similar data across different data sets and to interpret and explain the reason for the extraction.
[0026] 〔Exemplary Embodiment 2〕 The second exemplary embodiment of the present invention will be described in detail with reference to the drawings. Note that components having the same functions as those described in Exemplary Embodiment 1 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0027] (Configuration of Information Processing Apparatus) The configuration of the information processing apparatus 2 according to the present exemplary embodiment will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the configuration of the information processing apparatus 2. As shown in the figure, the information processing apparatus 2 includes a control unit 20 that comprehensively controls each part of the information processing apparatus 2, and a storage unit 21 that stores various data used by the information processing apparatus 2. Further, the information processing apparatus 2 includes an input unit 22 that receives input of various data to the information processing apparatus 2, and an output unit 23 for the information processing apparatus 2 to output various data.
[0028] Further, the control unit 20 includes a feature amount calculation unit (feature amount calculation means) 201, a conversion rule generation unit 202, a conversion unit (conversion means) 203, and a similar data extraction unit 204. Further, the storage unit 21 stores a data set 211, a conversion rule 212, and result data 213.
[0029] The feature amount calculation unit 201 calculates a feature amount related to the attribute for the data included in the data set associated with the attribute expressed in natural language. More specifically, the data set 211 stored in the storage unit 21 includes a first data set and a second data set, and the feature amount calculation unit 201 calculates a feature amount for the data included in each of these data sets.
[0030] The first data set may be, for example, a collection of data related to one company, one corporation, one group, or one user. Further, the data included in the first data set may be, for example, the number of sales or the sales amount of the products handled by the above-mentioned company or the like. The same applies to the second data set.
[0031] The above attributes are related to the data included in the above dataset and may be expressed in natural language, that is, anything that can be recognized by people in terms of content. For example, classifications such as the major classification, medium classification, and minor classification of data may be used as the attributes of the data, or if the data is a product or service, the product name or service name may be used as the attributes of the data. Also, for example, words obtained by morphological analysis of the description text of a product or service may be used as attributes, or the aggregated product name or review text may be used as attributes.
[0032] Regarding data related to users, the occupation, nature, characteristics, etc. of the user may be used as the attributes of the data. For example, attributes such as salaryman, brand-oriented, regular purchaser, etc. may be used. Such attributes can also be identified, for example, from the results of questionnaires for users.
[0033] Recently, a service called data enrichment that adds additional information related to the data to the data in the dataset has become available. The additional information added by such a service can also be used as attributes. The above-mentioned user attributes may also be identified by using external services in the same way as data enrichment.
[0034] Based on the relationship between two attributes, the conversion rule generation unit 202 generates a conversion rule for converting the feature quantity related to one attribute into the feature quantity related to the other attribute. The conversion rule generated by the conversion rule generation unit 202 is stored in the storage unit 21 as the conversion rule 212.
[0035] Based on the relationship between two attributes, the conversion unit 203 converts the feature quantity related to one attribute into the feature quantity related to the other attribute. Specifically, the conversion unit 203 performs the above conversion using the above-mentioned conversion rule 212.
[0036] The similar data extraction unit 204 extracts similar data similar to the data included in the second data set from among the plurality of data included in the first data set. Although details will be described later, the extraction of similar data is performed based on the first feature amount obtained by the conversion by the conversion unit 203 and the feature amount regarding the first attribute for the plurality of data included in the first data set. The similar data extracted by the similar data extraction unit 204 is stored in the storage unit 21 as the result data 213.
[0037] The similar data extraction unit 204 is not an essential configuration of the information processing apparatus 2. However, by providing the similar data extraction unit 204, there is an advantage that data similar to the data included in the second data set can be automatically extracted from among the plurality of data included in the first data set. For this reason, it is preferable that the information processing apparatus 2 includes the similar data extraction unit 204.
[0038] (Specific example of feature amount calculation) FIG. 4 is a diagram showing a specific example of feature amount calculation by the feature amount calculation unit 201. More specifically, FIG. 4 shows an example of calculating feature amounts from a source data set (second data set) D s and a target data set (first data set) D t .
[0039] The source data set D s is a data set in a table format in which a user ID, an item ID, and an amount used are associated. This source data set D s shows the usage history of a credit card. That is, the source data set D s shows the amount used of the credit card at the store specified by the item ID for the user specified by the user ID.
[0040] Also, the related data M s is data associated with the source data set D s , and the source data set D sIt includes natural language indicating the attributes of the data contained therein. Specifically, related data M s indicates, in natural language, the classification of the stores specified by each item ID contained in the source dataset D s and also indicates whether or not it corresponds to that classification by a numerical value of 0 or 1. For example, in related data M s for the store with item ID "i01_s", the attribute values of "retail" and "motorcycle shop" are "1", and the other attribute values are "0". This indicates that the store of "i01_s" is a retail store and a motorcycle shop, and not a toy store, hotel, or hospital.
[0041] On the other hand, one target dataset D t is a dataset in a table format in which user IDs, item IDs, and the number of purchases are associated. This target dataset D t shows the purchase history on the mail-order site. That is, the target dataset D t shows the number of purchases of the products specified by the item ID for the users specified by the user ID.
[0042] Also, related data M t is data associated with the target dataset D t and includes natural language indicating the attributes of the data contained in the target dataset D t . Related data M t shows, in natural language, the classification of the products specified by each item ID contained in the target dataset D t and also indicates whether or not it corresponds to that classification by a numerical value of 0 or 1. For example, in related data M t it is shown that the product with item ID "i01_t" is auto parts and vehicle parts and not a map or a toy. Note that hereinafter, the set of attribute names contained in related data M s is called the attribute set A s . Similarly, the set of attribute names contained in related data M t is called the attribute set A t .
[0043] The feature quantity calculation unit 201 calculates feature quantities regarding the attributes shown in the related data M, that is, the classification of the credit card usage stores, for the data included in the source data set D s Specifically, the feature quantity calculation unit 201 calculates, as a feature quantity (second feature quantity), a score indicating the degree to which the data included in the source data set D s corresponds to the classification of the credit card usage stores. More specifically, the feature quantity calculation unit 201 calculates a score indicating the degree to which the data included in the source data set D s corresponds to the classification of the credit card usage stores as a feature quantity.
[0044] Specifically, the first data included in the source data set D s has an item ID of "i01_s", and the related data M s indicates that the attributes of this item (specifically, the store) are "retail" and "motorcycle shop". Also, the second data included in the source data set D s has an item ID of "i03_s", and the related data M s indicates that the attribute of this item is "hotel". Therefore, the feature quantity calculation unit 201 calculates, as a feature quantity, a score indicating the degree to which the data included in the source data set D s corresponds to "retail", "motorcycle shop", and "hotel".
[0045] Also, when calculating the feature quantity, the feature quantity calculation unit 201 may perform weighting according to the data included in the source data set D s For example, in the source data set D s , weighting may be performed such that the larger the value of the usage amount, the larger the score. An example of such weighting is shown in Table T1 in FIG. 4. In Table T1, f(w) is the weight value. Which value of w is used as the value of any data included in the source data set D s may be determined in advance according to the data (such as users or items) included in the source data set D. The same applies to the function f(w), and any function may be determined in advance according to the data included in the source data set D.
[0046] The above score is obtained by obtaining a mapping φ that converts each data u included in the source data set D s into a vector of the attribute dimension of the related data M s . Similarly, the score of each data u' included in the target data set D s can be calculated by obtaining a mapping φ that converts the data u' into a vector of the attribute dimension of the related data M t . t t
[0047] Note that hereinafter, for matters common to the source data set D s and the target data set D t , the subscript of "Z" may be used instead of the subscripts of "s" and "t". For example, the source data set D s and the target data set D t may be denoted as the data set D Z . Also, the above mapping φ s and the mapping φ t may be denoted as the mapping φ Z .
[0048] The method of constructing the mapping φ Z is arbitrary. For example, the mapping φ Z shown by the following mathematical formula may be used.
[0049]
Equation
[0050] Also, g Z is a mapping that converts an item into a vector of the attribute dimension. The related data M s and M t Utilize the attribute matrix representing the attribute values shown as it is in a matrix to obtain g Z = M Z [i] may be used. Also, for example, a vector determined from a matrix obtained by attaching weights such as TFIDF or BM25 to the attribute matrix (for example, g Z (i) = M Z tfidf [i]) may be used.
[0051] Also, h is a mapping that converts a set into a vector of attribute dimensions. As h, a mapping for normalization such that the norm becomes 1 is set. For example, as shown in the following mathematical formula, the sum of f(w)g(i) may be used as h, or an average or the like may be applied.
[0052]
Equation
[0053] Also, through similar processing, a feature, that is, a score, as shown in the feature data F t is calculated from the target data set D t and the related data M t . The features of the data included in the target data set D t can also be represented as a feature vector.
[0054] (Generation method of conversion rule generation) The conversion rule generation unit 202 generates a conversion rule 212 that converts the features regarding the attributes of the source data set D s into the features regarding the attributes of the target data set D t . Since each of the above features is represented as a vector, hereinafter, these features are referred to as feature vectors.
[0055] For the above conversion, the feature vector φ s regarding the attributes of the source data set D s (u) may be converted into a feature vector regarding the attributes of the target data set D t by obtaining a mapping ψ. For example, as shown in the following formula, the matrix product of the matrix R to be described later and φ s (u) can be regarded as the mapping ψ(φ s (u)). Rφ s (u) becomes a feature vector (d t -dimensional feature) regarding the attributes of the target data set D t as shown in the following formula.
[0056]
Equation
[0057]
Equation
[0058] Also, the matrix R may be such that values below a certain threshold θ are converted to 0. In this case, ψ(φ s (u)) is represented by the following mathematical formula. By dividing the matrix R at the threshold θ and converting values below the threshold θ to 0, it is possible to prevent attributes with low relevance from affecting the score after conversion.
[0059]
Equation
[0060] For example, the conversion rule generation unit 202 may generate the matrix R based on the similarity of character strings. In this case, the conversion rule generation unit 202 may generate the matrix R shown in the following mathematical formula str . As the similarity of character strings, for example, Jaccard similarity, Levenshtein similarity, and similarity based on n-gram can be applied.
[0061]
Equation
[0062] [Number] As described above, the conversion rule generation unit 202 may generate a conversion rule 212 in which the higher the degree of similarity between the attributes of the source data set D s and the attributes of the target data set D t , the more strongly the score before conversion is reflected in the score after conversion.
[0063] According to the above configuration, the conversion rule 212 can be automatically generated. Further, the generated conversion rule 212 is such that the higher the degree of similarity between the attributes of the source data set D s and the attributes of the target data set D t , the more strongly the score before conversion is reflected in the score after conversion. Therefore, appropriate conversion according to the degree of similarity between attributes can be performed.
[0064] Also, when applying the optimal transport theory, the conversion rule generation unit 202 may generate a matrix R ot shown in the following mathematical formula. Note that the following mathematical formula is different from the mathematical formula used in a general optimal transport matrix in that normalization is performed in the column direction (j) on the right side of the first row. By this normalization, the scores of each attribute included in the attribute set A s of the source data set D s are distributed so that the sum of the scores becomes 1 in the attribute set A t of the target data set D t .
[0065] In this case, a transport matrix T that realizes mapping according to the similarity of words, such as Word Mover's Distance or Word Rotator's Distance, is used as the conversion rule 212. Note that an example of the generation of the matrix R ot will be described later with reference to FIG. 5. Also, the details of Word Mover's Distance and Word Rotator's Distance are described in, for example, the following literature. Matt Kusner et al., “From Word Embeddings To Document Distances”, 2015, PMLR 37:957-966 Sho Yokoi et al., “Word Rotator's Distance”, 2020, Association for Computational Linguistics, pp.2944-2960
[0066]
Number
[0067] According to the above configuration, the conversion rule can be automatically generated. Further, since the conversion rule is generated by the optimal transport theory, the strength of the relationship between the attributes of the source data set D s and the attributes of the target data set D t can be emphasized. For example, when the relationship between these attributes is strong, the score before conversion can be strongly reflected in the score after conversion, while when the relationship between these attributes is weak, the score before conversion can be made not to be reflected in the score after conversion. As a result, the relationship between the score after conversion and the attributes of the source data set D s before conversion becomes clear, so that the interpretability and explainability can be further enhanced.
[0068] Further, a combination of the various matrices described above with weights may be used as the matrix R for conversion. In this case, the matrix R is represented by the following mathematical formula.
[0069]
Number
[0070] In the example of FIG. 5, the conversion rule generation unit 202 generates an embedding representation for each attribute name included in the attribute set A s associated with the source data set D s and each attribute name included in the attribute set A t associated with the target data set D t . More specifically, the conversion rule generation unit 202 embeds each attribute name included in the attribute set A s and each attribute name included in the attribute set A t into the n-dimensional Euclidean space SP1. n may be set to about 300, for example.
[0071] In the Euclidean space SP1, words (attribute names) that are similar to each other are embedded at positions close to each other. For this reason, "toy" and "toy store" are embedded at close positions, and "motorcycle supplies", "motorcycle store", and "auto parts" are embedded at close positions. Such an embedding representation can be generated by, for example, fasttext or the like.
[0072] Next, the conversion rule generation unit 202 calculates the World Rotator's Distance for the above embedding expression and sets the optimal transport matrix T as the matrix R. The matrix R is a matrix that indicates the similarity between attributes as a numerical value from 0 to 1. Note that the closer the value is to 1, the higher the similarity. For example, in the matrix R shown in FIG. 5, the values for combinations of attributes of similar concepts such as "motorcycle supplies" and "motorcycle shop" are 1. On the other hand, the values for combinations of attributes of dissimilar concepts such as "motorcycle supplies" and "toy store" are 0. Such a matrix R is stored in the storage unit 21 as the conversion rule 212.
[0073] As in the example of FIG. 5, in the matrix R generated based on the optimal transport theory, combinations of similar concepts often become 1, combinations of dissimilar concepts often become 0, and intermediate numerical values are rare. Therefore, when generating conversion rules based on the optimal transport theory, the relationship between the attributes of the source data set D s and the attributes of the target data set D t is easy to understand. That is, when using the optimal transport matrix as the conversion rule 212, as described above, there is an advantage that the explicability and interpretability of the final classification result are improved. On the other hand, for example, when using the cosine similarity (R cos ) etc. as the conversion rule 212, many intermediate values between -1 and 1 will be output, so the explicability etc. may be lower than when using the optimal transport matrix.
[0074] (Specific example of conversion) FIG. 6 is a diagram showing an example of the conversion of feature amounts using the matrix R shown in FIG. 5. As shown in the figure, by taking the matrix product of the matrix R and the feature vector F s , the feature vector F s can be converted into the feature vector F t ' of the attributes of the target data set D s .
[0075] Specifically, the feature vector F sThe score of "Retail" in is "0.05", which is converted to the score of "Automobile Parts" "0.05" by the matrix R. Also, the score of "Motorcycle Shop" "0.55" is converted to the score of "Motorcycle Supplies" (0.55) by the matrix R, and the score of "Hotel" "0.4" is converted to the score of "Motorcycle Supplies" (0.1) by the matrix R. Thus, the feature vector F s ' the score of "Motorcycle Supplies" is "0.65".
[0076] Feature vector F s corresponds to the user (the user with the user ID "u01_s" in Figure 4), indicating that the user uses the credit card for "Retail", "Motorcycle Shop", and "Hotel" at the rates of 0.05, 0.55, and 0.4 respectively. s And the feature vector F
[0077] And the feature vector F s ' shows the inference result that the above user is interested in "Motorcycle Supplies", "Automobile Parts", and "Map" at the rates of 0.65, 0.05, and 0.3 respectively on the e-commerce site.
[0078] As described above, the feature amount calculation unit 201 may calculate a score indicating the degree to which the data included in the source data set D s corresponds to the attributes associated with the source data set D s . Then, the conversion unit 203 applies the conversion rule 212 generated based on the relationship between the attributes of the target data set D t and the attributes of the source data set D s to convert the score of the source data set D s into a score indicating the degree to which it corresponds to the attributes of the target data set D t .
[0079] According to the above configuration, the score of the source data set D s can be converted into a score indicating the degree to which it corresponds to the attributes of the target data set D tis converted into a score indicating the degree of correspondence to the attribute, and based on the magnitude of the value of this score, similar data can be easily extracted across the target data set D t and the source data set D s Moreover, regarding the reason for extraction, it can also be easily explained and interpreted based on the magnitude of the value of the score.
[0080] (Specific example of similar data extraction) The similar data extraction unit 204 extracts similar data similar to the data included in the source data set D t from among a plurality of data included in the target data set D s To extract similar data, a mapping such as the following g may be configured. The mapping g converts the set U s of users included in the given source data set D s into a subset 2 t of the set U t of users included in the target data set D Ut and converts the element u included in the source data set D s into an element C included in the set of users U t . It can also be said that the mapping g returns a data group of the target data set D s similar to the data of t .
[0081]
Number
[0082]
Number
[0083] And the similar data extraction unit 204 may perform clustering on this union set and extract the data u’ included in the cluster to which the data u belongs as similar data. Note that the method of this clustering is arbitrary, and for example, the K-Means method, Ward method, DBSCAN, etc. can also be applied.
[0084] Also, for example, instead of executing the extraction of similar data at once as described above, the similar data extraction unit 204 may perform clustering on the feature quantity φ t (u’) which is the set regarding the attributes associated with the data u’ of the target dataset
Number
[0085] An example applying this method is shown in FIG. 7. FIG. 7 is a diagram showing an example of extracting similar data. In this extraction example, first, for each user included in the target dataset D t , the feature quantity calculation unit 201 calculates a feature quantity vector based on the corresponding attributes (specifically, “motorcycle supplies” to “toys”). Then, based on the calculated feature quantity vector, the similar data extraction unit 204 clusters each user in the feature space SP2 as shown in [1] of the figure.
[0086] On the other hand, for the source dataset D sFor the user included in s (the user with user ID "u01_s"), as shown in [2] of the same figure, a feature vector F based on the corresponding attribute (specifically, "retail" to "hospital") is generated by the feature quantity calculation unit 201. s Also, as shown in [3] of the same figure, the feature vector F s is converted into a feature vector F
[0087] ' by the conversion unit 203. s Here, since the feature vector F t ' is represented by the attributes of the target data set D, it can be classified into any of the clusters on the feature space SP2 generated by the above-described process of [1].
[0088] Then, the similar data extraction unit 204 extracts each user included in the cluster to which the feature vector F s ' belongs as a user similar to the users included in the source data set D s . In the example of FIG. 7, as shown in [4] of the same figure, three users with user IDs "u01_t", "u05_t", and "u012_t" are extracted. The scores of the feature quantities of these users are similar to each other as shown in the table T2 of the same figure, and their scores are also similar to the feature vector F s of the users in the source data set D s '.
[0089] Note that, as in the above example, when the data is a user, similar users may be extracted by applying the audience expansion technique. The process of the audience expansion technique is to give scores to each user, sort them in descending order of the scores, and return the top users as similar users, which is similar to clustering. The details of the audience expansion technique are described in, for example, the following literature. Stephanie deWet et al., “Finding Users Who Act Alike: Transfer Learning for Expanding Advertiser Audiences”, July 2019, Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp.2251-2259 (Flow of the information processing method: Preparation stage) The flow of the information processing method executed by the information processing apparatus 2 will be described based on FIG. 8. FIG. 8 is a flowchart showing the flow of processing executed by the information processing apparatus 2 in the preparation stage before performing feature quantity conversion and similar data extraction.
[0090] In S21, the information processing apparatus 2 reads the first and second data sets. The first and second data sets are input via, for example, the input unit 22 and stored as the data set 21 in the storage unit 21.
[0091] The first data set includes a plurality of data, and a first attribute represented in natural language is associated with the first data set. The first data set corresponds to the above-described target data set. The first attribute may or may not be shown in the first data set. In the latter case, separately from the first data set, data indicating the first attribute (for example, the related data M t such as the data) may be read.
[0092] On the other hand, a second attribute different from the first attribute represented in natural language is associated with the second data set. The second data set corresponds to the above-described source data set. Since only the attribute is used in the preparation stage for the second data set, it is not necessary to read the data of the second data set in S21. In S21, for the second data set, data indicating the second attribute (for example, the related data M s such as the data) may be read.
[0093] In S22, the feature quantity calculation unit 201 calculates a feature quantity regarding the first attribute for each data of the first data set read in S21. For example, the feature quantity calculation unit 201 may calculate a feature vector F as shown in FIG. 4 t and the like.
[0094] In S23, the conversion unit 203 generates a conversion rule for converting the second feature quantity regarding the second attribute in the second data set read in S21 into the first feature quantity regarding the first attribute in the first data set also read in S21. For example, the conversion unit 203 may generate a matrix R shown in FIG. 5 as the conversion rule. The generated conversion rule is stored in the storage unit 21 as the conversion rule 212. Thus, the processing in the preparation stage is completed. Note that the processing of S22 and S23 may be performed simultaneously in parallel, or the processing of S23 may be performed first.
[0095] (Flow of information processing method: inference stage) FIG. 9 is a flowchart showing the flow of the information processing method performed after the processing in the above-described preparation stage. In the flow of FIG. 9, after conversion using the conversion rule 212 generated in the preparation stage, extraction (inference) of similar data is performed.
[0096] In S31, the information processing apparatus 2 reads the data of the second data set. The data is input, for example, via the input unit 22 and stored in the storage unit 21 as the data set 21. In S31, at least one data included in the second data set may be read. Note that if the reading of the data of the second data set has also been completed in S21 of FIG. 8, the processing of S31 is omitted.
[0097] In S32, the feature quantity calculation unit 201 calculates a second feature quantity regarding the second attribute for the data of the second data set read in S31. For example, the feature quantity calculation unit 201 may calculate a feature vector F as shown in FIG. 4 s and the like.
[0098] In S33, the conversion unit 203 reads the conversion rule 212 generated in S23 of FIG. 8. Then, in S34, the conversion unit 203 uses the conversion rule 212 read in S33 to convert the second feature amount calculated in S32 into the first feature amount related to the first attribute. For example, as in the example of FIG. 6, the conversion unit 203 multiplies the matrix R, which is the conversion rule 212, and the feature vector F, which is the second feature amount s to obtain the feature vector F s ', which is the first feature amount.
[0099] In S35, the similar data extraction unit 204 reads the feature amounts calculated in S22 of FIG. 8, that is, the feature amounts related to the first attribute for each data included in the first data set. Then, in S36, the similar data extraction unit 204 extracts similar data similar to the data included in the second data set from among the plurality of data included in the first data set based on the feature amounts read in S35 and the first feature amount generated by the conversion in S34. Further, the similar data extraction unit 204 stores the extracted data in the storage unit 21 as the result data 213.
[0100] From the above, the process of FIG. 9 ends. By changing the data read in S31, similar data can be extracted from the first data set for each data included in the second data set. Note that the processes of S31, S33, and S35 may be performed simultaneously in parallel, and the execution order of these processes may be changed as appropriate.
[0101] 〔Modification Example〕 In the above exemplary embodiment, an example in which similar data is extracted using the converted feature amount and the extracted similar data is stored as the result data 213 has been described. However, the use of the converted feature amount is arbitrary and is not limited to this example.
[0102] For example, after extracting similar data using the transformed feature quantities, a feature image may be generated that visualizes the features of the feature quantities of the extracted similar data. This will be described with reference to FIG. 10. FIG. 10 is a diagram showing an example of generating a feature image that visualizes the features of the feature quantities of similar data.
[0103] In the example of FIG. 10, a feature image IMG is generated from the feature quantities of three users shown in the table T2 shown in FIG. 7. Such a feature image IMG can be generated by calculating the average value of the values of each attribute of the three users and drawing the attribute character string according to the magnitude of the value. For example, such a feature image can be generated by using WordCloud.
[0104] Such a feature image can also be generated for the transformed feature quantities. For example, a feature image may be generated from the transformed feature quantities of the data in the source data set, and a feature image may be generated from the feature quantities of the data in the target data set, and they may be stored as result data 213. In this case, the relevance between the data in the target data set and the data in the source data set is represented by the feature image. Note that, in this case, there may be one piece of data each for the target data set and the source data set.
[0105] 〔Application Example〕 The above-described information processing apparatus 2 can be used for various applications. For example, the information processing apparatus 2 can extract similar users similar to user A from a plurality of users corresponding to each data included in the target data set using the source data set of a certain user A.
[0106] Here, the information processing apparatus 2 can acquire data related to similar users from the data included in the target data set. For example, when the target data set indicates the purchase history of an e-commerce site, the information processing apparatus 2 can acquire data indicating the products purchased by the similar users.
[0107] Then, the information processing apparatus 2 can present information (such as an advertisement) that recommends the purchase of a product indicated by the acquired data to user A. For example, when the source data set indicates the usage history of a credit card, the information processing apparatus 2 may cause an advertisement to be displayed on the user A's my page or the like at the credit card company. At this time, even if user A has never used the above mail-order site, the reason for recommending the purchase of the product can be explained based on the attributes of the source data set.
[0108] Also, according to the information processing apparatus 2, for example, using a source data set indicating the usage history of a credit card at a certain store, it is also possible to analyze how users who shop at that store respond to what attributes on the mail-order site. In other words, according to the information processing apparatus 2, it is possible to analyze what product attribute scores become high on the mail-order site. Such analysis results are useful for decisions such as determining products to be sold at the store.
[0109] In addition, the information processing apparatus 2 can also be used for various cross-industry analyses. For example, assume that Company A, a clothing retailer, is considering entering the automotive industry, which is another industry. In this case, for example, data related to Company A's trading companies can be used as the source data set, and data related to the trading companies and handled products of Company B, which sells automotive supplies, can be used as the target data set. Also, the attributes of the target data set may be the classification of the handled products. In this case, the information processing apparatus 2 can extract trading companies of Company B that are similar to Company A's trading companies.
[0110] Here, the attributes of the trading companies of Company B extracted indicate the classification of the products handled by those trading companies. Among such attributes, those with high scores can be said to have a high degree of relevance to both the clothing industry and the automotive industry. For example, if, as a result of performing the above analysis, Information Processing Device 2 has a high score for the attribute of "seat", the product to be recommended to Company A may be determined to be vehicle seats. Company A can then request the manufacturing of vehicle seats from its trading companies based on this recommendation. In this way, Information Processing Device 2 can also be used for various cross-industry analyses.
[0111] 〔Modification Example〕 The execution entity of each process described in each of the above exemplary embodiments is arbitrary and is not limited to the above examples. That is, an information processing system having functions similar to Information Processing Devices 1 and 2 can be constructed by a plurality of mutually communicable devices. Also, for example, each process described in each of the above exemplary embodiments can be realized in the form of SaaS (Software as a Service) or the like.
[0112] 〔Example of Realization by Software〕 Some or all of the functions of Information Processing Devices 1 and 2 may be realized by hardware such as an integrated circuit (IC chip), or may be realized by software.
[0113] In the latter case, Information Processing Devices 1 and 2 are realized by a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as Computer C) is shown in FIG. 11. Computer C includes at least one processor C1 and at least one memory C2. A program P for operating Computer C as Information Processing Devices 1 and 2 is recorded in memory C2. In Computer C, processor C1 reads and executes program P from memory C2, thereby realizing each function of Information Processing Devices 1 and 2.
[0114] As the processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating point number Processing Unit), PPU (Physics Processing Unit), microcontroller, or a combination thereof can be used. As the memory C2, for example, a flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0115] Note that the computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and temporarily storing various data. Also, the computer C may further include a communication interface for transmitting and receiving data to and from other devices. Further, the computer C may further include an input / output interface for connecting input / output devices such as a keyboard, mouse, display, and printer.
[0116] Also, the program P can be recorded on a non-transitory tangible recording medium M readable by the computer C. As such a recording medium M, for example, a tape, disk, card, semiconductor memory, or programmable logic circuit can be used. The computer C can obtain the program P via such a recording medium M. Also, the program P can be transmitted via a transmission medium. As such a transmission medium, for example, a communication network or broadcast wave can be used. The computer C can also obtain the program P via such a transmission medium.
[0117] 〔Supplementary Note 1〕 The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. For example, embodiments obtained by appropriately combining the technical means disclosed in the above-described embodiments are also included in the technical scope of the present invention.
[0118] [Supplementary Note 2] Some or all of the above-described embodiments may also be described as follows. However, the present invention is not limited to the aspects described below.
[0119] (Supplementary Note 1) An information processing apparatus comprising: a first data set including at least one data and having a first attribute represented in natural language associated therewith; and a second data set including a plurality of data and having a second attribute represented in natural language associated therewith, and feature quantity calculation means for calculating, for the data included in the second data set, a second feature quantity regarding the second attribute, and conversion means for converting the second feature quantity into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0120] According to the above configuration, it becomes possible to extract similar data across different data sets, namely, the first data set and the second data set, and it also becomes possible to explain and interpret the reason for the extraction.
[0121] (Supplementary Note 2) The information processing apparatus according to Supplementary Note 1, wherein the feature quantity calculation means calculates, as the second feature quantity, a second score indicating the degree to which the data included in the second data set corresponds to the second attribute associated with the second data set, and the conversion means applies a conversion rule generated based on the relationship between the first attribute and the second attribute to convert the second score into a first score indicating the degree to which the first attribute corresponds, and uses the converted score as the first feature quantity.
[0122] According to the above configuration, since the first score indicating the degree corresponding to the first attribute is used as the first feature amount, similar data can be easily extracted across the first dataset and the second dataset based on the magnitude of the value of the first score. Also, regarding the reason for the extraction, it can be easily explained and interpreted based on the magnitude of the value of the first score.
[0123] (Appendix 3) The information processing apparatus according to Appendix 2, comprising conversion rule generation means for generating a conversion rule such that the second score is more strongly reflected in the first score as the degree of similarity between the first attribute and the second attribute is higher.
[0124] According to the above configuration, a conversion rule can be automatically generated. Also, since the generated conversion rule is such that the second score is more strongly reflected in the first score as the degree of similarity between the first attribute and the second attribute is higher, a proper conversion according to the degree of similarity between the first attribute and the second attribute can be performed.
[0125] (Appendix 4) The information processing apparatus according to Appendix 2, comprising conversion rule generation means for generating the conversion rule from the first attribute and the second attribute by optimal transport theory.
[0126] According to the above configuration, a conversion rule can be automatically generated. Also, since the conversion rule is generated by optimal transport theory, the strength of the relationship between the first attribute and the second attribute can be emphasized, and the interpretability and explainability can be further enhanced.
[0127] (Appendix 5) The information processing apparatus according to any one of Appendices 1 to 4, comprising similar data extraction means for extracting similar data similar to the data included in the second dataset from among the plurality of data included in the first dataset based on the first feature amount of the data included in the second dataset generated by the conversion by the conversion means and the feature amount regarding the first attribute of the plurality of data included in the first dataset.
[0128] According to the above configuration, data similar to the data included in the first data set can be automatically extracted from among the plurality of data included in the second data set.
[0129] (Appendix 6) An information processing method, including: for data included in the second data set out of a first data set including at least one piece of data and associated with a first attribute expressed in natural language, and a second data set including a plurality of data and associated with a second attribute expressed in natural language, calculating a second feature amount regarding the second attribute; and converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0130] According to the above information processing method, it becomes possible to extract similar data across different data sets, namely, the first data set and the second data set, and it also becomes possible to explain and interpret the reason for the extraction.
[0131] (Appendix 7) A program that causes a computer to function as a feature amount calculation means for calculating a second feature amount regarding a second attribute for data included in the second data set out of a first data set including at least one piece of data and associated with a first attribute expressed in natural language, and a second data set including a plurality of data and associated with a second attribute expressed in natural language, and a conversion means for converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0132] According to the above program, it becomes possible to extract similar data across different data sets, namely, the first data set and the second data set, and it also becomes possible to explain and interpret the reason for the extraction.
[0133] [Appendix Item 3] Some or all of the above-described embodiments can also be expressed as follows. An information processing apparatus including at least one processor, the processor including at least one data, a first dataset associated with a first attribute expressed in natural language, and a second dataset including a plurality of data and associated with a second attribute expressed in natural language, the processor performing a feature amount calculation process for calculating a second feature amount regarding the second attribute for the data included in the second dataset, and a conversion process for converting the second feature amount into a first feature amount regarding the first attribute based on the relationship between the first attribute and the second attribute.
[0134] Note that this information processing apparatus may further include a memory, and a program for causing the processor to execute the feature amount calculation process and the conversion process may be stored in this memory. Further, this program may be recorded on a non-transitory tangible computer-readable recording medium.
Explanation of Signs
[0135] 1 Information processing apparatus 11 Feature amount calculation unit 12 Conversion unit 2 Information processing apparatus 201 Feature amount calculation unit 202 Conversion rule generation unit 203 Conversion unit 204 Similar data extraction unit 211 Dataset (first dataset, second dataset) 212 Conversion rule
Claims
1. A first dataset including at least one piece of data and associated with a first attribute represented in natural language, and among second datasets including a plurality of pieces of data and associated with a second attribute represented in natural language, for the data included in the second dataset, feature quantity calculation means for calculating a second feature quantity regarding the second attribute; conversion means for converting the second feature quantity into a first feature quantity regarding the first attribute based on the relationship between the first attribute and the second attribute, comprising: The feature quantity calculation means calculates, as the second feature quantity, a second score indicating the degree to which the data included in the second dataset corresponds to the second attribute associated with the second dataset; The conversion means applies a conversion rule generated based on the relationship between the first attribute and the second attribute, and uses, as the first feature quantity, the converted value obtained by converting the second score into a first score indicating the degree to which the converted value corresponds to the first attribute. An information processing apparatus.
2. The information processing apparatus according to claim 1, further comprising conversion rule generation means for generating a conversion rule such that the higher the degree of similarity between the first attribute and the second attribute, the stronger the second score is reflected in the first score.
3. The information processing apparatus according to claim 1, further comprising conversion rule generation means for generating the conversion rule from the first attribute and the second attribute by means of optimal transport theory.
4. Based on the first feature quantity regarding the data included in the second dataset generated by the conversion by the conversion means and the feature quantity regarding the first attribute of the plurality of pieces of data included in the first dataset, among the plurality of pieces of data included in the first dataset, similarity data extraction means for extracting similarity data similar to the data included in the second dataset. The information processing apparatus according to any one of claims 1 to 3.
5. At least one processor calculates, for the data included in the second dataset, a second feature quantity regarding the second attribute, among a first dataset including at least one piece of data and associated with a first attribute represented in natural language, and a second dataset including a plurality of pieces of data and associated with a second attribute represented in natural language; converting the second feature amount into a first feature amount related to the first attribute based on the relationship between the first attribute and the second attribute; in calculating the second feature amount, the at least one processor calculates, as the second feature amount, a second score indicating the degree to which the data included in the second data set corresponds to the second attribute associated with the second data set; an information processing method, wherein in the conversion, the at least one processor applies a conversion rule generated based on the relationship between the first attribute and the second attribute, and converts the second score into a first score indicating the degree to which the first attribute corresponds, and uses the converted score as the first feature amount. **Claim 6** A computer, among a first data set including at least one data and associated with a first attribute expressed in natural language, and a second data set including a plurality of data and associated with a second attribute expressed in natural language, for the data included in the second data set, a feature amount calculation means for calculating a second feature amount related to the second attribute, and functioning as a conversion means for converting the second feature amount into a first feature amount related to the first attribute based on the relationship between the first attribute and the second attribute; the feature amount calculation means calculates, as the second feature amount, a second score indicating the degree to which the data included in the second data set corresponds to the second attribute associated with the second data set; a program, wherein the conversion means applies a conversion rule generated based on the relationship between the first attribute and the second attribute, and converts the second score into a first score indicating the degree to which the first attribute corresponds, and uses the converted score as the first feature amount.
Citation Information
Patent Citations
Duplicate record detection system and, duplicate record detection program
JP2006163941A
Document classification program, document classification method, and document classification apparatus
JP2006301920A
Data conversion program, data conversion method, and data converter
JP2018055551A
Information providing method
JP2020038727A