Information processing apparatus, method, and program

The information processing apparatus addresses the challenge of analyzing data with changing variable names by converting and combining variable and value vectors, enabling efficient analysis of data with diverse column configurations.

JP2025091092APending Publication Date: 2025-06-18KK TOSHIBA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023206089
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Existing tabular data analysis technologies struggle to analyze data with changes, additions, or deletions of variable names, as they require re-learning or discarding non-common columns, limiting their ability to handle data with different column configurations.

Method used

An information processing apparatus comprising a data and correspondence acquisition unit, a vector generation unit, and a vector combination unit, which acquires variable names and values, converts them into vectors, and combines these vectors based on their correspondence, enabling analysis even when variable names change, add, or delete.

Benefits of technology

This solution allows for the analysis of data with varying column configurations without the need for re-learning or discarding data, reducing computational costs and enhancing the ability to process diverse data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091092000001_ABST
    Figure 2025091092000001_ABST
Patent Text Reader

Abstract

To provide a method that can analyze data in which variable names and values are associated even though the variable names are changed, added or reduced.SOLUTION: An information processing apparatus according to an embodiment includes a data and correspondence relationship acquisition unit, a vector generation unit, and a vector combination unit. The data and correspondence relationship acquisition unit acquires data including a variable name and a value associated with the variable name, and a correspondence relationship between the variable name and the value. The vector generation unit generates a variable name vector corresponding to each variable name and a value vector corresponding to the value associated with each variable name. The vector combination unit combines the variable name vector and the value vector based on the correspondence relationship.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an information processing apparatus, method, and program.

Background Art

[0002] With the development of IoT (Internet of Things) technology, information processing apparatuses have become capable of acquiring various types of data, such as multivariate data represented by tabular data. Here, multivariate data is data obtained by aggregating data of different natures and is used in many scenarios. For example, in a manufacturing site, multivariate data that aggregates manufacturing data related to the state of products, such as manufacturing conditions such as the names of materials and devices used during manufacturing and inspection data of manufactured products, is used. Also, in a medical site, multivariate data that aggregates medical data related to patients, such as patient attribute information and various test values, is used.

[0003] Furthermore, by analyzing these multivariate data, it is expected to obtain analysis results such as failures of products or devices, or candidates for diseases. For example, in a scenario where tabular data, which is a type of multivariate data, is analyzed, various machine learning techniques for classification or regression are used as tabular data analysis techniques. Here, tabular data has the values of each variable corresponding to one sample along the row direction, and stores one variable name and the values of each sample corresponding thereto along the column direction. Each variable includes a variable name and its value. This type of tabular data analysis technique can only analyze data with the same column configuration as the training data during learning during operation because it is premised that the column configuration of the tabular data is the same between learning and operation.

[0004] However, when analyzing tabular data, during operation, the column configuration of the tabular data may be different from that of the training data during learning. For example, when the tabular data is manufacturing data, due to changes in the manufacturing process, addition or reduction of inspection items, etc., the variables constituting the manufacturing data may change between the learning time and the operation time. In this case, the tabular data analysis technology cannot analyze the manufacturing data with the changed column configuration due to changes, additions, or deletions of variable names. Therefore, the tabular data analysis technology needs to re-learn using the training data after the column configuration is changed, or extract only the common columns from the manufacturing data before and after the change in the column configuration and perform the analysis. In the latter case, the remaining columns after extraction are not used for the analysis and are discarded (ignored).

[0005] Therefore, it is desired to enable the analysis of data with changes, additions, or deletions of variable names. In this way, when data with different column configurations can be analyzed, effects such as the need for data augmentation or re-learning of the training data can be expected to be eliminated.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] The problem to be solved by the present invention is to provide an information processing apparatus, method, and program that can analyze data in which variable names and values are associated even when changes, additions, or deletions of variable names occur.

Means for Solving the Problems

[0008] The information processing apparatus according to the embodiment includes a data and correspondence acquisition unit, a vector generation unit, and a vector combination unit. The data and correspondence acquisition unit acquires a variable name, a value associated with the variable name, and the correspondence between the variable name and the value. The vector generation unit converts each of the variable name and the value into a vector. The vector combination unit combines the vectors obtained from the variable name and the value based on the correspondence between the variable name and the value.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Modes for Carrying Out the Invention

[0010] Hereinafter, each embodiment will be exemplarily described with reference to the drawings. In the following description, tabular data stored in a csv file or the like will be described as an example of multivariate data in which variable names and values are associated. However, the present invention is not limited to this, and as multivariate data, file format data in which variable names and values are associated, such as json or yaml, may be used.

[0011] <First Embodiment> FIG. 1 is a block diagram showing an example of the configuration of an information processing apparatus according to the first embodiment. The information processing apparatus 10 includes a data and correspondence acquisition unit 1, a vector generation unit 4, and a vector combination unit 5.

[0012] Here, the data and correspondence relationship acquisition unit 1 receives the input of tabular data 201 which is a type of multivariate data. Here, the multivariate data is data obtained by aggregating data with different properties, and is collected for a plurality of samples. For example, the tabular data 201 may store data in a plurality of cells (0 ≦ i ≦ m - 1, 0 ≦ j ≦ n - 1) of m rows and n columns. For example, a plurality of variable names arranged in a plurality of columns (i = 0, 1 ≦ j ≦ n - 1) after the first column in the 0th row, a plurality of sample IDs arranged in a plurality of rows (1 ≦ i ≦ m - 1, j = 0) after the first row in the 0th column, and a plurality of values arranged in a plurality of cells (1 ≦ i ≦ m - 1, 1 ≦ j ≦ m - 1) after the first row and after the first column may be stored in association with each other. In such a case, the tabular data 201 has a plurality of variables in one sample along the row direction, and each of the plurality of variables has a variable name (column name) and a value associated with the variable name along the column direction. Note that the sample ID is not essential and may be omitted. For example, when the sample is anonymous like big data, the sample ID can be omitted. The plurality of variables may be independent of each other or may have a relationship with each other. As the value associated with the variable name, data in any data format can be appropriately used. For example, the value may be any of a discrete category value, text data, and a continuous value represented by a numerical value. Further, the data and correspondence relationship acquisition unit 1 acquires data 202 including a variable name and a value associated with the variable name, and a correspondence relationship 203 between the variable name and the value from the tabular data 201.

[0013] The vector generation unit 4 generates a variable name vector 207 corresponding to each variable name and a value vector 208 corresponding to the value associated with each variable name. For example, when the value is a category value, the value vector 208 is a vector corresponding to the category value. When the value is a category value or text data, the vector generation unit 4 may generate a value vector 208, which is a text vector, by performing text analysis on the value. Also, for example, when the value is a numerical value, the value vector 208 is a vector corresponding to the numerical value. In this case, the vector generation unit 4 may generate a value vector 208, which is the output of the neural network, by inputting the numerical value into the neural network. Also, for example, the vector generation unit 4 may generate a value vector 208 by linearly transforming the numerical value.

[0014] The vector combination unit 5 combines the variable name vector 207 and the value vector 208 based on the correspondence 203 between the variable name and the value. For example, the vector combination unit 5 may combine the variable name vector 207 and the value vector 208 by adding the variable name vector 207 and the value vector 208 related to the same variable name. Also, for example, the vector combination unit may combine the variable name vector 207 and the value vector 208 by arranging the respective elements of the variable name vector 207 and the value vector 208 related to the same variable name. Also, for example, the vector combination unit may input the variable name vector 207 and the value vector 208 related to the same variable name into the neural network and output a vector 209 obtained by combining the variable name vector 207 and the value vector 208. The vector 209 is a vector of each variable output in the number (number of columns) of variable names.

[0015] Next, the operation of the information processing apparatus configured as described above will be described with reference to the flowchart of FIG. 2 and the schematic diagrams of FIGS. 3 and 4.

[0016] (Step ST10: Acquisition of Data and Correspondence) As shown in FIGS. 2 and 3, the data and correspondence relationship acquisition unit 1 receives the input of tabular data 201. The tabular data 201 has three variables (numerical value 1, 100), (numerical value 2, 200), (category 1, X) in one sample along the row direction, and each of the plurality of variables has a variable name (column name) and a value corresponding to the variable name along the column direction. Further, the data and correspondence relationship acquisition unit 1 acquires data 202 including the variable name and the value corresponding to the variable name, and the correspondence relationship 203 between the variable name and the value from the tabular data 201.

[0017] (Step ST20: Vector generation) Based on the data 202, the vector generation unit 4 generates variable name vectors 207 corresponding to each of the variable names "numerical value 1", "numerical value 2", and "category 1". Further, the vector generation unit 4 generates value vectors 208 corresponding to the values "100", "200", and "X" associated with each variable name. When the input data has N variables (number of columns N), the vector generation unit 4 generates N variable name vectors 207 obtained from N variable names and N value vectors obtained from the values associated with each variable name. In FIG. 3, N = 3.

[0018] Regarding variable names, they can be regarded as text data. As a method for vectorizing variable names, there is a method of learning embedding vectors. In this method, a vocabulary list is created from the variable names appearing in the training data, and the embedding vectors corresponding to each vocabulary are learned. Also, the variable name can be tokenized, and the embedding vectors can be learned for each token. In this case, from one variable name, several token vectors corresponding to the tokens constituting the variable name can be obtained. The variable name vector 207 may be generated by averaging the token vectors obtained from each token constituting the variable name. Also, the variable name vector 207 may be generated by inputting the set of token vectors constituting the variable name into a neural network. By performing tokenization, it is possible to expect the effect that variable name vectors obtained from similar variable names containing the same token are similar to each other. Also, when an unknown variable name appears during inference, the embedding vectors learned from other variable names containing the tokens constituting the variable name may be utilized.

[0019] Regarding the value associated with the variable name, the method for obtaining the vector is changed according to the type of data. When the value is a categorical value or text data, it can be vectorized by the same process as the variable name. The embedding vectors to be learned may be common with those of the variable name. By sharing the embedding vectors with the variable name, it is possible to learn the embedding vectors of various tokens. For example, when a token not included in the variable names in the training data is included in the value, the embedding is learned from the tokens constituting the value, and the variable name vector can be generated even for an unknown input as a variable name. Also, the tabular data 201 may have text data stored in the value. In this case, the value may be vectorized using a text feature vector extractor, which is a natural language processing technique.

[0020] When the value associated with the variable name is a numerical value, vectorization is performed by a process different from that for categorical values or text data. For example, a value vector 208 may be generated by linearly transforming the numerical value. Specifically, for example, a value vector 208 may be generated by multiplying the numerical value by a weight matrix. When the output value vector 208 is of dimension D num since each variable stores a scalar value, the size of the weight matrix W num is 1×D num dimensions. Let the value vector 208 be v num and the original numerical value be x num When D num dimensional bias b num is added, the value vector v num is expressed as follows.

[0021] v num =x num ·W num +b num Here, the weight matrix W num and the bias b num are determined by learning. Similarly, a value vector 208 may be obtained from a numerical value by a neural network with a multi-layer transformation process. Also, vectorization of numerical values may be performed by the method described in the technical literature (Yury Gorishniy, Ivan Rubachev, Artem Babenko “On Embeddings for Numerical Features in Tabular Deep Learning” Advances in Neural Information Processing Systems 35, pages 24991-25004, 2022.). Alternatively, by treating the numerical value as a character string, vectorization may be performed in the same way as for categorical values.

[0022] (Step ST30: Vector combination) Based on the correspondence relationship 203 between the variable names and values in each variable, the vector combining unit 5 combines the variable name vector 207 and the value vector 208 to obtain a vector 209 corresponding to each variable. Here, when the input data to the information processing apparatus 10 is N variables, the vector combining unit 5 acquires and outputs N D-dimensional vectors 209. In the example of FIG. 3, when the tabular data 201 is three variables, three D-dimensional vectors 209 are output.

[0023] Specifically, the vector combining unit 5 takes two vectors, the variable name vector 207 and the value vector 208, as inputs, and outputs a D-dimensional vector 209 by the sum of both vectors. In this case, both the variable name vector 207 and the value vector 208 are D-dimensional. Note that this is not the only case, and the variable name vector 207 and the value vector 208 may have different dimensions. For example, represent the variable name vector 207 as v variable and represent the variable name vector v variable as D variable dimensions, represent the value vector 208 as v value and represent the value vector v value as D value dimensions. Also, represent the vector 209 of each variable as v. At this time, as shown below, a vector v obtained by arranging the elements of the variable name vector v variable and the elements of the value vector v value may be a D variable + D value dimensional vector corresponding to each variable.

[0024]

Number

[0025] Further, the vector combining unit 5 may linearly transform the vector v, which is the above D variable + D value dimensional vector, and use the obtained D-dimensional vector as the vector 209.

[0026] In any case, the vector combining unit 5 performs a combining process of obtaining each vector 209 corresponding to each variable from the variable name vector 207 and the value vector 208. By performing such a combining process, the correspondence between the variable name and the value of the input data can be explicitly shown in each vector 209.

[0027] As described above, according to the first embodiment, the data and correspondence acquisition unit 1 acquires data 202 including a variable name and a value corresponding to the variable name, and a correspondence 203 between the variable name and the value from the tabular data 201. The vector generation unit 4 generates a variable name vector 207 corresponding to each variable name and a value vector 208 corresponding to the value associated with each variable name. The vector combining unit 5 combines the variable name vector 207 and the value vector 208 based on the correspondence 203 between the variable name and the value. In this way, with the configuration of combining the variable name vector 207 corresponding to each variable name and the value vector 208 corresponding to the value associated with each variable name, even if a change, addition, or deletion of the variable name occurs in the data in which the variable name and the value are associated, the data can be made analyzable.

[0028] Regarding this effect, as a comparative example, a technical document (Zifeng Wang and Jimeng Sun “TransTab: Learning Transferable Tabular Transformers Across Tables” Advances in Neural Information Processing Systems 35, pages 2902-2915, 2022. (hereinafter referred to as [Wang22])) will be cited and explained. In the technical document [Wang22] of the comparative example, a technique has been proposed to make it possible to handle tabular data with different column configurations by converting the components of the tabular data into a set of vectors. In this type of comparative example, the column names (variable names) and category values of the tabular data are tokenized, and each token is vectorized. At that time, in the comparative example, information such as which of the column name and the category value each token is derived from and which column it belongs to is lost. Therefore, in the comparative example, the correspondence between the column name and the value is lost, and it becomes impossible to distinguish when the same category value is used in different columns. Therefore, in the comparative example, it is not possible to identify which column each category value belongs to, and there is an inconvenience that it cannot be processed appropriately.

[0029] On the other hand, according to the first embodiment, by obtaining the correspondence 203 between the variable name and the value, generating vectors corresponding to each of the variable name and the value, and combining the variable name vector 207 and the value vector 208 using the correspondence 203, it becomes possible to explicitly distinguish which variable name the value corresponds to. Supplementary, in the first embodiment, as shown in the upper and lower parts of FIG. 4, between two samples, a set 210 of vectors including the variable name vector 207 and the value vector 208 generated by the vector generation unit 4 is the same set as each other. Even in this case, in the first embodiment, between two samples, a set 211 including two vectors 209 output by the vector combination unit 5 is a different set 211 from each other. According to such a first embodiment, it becomes possible to distinguish and process each set 211 of vectors 209 obtained from two samples.

[0030] Therefore, in the case of tabular data using the same categorical values in multiple columns between two samples, in the comparative example, since the correspondence between the column names and the values is lost, different samples cannot be distinguished. To supplement, the comparative example corresponds to the configuration without the vector combining unit 5 in FIG. 4, so that the same set 210 can be obtained between two samples. In such a comparative example, the set 210 of each vector obtained from the two samples cannot be distinguished in the subsequent processing.

[0031] On the other hand, in the first embodiment, different samples can be distinguished in order to maintain the correspondence between the column names (variable names) and the values. That is, since the first embodiment has a configuration including the vector combining unit 5 in FIG. 4, different sets 211 can be obtained between two samples. In such a first embodiment, the set 211 of each vector 209 obtained from the two samples can be distinguished in the subsequent processing. In addition, in the first embodiment, the number of vectors constituting the set 211 is the same as the number of variables N (the number of columns N). On the other hand, in the technical document [Wang22] of the comparative example, the number of vectors constituting the set 210 is the number of tokens M obtained by token splitting (M≧2N). Therefore, according to the first embodiment, compared with the comparative example, the number of vectors to be processed in the subsequent processing can be reduced to less than half, so that the calculation cost of the subsequent processing can be reduced.

[0032] Further, according to the first embodiment, the value may be a categorical value, and the value vector 208 may be a vector corresponding to the categorical value. In this case, even when the value is a categorical value, the above-described effects can be obtained.

[0033] Further, according to the first embodiment, when the value is a categorical value or text data, the vector generation unit 4 may generate the value vector 208, which is a text vector, by text-analyzing the value. In this case, even when the value is a categorical value or text data, the above-described effects can be obtained.

[0034] Also, according to the first embodiment, the value may be a numerical value, and the value vector 208 may be a vector corresponding to the numerical value. Even in the case where the value is a numerical value, the above-described effects can be obtained.

[0035] Also, according to the first embodiment, the vector generation unit 4 may generate the value vector 208, which is the output of the neural network, by inputting the numerical value into the neural network. In this case, in addition to the above-described effects, the value vector 208 can be generated using the neural network.

[0036] Also, according to the first embodiment, the vector generation unit 4 may generate the value vector 208 by linearly transforming the numerical value. In this case, in addition to the above-described effects, the value vector 208 can be generated using linear transformation.

[0037] Also, according to the first embodiment, the data and correspondence relationship acquisition unit 1 acquires the data 202 and the correspondence relationship 203 from the tabular data 201 having a plurality of variables in one sample along the row direction and each of the plurality of variables having a variable name and a value along the column direction. Therefore, since the correspondence relationship 203 used for vector combination can be acquired from the tabular data 201, in addition to the above-described effects, the labor of the operator inputting the correspondence relationship 203 can be omitted.

[0038] Also, according to the first embodiment, the vector combination unit 5 combines the variable name vector 207 and the value vector 208 by the sum of the variable name vector 207 and the value vector 208 regarding the same variable name. Therefore, in addition to the above-described effects, when both the variable name vector 207 and the value vector 208 are D-dimensional, the combined vector 209 can be obtained as a D-dimensional vector.

[0039] Further, according to the first embodiment, the vector combining unit 5 may combine the variable name vector 207 and the value vector 208 regarding the same variable name by arranging the respective elements of the variable name vector 207 and the value vector 208. In this case, in addition to the above-described effects, the load of the process for obtaining the combined vector 209 can be reduced as compared with the case of obtaining the combined vector 209 by arithmetic processing.

[0040] Further, according to the first embodiment, the vector combining unit 5 may output a vector 209 obtained by combining the variable name vector 207 and the value vector 208 regarding the same variable name by inputting the variable name vector 207 and the value vector 208 into a neural network. In this case, in addition to the above-described effects, the combined vector 209 can be generated using the neural network.

[0041] <Modification Example of the First Embodiment> Subsequently, a modification example of the first embodiment will be described. This modification example can be similarly applied to the following respective embodiments.

[0042] FIG. 5 is a schematic diagram for explaining a modification example of the first embodiment. The same components as the above-described components are denoted by the same reference numerals and the detailed description thereof is omitted. Here, mainly, different parts will be described. The following respective embodiments will also omit the overlapping description in the same manner.

[0043] Here, the information processing apparatus 10 further includes a processing unit 6 at a subsequent stage of the vector combining unit 5. The processing unit 6 performs classification processing or regression processing based on the set 211 by inputting the set 211 of vectors 209 combined by the vector combining unit 5 into the neural network 61. Supplementary, the vector combining unit 5 outputs N D-dimensional vectors 209 corresponding to each variable. The processing unit 6 performs classification and regression of the tabular data 201 which is input data based on the N D-dimensional vectors 209 output from the vector combining unit 5.

[0044] Other configurations are the same as those in the first embodiment.

[0045] According to the above-described modification example, in addition to the effects of the first embodiment, classification processing or regression processing based on the set 211 of combined vectors 209 can be performed.

[0046] In this modification example, the processing unit 6 performs feature extraction using a neural network 61 that takes the set 211 of vectors 209 as input, and performs regression or classification from the obtained feature vectors. At this time, assuming that the configuration of the neural network 61 inputs an arbitrary number of D-dimensional vectors, it is possible to handle input data with an arbitrary number of variables. By describing the vector generation unit 4, the vector combination unit 5, and the subsequent classification and regression processing with differentiable operations, learning by the gradient descent method can be performed using training data with teacher labels. Also, self-supervised learning without using teacher labels may be applied by the method described in the above-mentioned technical document [Wang22].

[0047] When the neural network 61 for feature extraction has a multi-layer structure that outputs the same number of D-dimensional vectors as the input vectors in each layer as intermediate outputs, the intermediate outputs can be interpreted as the value vectors 208 in the vector combination unit 5, and the processing in the vector combination unit 5 may be performed in each layer of the neural network 61.

[0048] Also, by configuring the neural network 61 for feature extraction to be able to process an arbitrary number of vectors, data with different variable configurations can be utilized. For example, in a manufacturing site, it is assumed that some variables may be changed, added, or deleted due to changes in the manufacturing process, addition or reduction of inspection items, etc. In known technologies such as Patent Document 1, there is a constraint that the variable configuration of the input data is the same, so data before the configuration change cannot be utilized. However, when some variable configurations are the same, it is possible to increase the number of training data by using the data before the configuration change. For example, data with different variable configurations may be mixed as training data to train a machine learning model.

[0049] Also, in this modification example, data with an unknown variable configuration may be input to the neural network 61 during inference. However, the neural network 61 has been machine-learned using data with different variable configurations as training data. Such a neural network 61 can handle combinations of variables not included in the training data for variables included in the data with an unknown variable configuration among the variables included in the training data.

[0050] <Second Embodiment> Next, an information processing apparatus according to the second embodiment will be described.

[0051] The second embodiment is a modification example of the first embodiment and represents a specific example when the value corresponding to the variable name is a categorical value. When the value is a categorical value, the value vector 208 is a vector corresponding to the categorical value.

[0052] FIG. 6 is a block diagram showing an example of the configuration of an information processing apparatus according to the second embodiment. This information processing apparatus 10 further includes a token vectorization unit 2 and a variable name / categorical value specifying unit 3 as compared with the configuration shown in FIG. 1.

[0053] Here, the token vectorization unit 2 token-divides the variable name and the categorical value for a variable including the categorical value and the variable name associated with the categorical value among the data 202 acquired by the data and correspondence relationship acquisition unit 1, and generates a token vector 205 corresponding to each obtained token 204.

[0054] The variable name / categorical value specifying unit 3 specifies the variable from which the token vector 205 is derived and the variable name or categorical value from which the token vector 205 is derived for each token 204 obtained by the token vectorization unit 2.

[0055] Accordingly, based on the specified result 206, the vector generation unit 4 generates a variable name vector 207 from the token vector 205 derived from the variable name and a value vector 208 from the token vector 205 derived from the category value for each variable.

[0056] For example, the vector generation unit 4 may generate the value vector 208 by processing a set of token vectors 205 derived from the category value with a neural network. Also, the vector generation unit 4 may generate the value vector 208 by averaging the token vectors 205 derived from the category value.

[0057] Similarly, the vector generation unit 4 may generate the variable name vector 207 by processing a set of token vectors 205 obtained from the tokens 204 constituting the variable name with a neural network. Also, the vector generation unit 4 may generate the variable name vector 207 by averaging the token vectors obtained from the tokens constituting the variable name.

[0058] Other configurations are the same as those in the first embodiment.

[0059] Next, the operation of the information processing apparatus configured as described above will be described with reference to the flowchart of FIG. 7 and the schematic diagram of FIG. 8. Note that in FIG. 7, steps ST21 to ST23 surrounded by a broken line are the parts that are significantly changed from the operations shown in FIG. 2.

[0060] (Step ST10: Data and correspondence relationship acquisition) As shown in FIGS. 7 and 8, the data and correspondence relationship acquisition unit 1 receives an input of tabular data 201. The tabular data 201 has three variables (category 1, AB) (category 2, XY) in one sample along the row direction, and each of the plurality of variables has a variable name (column name) and a value corresponding to the variable name along the column direction. Note that the values corresponding to the variable names "category 1" and "category 2" are the category values "AB" and "XY".

[0061] In addition, the data and correspondence relationship acquisition unit 1 acquires data 202 including variable names and values corresponding to the variable names, and a correspondence relationship 203 between variable names and values from the tabular data 201.

[0062] (Step ST21: Token Vectorization) For variables including category values and variable names corresponding to the category values among the acquired data 202, the token vectorization unit 2 performs token splitting on the variable names and category values, and generates token vectors 205 corresponding to the obtained respective tokens 204. In FIG. 8, the variable names are "Category 1" and "Category 2", the category values are "AB" and "XY", each token 204 is "Category", "1", "Category", "2", "A", "B", "X", "Y", and the token vectors 205 are generated in the number corresponding to each token 204.

[0063] That is, in step ST21, as described in the first embodiment, the variable names and category values are token-split, and for each obtained token 204, vectorization is performed in units of tokens. Note that this is not the only way, and by combining all variable names and category values of the input data into one string data, token splitting may be performed on the combined string data. Also, after performing token splitting, the token vectorization unit 2 converts each token 204 into a learnable token embedding vector (token vector 205). At this time, the token embedding vectors may be common for variable names and category values. Further, the token vectorization unit 2 may perform the token splitting process independently for each variable name and category value without combining them into string data.

[0064] (Step ST22: Specifying Variable Names and Category Values) The variable name and category value identification unit 3 identifies, for each token 204, the variable from which the token vector 205 is derived, and the variable name or category value from which the token vector 205 is derived. To supplement, when the token vectorization unit 2 performs token splitting on the string data obtained by combining the variable name and the category value, it becomes unclear which variable each token vector 205 belongs to and whether it is from the variable name or the category value. On the other hand, in the subsequent vector generation unit 4, a variable name vector 207 and a value vector 208 are generated as vectors corresponding to the variable name and the category value respectively. Therefore, between the token vectorization unit 2 and the vector generation unit 4, the variable name and category value identification unit 3 identifies and records which variable each token vector 205 is derived from and whether it is from the variable name or the value.

[0065] (Step ST23: Vector Generation) Based on the result 206 identified in step ST23, the vector generation unit 4 generates a variable name vector 207 from the token vectors 205 derived from the variable name and a value vector 208 from the token vectors 205 derived from the category value for each variable. As for the process of generating a vector from a token vector, as described in the first embodiment, methods such as processing the set of token vectors 205 by a neural network and averaging the token vectors 205 can be appropriately used.

[0066] (Step ST30: Vector Combination) Based on the correspondence 203 obtained in step ST10, the vector combination unit 5 combines the variable name vector 207 and the value vector 208 generated in step ST23 to obtain a vector 209 corresponding to each variable.

[0067] According to the second embodiment as described above, the token vectorization unit 2 tokenizes the variable name and the category value for a variable including the category value and the variable name associated with the category value among the data 202 acquired by the data and correspondence acquisition unit 1, and generates a token vector 205 corresponding to each obtained token 204. The variable name / category value identification unit 3 identifies, for each token 204 obtained by the token vectorization unit 2, the variable from which the token vector 205 is derived and the variable name or category value from which the token vector 205 is derived. Based on the identified result 206, the vector generation unit 4 generates a variable name vector 207 from the token vectors 205 derived from the variable name and a value vector 208 from the token vectors 205 derived from the category value for each variable. Therefore, in addition to the effects described above, when the value associated with the variable name is a category value, the correspondence between the token vector 205 constituting the variable name or the category value and the variable name or the category value can be clarified.

[0068] Also, according to the second embodiment, the vector generation unit 4 may generate the value vector 208 by processing a set of token vectors 205 derived from the category value by a neural network. In this case, in addition to the effects described above, the value vector 208 can be generated using a neural network.

[0069] Also, according to the second embodiment, the vector generation unit 4 may generate the value vector 208 by averaging the token vectors 205 derived from the category value. In this case, in addition to the effects described above, the value vector 208 can be generated using averaging.

[0070] Also, according to the second embodiment, the vector generation unit 4 may generate the variable name vector 207 by processing a set of token vectors 205 obtained from the tokens 204 constituting the variable name by a neural network. In this case, in addition to the effects described above, the variable name vector 207 can be generated using a neural network.

[0071] Further, according to the second embodiment, the vector generation unit 4 may generate the variable name vector 207 by averaging the token vectors obtained from the tokens constituting the variable name. In this case, in addition to the effects described above, the variable name vector 207 can be generated using averaging.

[0072] <Third Embodiment> Next, an information processing apparatus according to the third embodiment will be described.

[0073] The third embodiment is a modification of the first embodiment and represents a specific example when the value associated with the variable name is a numerical value. When the value is a numerical value, the value vector 208 is a vector corresponding to the numerical value.

[0074] FIG. 9 is a block diagram showing an example of the configuration of an information processing apparatus according to the third embodiment. This information processing apparatus 10 has a configuration in which the vector generation unit 4 shown in FIG. 1 is divided into a variable name vector generation unit 41 and a numerical value vector generation unit 42.

[0075] Here, the variable name vector generation unit 41 generates a variable name vector 207 corresponding to each variable name 202a from the data 202 including the variable name 202a and the numerical value 202b associated with the variable name 202a.

[0076] The numerical value vector generation unit 42 generates a value vector 208 corresponding to the numerical value 202b associated with each variable name 202a from the data 202. The numerical value vector generation unit 42 may generate the value vector 208, which is the output of the neural network, by inputting the numerical value 202b into the neural network, for example. Further, the numerical value vector generation unit 42 may generate the value vector 208 by linearly transforming the numerical value 202b. Note that the present invention is not limited to this, and the numerical value vector generation unit 42 may generate the value vector 208 from the token vector 205 corresponding to the numerical value 202b in the same manner as in the second embodiment by regarding the numerical value 202b as a character string.

[0077] Other configurations are the same as those in the first embodiment.

[0078] Next, the operation of the information processing apparatus configured as described above will be described with reference to the flowchart of FIG. 10 and the schematic diagram of FIG. 11. In FIG. 10, steps ST20-1 and ST20-2 surrounded by a broken line are parts that are significantly changed from the operations shown in FIG. 2. Also, steps ST20-1 and ST20-2 may be executed in either order.

[0079] (Step ST10: Acquisition of Data and Corresponding Relationships) As shown in FIGS. 10 and 11, the data and corresponding relationship acquisition unit 1 receives the input of the tabular data 201. Further, the data and corresponding relationship acquisition unit 1 acquires data 202 including the variable name 202a and the value corresponding to the variable name 202a, and the corresponding relationship 203 between the variable name 202a and the value from the tabular data 201. Here, the value corresponding to the variable name 202a is the numerical value 202b.

[0080] (Step ST20-1: Generation of Variable Name Vector) The variable name vector generation unit 41 generates a variable name vector 207 corresponding to each variable name 202a from the data 202 including the variable name 202a and the numerical value 202b corresponding to the variable name 202a. Note that the process of generating the variable name vector 207 from the variable name 202a may be the method described in the first embodiment, or the method using the token vectorization unit 2, the variable name / category value specifying unit 3, and the vector generation unit 4 described in the second embodiment.

[0081] (Step ST20-2: Generation of Numerical Vector) The numerical vector generation unit 42 generates a value vector 208 corresponding to the numerical value 202b associated with each variable name 202a among the data 202. As the process of generating the value vector 208 from the numerical value 202b, the method described in the first embodiment can be appropriately used. Also, as the process of generating the value vector 208 from the numerical value 202b, by treating the numerical value 202b as a character string, a method using the token vectorization unit 2, the variable name / category value identification unit 3, and the vector generation unit 4 described in the second embodiment may be used.

[0082] (Step ST30: Vector combination) Based on the correspondence 203 acquired in step ST10, the vector combination unit 5 combines the variable name vector 207 generated in step ST20-1 and the value vector 208 generated in step ST20-2, and acquires a vector 209 corresponding to each variable.

[0083] As described above, according to the third embodiment, the variable name vector generation unit 41 generates a variable name vector 207 corresponding to each variable name 202a among the data 202 including the variable name 202a and the numerical value 202b associated with the variable name 202a. The numerical vector generation unit 42 generates a value vector 208 corresponding to the numerical value 202b associated with each variable name 202a among the data 202. Therefore, in addition to the effects described above, the variable name vector 207 and the value vector 208 can be generated in parallel.

[0084] <Fourth Embodiment> FIG. 12 is a block diagram showing an example of the hardware configuration of the information processing apparatus according to the fourth embodiment. The fourth embodiment is a specific example of the first to third embodiments, and is a form in which the information processing apparatus 10 is realized by a computer.

[0085] The information processing apparatus 10 includes, as hardware, a CPU (Central Processing Unit) 11, a RAM (Random Access Memory) 12, a program memory 13, an auxiliary storage device 14, and an input / output interface 15. The CPU 11 communicates with the RAM 12, the program memory 13, the auxiliary storage device 14, and the input / output interface 15 via a bus. That is, the information processing apparatus 10 of the present embodiment is realized by a computer having such a hardware configuration.

[0086] The CPU 11 is an example of a general-purpose processor. The RAM 12 is used by the CPU 11 as a working memory. The RAM 12 includes a volatile memory such as SDRAM (Synchronous Dynamic Random Access Memory). The program memory 13 stores programs for realizing each function of each part according to each embodiment in the computer. Further, as the program memory 13, for example, a ROM (Read-Only Memory), a part of the auxiliary storage device 14, or a combination thereof is used. The auxiliary storage device 14 stores data non-temporarily. The auxiliary storage device 14 includes a non-volatile memory such as an HDD (hard disc drive) or an SSD (solid state drive).

[0087] The input / output interface 15 is an interface for connecting to other devices. The input / output interface 15 is used, for example, for connecting to a keyboard, a mouse, and a display.

[0088] The program stored in the program memory 13 includes computer-executable instructions. When the program (computer-executable instructions) is executed by the CPU 11 which is a processing circuit, it causes the CPU 11 to execute a predetermined process. For example, when the program is executed by the CPU 11, it causes the CPU 11 to execute a series of processes described with respect to each part of FIGS. 1, 6, and 9. For example, the computer-executable instructions included in the program, when executed by the CPU 11, cause the CPU 11 to execute an information processing method. The information processing method may include each step corresponding to each function of each part described above. Further, the information processing method may appropriately include each step shown in FIGS. 2, 7, and 10. Further, instead of the program, a learned model stored in the program memory 13 may be executed.

[0089] As the learned model, for example, it may be a model for causing a computer to function so as to output each vector obtained by combining, with respect to the same variable name, a variable name vector corresponding to each variable name and a value vector corresponding to the value associated with each variable name, which is machine-learned based on multivariate data including a plurality of variable names and values associated with each of the plurality of variable names.

[0090] Also, for example, as the learned model, it may be a model that is machine-learned using teacher data in which multivariate data including a plurality of variable names and values associated with each of the plurality of variable names is associated with a correct label for classification processing or an output value for regression processing based on this multivariate data, and causes a computer to execute a process of receiving, as an input, the multivariate data at the current time point and a process of outputting the correct label or output value at the current time point based on the received multivariate data.

[0091] Also, as the learned model, for example, it may be a model for causing a computer to function so as to execute the processing by the neural network of each of the above-described embodiments.

[0092] A program or a trained model may be provided to an information processing apparatus 10, which is a computer, in a state stored in a computer-readable storage medium. In this case, for example, the information processing apparatus 10 further includes a drive (not shown) for reading data from the storage medium, and acquires a program from the storage medium. As the storage medium, for example, a magnetic disk, an optical disk (such as CD-ROM, CD-R, DVD-ROM, DVD-R), a magneto-optical disk (such as MO), a semiconductor memory, etc. can be appropriately used. The storage medium may be referred to as a non-transitory computer readable storage medium. Also, the program or the trained model may be stored in a server on a communication network, and the information processing apparatus 10 may download the program or the trained model from the server using the input / output interface 15.

[0093] The processing circuit that executes the program or the trained model is not limited to a general-purpose hardware processor such as the CPU 11, and a dedicated hardware processor such as an ASIC (Application Specific Integrated Circuit) may be used. The term processing circuit includes at least one general-purpose hardware processor, at least one dedicated hardware processor, or a combination of at least one general-purpose hardware processor and at least one dedicated hardware processor. In the example shown in FIG. 12, the CPU 11, the RAM 12, and the program memory 13 correspond to the processing circuit.

[0094] According to at least one of the embodiments described above, with respect to data in which variable names and values are associated, even if the variable names are changed, added, or deleted, the data can be made analyzable. This is the same for at least one of the above-described modifications.

[0095] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and the equivalent scope thereof.

Explanation of Signs

[0096] 10... Information processing device, 1... Data and correspondence acquisition unit, 2... Token vectorization unit, 3... Variable name / category value specification unit, 4... Vector generation unit, 5... Vector combination unit, 6... Processing unit, 61... Neural network, 201... Tabular data, 202... Data, 202a... Variable name, 202b... Numerical value, 203... Correspondence, 204... Token, 205... Token vector, 206... Specified result, 207... Variable name vector, 208... Value vector, 209... Vector, 210, 211... Set.

Claims

1. Data including a variable name and a value corresponding to the variable name, and a data and correspondence relationship acquisition unit that acquires the correspondence relationship between the variable name and the value, A vector generation unit that generates a variable name vector corresponding to each variable name and a value vector corresponding to the value associated with each variable name, A vector combination unit that combines the variable name vector and the value vector based on the correspondence relationship, An information processing apparatus comprising the above.

2. The value is a category value, The value vector is a vector corresponding to the category value, The information processing apparatus according to Claim 1.

3. For a variable including the category value and the variable name associated with the category value among the acquired data, tokenize the variable name and the category value, and generate a token vector corresponding to each obtained token, a token vectorization unit, For each of the tokens, a variable name / category value specifying unit that specifies the variable from which the token vector is derived and the variable name or the category value from which the token vector is derived, Further comprising: Based on the specified result, the vector generation unit generates the variable name vector from the token vector derived from the variable name and the value vector from the token vector derived from the category value for each variable, The information processing apparatus according to Claim 2.

4. The vector generation unit generates the value vector by processing a set of token vectors derived from the category value by a neural network, The information processing apparatus according to Claim 3.

5. The vector generation unit generates the value vector by averaging the token vectors derived from the category value, The information processing apparatus according to claim 3.

6. When the value is a categorical value or text data, the vector generation unit generates the value vector, which is a text vector, by performing text analysis on the value. The information processing apparatus according to claim 1.

7. The value is a numerical value. The value vector is a vector corresponding to the numerical value. The information processing apparatus according to claim 1.

8. The vector generation unit generates the value vector, which is the output of the neural network, by inputting the numerical value into the neural network. The information processing apparatus according to claim 7.

9. The vector generation unit generates the value vector by linearly transforming the numerical value. The information processing apparatus according to claim 7.

10. The data and correspondence acquisition unit has a plurality of variables in one sample along the row direction, and acquires the data and the correspondence from tabular data in which each of the plurality of variables has the variable name and the value along the column direction. The information processing apparatus according to any one of claims 1 to 9.

11. The vector generation unit generates the variable name vector by processing a set of token vectors obtained from tokens constituting the variable name by a neural network. The information processing apparatus according to any one of claims 1 to 9.

12. The vector generation unit generates the variable name vector by averaging token vectors obtained from tokens constituting the variable name. The information processing apparatus according to any one of claims 1 to 9.

13. The vector combining unit combines the variable name vector and the value vector by adding the variable name vector and the value vector for the same variable name, according to any one of claims 1 to 9.

14. The vector combining unit combines the variable name vector and the value vector by arranging the respective elements of the variable name vector and the value vector for the same variable name, according to any one of claims 1 to 9.

15. The vector combining unit inputs the variable name vector and the value vector for the same variable name into a neural network, and outputs a vector obtained by combining the variable name vector and the value vector, according to any one of claims 1 to 9.

16. A processing unit that performs classification processing or regression processing based on the set by inputting the set of the combined vectors into a neural network, The information processing apparatus according to any one of claims 1 to 9, further comprising the processing unit.

17. A data and correspondence relationship acquisition unit acquires data including a variable name and a value corresponding to the variable name, and a correspondence relationship between the variable name and the value, A vector generation unit generates a variable name vector corresponding to each variable name and a value vector corresponding to the value corresponding to each variable name, A vector combining unit combines the variable name vector and the value vector based on the correspondence relationship, An information processing method comprising:

18. A function of acquiring data including a variable name and a value corresponding to the variable name, and a correspondence relationship between the variable name and the value, A function of generating a variable name vector corresponding to each variable name and a value vector corresponding to the value corresponding to each variable name, A function of combining the variable name vector and the value vector based on the correspondence relationship, A program for realizing [it] on a computer.

Citation Information

Patent Citations

  • Interpretable tabular data learning using sequential sparse attention

    JP2022543393A