Database generation device, database generation method, and database generation program
By integrating multiple customer or consumer databases and using efficient connection databases and data generation technology, the problem of insufficient number of samples and items after database integration is solved, and comprehensive database generation with high accuracy and market compatibility is achieved.
Patent Information
- Application Number
- JP2021139603
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-08-30
AI Technical Summary
When multiple customer databases or consumer databases are integrated, the customer IDs or consumer IDs do not match, resulting in a small number of integrated database samples and limited items and metrics, which cannot be used as an effective comprehensive database.
By integrating the comprehensive database with other databases, a second comprehensive database is generated with an efficiently connected database, and through data extraction and generation, the number of samples is increased to match market statistics, and a comprehensive database with more samples and a wide range of items is generated.
It realizes the automatic generation of comprehensive databases corresponding to the number of samples and items, making them compatible with the real market and improving the accuracy and practicality of the database.
Smart Images

Figure 0007675431000001 
Figure 0007675431000002 
Figure 0007675431000003
Abstract
Description
[Technical field]
[0001] The present invention relates to the technical field of a database generation device, a database generation method, and a program for generating a database. More specifically, the present invention relates to a database generation device and a database generation method for integrating a plurality of different databases to generate an integrated database, and a program for the database generation device. [Background technology]
[0002] Generally, various companies manage one or more customer databases or consumer databases that each contain information about their customers or general consumers. These customer databases and consumer databases vary widely in the number of samples stored therein and in the items (indexes) used as databases, depending on their purposes, etc.
[0003] In addition, there are cases where it is necessary to integrate an external database with different attributes or items into a customer database or consumer database managed in-house, and generate a new customer database with a larger number of samples, etc. The following Patent Document 1 is an example of a prior art document that discloses a conventional technique for integrating such databases.
[0004] The conventional technology disclosed in Patent Document 1 has the objective of "providing a database integration device or the like that can execute join processing at a higher speed", and is configured as follows: "A reception unit of the database integration device receives a request to merge multiple pieces of data to be joined from a client, and a determination unit of the database integration device determines a database system that will execute the join processing based on join feasibility information that indicates whether each combination of database systems having databases each storing the data to be joined specified by the request can read the data to be joined from the database system to be combined and execute the join processing, and whether the data to be joined can be made to be read by the database system to be combined, a generation unit of the database integration device generates an execution plan for executing the join processing, and an execution unit of the database integration device transmits the request to the database system based on the execution plan." [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 6181250 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in general, when integrating multiple customer databases or consumer databases, there are cases where the customer ID or consumer ID whose data is included in one customer database or one consumer database does not match the customer ID or consumer ID whose data is included in another customer database or another consumer database. When integrating such customer databases or consumer databases, the only way was to integrate data on customer IDs or consumer IDs common to multiple customer databases and consumer databases. For this reason, the integrated database obtained as a result of the integration (an integrated database containing only data on the common customer ID or consumer ID) has a small number of samples and the items and indicators as a database are limited, resulting in a problem that only databases that cannot be used as an integrated database are generated. This problem leads to the problem that the more customer databases and consumer databases are integrated, the fewer customers and consumers that are common to each customer database and consumer database, making it useless as an integrated database.
[0007] The present invention has been made in consideration of the above-mentioned problems, and one example of the object of the present invention is to provide a database generation device and a database generation method, as well as a program for said database generation device, that are capable of automatically generating an integrated database with a large number of samples and a wide range of database items (indicators), even when integrating multiple databases. [Means for solving the problem]
[0016] In order to solve the above problem, the following claims are provided: 1 The invention described in An integrated database is obtained by integrating and expanding an integrated database related to product purchases using an integrating database, and further integrating another database that is different from the integrated database in at least one of the number of samples or items as a database into the integrated database related to product purchases. A database generating device, A database that may be expected to improve accuracy from the accuracy of the original database by integrating various items or indicators and includes a predetermined number of samples. Integrating a general-purpose connection database into the integrated database to generate a second integrated database. Ruto and the generated second integrated database. Based on the accuracy rate of The first step is accuracy. 2 Accuracy of the integrated database Based on the accuracy rate of The first step is accuracy. 1 When the accuracy is equal to or higher than the above Actually used for integration with integrated databases Extracting data of valid items from the second integrated database to generate an extracted second integrated database. Lottery extraction means; and the generated extracted second integrated database. Based on the accuracy rate of The first step is accuracy. 3 The accuracy is as above 2 and a data generation means for generating data when the accuracy is equal to or greater than the specified accuracy so as to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database, and a generation means for approximating data of the extracted second integrated database including the generated data to data of a market statistics database including statistical information of the actual market, to generate an extracted second integrated database with an increased number of samples. In order to solve the above-mentioned problems, the invention described in claim 7 is a database generating device that further integrates another database, which is an integrated database obtained by integrating and expanding an integrated database related to product purchases using an integrating database, and which differs from the integrated database related to product purchases in at least one of the number of samples or items as a database, into the integrated database related to product purchases. The database generating method is executed in the database generating device having an integrating means, an extracting means, a data generating means, and a generating means, and includes a step of integrating a general-purpose connecting database, which is a database whose accuracy may be expected to be improved from that of the original database by integrating it, into the integrated database by the integrating means to generate a second integrated database, and a step of extracting the second integrated database by extracting the second integrated database from the general-purpose connecting database, which is a database whose accuracy may be improved from that of the original database by integrating it, and which includes a general-purpose connecting database which includes various items or indicators and a predetermined number of samples. the extraction step of extracting, by the extraction means, data of valid items actually used in integration with the integrated database, from the second integrated database, to generate an extracted second integrated database, when a second accuracy, which is an accuracy based on the accuracy rate of the integrated database, is equal to or higher than the first accuracy, which is an accuracy based on the accuracy rate of the integrated database; the data generation step of generating, by the data generation means, data so as to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database, when a third accuracy, which is an accuracy based on the accuracy rate of the generated extracted second integrated database, is equal to or higher than the second accuracy; and the generation step of approximating, by the generation means, data of the extracted second integrated database including the generated data, to data of a market statistics database including statistical information of a real market, to generate an extracted second integrated database with an increased number of samples. In order to solve the above problem, the invention described in claim 8 provides a database generating device that further integrates another database, which is an integrated database obtained by integrating and expanding an integrated database related to product purchases using an integrating database, into the integrated database related to product purchases, and which is different from the integrated database in at least one of the number of samples or items as a database, and which is a database in which accuracy can be expected to be improved from that of the original database by integrating a general-purpose connection database that includes various items or indicators and a predetermined number of samples into the integrated database to generate a second integrated database, and a database generating device that generates a second integrated database based on the accuracy of the generated second integrated database, the database generating device further integrating another database, which is different from the integrated database in at least one of the number of samples or items as a database, into the integrated database related to product purchases, the database generating device further integrating a general-purpose connection database that includes various items or indicators and a predetermined number of samples into the integrated database, the database generating device further integrating a general-purpose connection database that includes ... When a certain second accuracy is equal to or higher than a first accuracy, which is an accuracy based on the accuracy rate of the integrated database, the device functions as an extraction means for extracting data of valid items actually used in integration with the integrated database from the second integrated database to generate an extracted second integrated database; when a third accuracy, which is an accuracy based on the accuracy rate of the generated extracted second integrated database, is equal to or higher than the second accuracy, the device functions as a data generation means for generating data to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database; and a generation means for approximating data of the extracted second integrated database including the generated data to data of a market statistics database including statistical information of real market, to generate an extracted second integrated database with an increased number of samples.
[0017] Claim 1, claim 7 or claim 8 According to the invention described in , No. 2. Integrated Database 2 The accuracy is the first 1 If the accuracy is higher than the accuracy, a second integrated database is generated. 3 The accuracy is 2 If the accuracy is higher than the specified accuracy, the number of samples in the extracted second integrated database is increased to match the number of samples in the integrated database, and then the number of samples is approximated to the data in the market statistics database to generate an extracted second integrated database with an increased number of samples. Thus, an integrated database that has the number of samples and items corresponding to the integrated database and also corresponds to the real market can be automatically generated as the extracted second integrated database with an increased number of samples.
[0018] In order to solve the above problem, the following claims are provided: 2 The invention described in claim 1 In the database generating device according to the 2 The accuracy is as above 1 When the accuracy is less than the above 3 The accuracy is as above 2When the precision is less than record The data of the valid items is extracted from the connection database. The above The first step to integration 2 The apparatus further includes an extraction means.
[0019] Claim 2 According to the invention described in claim 1 In addition to the effects of the invention described in 2 The accuracy is 1 When the accuracy is less than 3 The accuracy is 2 When the accuracy is below the required accuracy, data on valid items actually used in the integration with the integrated database is extracted from the connection database and used for the integration, so that a second integrated database with an increased sample number and higher accuracy can be automatically generated.
[0020] In order to solve the above problem, the following claims are provided: 3 The invention described in claim 1 or claims 2 In the database generating apparatus described in the above, the generated sample number increased extraction second integrated database Based on the accuracy rate of The first step is accuracy. 4 The accuracy is as above 3 When the precision is less than the limit, the data generating means is configured to regenerate the data to increase the number of samples in the extracted second integrated database.
[0021] Claim 3 According to the invention described in claim 1 or claims 2 In addition to the effects of the invention described in the above, the number of samples is increased and the second integrated database is extracted. 4 The accuracy is 3 When the accuracy is below the required accuracy, data for increasing the number of samples in the extracted second integrated database is regenerated, so that an extracted second integrated database with an increased number of samples and higher accuracy can be automatically generated.
[0024] In order to solve the above problem, the following claims are provided: 4 The invention described in claims 1 to 5 is3 In the database generation device described in any one of the above, the extraction of data for the effective items is configured to be performed using at least one of a principal component analysis method, a variable importance method, or a method using a SHapley Additive exPlanations (SHAP) library.
[0025] Claim 4 According to the invention described in claim 1 to claim 3 In addition to the effect of the invention described in any one of claims 1 to 5, data on effective items is extracted using at least one of the principal component analysis method, the variable importance method, or a method using the SHAP library, so that data on effective items that are more practical can be extracted.
[0028] In order to solve the above problem, the following claims are provided: 5 The invention described in claim 1 From claims 3 In the database generation device described in any one of the above, the device further includes a matching evaluation means for evaluating the degree of matching between the data contained in the generated sample number increased extraction second integrated database and the data contained in the connection database, and a notification means for notifying the degree of matching information indicating the evaluated degree of matching.
[0029] Claim 5 According to the invention described in claim 1 From claims 3 In addition to the effect of the invention described in any one of claims 1 to 5, the degree of match between the data contained in the generated sample number-increased extracted second integrated database and the data contained in the connection database is evaluated, and matching degree information indicating the evaluated degree of match is notified, so that the degree of match between the finally generated sample number-increased extracted second integrated database and the original connection database can be easily recognized.
[0030] In order to solve the above problem, the following claims are provided: 6 The invention described in claim 5In the database generation device described in the above, the matching evaluation means is configured to evaluate the matching degree using at least one of the mean / variance method, the histogram method, the method using statistical distribution of aggregated data, the S (Signal) / N (Noise) method, or an evaluation method using Cronbach's alpha coefficient.
[0031] Claim 6 According to the invention described in claim 5 In addition to the effects of the invention described above, the degree of match is evaluated using at least one of the mean / variance method, the histogram method, the statistical distribution method of aggregated data, the S / N method, or an evaluation method using Cronbach's alpha coefficient, so that the degree of match can be recognized more accurately. Effect of the Invention
[0032] As described above, according to the present invention, When the second precision of the second integrated database is equal to or greater than the first precision of the integrated database, an extracted second integrated database is generated. When the third precision of the extracted second integrated database is equal to or greater than the second precision, the number of samples in the extracted second integrated database is increased to match the number of samples in the integrated database, and then the number of samples is increased by approximating the data in the market statistics database. Generate the consolidated database.
[0033] Therefore, The second integrated database will have an increased number of samples and items corresponding to the integrated database and will also correspond to the real market. It can be generated automatically. [Brief description of the drawings]
[0034] [Figure 1] 1 is a block diagram showing a schematic configuration of a database generating device according to a first embodiment. [Diagram 2] 2 is a block diagram showing a schematic configuration of an extraction unit constituting the database generating device of the first embodiment. FIG. [Diagram 3] 5 is a flowchart showing a database generation process according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the contents of a database before a database generation process according to the first embodiment is executed. [Diagram 5] FIG. 4 is a diagram illustrating an example of the contents of a database after a database generation process according to the first embodiment is executed. [Figure 6] 13 is a flowchart showing a database generation process according to the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0035] Next, embodiments of the present invention will be described with reference to the drawings. Note that each embodiment described below is an embodiment in which the present invention is applied to a database generating device that integrates data from multiple different databases to generate a new integrated database.
[0036] (I) First embodiment First, a first embodiment of the present invention will be described with reference to Figs. 1 to 5. Fig. 1 is a block diagram showing the general configuration of a database generating device of the first embodiment, Fig. 2 is a block diagram showing the general configuration of an extraction unit constituting the database generating device, and Fig. 3 is a flowchart showing a database generating process of the first embodiment. Fig. 4 is a diagram illustrating an example of the contents of the database before the database generating process is executed, and Fig. 5 is a diagram illustrating an example of the contents of the database after the database generating process is executed. In Figs. 1 and 3, "database" is appropriately abbreviated as "DB."
[0037] As shown in FIG. 1, the database generation device S of the first embodiment is specifically realized by, for example, a personal computer, and is composed of a processing unit 1 consisting of a CPU or the like, a recording unit 2 consisting of an HDD (Hard Disk Drive) or an SSD (Solid State Drive) or the like, an operation unit 3 consisting of a keyboard, a mouse, etc., and a display 4 consisting of a liquid crystal display or the like.
[0038] The processing unit 1 is composed of an evaluation unit 10, an extraction unit 11, a generation unit 12, and an integration unit 13. The extraction unit 11 is further composed of a principal component analysis extraction unit 110, a variable importance extraction unit 111, and a SHAP extraction unit 112, as shown in FIG.
[0039] In this case, the evaluation unit 10, the extraction unit 11, the generation unit 12, and the integration unit 13 may be realized by a hardware logic circuit including a CPU constituting the processing unit 1, or may be realized in software by the CPU reading and executing a program corresponding to the database generation process of the first embodiment described later. Similarly, the principal component analysis extraction unit 110, the variable importance extraction unit 111, and the SHAP extraction unit 112 may be realized by a hardware logic circuit including a CPU constituting the extraction unit 11, or may be realized in software by the CPU reading and executing a program corresponding to the database generation process. Note that each of the above programs may be pre-recorded in the recording unit 2 and read by the CPU, or may be configured so that the CPU obtains and uses the program recorded in an external server device (not shown) via a network such as the Internet.
[0040] At this time, the evaluation unit 10 corresponds to an example of the "first evaluation means" of the present invention, an example of the "second evaluation means", an example of the "third evaluation means", an example of the "fourth evaluation means", an example of the "fifth evaluation means", an example of the "sixth evaluation means", an example of the "seventh evaluation means" and an example of the "eighth evaluation means", respectively, and the extraction unit 11 corresponds to an example of the "extraction means" and an example of the "second extraction means" of the present invention. Also, the generation unit 12 corresponds to an example of the "generation means" of the present invention, an example of the "data generation means" and an example of the "second generation means", respectively, and the integration unit 13 corresponds to an example of the "integration means" and an example of the "second integration means" of the present invention. Furthermore, the processing unit 1 corresponds to an example of the "matching degree evaluation means" of the present invention, and the display 4 corresponds to an example of the "notification means" of the present invention.
[0041] In the above configuration, the database generating device S is a database generating device that generates an integrated database 102 by integrating data of the main database 100 and data of the donor database 101 shown in Fig. 1. At this time, the data of the main database 100 and the donor database 101 to be integrated may be pre-recorded in the recording unit 2, or may be acquired via a network such as the Internet from an external server device (not shown) or the like every time the database generating process of the first embodiment is executed.
[0042] Here, the main database 100 in the first embodiment is, for example, a database of customers or general consumers that a company uses on a daily basis in sales and development work related to the product brand to which the product belongs, and is a database that belongs to the company or the relevant department. Such a main database 100 is basically a database that has a large number of samples (for example, more than tens of thousands of samples) and contains many items (indicators) related to the product brand, and also contains data on actual customers of the product. In contrast, the main database 100 does not contain much data (number of samples) for items (indicators) that are not directly related to the product brand.
[0043] In contrast to the main database 100 as described above, the donor database 101 of the first embodiment is a database that does not belong to the above-mentioned company, for example, created by an external research company or a department other than the above-mentioned department of the company. Such a donor database 101 has few items (indicators) related to specific products or product brands like the main database 100, and often does not have a large number of samples (for example, about 0 to 1,000 samples). However, the donor database 101 is a database that includes many items (indicators) that are not directly related to the above-mentioned product brands, such as items (indicators) related to the lifestyle of general purchasers (general purchasers including purchasers of products other than the above-mentioned products) and items (indicators) related to general values.
[0044] Then, the database generation device S integrates the data of the donor database 101 with the data of the main database 100 having the attributes described above, and diversifies the items (indicators) to generate an integrated database 102 that will be useful to the company.
[0045] More specifically, first, the recording unit 2 of the database generation device S temporarily records the data of the carefully selected donor database 103 and the sample generation carefully selected donor database 104 described later, which are generated in the database generation process of the first embodiment, and also records other data necessary for the database generation process, and outputs it to the processing unit 1 as necessary.
[0046] Meanwhile, the evaluation unit 10 of the processing unit 1 evaluates the accuracy of each database such as the main database 100 from the viewpoint of its accuracy rate, for example, by an evaluation method using a conventional cross validation method using a so-called confusion matrix. Here, regarding the accuracy rate, for example, when a purchaser whose data of a predicted purchase item is stored in the database as a sample actually purchases the predicted purchase item, the accuracy rate of the database including the sample is improved.
[0047] Next, the extraction unit 11 extracts effective indicators that are effective for generating the integrated database 102 from the items (indices) of the donor database 101.
[0048] Here, the method of extracting the above-mentioned effective index in the extracting unit 11 of the first embodiment will be described with particular reference to FIG.
[0049] The extraction of effective indices by the extraction unit 11 is performed by at least one of the principal component analysis extraction unit 110, the variable importance extraction unit 111, and the SHAP extraction unit 112 shown in Fig. 2. At this time, the principal component analysis extraction unit 110 extracts effective indices by a principal component analysis method similar to that used in the past. More specifically, the principal component analysis extraction unit 110 extracts items (indices) of principal components whose cumulative contribution rate is equal to or greater than a preset, changeable cumulative contribution rate threshold (e.g., 70%) and whose absolute value of the principal component loading amount is equal to or greater than a preset, changeable principal component loading amount threshold (e.g., 0.01) as effective indices.
[0050] On the other hand, the variable importance extraction unit 111 extracts effective indices by a conventional variable importance method. More specifically, the variable importance extraction unit 111 extracts items (indices) whose variable importance is equal to or greater than a variable importance threshold (e.g., 0.002) that is previously set so as to be changeable, as effective indices. The SHAP extraction unit 112 extracts effective indices by a conventional SHAP method. More specifically, the SHAP extraction unit 112 extracts items (indices) that fall within a SHAP threshold (e.g., top 20) previously set for the objective variable, as effective indices. At this time, if there are multiple items (e.g., product brands, etc.) that are objective variables, the SHAP extraction unit 112 adds them to the effective indices by sum integration (OR integration). It is preferable that the extraction result of the principal component analysis extraction unit 110, the variable importance extraction unit 111, or the SHAP extraction unit 112 is used as the extraction result of the extraction unit 11, for example, be previously set according to the attributes of the main database 100 or the attributes of the integrated database 102 to be generated.
[0051] Then, the effective indices output from at least one of the principal component analysis extraction unit 110, the variable importance extraction unit 111, and the SHAP extraction unit 112 are output as the extraction result by the extraction unit 11 through sum integration (OR integration). Then, the data of the effective indices as the extraction result is temporarily recorded in the recording unit 2 as the carefully selected donor database 103.
[0052] Next, returning to FIG. 1, in order to match the number of samples in the carefully selected donor database 103 with the number of samples in the main database 100 (for example, to make the number of samples in the carefully selected donor database 103 the same as the number of samples in the main database 100), the generation unit 12 of the processing unit 1 generates new data (samples) for the carefully selected donor database 103 using a new sample generation method using AI technology such as a conventional weight-back method or GAN (generative adversarial network) technology, in addition to the technology for which a patent application is pending by the inventors of the present invention (Patent Application No. 2020-085546), and adds this to the carefully selected donor database 103 to generate a sample generated carefully selected donor database 104, which is temporarily recorded in the recording unit 2.
[0053] As a result, the integration unit 13 integrates the data of the recorded sample generation carefully selected donor database 104 and the data of the original main database 100 in a conventional manner to generate the integrated database 102 of the first embodiment. In such an integrated database 102, the characteristics (advantages) of the donor database 101 are applied to the main database 100, thereby compensating for the shortcomings of the main database 100. As a result, an integrated database 102 that is extremely useful for the business activities of the company to which the main database 100 belongs (i.e., an integrated database with a large number of samples and a wide range of database items (indicators)) is automatically obtained.
[0054] The operations required to execute each of the above-mentioned functions are executed by the operation unit 3, and operation signals corresponding to the operations are output to the processing unit 1. In response to the operation signals, the processing unit 1 executes the series of functions described above. Information required to execute the functions is displayed, for example, on the display 4, and presented to the operator of the database generating device S.
[0055] Next, the database generation process executed in the database generation device S of the first embodiment will be specifically described with reference to FIGS.
[0056] The database generation process of the first embodiment executed by the database generation device S having the above-mentioned functions starts, for example, when a power switch (not shown) of the database generation device S is turned on.
[0057] When the database generation process is started, first, data of the main database 100 and data of the donor database 101 are acquired in the database generation process S. Next, the evaluation unit 10 of the processing unit 1 evaluates the accuracy of the main database 100 by the above-mentioned evaluation method based on the acquired data of the main database 100, and temporarily records the evaluation result as "Evaluation A" in the recording unit 2 (step S1).
[0058] Next, in parallel with step S1, the evaluation unit 10 evaluates the accuracy of the donor database 101 by the above-mentioned evaluation method based on the acquired data of the donor database 101, and temporarily records the evaluation result in the recording unit 2 as "evaluation B" (step S2). Next, the extraction unit 11 of the processing unit 1 extracts data corresponding to the effectiveness index from the data of the donor database 101 by the above-mentioned extraction method, generates the carefully selected donor database 103 using the extracted data, and temporarily records it in the recording unit 2 (step S3). After that, the evaluation unit 10 evaluates the accuracy of the carefully selected donor database 103 by the above-mentioned evaluation method based on the generated data of the carefully selected donor database 103, and temporarily records the evaluation result in the recording unit 2 as "evaluation C" (step S4).
[0059] Next, the processing unit 1 judges whether the evaluation C recorded in the recording unit 2 is equal to or greater than the evaluation B (step S5). If the judgment in step S5 indicates that the evaluation C is less than the evaluation B (step S5: NO), the extraction of the effective index in step S3 is deemed insufficient, and the process returns to step S3 again, and the extraction unit 11 re-extracts the effective index. On the other hand, if the judgment in step S5 indicates that the evaluation C is equal to or greater than the evaluation B (step S5: YES), the generation unit 12 of the processing unit 1 then performs sample generation (data generation) for the carefully selected donor database 103 using the above-mentioned generation method, generates the sample generated carefully selected donor database 104, and temporarily records it in the recording unit 2 (step S6).
[0060] Then, the integration unit 13 of the processing unit 1 integrates the recorded data of the sample generation carefully selected donor database 104 and the data of the original main database 100 in a conventional manner to generate an integrated database 102, which is temporarily recorded in the recording unit 2 (step S7). At this time, the integrated database 102 may be stored in an external server device (not shown). Next, the evaluation unit 10 evaluates the accuracy of the integrated database 102 by the evaluation method described above based on the recorded data of the integrated database 102, and temporarily records the evaluation result as "Evaluation D" in the recording unit 2 (step S8).
[0061] Next, the processing unit 1 judges whether the evaluation D recorded in the recording unit 2 is equal to or greater than the evaluation A (see step S1 above) (step S9). If the judgment in step S9 indicates that the evaluation D is less than the evaluation A (step S9: NO), the extraction of the effective index in step S3 included in the generation process of the current integrated database 102 is insufficient, and the process returns to step S3 again, where the extraction unit 11 further extracts the effective index. On the other hand, if the judgment in step S9 indicates that the evaluation D is equal to or greater than the evaluation A (step S9: YES), the processing unit 1 then evaluates the degree of match between the data in the integrated database 102 and the data in the donor database 101 at that time (step S10).
[0062] Here, the evaluation of the degree of match performed in step S10 is to evaluate the degree to which the data of the integrated database 102, which is the result of integrating the generalized donor database 101 with the company's own main database 100, matches the data of the donor database 101, that is, whether the database is more versatile. Specifically, the evaluation method in step S10 is the same as in the past, for example, at least one of the mean / variance method, the histogram method, the statistical distribution utilization method of the collected data, the S / N method, or the evaluation method using Cronbach's alpha coefficient. In this case, when the degree of match is determined using, for example, the mean / variance method, the degree of match is higher as the average value and the variance range match. The evaluation result of the match is then displayed (output) using, for example, the display 4, or is provided to the person in charge of the company to which the main database 100 belongs, together with the data of the integrated database 102 recorded in the recording unit 2 (step S11). If it is desired to increase versatility by, for example, increasing the degree of matching based on the evaluation result of the degree of matching in step S10 to bring the attributes of the integrated database 102 closer to the attributes of the donor database 101, it is preferable to perform the database generation process with stricter standards for determining whether the data is true or false in relation to the data in the donor database 101. On the other hand, if it is desired to increase the accuracy of the evaluation value in the evaluation unit 10 rather than the degree of matching, it is preferable to, for example, make the standards for determining the objective variables in each database stricter.
[0063] Thereafter, the processing unit 1 judges whether or not to end the database generation process of the first embodiment, for example, by an end operation on the operation unit 3 (step S12). If the judgment in step S12 is that the database generation process is to be ended (step S12: YES), the processing unit 1 ends the database generation process as is. On the other hand, if the judgment in step S12 is that the database generation process is to be continued, for example, for another main database 100 or another donor database 101 (step S12: NO), the processing unit 1 returns to the above steps S1 and S2 and continues the above-described process for the other main database 100 or the other donor database 101.
[0064] Next, a comparison between the main database 100 and the integrated database 102 as a result of the database generation process of the first embodiment will be specifically described with reference to Figures 4 and 5. Figures 4 and 5 show the results of the database generation process of the first embodiment being executed on the main databases 100 of not only one company but multiple companies.
[0065] First, as shown in Fig. 4, the main database 100 belonging to a certain company A records (accumulates) data such as attributes and whether or not the customer participated in a campaign implemented by company A in association with an ID indicating each customer, etc. In this case, although the participation in a campaign in particular is data unique to company A, data such as the customer's general movement history is not included (see the hatched portion in Fig. 4).
[0066] On the other hand, in the database generation process of the first embodiment described above, the donor database 101 of the first embodiment is applied to the main database 100 of Company A. The donor database 101 used in this case contains data indicating general lifestyles or values, such as the movement history and service usage history, as samples. When the database generation process of the first embodiment using such a donor database 101 is executed on the main database 100, the integrated database 102 obtained as a result may contain, as an example, data such as the movement history, for which data (samples) were not obtained because they are less relevant to the business activities of Company A, as shown in FIG. 5. As a result, the integrated database 102 that is extremely useful for the business activities of Company A is automatically obtained.
[0067] As described above, according to the database generation process by the database generation device S of the first embodiment, the accuracy of the donor database 101 is evaluated as B, the accuracy of the carefully selected donor database 103 is evaluated as C, and when evaluation C≧evaluation B, a sample generated carefully selected donor database 104 is generated and integrated with the main database 100 to generate the integrated database 102 (see steps S1 to S7 in FIG. 3). Therefore, the sample generated carefully selected donor database 104 generated based on the evaluation results of the accuracy of the donor database 101 and the accuracy of the carefully selected donor database 103 is integrated with the main database 100 to generate the integrated database 102, so that the integrated database 102 with a large number of samples and a wide range of items (indicators) as a database can be automatically generated.
[0068] According to a simulation by the inventors of the present invention, the integrated database 102 (containing the same number of samples as the samples in the main database and the number of variables being the sum of the number of variables in the main database 100 and the number of variables in the donor database 101) obtained by integrating the main database 100 (with a correct answer rate for evaluation B of 80% or more) related to general values, which contains samples in the thousands and has a larger number of variables than the main database 100, and is related to product brands, using the database generation process of the first embodiment, has a correct answer rate for evaluation D that is higher than that of the original main database 100 and approaches that of the donor database 101. From these, it can be seen that the database generation process of the first embodiment makes it possible to automatically generate an integrated database 102 that not only contains a large number of samples and a wide variety of items (indicators) as a database, but also has a dramatically improved accuracy (correct answer rate) compared to the original main database 100.
[0069] Furthermore, when evaluation C is less than evaluation B, data on the validity indicators is re-extracted and the carefully selected donor database 103 is re-generated, and the accuracy of the re-generated carefully selected donor database 103 is re-evaluated, so that a more accurate integrated database 102 can be automatically generated.
[0070] Furthermore, when the accuracy (rating D) of the generated integrated database 102 is less than the accuracy (rating A) of the original main database 100, data for valid indicators is re-extracted and the carefully selected donor database 103 is re-generated, and the accuracy of the re-generated carefully selected donor database 103 is re-evaluated (see steps S8 and S9 in Figure 3), so that an even more accurate integrated database 102 can be automatically generated.
[0071] Furthermore, if the above evaluation D is higher than the above evaluation A, the database generation process of the first embodiment is terminated and the contents of the integrated database 102 are confirmed (see step S9 in FIG. 3: YES), so that an integrated database 102 with higher accuracy than the main database 100 can be automatically generated.
[0072] Furthermore, since each evaluation by the evaluation unit 10 is performed using a cross-validation method that uses a confusion matrix, the accuracy of each database can be accurately evaluated.
[0073] Furthermore, since the extraction of effective indicators by the extraction unit 11 is performed using at least one of the principal component analysis method, the variable importance method, or a method using the SHAP library, it is possible to extract data on effective indicators that are more practical.
[0074] In addition, the degree of match between the data contained in the generated integrated database 102 and the data contained in the donor database 101 is evaluated and the evaluated degree of match is output (see steps S10 and S11 in Figure 3), so that the degree of match between the finally generated integrated database 102 and the original donor database 101 can be easily recognized.
[0075] Furthermore, since the evaluation of the degree of match in step S10 is performed using at least one of the mean / variance method, the histogram method, the statistical distribution method of aggregated data, the S / N method, or an evaluation method using Cronbach's alpha coefficient, the degree of match can be recognized more accurately. (II) Second embodiment
[0076] Next, a second embodiment which is another embodiment of the present invention will be described with reference to Fig. 6. Fig. 6 is a flowchart showing a database generation process of the second embodiment.
[0077] In the database generation process of the first embodiment described above, the main database 100 and the donor database 101 are integrated to generate the integrated database 102. In contrast, in the database generation process of the second embodiment described below, the integrated database 102 generated above is further expanded in a similar manner to the database generation process of the first embodiment to generate a database applicable to various markets (including so-called virtual markets on the Internet).
[0078] Incidentally, the hardware configuration of the database generation process of the second embodiment is basically the same as the hardware configuration of the database generation device S of the first embodiment, so in the following explanation, the same members as those in the database generation device S are given the same member numbers and detailed explanations are omitted. Furthermore, among the database generation process of the second embodiment, the same processes as those in the database generation process of the first embodiment described above are given the same step numbers and detailed explanations are omitted.
[0079] As shown in FIG. 6, the database generation process of the second embodiment executed in the database generation device of the second embodiment is started, similar to the database generation process of the first embodiment, for example, from the timing when the power switch of the database generation device of the second embodiment is turned on.
[0080] When the database generation process is started, first, data of the integrated database 102 generated by the database generation process of the first embodiment is obtained. At this time, the integrated database 102 provided to the database generation process of the second embodiment may be one in which one main database 100 and one donor database 101 are integrated by the database generation process of the first embodiment, or may be an integrated database generated by integrating one or more main databases 100 and one or more donor databases 101 by successively repeating the database generation process of the first embodiment multiple times.
[0081] Next, the evaluation unit 10 of the processing unit 1 of the second embodiment evaluates the accuracy of the integrated database 102 based on the acquired data of the integrated database 102 using an evaluation method similar to that of the database generation process of the first embodiment, and temporarily records the evaluation result as "evaluation a" in the recording unit 2 of the second embodiment (step S20).
[0082] Next, the integration unit 13 of the processing unit 1 of the second embodiment integrates the data of the integrated database 102 and the data of the connection database 124 of the second embodiment in a manner similar to that of the conventional method, generates a high-precision integrated database 120, and temporarily records it in the recording unit 2 (step S21).
[0083] Here, the connection database 124 is a database that may function as a sort of "glue" for connecting and integrating two databases, or that may be expected to improve the accuracy of the original database by integrating using the connection database, and is a highly versatile database that includes various items (indicators) and a predetermined number of samples.
[0084] Next, the evaluation unit 10 evaluates the accuracy of the high-precision integrated database 120 using the evaluation method described above based on the data of the high-precision integrated database 120 that has been generated and recorded, and temporarily records the evaluation result in the recording unit 2 as "evaluation b" (step S22).
[0085] Next, the processing unit 1 judges whether the evaluation c recorded in the recording unit 2 is equal to or greater than the evaluation b (step S23). If the judgment in step S23 indicates that the evaluation c is less than the evaluation b (step S23: NO), in order to improve the accuracy of the connection database 124, the extraction unit 11 of the processing unit 1 extracts data corresponding to effective indicators from the data of the connection database 124 at that time by the above-mentioned extraction method, and generates a new connection database 124 (with carefully selected items (indicators)) using the extracted data and temporarily records it in the recording unit 2 (step S27). This new connection database 124 is then subjected to the process in the above-mentioned step S21.
[0086] On the other hand, in the judgment of step S23, if the evaluation c is equal to or higher than the evaluation b (step S23: YES), the extraction unit 11 then extracts data corresponding to the effective index from the data of the recorded high-precision integrated database 120 by the above-mentioned extraction method in order to improve the accuracy of the high-precision integrated database 120, generates the carefully selected high-precision integrated database 121 using the extracted data, and temporarily records it in the recording unit 2 (step S24). Here, the generation of the carefully selected high-precision integrated database 121 (step S24) includes the generation of items (indexes) corresponding to a predetermined virtual market whose attributes or characteristics are similar to those of the integrated database 102, and the generation of a model corresponding thereto. After that, the evaluation unit 10 evaluates the accuracy of the carefully selected high-precision integrated database 121 by the above-mentioned evaluation method based on the data of the generated carefully selected high-precision integrated database 121, and temporarily records the evaluation result as "evaluation c" in the recording unit 2 (step S25).
[0087] Next, the processing unit 1 judges whether the evaluation c recorded in the recording unit 2 is equal to or greater than the evaluation b (see step S22) (step S26). If the judgment in step S26 finds that the evaluation c is less than the evaluation b (step S26: NO), it is determined that the extraction of effective indicators in the connection database 124 in step S27 was insufficient, and the process returns to step S27 again, where the extraction unit 11 further extracts effective indicators and provides them to the subsequent step S21. On the other hand, if the judgment in step S26 finds that the evaluation c is equal to or greater than the evaluation b (step S26: YES), then the generation unit 12 of the processing unit 1 generates samples (data generation) for the carefully selected high-precision integrated database 121 using the generation method described above (step S28).
[0088] Next, the processing unit 1 uses a predetermined market statistics database 122, which is a market statistics database containing statistical information in a real (non-virtual) market and has attributes or characteristics similar to those of the integrated database 102, to approximate the data of the carefully selected, high-precision integrated database 121 after the sample is generated to the data of the market statistics database 122 (step S29), and generates a sample generated carefully selected, high-precision integrated database 123 using the approximated data and temporarily records it in the recording unit 2 (step S30).
[0089] Then, based on the data of the generated sample generated carefully selected high-precision integrated database 123, the evaluation unit 10 evaluates the accuracy of the sample generated carefully selected high-precision integrated database 123 using the evaluation method described above, and temporarily records the evaluation result in the recording unit 2 as ``evaluation d'' (step S31).
[0090] Next, the processing unit 1 judges whether or not the evaluation d recorded in the recording unit 2 is equal to or greater than the evaluation c (see step S25 above) (step S32). If the judgment in step S32 is that the evaluation d is less than the evaluation c (step S32: NO), it is determined that the accuracy of the processes such as sample generation and approximation to the data of the market statistics database 122 in steps S28 to S30 was insufficient, and the processing unit 1 returns to step S28 again to repeat the subsequent processes.
[0091] On the other hand, if it is determined in step S32 that the evaluation d is equal to or greater than the evaluation c (step S32: YES), then the processing unit 1 evaluates the degree of match between the data in the sample generation carefully selected high-precision integrated database 123 at that time and the data in the connection database 124 and outputs the evaluation result in the same manner as in steps S10 and S11 in the database generation process of the first embodiment.
[0092] Thereafter, the processing unit 1 judges whether or not to end the database generation process of the second embodiment, for example, by an end operation by the operation unit 3 (step S33). If the judgment in step S33 is that the database generation process is to be ended (step S33: YES), the processing unit 1 ends the database generation process as is. On the other hand, if the judgment in step S33 is that the database generation process is to be continued, for example, with another integrated database 102 as the target (step S33: NO), the processing unit 1 returns to the above step S20 and continues the above-described process with the other integrated database 102 as the target.
[0093] The database generation process of the second embodiment described above can also provide the same effects as the database generation process of the first embodiment.
[0094] That is, the accuracy of the integrated database 102 is rated a, the accuracy of the high-precision integrated database 120 is rated b, and when rated b≧rated a, a carefully selected high-precision integrated database 121 is generated, and when rated c≧rated b, the accuracy of the carefully selected high-precision integrated database 121 is rated c, and a sample generated carefully selected high-precision integrated database 123 is generated by approximating it to the data of the market statistics database 122 (see steps S20 to S30 in FIG. 6). Thus, a sample generated carefully selected high-precision integrated database 123 that has the number of samples and items corresponding to the integrated database 102 and also corresponds to the real market can be automatically generated.
[0095] Furthermore, when evaluation b<evaluation a (see step S23 in FIG. 6: NO) or when evaluation c<evaluation b (see step S26 in FIG. 6: NO), data on valid items is extracted from the connection database 124 and is made available for integration with the integrated database 102 (see step S27 in FIG. 6), making it possible to automatically generate a more highly accurate sample generation and carefully selected high-precision integrated database 123.
[0096] Furthermore, when the accuracy evaluation d of the sample generated, carefully selected, high-precision integrated database 123 is less than the evaluation c (see step S32: NO in Figure 3), sample generation (data generation) is executed again as step S28, so that a sample generated, carefully selected, high-precision integrated database 123 with even higher accuracy can be automatically generated.
[0097] Furthermore, the degree of match between the data contained in the generated sample generation carefully selected high-precision integrated database 123 and the data contained in the connection database 124 is evaluated (see step S10 in Figure 6), and matching information indicating the evaluated degree of match is notified (see step S11 in Figure 6), so that the degree of match between the finally generated sample generation carefully selected high-precision integrated database 123 and the original connection database 124 can be easily recognized. [Industrial Applicability]
[0098] As described above, the present invention can be used in the field of database integration, and particularly when applied to the field of integrating databases with different numbers of samples and / or items (indicators), it can produce particularly significant effects. [Explanation of symbols]
[0099] 1 Processing section 2 Recording section 3 Control section 4. Display 10 Evaluation Section 11 Extraction part 110 Principal component analysis extraction part 111 Variable Importance Extraction Unit 112 SHAP extraction part 12 Generation part 13 Integration Department 100 Main Database 101 Donor Database 102 Integrated Database 103 Carefully Selected Donor Database 104 Sample Generation Carefully Selected Donor Database 120 High-precision integrated database 121 carefully selected high-precision integrated databases 122 Market Statistics Database 123 Sample Generation Carefully Selected High-Precision Integrated Database 124 Connection Database S Database generator
Claims
1. A database generation device which further integrates another database, which is an integrated database obtained by integrating and expanding an integrated database related to product purchases using an integrating database, and which differs from the integrated database related to product purchases in at least one of the number of samples or items as a database, comprising: an integration means for integrating a general-purpose connection database, which is a database that may be expected to improve accuracy from the accuracy of the original database by integration, and which includes various items or indicators and a predetermined number of samples, into the integrated database to generate a second integrated database; an extraction means for extracting data of effective items actually used for integration with the integrated database from the second integrated database to generate an extracted second integrated database when a second accuracy based on the accuracy of the generated second integrated database is equal to or greater than a first accuracy based on the accuracy of the integrated database; a data generating means for generating data so as to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database when a third accuracy, which is an accuracy based on a rate of correct answers of the generated extracted second integrated database, is equal to or greater than the second accuracy; a generating means for approximating data of the extracted second integrated database including the generated data to data of a market statistics database including statistical information of a real market, thereby generating an extracted second integrated database with an increased number of samples; A database generating device comprising:
2. 2. The database generating device according to claim 1, A database generation device further comprising a second extraction means for extracting data of the valid items from the connection database and providing it for the integration when the second accuracy is less than the first accuracy or when the third accuracy is less than the second accuracy.
3. 3. The database generating device according to claim 1, A database generation device characterized in that when a fourth accuracy, which is an accuracy based on the accuracy rate of the generated sample number-increasing extracted second integrated database, is less than the third accuracy, the data generation means regenerates the data to increase the number of samples in the extracted second integrated database.
4. 4. The database generating device according to claim 1, A database generating device, characterized in that the extraction of data on the effective items is performed using at least one of a principal component analysis method, a variable importance method, and a method using a SHAP (SHapley Additive exPlanations) library.
5. 4. The database generating device according to claim 1, a matching degree evaluation means for evaluating a matching degree between the data included in the generated sample number-increased extraction second integrated database and the data included in the connection database; a notification means for notifying a match degree information indicating the evaluated match degree; The database generating device further comprises:
6. 6. The database generating device according to claim 5, The database generation device is characterized in that the matching evaluation means evaluates the matching degree using at least one of the mean / variance method, the histogram method, the statistical distribution utilization method of aggregated data, the S (Signal) / N (Noise) method, or an evaluation method using Cronbach's alpha coefficient.
7. A database generating device which further integrates another database into the integrated database relating to product purchases, the other database having at least one of the number of samples or items as a database different from that of the integrated database obtained by integrating and expanding an integrated database relating to product purchases using an integrating database, the database generating method being executed in the database generating device which has an integration means, an extraction means, a data generating means, and a generating means, an integration step of integrating a general-purpose connection database, which is a database that may be expected to improve accuracy from the accuracy of the original database by integration, and which includes various items or indicators and a predetermined number of samples, into the integrated database by the integration means to generate a second integrated database; an extraction step of extracting, by the extraction means, data of effective items actually used for integration with the integrated database from the second integrated database to generate an extracted second integrated database when a second accuracy based on the accuracy of the generated second integrated database is equal to or greater than a first accuracy based on the accuracy of the integrated database; a data generating step of generating data by the data generating means so as to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database when a third accuracy, which is an accuracy based on a correct answer rate of the generated extracted second integrated database, is equal to or greater than the second accuracy; a generating step of approximating data of the extracted second integrated database including the generated data by the generating means to data of a market statistics database including statistical information of a real market, thereby generating an extracted second integrated database with an increased sample number; A database generating method comprising:
8. A computer included in a database generating device which further integrates another database, which is an integrated database obtained by integrating and expanding an integrated database related to product purchases using an integrating database, and which is different from the integrated database related to product purchases in at least one of the number of samples or items as a database, an integration means for integrating a general-purpose connection database including various items or indicators and a predetermined number of samples into the integrated database, the general-purpose connection database being a database that may be expected to improve accuracy from the accuracy of the original database through integration, to generate a second integrated database; an extraction means for extracting data of effective items actually used for integration with the integrated database from the second integrated database to generate an extracted second integrated database when a second accuracy, which is an accuracy based on a correct answer rate of the generated second integrated database, is equal to or higher than a first accuracy, which is an accuracy based on a correct answer rate of the integrated database; a data generating means for generating data so as to increase the number of samples in the extracted second integrated database to match the number of samples in the integrated database when a third accuracy, which is an accuracy based on a correct answer rate of the generated extracted second integrated database, is equal to or greater than the second accuracy; and a generating means for approximating data of the extracted second integrated database including the generated data to data of a market statistics database including statistical information of a real market, thereby generating an extracted second integrated database with an increased number of samples; A database generating program that functions as a database generating program.
Citation Information
Patent Citations
Fixing structure of resin molding
JP1986081250A
Table classification device, table classification method, and table classification program
JP2010039593A
Similarity calculation program and similarity calculation device
JP2011154540A
Database binding apparatus, database binding method, and database binding program
JP2019159837A