Automatic data linking across datasets

By projecting multiple databases into an equal-dimensional feature space and calculating the link matrix, the nearest neighbors of the datasets are identified and aggregated, solving the problem of insufficient feature utilization among datasets in cross-modal analysis and improving the performance and analysis efficiency of the computer system.

CN115769194BActive Publication Date: 2026-07-28INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2021-06-24
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively link multiple disjoint datasets of studies or patterns that do not have overlapping features, resulting in inefficient data utilization in cross-modal analysis. Furthermore, existing methods primarily rely on metadata and directed transfers of datasets, lacking the ability to perform cross-dataset analysis when there are no common attributes.

Method used

By generating baseline predictions and using a dimensionality reduction algorithm to project multiple databases into an equal-dimensional feature space, a link matrix is ​​calculated to identify nearby neighbor samples, and data is aggregated for cross-domain links to generate a user interface design to optimize computer system performance.

Benefits of technology

It enables cross-modal analysis in the absence of common features, improves the performance of datasets and the ability to analyze the influence of features, and optimizes the user interface design and performance of the computer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769194B_ABST
    Figure CN115769194B_ABST
Patent Text Reader

Abstract

A computer system, a computer program product, and a method are provided. The method includes generating a baseline prediction for each database of a plurality of databases, wherein the generated baseline prediction includes data points for each performance characteristic of a respective database of the plurality of databases; generating a linkage matrix in an equal-dimensional characteristic space using a dimension reduction algorithm based on measured differences between actual performance data points of each respective database and respective data points associated with the generated baseline prediction for each database of the plurality of databases; linking one or more databases of the plurality of databases based on measured differences within the generated linkage matrix reaching a predetermined threshold of measured performance; and generating a design for the computer system that is displayable through a user interface using the linked databases. The generated design is subsequently used for construction.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] This invention generally relates to the field of data extraction technology, and more specifically to data-driven insights for strategy formulation.

[0002] Data extraction is the process of retrieving data from various sources. Typically, data extraction is performed in the course of efforts to further process the extracted data, migrate the data to a database (such as a data warehouse or data lake), or further analyze the extracted data. Data extraction is the action or process of retrieving data from a data source for further data processing or data storage. After importing into an intermediate extraction environment, there is usually data transformation, and possibly metadata addition before exporting to another stage of the data workflow. Data-driven data extraction is a method for making strategic decisions based on data analysis and interpretation.

[0003] Data collection is the process of gathering and measuring information about variables of interest in an established, systematic manner that enables researchers to answer stated research questions, test hypotheses, and evaluate results. Data collection is an integral part of research in all fields, including the physical and social sciences, humanities, and business. The goal of all data collection is to capture quality evidence that allows analysis to lead to persuasive and credible answers to the questions posed, and to ensure that the focus of accurate and honest data collection remains consistent.

[0004] Data processing is the collection and manipulation of data items to produce meaningful information. In this sense, it can be considered a subset of information processing, as well as any change in information that is detectable to an observer. Automated data processing can refer to electronic data processing, which uses relatively simple, repetitive activities to process large amounts of similar information. Summary of the Invention

[0005] Embodiments of the present invention provide a computer system, a computer program product, and a method, comprising: generating a baseline prediction for each database in a plurality of databases, wherein the generated baseline prediction includes data points for each performance feature of the respective database in the plurality of databases; and generating a link matrix using a dimension reduction algorithm in an equal-dimensional feature space based on the difference between the actual performance data points of each respective database and the corresponding data points associated with the baseline prediction generated for each database in the plurality of databases.

[0006] Based on the difference in measurements within the generated link matrix reaching a predetermined threshold for measurement performance, one or more databases from the plurality of databases are linked; and the linked databases are used to generate a design for a computer system that can be displayed via a user interface. The generated design is then used to construct the computer system. Attached Figure Description

[0007] Figure 1 It is a functional block diagram depicting an environment having a computing device connected to or communicating with another computing device according to at least one embodiment of the present invention;

[0008] Figure 2 This is a flowchart illustrating the operational steps for generating an interpretation of an overlapping data link dataset according to at least one embodiment of the present invention;

[0009] Figure 3 This is a collection of examples illustrating cross-domain links according to at least one embodiment of the present invention;

[0010] Figure 4 This is a sample diagram illustrating the operational steps for generating a cross-domain interpretation according to at least one embodiment of the present invention;

[0011] Figure 5 This is a flowchart 500 illustrating the operational steps of performing cross-domain modeling on a linked dataset according to at least one embodiment of the present invention;

[0012] Figure 6 This is an illustration of an algorithm for selecting cross-domain neighbors according to at least one embodiment of the present invention;

[0013] Figures 7A-7E It is a collection of examples of outputs linked across datasets according to at least one embodiment of the present invention; and

[0014] Figure 8 An embodiment of the invention is depicted. Figure 1 A block diagram of the components of the computing system within the computing display environment. Detailed Implementation

[0015] Various embodiments of the present invention recognize the need to improve current data collection techniques aimed at leveraging two or more disjoint studies or patterns that do not have overlapping features, thereby improving the performance of each pattern and enabling cross-modal interpretation by linking these patterns at the data level. Embodiments of the present invention provide systems, methods, and computer program products for improving existing cross-modal linking. Currently, data is the most valuable asset to have in the growing trend of extracting data-driven insights for strategy formulation and process automation. Consequently, particularly in low-constraint settings, significant resources are allocated to collecting such data. The aggregated use of data obtained from multiple specific and ad hoc instances helps to encode better generalization during cross-modal analysis. Current cross-modal linking focuses on computer vision or image modalities, metadata dependent on the dataset, and a one-way transfer of data. Furthermore, existing data linking techniques are motivated by the need for entity reconciliation that requires common attributes across datasets or modalities. Some of the techniques commonly used include data augmentation, synthetic data generation, transfer learning (which includes domain adaptation), multi-task learning, and a few demonstration learning, record linking, and entity reconciliation and servicing. Various embodiments of the present invention improve upon current cross-dataset utilization services by linking disjoint datasets that have no common attributes across the datasets; and the procedure achieves this by projecting these datasets into an equal-dimensional feature space, which also serves as a dimensionality reduction, and computing a link matrix containing distance metrics between samples across different datasets, which allows cross-modal analysis to occur even in the absence of common features. Embodiments of the present invention retrieve input from at least one database, generate predictions as a model for each database, which are used as a baseline to compare performance achieved after cross-dataset linking, and then obtain data-level links between these databases by projecting the databases into an equal-dimensional feature space, where a link matrix is ​​computed to obtain neighbor samples across the databases. The model is executed on the linked databases, and the prediction performance is compared to the baseline.

[0016] Figure 1 This is a functional block diagram of a computing environment 100 according to an embodiment of the present invention. The computing environment 100 includes a computing device 102 and a server computing device 108. The computing device 102 and server computing device 108 may be a desktop computer, a laptop computer, a dedicated computer server, a smartphone, or any other computing device known in the art. In some embodiments, the computing device 102 and server computing device 108 may represent computing devices that utilize multiple computers or components to act as a single seamless resource pool when accessed via network 106. Generally, the computing device 102 and server computing device 108 may represent any electronic device or combination of electronic devices capable of executing machine-readable program instructions, as per [reference to...]. Figure 8 More detailed description.

[0017] Computing device 102 may include program 104. Program 104 may be a standalone program on computing device 102. In another embodiment, program 104 may be stored on server computing device 108. In this embodiment, program 104 retrieves input from at least one database, generates predictions as models for each database, links model predictions to generate a baseline, compares model predictions in an equal-dimensional feature space using a dimensionality reduction algorithm, calculates a link matrix based on the distance between database samples in the reduced feature space, and measures performance improvement and impact analysis by performing cross-database modeling on the linked databases. For example, program 104 receives input data from multiple modalities, where each modality includes specific features such as nutrition, transportation, and healthcare; generates performance predictions based on the inputs received for each modality; projects these modalities into an equal-dimensional feature space using a dimensionality reduction technique; evaluates a distance metric for each pair of samples extracted from different datasets or modalities; identifies the nearest neighbors of samples across modalities using the distance metric; aggregates extracted data from nearest neighbors and links them to the original feature space; generates output modeling and predictions based on the linked datasets; and measures performance improvement and cross-modal impact analysis for common features.

[0018] Network 106 may be a local area network (“LAN”), a wide area network (“WAN”) (such as the Internet), or a combination of both; and it may include wired, wireless, or fiber optic connections. Typically, network 106 may be any combination of connections and protocols that support communication between computing device 102 and server computing device 108 (specifically, program 104 according to a preferred embodiment of the invention).

[0019] Server computing device 108 may include program 104 and may communicate with computing device 102 via network 106.

[0020] Figure 2 This is a flowchart illustrating the operational steps of measuring performance prediction using a dimensionality reduction algorithm according to at least one embodiment of the present invention.

[0021] In step 202, program 104 retrieves input from at least one database. In this embodiment, program 104 receives opt-in / opt-out permission from the user to gain access to a specific modality, which allows program 104 to receive input from that modality. For example, program 104 receives user transformation data from a transformation modality or transformation database.

[0022] In step 204, program 104 performs dimensionality reduction. In this embodiment, program 104 reduces the dimensionality of the retrieved input by projecting the two datasets onto an isodimensional space, which enables head-to-head comparisons of samples from different datasets that potentially have different types and numbers of features. Isodimensionality is applied to objects that are nearly identical in size or expand in multiple directions and across any number of dimensions. In this embodiment, program 104 performs dimensionality reduction by utilizing the following as a function:

[0023] (1)

[0024] (2)

[0025] For equations (1) and (2), D1 and D2 represent disjoint databases and D 12 and D 21 Indicates the database to which the link is located. The dimensions of the feature space are represented by K, L, and R, which represent D1, D2, and D3, respectively. 12 Or D 21 The dimension of the feature space, and because of the feature dimension, R is equal to or less than K or L. In this embodiment, D 12 and D 21 Both are linked datasets, but the applications of the links differ; in D 12 The baseline dataset is D1, and more data is extracted from D2 and linked to baseline D1; in D... 21 The baseline dataset is D2, and more information is extracted from D1 and linked to baseline D2. In this embodiment, procedure 104 utilizes equations (1) and (2) to perform dimensionality reduction, which projects at least two datasets into an equal-dimensional feature space that allows head-to-head comparisons of samples across datasets. For example, procedure 104 performs dimensionality reduction based on principal component analysis using the first R feature vectors; feature importance scores using the first R features in both positive and negative directions; and a dense autoencoder using the output of an R-dimensional encoder.

[0026] In step 206, program 104 calculates the link matrix. In this embodiment, program 104 calculates the proximity link matrix based on the performed dimensionality reduction. In this embodiment, program 104 calculates the link matrix to link samples across two domains, and the matrix contains the distance between each pair of samples. In this embodiment, program 104 calculates the link matrix by utilizing the following as a function:

[0027] (3)

[0028] (4)

[0029] Regarding equations (3) and (4), i is a variable representing the range from 1 to N (the index of sample d in D1); j is a variable representing the range from 1 to M (the index of sample D2); and D ij Indicating D1 With D2 The Euclidean distance between them. In this embodiment, a domain is defined as a specific cluster of information with unique characteristics. For example, procedure 104 calculates a link matrix that details the distance between the transport domain and the nutrient domain. In another embodiment, procedure 104 can use the calculated link matrix to identify nearest neighbors of samples across datasets and link information between these nearest neighbors. In yet another embodiment, procedure 104 integrates domain expert knowledge within the link matrix calculation.

[0030] In step 208, program 104 selects cross-dataset neighbors. In this embodiment, program 104 selects nearest neighbors for each sample by soft-matching across the dataset based on the link matrix of the nearest cross-domain neighbors. In this embodiment, program 104 defines the dataset and domain as interchangeable. In this embodiment, program 104 selects cross-domain neighbors by utilizing the following as a function:

[0031] (5)

[0032] (6)

[0033] Regarding equations (5) and (6), Indicates multiple neighbors and and This indicates the selected cross-domain neighbor. The neighbors in the middle contain samples from D2, and similarly, in The neighbors in the dataset include samples from D1. In this embodiment, procedure 104 utilizes equations (5) and (6) by soft-matching cross-domain neighbors specific to each sample and each cross-domain neighbor. For example, procedure 104 selects domain neighbors based on common features, weighted features based on t-scores, and average features from across the domains. In another embodiment, procedure 104 uses this selection method to select cross-domain neighbors that do not share common data points or common features. In yet another embodiment, procedure 104 selects cross-domain neighbors based on feature scaling of the standard deviation of the importance scores of the variables, which affects the dataset distance to the other dataset.

[0034] In step 210, program 104 aggregates the extracted data. In this embodiment, program 104 extracts data from selected cross-domain neighbors, averages the extracted data based on features, and aggregates the average to generate aggregated data extract. In this embodiment, program 104 aggregates the extracted data using the following formula:

[0035] (7)

[0036] (8)

[0037] Regarding equations (7) and (8), and This represents the aggregated data extraction, along with other variables represented in the preceding equations. For example, procedure 104 extracts data from multiple transport domains, averages the data based on distance, and aggregates the average distances to generate a selected cross-domain aggregated distance. In another embodiment, procedure 104 aggregates the extracted data even if there are no common data points or features.

[0038] In step 212, program 104 performs cross-dataset linking. In this embodiment, program 104 performs cross-domain linking by concatenating the original data and aggregating the extracted data. In this embodiment, program 104 performs cross-domain linking to measure performance improvements of features across domains or datasets and to perform cross-domain impact analysis. In this embodiment, program 104 performs cross-domain linking to achieve performance improvements in each domain by utilizing link information from selected cross-domain neighbors. In this embodiment, program 104 provides a cross-domain interpretation that can analyze the impact of features in one domain on the outputs of different domains. In this embodiment, program 104 utilizes large, unlabeled domains by performing cross-domain linking. For example, program 104 performs cross-domain linking on the transportation domain, measures the transportation outputs of the linked cross-domains, and analyzes the impact of user interventions on the measured outputs based on the calculation of the link matrix. In another embodiment, program 104 applies cross-domain linking in a data-driven manner in low or limited resource settings (e.g., less precise acquisition equipment, unsatisfactory domain experts, limited data availability, and computational resources). In another embodiment, program 104 measures the performance improvement of one or more linked databases and measures the effect of previously unmeasured residual features based on the performance improvement of features measured within one or more linked databases.

[0039] In step 214, program 104 generates a design for a user interface. In this embodiment, program 104 generates the design for the user interface based on the measured performance improvement compared to a baseline. In this embodiment, program 104 generates a design for the user interface that allows the user to identify improved scores for specific features of computing device 102 based on cross-dataset links. In this embodiment, program 104 generates a design for a user interface including a tactile sensor that receives input from the user. In this embodiment, program 104 receives input via the tactile sensor and optimizes the generated design for the user interface. In this embodiment, program 104 optimizes the generated design for the user interface by ranking the aggregated extracted data according to at least one category, where the category can be a sample number, identified location, dataset, size, etc. For example, program 104 generates a smart car user interface that displays the measured performance of features of the smart car based on cross-domain links between a transportation dataset and a nutrition dataset. In another embodiment, program 104 receives input from a tactile sensor of a computing device for linked databases among multiple databases; generates a design for a first computer system based on the input received from the tactile sensor of the computing device; modifies the generated design for the first computer system based on the input received from the tactile sensor of the computing device; optimizes the generated design for the first computer system based on the modified design; displays the optimized design for the first computer system based on a combination of the modified design, the input received from the tactile sensor of the computing device, and the linked databases; and constructs a second computing system based on the optimized design of the computer system that can be displayed through a user interface.

[0040] In step 216, program 104 constructs a computer system. In this embodiment, program 104 constructs the computer system based on cross-domain links and the design of the generated user interface. In this embodiment, program 104 constructs the computer system to improve the performance of features of computing device 102. In another embodiment, program 104 constructs the computer system based on input received from tactile sensors from the design of the generated user interface. In yet another embodiment, program 104 constructs the computer system based on optimized cross-domain links combined with input received from tactile sensors from the design of the generated user interface, and the constructed computer system improves the performance of features of computing device 102.

[0041] Figure 3 Several examples of cross-domain connections according to at least one embodiment of the present invention are illustrated.

[0042] Example 302 depicts an example of modeling independently without using cross-domain links. In Example 302, program 104 analyzes the data generated by D1 and D2. NAt least one dataset is represented. In Example 302, program 104 generates a dataset consisting of M1 and M... N At least one model is represented. In Example 302, procedure 104 predicts the outcome based on the model associated with O1 and O. N The result is indicated.

[0043] Example 304 illustrates an example of a cross-domain link before the model. In Example 304, Program 104 analyzes the cross-domain links between D1 and D2. N This represents at least two databases. In Example 304, program 104 performs a cross-domain join based on the analysis of the databases. In Example 304, program 104 generates a model represented by M based on the cross-domain join. In Example 304, program 104 predicts the model represented by O based on the generated model. new (If the link has a single result) or O 1:N (If each dataset has different results) represents the result.

[0044] Example 306 illustrates an example of cross-domain linking after a model. In Example 306, Program 104 analyzes the cross-domain linking by D1 and D2. N The representation refers to at least two databases. In Example 306, program 104 generates a dataset consisting of M1 and M2 based on the analysis of at least two datasets. N At least two models are represented. The outputs of these models are used as a baseline to compare the performance improvement of the measurement achieved using cross-domain linking. In Example 306, procedure 104 performs cross-domain linking based on the generated models. In Example 306, procedure 104 predicts the results based on the cross-domain linking.

[0045] Figure 4 This is a sample diagram 400 illustrating the operational steps for generating a cross-domain interpretation according to at least one embodiment of the present invention.

[0046] In step 402, program 104 analyzes at least two datasets. In this embodiment, program 104 analyzes at least two datasets represented by D1 and D2. In this embodiment, program 104 analyzes at least two disjoint datasets by tracking the sets of elements divided into multiple disjoint subsets and providing near-constant-time operations to add new sets, merge existing sets, and determine whether elements are in the same set.

[0047] In step 404, program 104 generates at least two predictions. In this embodiment, program 104 generates at least two domain-specific predictions, and these predictions may be based on a model generated from the analysis of disjoint datasets.

[0048] In step 406, the program performs data-level cross-domain links. In this embodiment, program 104 performs data-level cross-domain links as disclosed in the previous graph. In this embodiment, program 104 performs data-level cross-domain links by retrieving data from at least two disjoint datasets; performing dimensionality reduction; calculating the link matrix; selecting cross-domain neighbors; and aggregating the extracted data.

[0049] In step 408, program 104 generates at least two cross-domain datasets. In this embodiment, program 104 generates datasets based on data-level cross-domain links from D... 12 and D 21 This represents at least two cross-domain datasets. In this embodiment, the cross-domain datasets are generated from linked disjoint datasets.

[0050] In step 410, program 104 generates at least two cross-domain datasets and predictions. In this embodiment, program 104 generates at least two cross-domain datasets and predictions based on the analysis of the generated cross-domain datasets and M. 12 and M 21 This represents their associated prediction models. In this embodiment, program 104 uses D 12 and D 21 As data and their models M 12 and M 21 Generate predictions for at least two domains. Compare the prediction performance with the baseline output.

[0051] In step 412, the program generates a cross-domain interpretation. In this embodiment, program 104 generates a cross-domain interpretation based on the generated cross-domain prediction. In this embodiment, program 104 generates a cross-domain interpretation based on the output of analysis of disjoint datasets and analysis of the generated cross-domain prediction. For example, program 104 generates a cross-domain interpretation of transportation data for a linked transportation dataset. In another embodiment, program 104 generates a domain-specific interpretation based on the generated domain-specific prediction.

[0052] Figure 5 This is a flowchart 500 illustrating the steps of performing cross-domain modeling on a linked dataset according to at least one embodiment of the present invention.

[0053] In step 502, procedure 104 generates a baseline. In this embodiment, procedure 104 generates a baseline based on domain-specific predictions. In this embodiment, the domain-specific predictions correspond to each dataset, thus making the predictions dataset-specific. For example, procedure 104 generates a driving speed baseline based on a transportation dataset.

[0054] In step 504, program 104 performs dimensionality reduction. In this embodiment, program 104 projects these datasets into an equal-dimensional feature space and performs dimensionality reduction. In this embodiment, program 104 projects these datasets into an equal-dimensional feature space to compare the distances between samples in different datasets that were originally represented by feature spaces of different types and dimensions.

[0055] In step 506, procedure 104 identifies important features in each dataset. In this embodiment, procedure 104 identifies important features based on the t-score calculated for each dataset after modeling to utilize both positive and negative features. In this embodiment, procedure 104 calculates the t-score for each dataset by evaluating the significance or importance score of each feature in the resulting prediction.

[0056] In step 508, procedure 104 matches samples across datasets. In this embodiment, procedure 104 matches samples across datasets using at least one machine learning technique to provide unsupervised dimensionality reduction. In this embodiment, procedure 104 matches samples across datasets using principal component analysis and an autoencoder, wherein principal component analysis and an autoencoder provide unsupervised dimensionality reduction that can compute the link matrix using a distance metric.

[0057] In step 510, program 104 calculates the link matrix. In this embodiment, program 104 calculates the link matrix as discussed in the previous figures. In this embodiment, program 104 calculates the link matrix by calculating pairwise distances between samples. In this embodiment, program 104 calculates the link matrix in response to performing dimensionality reduction.

[0058] In step 512, procedure 104 selects nearest neighbors across databases. In this embodiment, procedure 104 uses the distance metric in the link matrix to select samples in other databases that are nearest neighbors to the baseline database; that is, samples with a lower distance metric are assumed to be nearest neighbors. In this embodiment, procedure 104 applies a soft-neighbor database selection method. In this embodiment, procedure 104 uses a computed association matrix and data identified in the neighbor databases to aggregate and concatenate each sample to select a neighbor database.

[0059] In step 514, procedure 104 generates cross-domain modeling. In this embodiment, procedure 104 generates cross-domain modeling based on linked datasets. In this embodiment, procedure 104 generates cross-domain modeling and provides result predictions for measurement comparison with baseline performance without links. In this embodiment, procedure 104 generates cross-domain modeling to analyze the cross-dataset impact of features.

[0060] In step 516, procedure 104 integrates domain knowledge. In this embodiment, procedure 104 integrates domain knowledge within the generated cross-domain model based on expert knowledge. In this embodiment, procedure 104 aggregates domain knowledge within the generated cross-domain model based on expert knowledge, and this expert knowledge is integrated into the link matrix calculation and nearest neighbors are selected.

[0061] Figure 6 This is an illustration of an algorithm 600 for selecting cross-domain neighbors according to at least one embodiment of the present invention.

[0062] In this embodiment, program 104 uses algorithm 600 to select cross-domain neighbors and extract aggregated link information. In this embodiment, algorithm 600 is described by the following equation:

[0063] (9)I - ←argmin(d i1 ,d i2 ,…,d iM );

[0064] (5) (7)

[0067]

[0068] (10)I - ←argmin(d 1j ,d 2j ,…,d Nj );

[0069] (6) (8)

[0072]

[0073] Regarding equations (9) and (10), C represents cross-domain neighbors, A represents aggregated neighbor data, U represents cross-domain links, S represents dimensionality reduction, Q represents the number of nearby neighbors, and This indicates a cascading operation. In this embodiment, algorithm 600 uses equations (5) and (7) to achieve the output of cross-domain links across a specific dataset. In this embodiment, the algorithm uses equations (6) and (8) to achieve the output of cross-domain links for different specific datasets. In response to achieving the output of cross-domain links, and in this embodiment, algorithm 600 calculates a measured performance improvement for the generated cross-domain linked datasets. In this embodiment, algorithm 600 calculates the linked dataset D as follows. 12 and D 21 :

[0074] (11)

[0075] (12)

[0076] Regarding equations (11) and (12), D 12 and D 21 The dataset represents the cross-domain linked dataset. In this embodiment, algorithm 600 uses equations (11) and (12) to select the cross-domain linked dataset. In this embodiment, algorithm 600 selects the cross-domain linked dataset, and the dimension of the feature space in the linked dataset becomes some of the individual dimensions of the disjoint datasets K and L, as follows:

[0077] (13)

[0078] Regarding formula (13), This indicates that the feature space of the linked dataset is the union of the feature spaces of the disjoint datasets; that is, D1 has the original K dimensions, and D2 has the L-dimensional feature space. Therefore, the linked datasets D... 12 and D 21 It will be the sum of these feature dimensions.

[0079] Figure 7 is a collection of examples of outputs from cross-dataset links according to at least one embodiment of the present invention.

[0080] Figure 7AExample 700 is an implementation of cross-dataset linking. In this embodiment, Example 700 identifies a country, determines a number of samples associated with the identified country, extracts a percentage of data from the determined number of samples, and displays a number of features. In this embodiment, Example 700 displays similar information from the cross-domain linked dataset to verify the cross-dataset linking. In this embodiment, Example 700 displays data from a dataset that identifies the country as Ethiopia, determines that 328 samples are associated with Ethiopia, extracts 2.43% of deaths from these 328 samples, and displays 67 features. In this embodiment, Example 700 displays data from disjoint datasets identifying countries such as Burkina Faso, Ghana, Kenya, Nigeria, and Ethiopia; it is determined that 10,364 samples are associated with Burkina Faso, 4,294 samples with Ghana, 14,949 samples with Kenya, 20,192 samples with Nigeria, and 7,193 samples with Ethiopia; 4.00% mortality is extracted from the 10,364 samples, 3.98% mortality from the 4,294 samples, 3.09% mortality from the 14,949 samples, 6.02% mortality from the 20,192 samples, and 4.74% mortality from the 7,193 samples; and 384 features are displayed for Burkina Faso, 319 features for Ghana, 316 features for Kenya, 368 features for Nigeria, and 329 features for Ethiopia. In this embodiment, Example 700, for example, shows a performance monitoring and accountability dataset linked across domains with demographic and health survey datasets, which are collected to understand maternal and infant health.

[0081] Figure 7B Example 702 is an example of measuring the performance improvement of cross-dataset linking before linking. In this embodiment, Example 702 identifies the classifier, determines the type of dimensionality reduction, and calculates the percentage performance improvement for separate domain datasets, cross-domain linked datasets, and random datasets. In this embodiment, Example 702 identifies the logistic regression classifier; determines the type of feature importance for which dimensionality reduction is performed; and calculates a 50.03% improvement percentage for the separate dataset, a 55.31% improvement percentage for the linked dataset, and a 51.39% improvement percentage for the random dataset. In this embodiment, Example 702 illustrates three types of dimensionality reduction, feature importance, principal component analysis, and autoencoders.

[0082] Figure 7CExample 704 is an example of measuring the performance improvement of linking across datasets after linking. In this embodiment, Example 704 identifies the classifier, determines the type of dimensionality reduction, and calculates the percentage performance improvement for separate domain datasets, cross-domain linked datasets, and random datasets. In this embodiment, Example 704 identifies the random field classifier; determines the type of autoencoder that performs dimensionality reduction; and calculates a 67.23% improvement percentage for separate datasets and a 99.99% improvement percentage for linked datasets. In this embodiment, Example 704 shows three types of dimensionality reduction, feature importance, principal component analysis, and autoencoders.

[0083] Figure 7D Example 706 is a validation of cross-domain linking. In this embodiment, Example 706 validates the percentage improvement of a dataset linked across multiple countries using cross-domain linking. In this embodiment, Example 706 identifies a random field classifier; determines the type of principal component analysis to perform dimensionality reduction; and calculates a 47.00% improvement percentage for the separate dataset, a 63.49% improvement percentage for Ethiopia, a 64.86% improvement percentage for Burkina Faso, a 74.12% improvement percentage for Ghana, a 56.25% improvement percentage for Kenya, and a 68.89% improvement percentage for Nigeria. In this embodiment, Example 706 shows the numerical difference in performance improvement of the separate dataset compared to the performance improvement of the cross-domain linked dataset.

[0084] Figure 7E Example 708 is a bar chart validation of cross-domain links. In this embodiment, example 708 depicts the validation in a series of bar charts. In this embodiment, example 708 provides performance improvements measured in relation to features of each separate domain. In this embodiment, example 708 maintains an x-axis for features and a y-axis for importance scores; for dataset 1, parity is not the feature being analyzed, but size is a feature with an improvement score of 0.19; for dataset 2, parity is a feature with an improvement score of 0.13, but size is not a feature being analyzed; and for the linked dataset 12, parity is a feature with an improvement score of 0.13, and size is a feature with an improvement score of 0.06. In this embodiment, example 708 also provides performance improvements measured in relation to features of cross-domain linked datasets. In this embodiment, example 708 validates the percentage improvement and depicts the validation via a bar chart. In this embodiment, example 708 illustrates a bar chart showing D1, D2, and D... 12 The top ten key features used for accurate outcome prediction and their improved scores, which are carried by specific features to determine the relative weight of the outcome.

[0085] Figure 8 An embodiment of the invention is depicted. Figure 1 A block diagram of the components of the computing system within the 800 computing system. It should be understood that... Figure 8 This illustration provides only one possible implementation and does not imply any limitation regarding the environment in which different embodiments may be implemented. Many modifications can be made to the depicted environment.

[0086] In specific embodiments of the invention, the programs described herein are identified based on applications for which they are implemented. However, it should be understood that any particular programming terminology used herein is for convenience only, and therefore the invention should not be limited to use only in any particular application identified and / or implied by such terminology.

[0087] The computing system 800 includes a communication structure 802 that provides communication between a cache 816, a memory 806, a permanent storage device 808, a communication unit 810, and an input / output (I / O) interface 812. The communication structure 802 can be implemented in any architecture designed to transfer data and / or control information between processors (such as microprocessors, communication and network processors), system memory, peripheral devices, and any other hardware components within the system. For example, the communication structure 802 can be implemented using one or more buses or crossbar switches.

[0088] Memory 806 and persistent storage device 808 are computer-readable storage media. In this embodiment, memory 806 includes random access memory (RAM). Typically, memory 806 may include any suitable volatile or non-volatile computer-readable storage medium. Cache 816 is a fast memory that enhances the performance of one or more computer processors 804 by storing recently accessed data and data close to the accessed data from memory 806.

[0089] Program 104 may be stored in persistent storage device 808 and memory 806 for execution by one or more of the respective computer processors 804 via cache 816. In one embodiment, persistent storage device 808 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage device 808 may include a solid-state drive, a semiconductor storage device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.

[0090] The media used in persistent storage device 808 can also be removable. For example, a removable hard disk drive can be used in persistent storage device 808. Other examples include optical discs and disks, thumb drives and smart cards, which are inserted into the drive to be transferred to another computer-readable storage medium that is also part of persistent storage device 808.

[0091] In these examples, communication unit 810 provides communication with other data processing systems or devices. In these examples, communication unit 810 includes one or more network interface cards. Communication unit 810 can provide communication via one or both physical and wireless communication links. Program 104 can be downloaded to permanent storage device 808 via communication unit 810.

[0092] One or more I / O interfaces 812 allow data input and output to other devices that can be connected to mobile devices, approval devices, and / or server computing devices 108. For example, I / O interface 812 can provide connectivity to external devices 818, such as keyboards, keypads, touchscreens, and / or other suitable input devices. External devices 818 may also include portable computer-readable storage media, such as thumb drives, portable optical discs or disks, and memory cards. Software and data used to implement embodiments of the invention, such as program 104, can be stored on such portable computer-readable storage media and can be loaded onto permanent storage device 808 via one or more I / O interfaces 812. One or more I / O interfaces 812 are also connected to display 820.

[0093] The display 820 provides a mechanism for displaying data to the user and can be, for example, a computer monitor.

[0094] The present invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.

[0095] Computer-readable storage media can be any tangible device capable of retaining and storing instructions used by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital universal disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as protrusions in slots or punch cards having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0096] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.

[0097] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Smalltalk, C++, etc.) and conventional procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may be personalized to perform aspects of this invention by utilizing the state information of the computer-readable program instructions.

[0098] The various aspects of the present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0099] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in the blocks or blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that directs a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium storing the instructions includes an article of manufacture containing instructions that implement aspects of the functions / actions specified in the blocks or blocks of a flowchart and / or block diagram.

[0100] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in the blocks or blocks of a flowchart and / or block diagram.

[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function(s). In some alternative implementations, the functions marked in the blocks may occur in a different order than indicated in the figures. For example, depending on the functions involved, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0102] Various embodiments of the invention have been described for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or technical improvements superior to those found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method, comprising: Prediction is performed based on a model generated for each database in a plurality of databases, and a baseline prediction is generated for each database in the plurality of databases, wherein the generated baseline prediction includes samples of the performance of each feature for the respective database in the plurality of databases; Based on a dimension reduction algorithm that projects different databases into an equal-dimensional feature space, the difference between the actual performance samples of each feature of the first database in the multiple databases and the baseline predictions of each feature performance included in the second database in the multiple databases are calculated, and a link matrix containing the calculated differences is generated. Based on the differences within the generated link matrix reaching a predetermined threshold for measured performance, link one or more databases from the plurality of databases; as well as Generate a design for a computer system that can be displayed via a user interface, the design representing performance improvements for each feature of the linked database. Specifically, based on a dimensionality reduction algorithm that projects different databases into an equal-dimensional feature space, the difference between the actual performance samples of each feature of the first database in the multiple databases and the baseline prediction samples of each feature performance included in the second database in the multiple databases is calculated, and a link matrix containing the calculated differences is generated, including: The actual performance samples of each feature of the first database in the plurality of databases are compared with the samples contained in the baseline prediction generated for the second database in the plurality of databases; The dimensionality reduction algorithm is used to perform head-to-head comparisons of samples, comparing samples of the actual performance of each feature of the first database in the plurality of databases with samples included in the corresponding baseline predictions of each of the remaining databases in the plurality of databases. The dimensionality reduction algorithm is used to perform head-to-head comparisons of samples, comparing samples of the actual performance of each feature in the first database among the plurality of databases with samples of the actual performance of each feature in each of the remaining databases among the plurality of databases; and Calculate the difference between samples of the actual performance characteristics of each database in the plurality of databases and the corresponding samples of the baseline prediction generated for that database.

2. The computer-implemented method according to claim 1, wherein, Based on the differences within the generated link matrix reaching a predetermined threshold for measured performance, linking one or more databases from the plurality of databases includes: Linking non-common performance features from one of the plurality of databases to another database among the plurality of databases, wherein linking the non-common performance features includes: The calculated difference between the predicted performance and the actual performance of each feature in one of the multiple databases is compared with the corresponding calculated difference in each of the remaining databases. Evaluate whether the difference between the predicted performance and the actual performance of each feature in each calculation meets the performance threshold requirement, and Link at least one feature of one database to another database among the plurality of databases to achieve the performance threshold requirement.

3. The computer-implemented method according to claim 1, wherein, Based on a dimensionality reduction algorithm in an equal-dimensional feature space, the difference between the actual performance samples of each feature of the first database in the multiple databases and the samples included in the baseline prediction generated for the second database in the multiple databases is calculated, and a link matrix containing the calculated differences is generated, including: Determine the Euclidean distance between a first sample in a first database and a second sample in a second database among the plurality of databases, wherein the Euclidean distance is the square root of the summation function of the squares of the difference between the first sample in the first database and the second sample in the second database among the plurality of databases.

4. The computer-implemented method according to claim 2 further includes: In response to determining that the difference between the calculated predicted performance and the actual performance for the feature does not meet the performance threshold requirement, samples of the actual performance of each feature in each of the plurality of databases are cascaded with samples from each of the plurality of databases included in the generated baseline prediction.

5. The computer-implemented method according to claim 1, further comprising: The characteristic performance improvement of the one or more linked databases is measured by comparing it with the baseline prediction generated for the second database; as well as The performance improvements are based on the characteristics measured within the database of the one or more links, and the effects of measuring the remaining characteristics that were not previously measured.

6. The computer-implemented method according to claim 1, wherein, Generating designs for computer systems that can be displayed through a user interface includes: Input is received from the tactile sensors of the computing device for a linked database among the plurality of databases; An improved design for generating each feature of a linked database representing a first computer system based on input received from the tactile sensors of a computing device; An improved design for the performance of each feature of the linked database for the first computer system is generated based on input received from the tactile sensors of the computing device. The improved design for the performance of each feature of the generated linked database for the first computer system is optimized based on the modified design of the improved performance of each feature of the generated linked database for the first computer system. Based on the modified design, input received from the tactile sensors of the computing device, and the linked database, an optimized, generated design for the first computer system representing each feature of the linked database is displayed, improving performance. A second computing system is constructed based on the optimized design of the computer system, which can be displayed through the user interface.

7. A computer program product, comprising: One or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions comprising: Program instructions for performing predictions based on models generated for each database in a plurality of databases and generating baseline predictions for each database in the plurality of databases, wherein the generated baseline predictions include samples of the performance of each feature for the respective database in the plurality of databases; Program instructions for calculating the difference between samples of the actual performance of each feature of the first database in the plurality of databases and samples of the performance of each feature in the baseline prediction generated for the second database in the plurality of databases, based on a dimensionality reduction algorithm that projects different databases into an equal-dimensional feature space, and generating a link matrix containing the calculated differences. Program instructions for linking one or more databases among the plurality of databases to reach a predetermined threshold of measured performance based on the differences within the generated link matrix; and Program instructions for generating designs for computer systems that can be displayed via a user interface, the designs representing improvements to the performance of each feature of the linked database. The program instructions for calculating the difference between samples of the actual performance of each feature of the first database in the plurality of databases and samples of the performance of each feature included in the baseline prediction generated for the second database in the plurality of databases, based on a dimensionality reduction algorithm that projects different databases into an equal-dimensional feature space, and generating a link matrix containing the calculated differences, include: Program instructions for comparing samples of the actual performance of each feature of the first database in the plurality of databases with samples contained in the baseline prediction generated for the second database in the plurality of databases; Program instructions for performing head-to-head comparisons of samples using the dimensionality reduction algorithm to compare samples of actual performance of each feature of the first database in the plurality of databases with samples included in the corresponding baseline predictions of each of the remaining databases in the plurality of databases. Program instructions for performing head-to-head comparisons of samples using the dimensionality reduction algorithm, comparing samples of actual performance for each feature of the first database in the plurality of databases with samples of actual performance for each feature of each of the remaining databases in the plurality of databases; and Program instructions for calculating the difference between samples of the actual performance characteristics of each of the plurality of databases and corresponding samples of the baseline prediction generated for that database.

8. The computer program product according to claim 7, wherein, The program instructions for linking one or more databases from the plurality of databases to reach a predetermined threshold for measured performance based on the differences within the generated link matrix include: Program instructions for linking a non-common performance feature from one of the plurality of databases to another database among the plurality of databases, wherein linking the non-common performance feature includes: Program instructions for comparing the calculated difference between the predicted performance and the actual performance of each feature in one of the plurality of databases with the corresponding calculated difference in each of the remaining databases. The program instructions used to evaluate whether the difference between each calculated performance and the actual performance of each feature meets the performance threshold requirement, and Program instructions for linking at least one feature of one database to another database among the plurality of databases to achieve the performance threshold requirement.

9. The computer program product of claim 7, comprising program instructions for calculating, based on a dimensionality reduction algorithm in an equal-dimensional feature space, the difference between samples of the actual performance of each feature of a first database in the plurality of databases and samples included in a baseline prediction generated for a second database in the plurality of databases, and generating a link matrix containing the calculated differences, comprising: Program instructions for determining the Euclidean distance between a first sample of a first database and a second sample of a second database in the plurality of databases, wherein the Euclidean distance is the square root of the summation function of the squares of the subtraction of the first sample of the first database and the second sample of the second database in the plurality of databases.

10. The computer program product according to claim 8, wherein, The program instructions stored on the one or more computer-readable storage media further include: In response to program instructions for determining that the difference between the calculated predicted performance and the actual performance for the feature does not meet a performance threshold requirement, program instructions for cascading samples of the actual performance of each feature from each of the plurality of databases with samples from each of the plurality of databases included in the generated baseline prediction.

11. The computer program product according to claim 7, wherein, The program instructions stored on the one or more computer-readable storage media further include: Program instructions for measuring the characteristic performance improvement of the one or more linked databases by comparing them with baseline predictions generated for a second database; and Program instructions for measuring the effect of performance improvements on previously unmeasured features based on features measured within the one or more linked databases.

12. The computer program product according to claim 7, wherein, Program instructions used to generate designs for computer systems that can be displayed through a user interface include: Program instructions for receiving input from the tactile sensors of a computing device for a linked database among the plurality of databases; Program instructions for generating an improved design for each feature of a linked database for a first computer system, based on input received from tactile sensors of a computing device. Program instructions for modifying the design of an improved performance of each feature of a database representing a linked database for a first computer system, based on input received from the tactile sensors of the computing device. Program instructions for optimizing the improved design of each feature performance of the generated linked database representation for the first computer system based on the improved design of each feature performance of the generated database representation for the first computer system. Program instructions for displaying an optimized, generated design for a first computer system representing each feature of the linked database, based on a combination of the modified design, input received from the tactile sensors of the computing device, and the linked database; and Program instructions for constructing a second computing system based on an optimized design of a computer system that can be displayed through the user interface.

13. A computer system, comprising: One or more computer processors; One or more computer-readable storage media; as well as Program instructions stored on the one or more computer-readable storage media and executable by at least one of the one or more processors, the program instructions comprising: Program instructions for performing predictions based on models generated for each database in a plurality of databases and generating baseline predictions for each database in the plurality of databases, wherein the generated baseline predictions include samples of the performance of each feature for the respective database in the plurality of databases; Program instructions for calculating the difference between samples of the actual performance of each feature of the first database in the plurality of databases and samples of the performance of each feature in the baseline prediction generated for the second database in the plurality of databases, based on a dimensionality reduction algorithm that projects different databases into an equal-dimensional feature space, and generating a link matrix containing the calculated differences. Program instructions for linking one or more databases among the plurality of databases to reach a predetermined threshold of measured performance based on the differences within the generated link matrix; and Program instructions for generating designs for computer systems that can be displayed via a user interface, the designs representing improvements to the performance of each feature of the linked database. The program instructions for calculating the difference between samples of the actual performance of each feature of the first database in the plurality of databases and samples of the performance of each feature included in the baseline prediction generated for the second database in the plurality of databases, based on a dimensionality reduction algorithm that projects different databases into an equal-dimensional feature space, and generating a link matrix containing the calculated differences, include: Program instructions for comparing samples of the actual performance of each feature of the first database in the plurality of databases with samples contained in the baseline prediction generated for the second database in the plurality of databases; Program instructions for performing head-to-head comparisons of samples using the dimensionality reduction algorithm to compare samples of actual performance of each feature of the first database in the plurality of databases with samples included in the corresponding baseline predictions of each of the remaining databases in the plurality of databases. Program instructions for performing head-to-head comparisons of samples using the dimensionality reduction algorithm, comparing samples of actual performance for each feature of the first database in the plurality of databases with samples of actual performance for each feature of each of the remaining databases in the plurality of databases; and Program instructions for calculating the difference between samples of the actual performance characteristics of each of the plurality of databases and corresponding samples of the baseline prediction generated for that database.

14. The computer system according to claim 13, wherein, The program instructions for linking one or more databases from the plurality of databases to reach a predetermined threshold for measured performance based on the differences within the generated link matrix include: Program instructions for linking a non-common performance feature from one of the plurality of databases to another database among the plurality of databases, wherein linking the non-common performance feature includes: Program instructions for comparing the calculated difference between the predicted performance and the actual performance of each feature from one of the plurality of databases with the corresponding calculated difference in each of the remaining databases. The program instructions used to evaluate whether the difference between each calculated performance and the actual performance of each feature meets the performance threshold requirement, and Program instructions for linking at least one feature of one database to another database among the plurality of databases to achieve the performance threshold requirement.

15. The computer system according to claim 13, wherein, The program instructions for calculating the difference between samples of the actual performance of each feature of a first database in a plurality of databases and samples included in the baseline prediction generated for a second database in a plurality of databases, based on a dimensionality reduction algorithm in an equal-dimensional feature space, and generating a link matrix containing the calculated differences, include: Program instructions for determining the Euclidean distance between a first sample of a first database and a second sample of a second database in the plurality of databases, wherein the Euclidean distance is the square root of the summation function of the squares of the subtraction of the first sample of the first database and the second sample of the second database in the plurality of databases.

16. The computer system according to claim 14, wherein, The program instructions stored on the one or more computer-readable storage media further include: In response to program instructions for determining that the difference between the calculated predicted performance and the actual performance for the feature does not meet a performance threshold requirement, program instructions for cascading samples of the actual performance of each feature from each of the plurality of databases with samples from each of the plurality of databases included in the generated baseline prediction.

17. The computer system according to claim 13, wherein, The program instructions stored on the one or more computer-readable storage media further include: Program instructions for measuring the characteristic performance improvement of the one or more linked databases by comparing them with baseline predictions generated for a second database; and Program instructions for measuring the effect of performance improvements on previously unmeasured features based on features measured within the one or more linked databases.