Sorting a dataset based on data attributes

By extracting target data fields and attributes from process documents and data usage documents using a computer system, generating a meta-dataset, and evaluating the similarity between the meta-dataset and the target meta-dataset, this approach solves the problem of low efficiency in dataset value assessment and ranking in existing technologies, and achieves efficient utilization and accurate matching of datasets.

CN114647627BActive Publication Date: 2025-11-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111423989.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-17
Filing Date
2021-11-26
Publication Date
2025-11-18
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively assess and rank the value of multiple datasets, especially when the content and intended use of the datasets are unclear, leading to inefficient use of the datasets.

Method used

The target data fields and attributes are extracted from process documents and data usage documents using a computer system to generate a meta-dataset. The similarity between the meta-dataset and the target meta-dataset is evaluated, and the meta-dataset is ranked based on the comparison attribute scores.

Benefits of technology

It enables efficient sorting of datasets based on their content and expected usage requirements, matching users' data needs and improving the utilization efficiency and accuracy of dataset value assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114647627B_ABST
    Figure CN114647627B_ABST
Patent Text Reader

Abstract

Sorting a set of data sets using a computer includes determining a target data field set from a process document set indicating user data field preferences. A target data set attribute set from the data use document set indicates user data range preferences. A plurality of metadata sets of the associated plurality of data sets having a field suitability value exceeding a predetermined suitability threshold are determined by the computer. The FSV represents a degree of similarity between a field set associated with the data set and the target data field set. The computer evaluates the metadata set with respect to the target attribute and generates a comparison attribute score for each candidate data set. The comparison attribute score indicates a degree of likelihood that the associated data set will have content exhibiting the target data set attribute. The computer candidate data set is based on the comparison attribute score.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates generally to the field of data set analysis, and more specifically to computer data set evaluation. BACKGROUND

[0002] A data set is a collection of data that can be used by various computer systems to provide answers to questions about many real-world and simulated situations. Often, a data set includes information about past transactions or other historical information from which predictions can be made about similar current and future transactions. In some domains, data sets are produced by user systems as a byproduct of system operation and held for future use. In other domains, data sets, especially large or custom data sets, can be provided for a fee to users by third parties. Artificial intelligence (AI) systems can identify patterns within the data contained in a data set to reveal trends that are often difficult to predict in other ways. Because data sets can vary widely in content, some data sets will be more useful than others to certain users.

[0003] The value of a data set can vary from use case to use case. If the intended use of the data is known, the value of the data set can be evaluated and ranked against evaluated data sets. SUMMARY

[0004] According to one embodiment, a computer-implemented method of ranking a plurality of data sets according to data set attributes includes identifying, by a computer, a target set of data fields from a set of process documents, the process documents indicating data field preferences of a user. The computer identifies a target set of data set attributes from a set of data use documents, and the data use documents indicating data range preferences of the user. The computer generates a set of meta data sets for the associated plurality of data sets. The computer determines candidate data sets having a field suitability value that exceeds a predetermined suitability threshold, and the field suitability value representing a degree of similarity between a set of fields associated with the data set and the target set of data fields. The computer evaluates the associated meta data set of each candidate data set with respect to the target attributes. The computer generates a comparison attribute score for each candidate data set, the comparison attribute score indicating a degree of likelihood that the associated data set will have content that exhibits the target data set attributes. The computer generates a list of the candidate data sets ordered by the comparison attribute scores.

[0005] According to aspects of the invention, the data use documents include information in a format selected from a list consisting of: business process execution language (BEPL) and unified modeling language (UML).

[0006] According to an aspect of the invention, data target attributes are extracted from elements of the process document, said elements being selected from a list consisting of class diagrams, activity diagrams, sequence diagrams, and component diagrams. According to an aspect of the invention, a candidate dataset with the highest comparison attribute score is designated as the selected dataset. According to an aspect of the invention, a set of search parameters for the search to be performed on said selected dataset is established; and historical usage fields in the metadata set associated with the selected dataset for the search are updated using search context values ​​representing aspects of the search parameters. According to an aspect of the invention, sorting is based at least in part on historical usage field values. According to an aspect of the invention, comparison attribute scores are based at least in part on desirability values ​​associated with each target dataset attribute in the target dataset attributes. According to an aspect of the invention, the metadata set includes information selected from a list consisting of: domain, gender, age group, geographic distribution, demographic distribution, statistical range of values, and context of applicability.

[0007] According to another embodiment, a system for sorting multiple datasets includes: a computer system including a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to: identify a target data field set from a set of process documents, the process documents indicating a user's data field preferences; identify a target dataset attribute set from a set of data usage documents, the data usage documents indicating the user's data range preferences; generate multiple meta-datasets for the multiple associated datasets; determine candidate datasets having field fitness values ​​exceeding a predetermined fitness threshold, the field fitness values ​​representing the degree of similarity between the set of fields associated with the dataset and the target data field set; evaluate the associated meta-datasets of each candidate dataset with respect to the target attributes, and generate a comparative attribute score for each candidate dataset by the computer, the comparative attribute score indicating the degree to which the associated dataset would have content that displays the attributes of the target dataset; and generate a list of the candidate datasets sorted according to the comparative attribute scores.

[0008] According to another embodiment, a computer program product for sorting multiple datasets includes a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to: identify a set of data target attributes indicative of a user's data field preferences from a set of process documents; identify a set of dataset target attributes from a set of data usage documents indicative of the user's data range preferences; generate multiple meta-datasets for the associated multiple datasets; determine the top k candidate datasets with field fitness values ​​exceeding a predetermined fitness threshold; evaluate the associated meta-datasets of each candidate dataset with respect to the target attributes, and generate a comparative attribute score for each candidate dataset; and sort the candidate datasets at least in part based on the comparative attribute scores.

[0009] The value of a given dataset can be based on a variety of factors, including the content of the dataset's recorded fields and the range of information it contains. For example, many data analysis systems require certain kinds of information (e.g., certain fields) to provide meaningful output, and a dataset with a larger amount of appropriate information (e.g., a higher number of desired data fields) is better than a dataset with fewer required data fields. Similarly, data analysis systems need data that is appropriate for the questions presented to the system in order to provide meaningful output, and the more relevant a given dataset is to its intended use case (e.g., the intended question to be asked), the higher its value.

[0010] The present invention matches user data requirements (including target data fields and target dataset attributes), including data requirements of users with business applications that should match metadata derived from the data in the dataset. According to various aspects of the invention, the metadata should represent the dataset content, describing the demographic and statistical characteristics of the data content.

[0011] The present invention relates data to domain and meaning through various methods, including ontology usage and key-value pair usage.

[0012] The present invention first selects a set of scores for providing datasets based on the target dataset requirements and their matching with metadata. This set of scores allows enterprises to assess which dataset is more suitable for their requirements.

[0013] According to various aspects of the invention, the derived metadata includes: statistical characteristics (e.g., distribution type, mean, variance, and correlation characteristics, any correlation; and whether it is time series data); various fields and their associated meanings / semantics (e.g., in a loan approval dataset, "spouse" is similar to "wife" and "husband"); if the "CSV" file and the associated pattern are known, various meanings suitable for that pattern (e.g., fields associated with opening new marketing channels may have certain meanings different from similarly named fields used to identify fields associated with sporting events) can be recorded as metadata; personal identification information (e.g., email, phone number, address / contact details) when granted use based on the consent and permission of the identified individual; fields related to previous dataset use (e.g., other datasets used by the dataset are identified through historical mining of dataset use); the derived metadata also includes information about the content representation, such as domain, gender, age group, geographic distribution (this may indicate an applicable age group, bank domain, or region, etc.).

[0014] This invention determines the value of a dataset based on its content (e.g., characterized by dataset metadata). According to this invention, the metadata includes descriptive information indicating the content-based characteristics of the dataset. This invention identifies data requirements for the business. This invention ranks the datasets and provides relevance scores based on the attribute and range values ​​of each metadata element and the range of those values. This invention develops and derives a systematic method for determining the value of a dataset based on business requirements and data content. This invention uses these scores and derives the ranking of each sub-aspect of the metadata. This invention uses dataset values ​​to compare two datasets relative to business requirements. This invention implements a dataset search mechanism based on data content. This invention uses the history of data usage in different contexts to generate metadata and uses it to identify business contexts during search events. This invention searches a corpus of datasets based on input from a set of business requirements; and ranks the results according to the best match for suitability. According to an aspect of the invention, target data fields (e.g., data fields supporting desired business processes) are defined using various diagrams in a standard format (e.g., a Business Process Execution Language (BPEL) that can provide extractable activities, actors, sequencing / sequences, and diagrams of related software engineering artifacts, and a Unified Modeling Language (UML)). According to an aspect of the invention, UML documentation may include class diagrams, activity diagrams, sequence diagrams, and component diagrams.

[0015] According to an aspect of the invention, activities from a BPEL diagram can be matched with a UML activity diagram and used to extract class-level components. According to an aspect of the invention, class-level components can provide all the requirements for a field.

[0016] This invention can derive business requirements. This invention can evaluate datasets and metadata. This invention can sort datasets, data values, and data sub-aspects. This invention can help determine whether a given dataset is relevant to providing information on how to open mobile or online business channels for a business, using questionnaires (or other data usage documents indicating data requirements).

[0017] Since datasets with content possessing certain attributes may be more useful than other datasets, aspects of the system (including, for example, user data requirement questionnaires and other data usage documents) help us identify what users need as content. Aspects of the present invention identify the relevant context of the data, indicating which datasets would be a good match for various user objectives (e.g., using coupons to promote new tea products, etc.).

[0018] According to some aspects of the present invention, PDAM includes a "discovery unit" for locating activity graphs and class graph locators. According to some aspects of the present invention, PDAM includes an entity extractor and an activity extractor.

[0019] According to some aspects of the invention, terminology attributes can be used interchangeably with word aspects. According to aspects of the invention, CAAM includes an ontology mapping engine. According to aspects of the invention, CAAM includes aspects for determining whether a dataset matches business requirements. According to aspects of the invention, HULUM 124 includes a data usage metadata extractor and a dataset history usage log, which indicates historical dataset usage and identifies other datasets used with the selected dataset. Attached Figure Description

[0020] These and other objects, features, and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments thereof, which is taken in conjunction with the accompanying drawings. The various features of the drawings are not to scale, as the illustrations are provided for clarity and to assist those skilled in the art in understanding the invention in conjunction with the detailed description. The drawings are shown below:

[0021] Figure 1 This is a schematic block diagram illustrating an overview of a computer-implemented method for sorting multiple datasets according to an embodiment of the present invention, based on dataset content and desired data attributes.

[0022] Figure 2 This is a system method illustrating a computer-implemented method for sorting multiple datasets according to the present invention (using, for example...). Figure 1 The flowchart shown is for the system implementation.

[0023] Figure 3A yes Figure 1 Alternative views of aspects of the system shown.

[0024] Figure 3B This is an embodiment of the invention for providing a collection of sorted datasets, such as... Figure 1 The diagram shows a schematic representation of aspects of the system.

[0025] Figure 4 yes Figure 1 The diagram shows a schematic overview of the system, in which aspects of the system are arranged in multiple stages.

[0026] Figure 5 yes Figure 1 The system shown is an alternative view, in which aspects of the system are arranged according to a workflow outline that includes a list of methods and related details.

[0027] Figure 6 This is a schematic representation of aspects of the "datavalue" and "dataranking" entries generated according to embodiments of the present invention.

[0028] Figure 7 This is an exemplary business data usage questionnaire and related sample answers according to an embodiment of the present invention.

[0029] Figure 8 This is a schematic block diagram illustrating a computer system according to an embodiment of the present disclosure, which may be wholly or partially incorporated. Figure 1 In one or more computers or devices shown, and with Figure 1 The systems and methods shown collaborate.

[0030] Figure 9 A cloud computing environment according to an embodiment of the present invention is shown.

[0031] Figure 10 An abstract model layer according to an embodiment of the present invention is shown. Detailed Implementation

[0032] The following description, provided with reference to the accompanying drawings, is intended to aid in a comprehensive understanding of exemplary embodiments of the invention as defined by the claims and their equivalents. It includes various specific details to aid understanding, but these details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Furthermore, for clarity and brevity, descriptions of well-known functions and structures may be omitted.

[0033] The terms and words used in the following description and claims are not limited to their bibliographical meaning, but are merely used to enable a clear and consistent understanding of the invention. Therefore, it will be apparent to those skilled in the art that the following description of exemplary embodiments of the invention is for illustrative purposes only and is not intended to limit the invention as defined by the appended claims and their equivalents.

[0034] It should be understood that the singular forms “one,” “an,” and “the” include plural referents unless the context clearly indicates otherwise. Thus, for example, referring to “one participant” includes referring to one or more such participants unless the context clearly indicates otherwise.

[0035] Now, referring to the accompanying drawings, and especially to... Figure 1 and Figure 2 This provides an overview of a system 100 for a computer-implemented method that sorts multiple datasets based on their contents, executed by a server computer 102 having optionally shared storage 104. Continue to refer to... Figure 1 The server computer 102 communicates with the source of the process document 106 (e.g., BPEL, UML diagrams, etc.) that indicates the desired dataset data fields. The server computer 102 includes a Process Document Analysis Module (PDAM) 112, which uses known UML processing and evaluation tools (including diagram identifiers and other similar UML content extractors) and a BPEL reader to examine and mine the document to identify target data fields. These target data fields provide information about the data format most compatible with the needs of a given user. As an example, a user can provide a document written in Business Process Execution Language (BPEL), and this format can indicate various activities, actors, and process sequences that are operationally important to the user's business, and these aspects can be extracted to help understand the user's data needs. As another example, a user can provide a document presented using Unified Modeling Language (UML) or a similar modeling language, and this format can provide insights into software artifacts operationally important to the user's processing system, including class diagrams, activity diagrams, and sequence diagrams. Activities from the BPEL document can be matched with UML diagrams and used to extract class-level components of the user's system. The server computer 102 uses the extracted class-level components to determine data field requirements.

[0036] Server computer 102 also communicates with the source of data use document 108 (e.g., a data request questionnaire) that indicates the desired attributes of the dataset. Server computer 102 also communicates with the source of one or more datasets 110.

[0037] Server computer 102 includes a Data Usage Document Analysis Module (DUDAM) 114 to evaluate data usage documents to identify target dataset attributes. Users can provide information about intended data usage in various ways (e.g., through detailed questionnaire responses, providing expected question sets, specifying topics of interest, etc.), and this information indicates what range of data content will best suit the user's needs. According to aspects of the invention, this information provides input about business requirements, allowing matching data content to be identified when encountered in various datasets. For example, if a user wants insights into marketing a specific product in a given region, a dataset containing information about the product's sales in that region is likely more valuable than a dataset that only includes sales information for that product in that different region. According to aspects of the invention, general sales information about the product may also be valuable to the user, and user data usage request queries can be structured to collect this level of detail. According to aspects of the invention, a wide variety of preferences regarding the scope of data can be collected from the user to train system 100 on user data usage preferences.

[0038] Server computer 102 includes a metadata generation module (MGM) 116 that uses (e.g., an ontology engine, statistical data processor, or similar known tools) to extract and generate metadata identifying dataset attributes. For example, server computer 102 can use a domain-specific ontology that includes machine-readable statements about domains (e.g., descriptions of various domain concepts and their relationships) to assign meaning to fields in a given dataset. MGM 116 can also receive simple key-value pairs indicating the meaning of fields. Other meaning assignment schemes chosen based on the judgment of those skilled in the art are also sufficient. The generated metadata can include a variety of useful information about the data contained in the given dataset. The derived metadata can include, for example, statistical characteristics of the data, such as the type of distribution, mean, and variance of the data values. The derived metadata can also include indications of time series data and other similar and various other data field correlations. According to aspects of the invention, the derived metadata can also include information about how the data has been previously used and what other queries it supports (e.g., derived through known data mining techniques). Exported metadata can also (when with confirmed consent from the original content provider) provide personally identifiable information that, depending on the provided confirmed consent, can be used for product marketing and opening new marketing channels or otherwise permitted. Exported metadata can include information about the demographic representation of the content found within the data, including subject areas and clustering by gender, age group, geographic distribution, etc. The exported metadata for a given dataset presents a summary of the dataset's content and provides indications of how well the dataset is suited for specific data uses. For example, exported metadata can indicate that a given dataset is well-suited for answering questions about a particular area, certain demographic ranges, geographic relevance, etc. The better a dataset is suited for a given data use, the more value it holds for users with those data use objectives.

[0039] Server computer 102 includes a Dataset Field Suitability Assessment Module (DFSAM) 118, which identifies datasets with Field Suitability Values ​​(FSVs) exceeding a predetermined suitability threshold. The FSV is calculated by comparing fields indicated by derived metadata of a given dataset 110 with target data fields determined by PDAM 112 to determine the number of matches between fields included in the dataset and preferred target data fields. The FSV indicates the degree of similarity between dataset fields and target data fields, which can be measured, for example, by the number of classification labels having a semantic similarity greater than 85% to the target data fields, or some other value selected based on the judgment of someone skilled in the art. To improve downstream computational efficiency, (DFSAM) 118 identifies the top k candidate datasets with FSVs greater than the suitability threshold and designates these candidate datasets as comparison datasets.

[0040] Server computer 102 includes a Comparative Attribute Evaluation Module (CAAM) 120 that compares the dataset metadata of the comparison datasets identified by (DFSAM) 118 to generate a Comparative Attribute Score (CASV) for each comparison dataset. This CASV represents the degree to which each associated comparison dataset has the likelihood of displaying the target dataset's attributes. CASV is determined, for example, by identifying multiple dataset attributes (or some other values ​​selected based on the judgment of someone skilled in the art) that have an attribute semantic similarity greater than 85% with the target dataset attribute. Server computer 102 includes a Candidate Dataset Ranking Module (CDRM) 122 that ranks the candidate dataset metadata based on the target attribute and generates a ranked list of candidate datasets indexed by the score values. Note that various dataset attributes may have different influence weights when applied to different data use documents 108, and these various attribute influence weights can be represented as dataset attribute necessity values ​​associated with various fields or other attributes included in the determined metadata. Server computer 102 includes a highest-ranked dataset selector that designates the comparison dataset with the highest comparative attribute score as the selected dataset.

[0041] As an example, according to aspects of the invention, the evaluation of metadata for two comparative datasets can reveal that the datasets possess attributes that reflect the target dataset (e.g., "dataset value range" and "dataset completeness"). If the "dataset value range" attribute has a higher user indication (e.g., more useful to a given user, via a data usage document) necessity value than the "dataset completeness" attribute, the dataset with a higher "dataset value range" score (e.g., wider value range) will be ranked as more suitable to meet the needs and preferences of the associated user than the dataset with a lower value range score (e.g., smaller value range). In the same example, the dataset with a higher "dataset completeness" score may not be ranked as more suitable because the "dataset completeness" attribute is less important than the "dataset value range". In this example, a relatively high score for the low-weighted "dataset completeness" attribute compared to other datasets is insufficient to ensure a high ranking of the associated dataset. However, in this example, if the dataset associated with a relatively high “dataset integrity” attribute score is presented as a set of attribute scores with an average attribute score higher than the average attribute score of other compared datasets, then that dataset can still be ranked higher by CDRM 122.

[0042] Server computer 102 also includes a Historical Usage Log Update Module (HULUM) 124, which updates historical data fields to assess the future use of the selected dataset with greater accuracy provided by historical context. According to an aspect of the invention, the historical usage field in the metadata set is associated with a dataset selected for searching using search context values ​​representing aspects of search parameters, and is updated whenever the dataset is selected using the search parameters. According to an aspect of the invention, note that data and its historical usage in different business applications can be tracked (e.g., via HULUM 124 and the selected dataset metadata and historical usage update module 126), and used to develop rich metadata (such as…). Figure 4 (Illustrated in section 440). For example, ongoing data usage by a business application can be used to identify associated data usage domains and frequency of use. This can be added back to the dataset as metadata (e.g., via HULUM 124 and the selected dataset metadata and historical usage update module 126), and future searches can be based on this ever-expanding set of metadata content. It is known to use dataset content to search datasets. According to an aspect of the invention, minor aspects (e.g., target dataset attributes) are included as dataset search criteria. For example, a minor aspect referred to as “number of empty set entries” can capture the number of empty (e.g., empty set) fields present in each record of a given dataset. According to an aspect of the invention, having minor aspects identified within the metadata allows a user to indicate a preference for datasets that display that aspect (e.g., indicating a high attribute requirement value). For example, if a given user’s requirement indicates a preference for datasets with a low number of empty set record entries, then CDRM 122 will prioritize datasets with a relatively low number of empty set record entries over datasets with more empty set record entries (e.g., more suitable for the user and more likely to meet the user’s data needs and preferences). According to an aspect of the invention, certain target dataset attributes (e.g., minor aspects) can also be directly identified as required for identifying a dataset as a selected dataset.

[0043] Now for specific reference Figure 2 And generally with reference to other accompanying drawings, according to various aspects of the present invention, a method for sorting multiple datasets based on dataset content and desired data attributes. At block 202, server computer 102 uses PDAM 112 to identify a set of target data fields from a collection of process documents, using graph identifiers, a BPEL reader, and UML evaluation tools (as described above), to examine and mine documents to identify target data fields.

[0044] At box 204, server computer 102 identifies the target dataset attribute set from the data usage document collection (as described above) via the Data Usage Document Analysis module DUDAM 114. At box 206, server computer 102 generates multiple metadata sets and associated multiple datasets via the Metadata Generation module (MGM) 116. At box 208, server computer 102 determines candidate datasets with field fitness values ​​exceeding a predetermined fitness threshold via DFSAM 118; the field fitness value (FSV) represents the degree of similarity between the field set associated with the dataset (via exported metadata information) and the target data field set. At box 210, server computer 102 determines the likelihood that each comparison dataset exhibits the target dataset attributes via CAAM 120.

[0045] At box 212, server computer 102 sorts candidate datasets at least in part based on comparative attribute score values ​​via CDRM 122. At boxes 214 and 216, server computer 102 establishes a set of search parameters to be performed on the selected dataset via selected dataset metadata and historical usage update module 126; and updates the historical usage field in the metadata set associated with the selected dataset to perform the search using search context values ​​representing aspects of the search parameters. At box 218, server computer 102 represents the selected dataset 218 via selected dataset renderer 128. According to aspects of the invention, the search context value may be a digital code providing information about domains in which a particular dataset has been previously used. The search context value may also be an unstructured text string and may represent other aspects of previous use of the provided dataset (including other datasets used collaboratively).

[0046] Now for reference Figure 3AThe diagram illustrates a high-level overview 310 of system 100. Specifically, business requirements, datasets, and metadata are provided as input to a data value engine for processing. According to aspects of the invention, the data value engine provides sorted datasets, data values, and aspect rankings as output. According to aspects of the invention, metadata comprises information about data in a given dataset (e.g., a domain associated with the given dataset). Metadata can be stored with the data in different forms. In many object storage arrangements, data is stored as objects, and metadata is stored in key-value pairs associated with the data objects. Metadata is primarily identified from within the data itself or manually using input from data experts (e.g., using automated mechanisms, such as analytical algorithms or similar routines selected by those skilled in the art), who provide additional insights and information about various data objects. According to aspects of the invention, a score is a numerical value representing the relative importance of aspects (e.g., attributes or features) in the metadata. Aspects are data attributes added to the dataset through automation or as part of input provided by domain experts. If two datasets are available, the dataset with the higher score is better for the specific requirements of the application. According to aspects of the invention, we preferably generate scores and rank the datasets based on the presence of content field attributes and preferred dataset attributes (e.g., minor aspects). The ranking identifies the relative importance of a given minor aspect to a given dataset across all features. The ranking also determines the relative placement of the various aspects when the dataset needs to be tailored to the data of a given user. For example, a dataset score for the "null" attribute would indicate that there are many empty sets of records in the associated dataset. Furthermore, server computer 102 will use the scores for each attribute and rank the attributes across the various comparison datasets and within each dataset (e.g., via CDRM 122).

[0047] Now for reference Figure 3B The diagram illustrates a schematic representation 320 of an example of a system 100 in use. Specifically, requests for certain information (represented by questions or a set of questions arranged in a questionnaire, and other data requirements) are passed to the data value engine. Several datasets (e.g., “HR Data,” “Customer Dataset,” and “Click Analysis”) and associated dataset metadata are also provided to the data value engine. The data value engine processes the input, the datasets provided based on a suitability assessment, and provides a list of datasets ranked according to the determined suitability. In the example shown, the “HR Data” dataset is the top-ranked dataset with a defined data value of 50; the “Click Analysis” dataset is the middle-ranked dataset with a defined data value of 46; and the “Customer” dataset is the lowest-ranked dataset with a defined data value of 35.

[0048] Now for reference Figure 4This section will discuss a schematic overview of System 100, illustrating aspects of the system arranged in multiple phases. Specifically, Phase 1 410 represents an aspect of an embodiment of the invention, collectively referred to as "Phase 1: Business Documentation and Process Analysis Engine," where BPEL documents, implementation artifacts, UML diagrams, and various component diagrams are processed for entity and activity extraction. The discovery units associated with Phase 1 410 include activity diagram locators and class diagram locators, adapted to identify fields necessary to support system activities, as represented in the process documents provided as input, based on established practices and requirements of a given user. Phase 2 420 represents an aspect of an embodiment of the invention, collectively referred to as "Phase 2: Dataset Value Assessment Engine," where various field requirements and desired dataset characteristics (including target fields identified in Phase 1 410 and dataset target attributes (e.g., dataset minor aspects) identified in Phase 3 430 (described more fully below)) are compared with the metadata of the respective provided datasets using known NLP, machine learning comparisons, and other computerized analysis methods. Dataset suitability values ​​are determined for each dataset, and the datasets are ranked according to these values. Phase 3, 430, collectively referred to as "Phase 3: Business Interactive Dataset Recommendation Engine," involves passing various business requirement questions, associated answers, and relevant system artifact mappings to the dataset evaluation engine 420 for use as described above. Phase 4, 440, collectively referred to as "Phase 4: Data Usage History," involves passing past records of dataset usage and extracted metadata describing that usage to Phase 2, 420, for supplementary consideration when determining dataset suitability values. Specifically, the output of Phase 4, 440 provides a correlated increase in historical perspective and score accuracy by allowing the evaluation engine of Phase 2, 420, to include metadata on past dataset usage and historical score values. This phase provides the system with an ever-increasing perspective in the case of repeated use, allowing System 100 to become more accurate with increased use.

[0049] Now for reference Figure 5 The diagram illustrates an alternative view of system 100, arranged according to an exemplary workflow outline 500, showing aspects of the system. Specifically, business questionnaire information and business process information are passed from the business owner to an activity identification stage, where target data aspects and required system classes are identified. A business metric-data converter then provides the required aspects from business and data fields to data field identifiers and generates aspect assessments. These aspect assessments are passed to a dataset assessment stage, where a dataset sorter provides a sorting of the datasets. This information is then passed back to the business owner as output.

[0050] Now for reference Figure 6The diagram 600 illustrates an example embodiment of "datavalue" and "dataranking" entries generated according to embodiments of the invention. Specifically, according to aspects of the invention, the entries provide an indication of a set of JSON-formatted key-value pairs useful for identifying and comparing dataset values ​​and associated dataset rankings. Note that other formats may be chosen based on the judgment of those skilled in the art.

[0051] Now for reference Figure 7 An exemplary questionnaire 700 (and sample answers) is shown regarding business needs for accounting solutions used to assess account churn. Server computer 102 collects and processes the answers (e.g., via DUDAM 114). The types of information that the associated business may prefer to collect are reflected in the answers provided in response to the questionnaire questions. Answers provided by users associated with the business are used to determine the attributes of the target dataset.

[0052] Regarding flowcharts and block diagrams, the flowcharts and block diagrams in the accompanying drawings of this disclosure illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagram and / or flowchart illustrations, and combinations of blocks in the block diagram and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0053] refer to Figure 8The system or computer environment 1000 includes a computer diagram 1010 shown in the form of a general-purpose computing device. The methods of the present invention can be implemented, for example, in a program 1060, including program instructions implemented on a computer-readable storage device or computer-readable storage medium, such as memory 1030, and more specifically, computer-readable storage medium 1050. This memory and / or computer-readable storage medium includes non-volatile memory or non-volatile storage. For example, memory 1030 may include storage medium 1034 such as RAM (random access memory) or ROM (read-only memory), and cache memory 1038. Program 1060 can be executed by processor 1020 of computer system 1010 (to execute program steps, code, or program code). Additional data storage devices may also be implemented as a database 1110 including data 1114. Computer system 1010 and program 1060 are general representations of a computer and a program that may be local to the user or provided as a remote service (e.g., as a cloud-based service), and may be provided using a website accessible through communication network 1200 (e.g., interacting with a network, the Internet, or a cloud service), as further exemplified. It should be understood that computer system 1010 herein also generally refers to a computer device or a computer included in a device such as a laptop or desktop computer, or one or more servers, either alone or as part of a data center. The computer system may include network adapter / interface 1026 and input / output (I / O) interfaces 1022. I / O interfaces 1022 allow input and output of data with external devices 1074 that can be connected to the computer system. Network adapter / interface 1026 can provide communication between the computer system and a network generally represented as communication network 1200.

[0054] Computer 1010 can be described in the general context of executable instructions of a computer system, such as program modules executed by the computer system. Typically, program modules can include routines, programs, objects, components, logic, data structures, etc., that perform specific tasks or implement specific abstract data types. Method steps and system components and techniques can be implemented in modules of program 1060 that perform tasks for each step of the method and system. These modules are generally represented in the diagram as program modules 1064. Program 1060 and program modules 1064 can execute specific steps, routines, subroutines, instructions, or code of a program.

[0055] The methods disclosed herein can run locally on a device such as a mobile device, or as a service on a server 1100, which may be remote and accessible via the communication network 1200. The program or executable instructions may also be provided as a service by a provider. The computer 1010 can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via the communication network 1200. In a distributed cloud computing environment, program modules can reside in local and remote computer system storage media, including memory storage devices.

[0056] Computer 1010 may include a variety of computer-readable media. Such media may be any available media accessible by computer 1010 (e.g., a computer system or server) and may include volatile and non-volatile media, as well as removable and non-removable media. Computer memory 1030 may include additional computer-readable media in the form of volatile memory, such as random access memory (RAM) 1034 and / or cache 1038. Computer 1010 may also include other removable / non-removable, volatile / non-volatile computer storage media, such as portable computer-readable storage media 1072 in one example. In one embodiment, computer-readable storage media 1050 may be provided for reading from and writing to non-removable, non-volatile magnetic media. Computer-readable storage media 1050 may be implemented, for example, as a hard disk drive. Additional memory and data storage devices may be provided, for example, as a storage system 1110 (e.g., a database) for storing data 1114 and communicating with processing unit 1020. The database may be stored on or part of server 1100. Although not shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media may be provided. In such an example, each may be connected to bus 1014 via one or more data media interfaces. As will be further described below, memory 1030 may include at least one program product that may include one or more program modules configured to perform the functions of embodiments of the present invention.

[0057] For example, the methods(s) described in this disclosure may be embodied in one or more computer programs, generally referred to as program 1060, and may be stored in memory 1030 within computer-readable storage medium 1050. Program 1060 may include program module 1064. Program module 1064 may generally perform the functions and / or methods of embodiments of the invention as described herein. One or more programs 1060 are stored in memory 1030 and can be executed by processing unit 1020. As an example, memory 1030 may store operating system 1052, application(s) 1054, other program modules, and program data on computer-readable storage medium 1050. It will be understood that program 1060, operating system 1052, and application(s) 1054 stored on computer-readable storage medium 1050 may similarly be executed by processing unit 1020. It should also be understood that application 1054 and program(s) 1060 are generally shown and may include all or part of one or more applications and programs discussed in this disclosure, or vice versa, that is, application 1054 and program 1060 may be all or part of one or more applications or programs discussed in this disclosure.

[0058] One or more programs may be stored in one or more computer-readable storage media, such that the program is contained in and / or encoded in the computer-readable storage media. In one example, the stored program may include program instructions for execution by a processor or a computer system having a processor, to perform a method or cause the computer system to perform one or more functions.

[0059] Computer 1010 can also communicate with one or more external devices 1074, such as a keyboard, pointing device, display 1080, etc.; one or more devices that enable a user to interact with computer 1010; and / or any device that enables computer 1010 to communicate with one or more other computing devices (e.g., network interface card, modem, etc.). This communication can occur via input / output (I / O) interface 1022. Furthermore, computer 1010 can also communicate with one or more networks 1200 (such as a local area network (LAN), a general area network (WAN), and / or a public network (e.g., the Internet)) via network adapter / interface 1026. As shown, network adapter 1026 communicates with other components of computer 1010 via bus 1014. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with computer 1010. Examples include, but are not limited to: microcode, device drivers 1024, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0060] It should be understood that a computer or a program running on computer 1010 can communicate with a server implemented as server 1100 via one or more communication networks implemented as communication network 1200. Communication network 1200 may include transmission media and network links, including, for example, wireless, wired, or fiber optic connections, as well as routers, firewalls, switches, and gateway computers. Communication network may include connections such as wired, wireless communication links, or fiber optic cables. Communication network may represent a global collection of networks and gateways (such as the Internet) that communicate with each other using various protocols such as Lightweight Directory Access Protocol (LDAP), Transmission Control Protocol / Internet Protocol (TCP / IP), Hypertext Transfer Protocol (HTTP), Wireless Application Protocol (WAP), etc. Networks may also include many different types of networks, such as intranets, local area networks (LANs), or wide area networks (WANs).

[0061] In one example, the computer can use a network that can use the Internet to access websites on the Web (World Wide Web). In one embodiment, the computer 1010, including a mobile device, can use a communication system or network 1200, which may include the Internet or a Public Switched Telephone Network (PSTN), such as a cellular network. The PSTN may include telephone lines, fiber optic cables, transmission links, cellular networks, and communication satellites. The Internet facilitates many search and text messaging technologies, such as sending queries to search engines using a cellular phone or laptop via text messaging (SMS), multimedia messaging service (MMS) (as opposed to SMS), email, or a web browser. Search engines can retrieve search results, i.e., links to websites, documents, or other downloadable data corresponding to the query, and similarly, provide the search results to the user via the device as, for example, web pages containing the search results.

[0062] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.

[0063] Computer-readable storage media can be tangible devices capable of retaining and storing instructions used by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0064] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.

[0065] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages ​​(including object-oriented programming languages ​​such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of this invention, electronic circuits, including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions to personalize the electronic circuits by utilizing the status information of the computer-readable program instructions.

[0066] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0067] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create parts for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0068] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0069] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function(s). In some alternative embodiments, the functions indicated in the blocks may occur in a non-linear order as shown in the figures. For example, two blocks shown consecutively may actually be implemented as a single step, executed simultaneously, substantially simultaneously, with partial or complete time overlap, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0070] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings set forth herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0071] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage devices, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0072] The features are as follows:

[0073] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power, such as server time and network storage, as needed, without requiring manual interaction with the service provider.

[0074] Wide Area Network (WAN) Access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0075] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Location independence has significance because consumers typically do not control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0076] Rapid Flexibility: In some cases, the ability to scale outwards and inwards quickly and flexibly can be provided. For consumers, the available capacity often appears unlimited and can be purchased in any quantity at any time.

[0077] Measurement services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the providers and consumers of the services being utilized.

[0078] The service model is as follows:

[0079] Software as a Service (SaaS): The capability offered to consumers is the ability to use the provider's applications running on cloud infrastructure. Applications can be accessed from a variety of client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage devices, or even individual application capabilities, with possible exceptions such as limited user-specific application configuration settings.

[0080] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications onto cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage devices, but they have control over the deployed applications and the configuration of any application hosting environment.

[0081] Infrastructure as a Service (IaaS): This provides consumers with the capability to deliver processing, storage, networking, and other basic computing resources that enable them to deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0082] The deployment model is as follows:

[0083] Private cloud: Cloud infrastructure operated solely by an organization. It can be managed by the organization or a third party and can exist inside or outside a building.

[0084] Community cloud: Cloud infrastructure shared by several organizations and supporting a specific community with shared concerns (e.g., tasks, security requirements, policies, and compliance considerations). It can be managed by an organization or a third party and can exist on-site or off-site.

[0085] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.

[0086] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain a single entity but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursting for load balancing between clouds).

[0087] Cloud computing environments are service-oriented, focusing on statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is the infrastructure of a network of interconnected nodes.

[0088] Now for reference Figure 9 The diagram illustrates an illustrative cloud computing environment 2050. As shown, the cloud computing environment 2050 includes one or more cloud computing nodes 2010 that can communicate with local computing devices used by cloud computing consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 2054A, a desktop computer 2054B, a laptop computer 2054C, and / or an automotive computer system 2054N. The nodes 2010 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 2050 to provide cloud consumers with infrastructure, platforms, and / or software-as-a-service that eliminates the need for them to maintain resources on their local computing devices. It should be understood that... Figure 9 The types of computing devices 2054A-N shown are intended to be illustrative only, and the computing node 2010 and cloud computing environment 2050 can communicate with any type of computing device on any type of network and / or network-addressable connection (e.g., using a web browser).

[0089] Now for reference Figure 10 This demonstrates the 2050 cloud computing environment ( Figure 9 This provides a set of functional abstractions. It should be understood beforehand that... Figure 10 The components, layers, and functions shown are for illustrative purposes only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0090] The hardware and software layer 2060 includes hardware and software components. Examples of hardware components include: a mainframe 2061; a server 2062 based on a RISC (Reduced Instruction Set Computer) architecture; a server 2063; a blade server 2064; a storage device 2065; and network and networking components 2066. In some embodiments, software components include network application server software 2067 and database software 2068.

[0091] The virtualization layer 2070 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 2071; virtual storage device 2072; virtual network 2073, including virtual private network; virtual application and operating system 2074; and virtual client 2075.

[0092] In one example, the management layer 2080 may provide the following functionalities: Resource Provisioning 2081 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and Pricing 2082 provides cost tracking when utilizing resources in the cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User Portal 2083 provides access to the cloud computing environment for consumers and system administrators. Service Level Management 2084 provides cloud resource allocation and management to ensure that required service levels are met. Service Level Agreement (SLA) Planning and Fulfillment 2085 provides pre-scheduling and procurement of cloud resources, where future needs are anticipated according to the SLA.

[0093] The workload layer 2090 provides examples of functionalities that can be leveraged in a cloud computing environment. Examples of workloads and functionalities that can be provided from this layer include: mapping and navigation 2091; software development and lifecycle management 2092; virtual classroom education delivery 2093; data analysis and processing 2094; transaction processing 2095; and automated methods for sorting multiple datasets based on dataset content and desired data attributes 2096.

[0094] Various embodiments of the invention have been described and presented for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Similarly, the examples of features or functions of the embodiments of this disclosure described herein, whether used in the description of specific embodiments or listed as examples, are not intended to limit the embodiments of this disclosure described herein, or to restrict the disclosure to the examples described herein. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technical techniques in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method for sorting multiple datasets based on dataset attributes, comprising: The computer identifies the target set of data fields from a set of process documents, which indicate the user's data field preferences; The computer identifies a set of target dataset attributes from a set of data usage documents, the data usage documents indicating the user’s data scope preferences, and the attributes include data attributes that are obtained automatically or added to the set of target dataset attributes as part of input provided by domain experts; Multiple meta-datasets are generated by computers for multiple related datasets; The computer determines candidate datasets with field fitness values ​​that exceed a predetermined fitness threshold, the field fitness values ​​representing the degree of similarity between the set of fields associated with the dataset and the target data set of fields; The computer evaluates the associated meta-dataset for each candidate dataset regarding the target attribute, and the computer generates a comparative attribute score for each candidate dataset, the comparative attribute score indicating the degree to which the associated dataset will have content that demonstrates the attribute of the target dataset; as well as The computer generates a list of candidate datasets sorted according to the scores of the comparison attributes.

2. The method of claim 1, wherein the data usage document includes information in a format selected from a list consisting of Business Process Execution Language (BEPL) and Unified Modeling Language (UML).

3. The method of claim 1, wherein the data target attribute is extracted from elements of the process document, the elements being selected from a list consisting of class diagrams, activity diagrams, sequence diagrams, and component diagrams.

4. The method of claim 1, further comprising designating the candidate dataset with the highest comparison attribute score as the selected dataset.

5. The method of claim 4, further comprising: establishing a set of search parameters for a search to be performed on the selected dataset; and updating a history usage field in the metadata set associated with the dataset selected for the search using search context values ​​representing aspects of the search parameters.

6. The method of claim 5, wherein the sorting is based at least in part on the historical usage field value.

7. The method of claim 1, wherein the comparison attribute score is based at least in part on an associated desirability value associated with each of the target dataset attributes.

8. The method of claim 1, wherein the metadata set comprises information selected from a list consisting of: domain, gender, age group, geographic distribution, demographic distribution, statistical range of values, and context of applicability.

9. A system for sorting multiple datasets based on dataset attributes, the system include: A computer system includes a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a computer to cause the computer to perform the method according to any one of claims 1-8.

10. A computer program product for sorting multiple datasets according to dataset attributes, the computer program product comprising program instructions executable by a computer to cause the computer to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • System and method for determining and optimizing resources of a data processing system utilized by a service request

    US20080317217A1

  • Data matching

    US20140156652A1

  • Systems and methods of data analytics

    US20140337320A1

  • Automatic database analysis

    US20190108230A1

  • Data processing simulator

    US20190377557A1