Assembly line generation method and related equipment
By grouping data cleaning operators by attribute groups and performing topological sorting, a pipeline supporting parallel computing is generated, which solves the problems of complexity and time consumption in data cleaning tasks and realizes the generation of efficient and accurate data cleaning pipelines.
Patent Information
- Application Number
- CN202410605239.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
In the absence of comprehensive prior data labels and unified standardized processes, existing technologies make data cleaning tasks complex and time-consuming, and there is an urgent need to improve the efficiency of automated processing and shorten processing time.
By extracting data from fine-grained data cleaning operators, multiple attribute groups supporting parallel computing are generated. The intra-group sorting of operator groups is determined based on the topological sorting results of the attribute groups, thereby achieving coarse-grained pipeline recommendation and optimizing the operator execution order.
It significantly reduces the number of operator combinations in the search, improves pipeline generation efficiency, shortens execution time, and ensures high-precision data cleaning results.
Smart Images

Figure CN120973774A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular to a pipeline generation method, system, computing device cluster, computer-readable storage medium, and computer program product. Background Technology
[0002] In computing, a pipeline divides the entire workflow into a series of consecutive stages or tasks, achieving efficient production or processing by using the output of each stage as the input of the next. Each stage focuses on a specific task and passes its results to the next stage, allowing the entire process to continue continuously.
[0003] Pipelines can be used in various large-scale data processing scenarios. For example, training large language models (LLMs) typically requires massive amounts of data, such as gigabytes (GB) or even terabytes (TB). Data quality significantly impacts the effectiveness of LLM. One way to improve data quality is through data cleaning; LLM can achieve or even surpass the performance of using proprietary datasets using only finely cleaned data.
[0004] However, real-world data often has shortcomings, especially in the absence of comprehensive prior data labeling and standardized processes, making data cleaning a complex and time-consuming task. The industry urgently needs a pipeline generation method to support automated processing, improve efficiency, and shorten processing time. Summary of the Invention
[0005] This application provides a pipeline generation method. This method extracts multiple attribute groups supporting parallel computation and their ordering by extracting fine-grained data cleaning operators, thereby achieving coarse-grained pipeline recommendation. Furthermore, the method determines the intra-group order of operators corresponding to the attribute groups based on the topological sorting results, resulting in a more accurate operator execution order and improved repair accuracy. Because this method first achieves coarse-grained pipeline recommendation by grouping attributes and then intelligently orchestrates the intra-group operator execution order, it significantly reduces the number of operator combinations that need to be searched, improving pipeline generation efficiency. Moreover, while maintaining accuracy, parallelism shortens execution time and improves execution efficiency.
[0006] Firstly, this application provides a pipeline generation method. This method can be executed by a pipeline generation system. A pipeline generation system is a tool used to generate pipelines (such as data cleaning pipelines). The pipeline generation system can be a software system, which can be a standalone software system or a plugin or functional module integrated into other software systems. The software system can be provided to the user in the form of a software package, which the user can deploy on computing devices, such as local computing devices or private clouds. The software system can also be provided to the user as a cloud service; for example, cloud providers can provide interfaces with pipeline generation capabilities for easy invocation. In some examples, the pipeline generation system can be a hardware system, such as a cluster of computing devices with pipeline generation capabilities. When the computing device cluster runs, it executes the pipeline generation method of this application.
[0007] Specifically, the pipeline generation system can acquire multiple data cleaning operators for pipeline generation, as well as the attribute relationships of the data. The pipeline orchestrates these data cleaning operators so that they execute according to the orchestration results. The data cleaning operators remove or repair abnormal data in the dataset. The pipeline generation system can extract attribute sets for the multiple data cleaning operators, including attributes of the input data and attributes of the output data. Then, the pipeline generation system can group the attributes in the attribute set according to the attribute relationships, obtaining multiple attribute groups supporting parallel computation and their order. Next, based on the topological sorting result of at least one attribute group, the pipeline generation system determines the intra-group order of the operator group corresponding to at least one attribute group. Based on the order of the multiple attribute groups and the intra-group order of the operator group corresponding to at least one attribute group, the pipeline generation system orchestrates the multiple data cleaning operators to obtain the pipeline.
[0008] In this method, the pipeline generation system extracts fine-grained data cleaning operators into multiple attribute groups that support parallel computing and obtains the ranking of these attribute groups, thereby achieving coarse-grained pipeline recommendation. Furthermore, this method determines the intra-group ranking of the operator groups corresponding to the attribute groups based on the topological ranking results of the attribute groups, which can obtain a more accurate operator execution order and improve repair accuracy. Because this method first achieves coarse-grained pipeline recommendation by grouping attributes and then intelligently orchestrates the intra-group operator execution order, it significantly reduces the number of operator combinations that need to be searched, solves the combinatorial explosion problem and the efficiency problem of the heuristic search stage in related technologies, and improves pipeline generation efficiency. Moreover, the data cleaning pipeline generated by this method considers the dependencies and order between operators, and has high accuracy. While ensuring accuracy, parallelism shortens the execution time and improves execution efficiency.
[0009] In some possible implementations, the pipeline generation system can acquire the contribution of multiple data cleaning operators, which indicates the data repair capability of these operators. For example, the contribution could be the repair rate of the data cleaning operators, estimated based on existing operators and representative samples, representing the number of data items that can be repaired by each operator per column of data. Accordingly, the pipeline generation system can determine the intra-group ranking of the operator group corresponding to at least one attribute group based on the topological sorting result of at least one attribute group and the contribution of the multiple data cleaning operators.
[0010] This method improves the accuracy of sorting and enhances the precision of repair by sorting the data cleaning operators based on the topological sorting results and incorporating their contribution.
[0011] In some possible implementations, the pipelined generation system can also perform dependency detection on the connected subgraphs corresponding to at least one of multiple attribute groups. When cyclic dependencies (also known as conflict cycles) exist, the pipelined generation system can disconnect the incoming edges of the target attribute column in the connected subgraph to obtain an acyclic attribute graph. The target attribute column is the attribute column in the connected subgraph that meets the quality requirements, such as the highest quality attribute column. The pipelined generation system can then perform topological sorting based on the acyclic attribute graph to obtain the topological sorting results for at least one attribute group.
[0012] This method obtains an acyclic attribute graph by detecting and cutting the cyclic dependencies of the connected subgraphs corresponding to the attribute groups. Topological sorting based on this acyclic attribute graph can improve the accuracy of the sorting.
[0013] In some possible implementations, the pipeline generation system can determine the target attribute column from the connected subgraph based on the quality of the attribute columns recorded in the knowledge base. The target attribute column includes the attribute columns with the highest quality. Alternatively, the pipeline generation system can obtain the target attribute column from the connected subgraph based on the primary key attribute set. The target attribute column is the attribute column included in the master data of the primary key attribute set. Here, the master data can be data shared between various systems, such as data to be shared between operational / transactional application systems and analytical systems. The primary key attribute set includes at least one primary key. The pipeline generation system can use a natural key mining algorithm to mine unique and meaningful fields or combinations of fields from the dataset (the set of data to be cleaned) as candidate keys, and then determine the primary key from the candidate keys.
[0014] This method can determine the highest quality attribute column based on the quality of the attribute columns recorded in the database, or determine the attribute columns included in the master data in the primary key attribute set as the highest quality attribute column, and use this as the basis to cut off circular dependencies, thus ensuring the accuracy of topological sorting.
[0015] In some possible implementations, the pipeline generation system can identify the semantics of multiple data cleaning operators. The semantics of a data cleaning operator can be its physical meaning, typically a functional description in natural language. Then, based on the semantics of the multiple data cleaning operators and the pipeline, the pipeline generation system generates an interpretation of the data cleaning operators in the pipeline using a pipeline interpretation template.
[0016] This method enhances the interpretability of the pipeline by generating interpretations of data cleaning operators in the pipeline based on the pipeline interpretation template through the semantics of data cleaning operators, and helps users understand the role of each node (or operator) in the pipeline.
[0017] In some possible implementations, the pipelined generation system can construct a directed dependency graph corresponding to an attribute set based on attribute associations and the attribute set itself. In this graph, vertices represent attributes within the attribute set, and edges represent dependencies between those attributes. The pipelined generation system can then perform connectivity analysis on the directed dependency graph to obtain multiple connected subgraphs. Based on these connected subgraphs, the system can group the attributes within the attribute set to obtain multiple attribute groups that support parallel computation.
[0018] This method constructs a dependent directed graph and performs connectivity analysis to obtain a connected subgraph. Based on the connected subgraph, attributes can be reasonably grouped, which helps to improve the generation and execution efficiency of the pipeline.
[0019] In some possible implementations, the pipelined generation system can group attributes belonging to the same connected subgraph into the same attribute group, resulting in multiple attribute groups. The pipelined generation system can then adjust at least one of these attribute groups based on the parallelism and the number of operators corresponding to the multiple connected subgraphs.
[0020] This method optimizes the attribute grouping results by combining the degree of parallelism and the number of operators (operator size or operator magnitude) corresponding to the connected subgraph, thereby further improving the accuracy of grouping and meeting the requirements of parallelism.
[0021] Secondly, this application provides a production line system. The system includes:
[0022] An attribute grouping subsystem is used to obtain multiple data cleaning operators for generating the pipeline, and to obtain the attribute associations of the data. The pipeline is used to orchestrate the multiple data cleaning operators so that the multiple data cleaning operators are executed according to the orchestration result. The data cleaning operators are used to remove or repair abnormal data in the dataset.
[0023] The attribute grouping subsystem is also used to extract the attribute set of the plurality of data cleaning operators, the attribute set including the attributes of the input data of the plurality of data cleaning operators and the attributes of the output data of the plurality of operators;
[0024] The attribute grouping subsystem is further configured to group the attributes in the attribute set according to the attribute association relationship, thereby obtaining multiple attribute groups that support parallel computing and the sorting of the multiple attribute groups.
[0025] An operator arrangement subsystem is used to determine the intra-group order of the operator group corresponding to the at least one attribute group based on the topological sorting result of at least one attribute group among the plurality of attribute groups;
[0026] The operator orchestration subsystem is further configured to orchestrate the multiple data cleaning operators according to the sorting of the multiple attribute groups and the intra-group sorting of the operator groups corresponding to the at least one attribute group, thereby obtaining a pipeline.
[0027] In some possible implementations, the operator orchestration subsystem is also used for:
[0028] Obtain the contribution of the plurality of data cleaning operators, wherein the contribution is used to indicate the data repair capability of the plurality of data cleaning operators;
[0029] The operator orchestration subsystem is specifically used for:
[0030] Based on the topological sorting result of at least one of the multiple attribute groups and the contribution of the multiple data cleaning operators, the intra-group sorting of the operator group corresponding to the at least one attribute group is determined.
[0031] In some possible implementations, the operator orchestration subsystem is also used for:
[0032] Dependency detection is performed on the connected subgraph corresponding to at least one of the plurality of attribute groups;
[0033] When a circular dependency exists, disconnect the incoming edges of the target attribute column in the connected subgraph to obtain an acyclic attribute graph, where the target attribute column is the attribute column in the connected subgraph that meets the quality requirements.
[0034] A topological sort is performed based on the acyclic attribute graph to obtain the topological sorting result of the at least one attribute group.
[0035] In some possible implementations, the operator orchestration subsystem is also used for:
[0036] Based on the quality of the attribute columns recorded in the knowledge base, the target attribute columns are determined from the connected subgraph, wherein the target attribute columns include the attribute columns with the highest quality; or...
[0037] Based on the primary key attribute set, the target attribute column is obtained from the connected subgraph, where the target attribute column is the attribute column included in the primary key attribute set.
[0038] In some possible implementations, the attribute grouping subsystem is also used for:
[0039] Identify the semantics of the multiple data cleaning operators;
[0040] The system also includes:
[0041] An interpretation generation subsystem is used to generate interpretations of the data cleaning operators in the pipeline based on the semantics of the plurality of data cleaning operators and the pipeline, using a pipeline interpretation template.
[0042] In some possible implementations, the attribute grouping subsystem is specifically used for:
[0043] Based on the attribute associations and the attribute set, a dependency directed graph corresponding to the attribute set is constructed. The vertices in the dependency directed graph represent the attributes in the attribute set, and the edges in the dependency directed graph represent the dependencies of the attributes in the attribute set.
[0044] Connectivity analysis is performed on the dependent directed graph to obtain multiple connected subgraphs;
[0045] Grouping the attributes in the attribute set according to the multiple connected subgraphs yields multiple attribute groups that support parallel computation.
[0046] In some possible implementations, the attribute grouping subsystem is specifically used for:
[0047] The attributes belonging to the same connected subgraph in the multiple connected subgraphs are grouped into the same attribute group to obtain multiple attribute groups;
[0048] Adjust at least one attribute group among the multiple attribute groups based on the parallelism and the number of operators corresponding to the multiple connected subgraphs.
[0049] Thirdly, this application provides a computing device cluster. The computing device cluster includes at least one computing device, which includes at least one processor and at least one memory. The at least one processor and the at least one memory communicate with each other. The at least one processor is configured to execute instructions stored in the at least one memory to cause the computing device or the computing device cluster to perform the method as described in the first aspect or any implementation thereof.
[0050] Fourthly, this application provides a computer-readable storage medium storing instructions that instruct a computing device or a cluster of computing devices to perform the method described in the first aspect or any implementation thereof.
[0051] Fifthly, this application provides a computer program product containing instructions that, when run on a computing device or a cluster of computing devices, causes the computing device or cluster of computing devices to perform the method described in the first aspect or any implementation thereof.
[0052] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0053] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.
[0054] Figure 1 This application provides a schematic diagram of a data cleaning pipeline;
[0055] Figure 2 A flowchart for generating a fine-grained cleaning operator sequence is provided in this application;
[0056] Figure 3 This application provides a schematic diagram of the architecture of a pipeline production system;
[0057] Figure 4 A flowchart of a pipeline generation method provided in this application;
[0058] Figure 5 This application provides a flowchart for generating a data cleaning operator based on a representative sample of user-given abnormal data;
[0059] Figure 6 A flowchart illustrating the various stages of pipeline generation provided in this application;
[0060] Figure 7 A flowchart illustrating an attribute group recommendation process provided in this application;
[0061] Figure 8 A schematic diagram of a dependent directed graph provided in this application;
[0062] Figure 9 A schematic diagram illustrating an attribute group adjustment provided in this application;
[0063] Figure 10 A flowchart illustrating a pipeline-level recommendation process provided for this application;
[0064] Figure 11 A schematic diagram illustrating data clustering provided in this application;
[0065] Figure 12 This application provides a schematic diagram of the intra-group sorting of an operator group;
[0066] Figure 13 A schematic diagram illustrating the generation of an interpretable pipeline provided in this application;
[0067] Figure 14 A schematic diagram of an interpretable production line provided for this application;
[0068] Figure 15 A schematic diagram of the structure of a computing device provided in this application;
[0069] Figure 16 This application provides a schematic diagram of the structure of a computing device cluster;
[0070] Figure 17 This application provides a schematic diagram of another computing device cluster structure.
[0071] Figure 18 This is a schematic diagram of another computing device cluster provided in this application. Detailed Implementation
[0072] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0073] First, some technical terms involved in the embodiments of this application will be introduced.
[0074] Data cleaning is the process of detecting and correcting (or deleting) corrupt or inaccurate records from a recordset, database table, or database. Specifically, data cleaning may include identifying incomplete, incorrect, inaccurate, or irrelevant portions of data and then performing replacement, modification, or deletion operations.
[0075] Data cleaning operators are a set of tools or functions used in the data cleaning process to remove or repair outlier data in a dataset. Outlier data can include, but is not limited to, missing data, duplicate data, or erroneous data. The tools used in the data cleaning process can include custom expressions or model operators, where model operators are computational units within a model used to implement specific functions. Data cleaning operators can obtain high-quality data by performing specific cleaning tasks, such as identifying and correcting erroneous data. Data cleaning operators can be categorized according to their cleaning granularity. For example, data cleaning operators for database tables can include cell-level data cleaning operators and column-level data cleaning operators.
[0076] To facilitate understanding, this application provides relevant examples for illustration.
[0077] Cell-level data cleaning operators can include the following operators:
[0078] Replace"British"with'English'
[0079] Replace "film-maker." with 'filmmaker'
[0080] Column-level data cleaning operators can include the following operators:
[0081] FORMAT_DATE(date,'yyyy / MM / DD')
[0082] UNIX_TIMESTAMP(unix,'yyyy / MM / DD')
[0083] A data cleaning pipeline is a pipeline that implements data cleaning, representing the data cleaning process. The pipeline orchestrates multiple data cleaning operators so that they are executed according to the orchestration. The data cleaning process can consist of a series of cleaning nodes, each containing a series of fine-grained cleaning operators (or simply operators), with the operators within a node having a specific order. The execution order of the cleaning nodes also also matters. This application provides an example of a data cleaning pipeline, such as... Figure 1As shown, Date, Gender, FnameToGender, and AreacodeToState are cleaning nodes. Among them, the operators in the Date cleaning node are used to modify the date into a specified format. Similarly, the operators in the Gender cleaning node are used to modify the gender into a specified format. For example, the data cleaning operator Replace “男” with ‘M’ in the Gender cleaning node is used to replace “男” in the cell with “M”, the data cleaning operator Replace “女” with ‘F’ is used to replace “女” in the cell with “F”, and the data cleaning operator Replace “Female” with ‘F’ is used to replace “Female” in the cell with “F”.
[0084] Data in the real world often has diverse defects. In the absence of comprehensive prior data labels and a unified standardization process, data cleaning is often a complex and time-consuming task. For this reason, related technologies provide a repair system for various data error types. This repair system supports allowing users to add existing data cleaning methods through a unified interface, and the data cleaning methods form a parameterized data cleaning library. This repair system can generate a series of intermediate results, which can be operators and parameters under a series of established goals, and then convert the above intermediate results into an intermediate language, and then use tree search and pruning to heuristically search for the intermediate results to obtain a data cleaning pipeline. This data cleaning pipeline can be a sequence of fine-grained cleaning operators. Figure 2 A flowchart for generating a sequence of fine-grained cleaning operators is shown. In order to evaluate the generated data cleaning pipeline, quality evaluation metric definitions can be carried out first, such as defining quality evaluation indicators or quality evaluation models. Among them, the quality evaluation indicators can include repair accuracy. Then a candidate set of cleaning operations can be generated. The candidate set of cleaning operations can include a series of cleaning operation operators (data cleaning operators) to be followed. A sequence of cleaning operators is generated through a search algorithm, and the sequence of cleaning operators is evaluated based on the defined quality evaluation indicators. If the quality evaluation indicators do not meet the requirements, continue to search to generate a new sequence of cleaning operators and conduct quality evaluation until the quality evaluation indicators meet the requirements, and finally obtain a sequence of fine-grained cleaning operators that meet the quality evaluation indicators. This sequence of cleaning operators can be used to clean large-scale data and output the cleaned data.
[0085] The above repair system for various data error types provides significant flexibility and adaptability. However, the combinatorial explosion problem that may occur during the data repair process and the efficiency problem in the heuristic search stage are the main challenges faced by the repair system.
[0086] In view of this, this application provides a pipeline generation method. This method can be executed by a pipeline generation system. A pipeline generation system is a tool used to generate pipelines, such as a data cleaning pipeline. The pipeline generation system can be a software system, which can be a standalone software system or a plugin or functional module integrated into other software systems. The software system can be provided to the user as a software package, which the user can deploy on computing devices, such as local computing devices or private clouds. The software system can also be provided to the user as a cloud service; for example, public cloud providers can provide application programming interfaces (APIs) with pipeline generation capabilities for user use. In some examples, the pipeline generation system can be a hardware system, such as a cluster of computing devices with pipeline generation capabilities. When the computing device cluster runs, it executes the pipeline generation method of this application.
[0087] Specifically, the pipeline generation system can acquire multiple data cleaning operators for pipeline generation, as well as the attribute relationships of the data. The pipeline orchestrates these data cleaning operators so that they execute according to the orchestration results. Data cleaning operators remove or repair outlier data in the dataset. The pipeline generation system can extract attribute sets for the multiple data cleaning operators, including attributes of the input data and attributes of the output data. Then, the pipeline generation system can group the attributes in the attribute set according to the attribute relationships, obtaining multiple attribute groups supporting parallel computation and their order. Next, the pipeline generation system can determine the intra-group order of the operator group corresponding to at least one attribute group based on the topological sorting result of at least one attribute group. Finally, the pipeline generation system orchestrates the multiple data cleaning operators based on the order of the attribute groups and the intra-group order of the operator group corresponding to at least one attribute group to obtain the pipeline.
[0088] This method extracts fine-grained data cleaning operators upwards to obtain multiple attribute groups supporting parallel computing and their ordering, thereby achieving coarse-grained pipeline recommendation. Furthermore, based on the topological sorting of attribute groups, the method determines the intra-group order of operators corresponding to those attribute groups, resulting in a more accurate operator execution order and improved repair accuracy. Because this method first achieves coarse-grained pipeline recommendation by grouping attributes and then intelligently orchestrates the intra-group operator execution order, it significantly reduces the number of operator combinations to be searched, solving the combinatorial explosion problem and efficiency issues in the heuristic search phase in related technologies, thus improving pipeline generation efficiency. Moreover, the data cleaning pipeline generated by this method considers the dependencies and order of operators, exhibiting high accuracy. By maintaining accuracy and employing parallelism, it shortens execution time and improves execution efficiency.
[0089] To make the technical solution of this application clearer and easier to understand, the architecture of the pipeline production system of this application will be described below with reference to the accompanying drawings.
[0090] See Figure 3 The diagram illustrates the architecture of a pipelined generation system 10, which includes an attribute grouping subsystem 100 and an operator arrangement subsystem 200. Further, the pipelined generation system 10 may also include an interpretation generation subsystem 300. Similar to the pipelined generation system 10, the attribute grouping subsystem 100, operator arrangement subsystem 200, and interpretation generation subsystem 300 can be implemented in software or hardware. The attribute grouping subsystem 100, operator arrangement subsystem 200, and interpretation generation subsystem 300 will be described below.
[0091] The attribute grouping subsystem 100 is used to acquire multiple data cleaning operators for generating the pipeline, as well as the attribute relationships of the data, and then extract attribute sets from the multiple data cleaning operators. The pipeline orchestrates the multiple data cleaning operators so that they are executed according to the orchestration result. The data cleaning operators remove or repair abnormal data in the dataset. The attribute sets include attributes of the input data and the output data of the multiple data cleaning operators. For example, the attributes of the data cleaning operators may include "zip", "areacode", "state", etc. The attribute grouping subsystem 100 is also used to group the attributes in the attribute set according to the attribute relationships, obtaining multiple attribute groups supporting parallel computing and the order of the multiple attribute groups. The multiple attribute groups supporting parallel computing can also be called parallel attribute blocks, or simply parallel blocks. Extracting parallel blocks from the operators enables coarse-grained pipeline recommendation.
[0092] The operator orchestration subsystem 200 is used to determine the intra-group order of the operator groups corresponding to at least one attribute group based on the topological sorting result of at least one attribute group among multiple attribute groups. The operator orchestration subsystem 200 is also used to orchestrate multiple data cleaning operators based on the sorting of multiple attribute groups and the intra-group order of the operator groups corresponding to at least one attribute group, thereby obtaining a pipeline. This enables intelligent orchestration of operator execution order based on the topological sorting result of attribute groups, improving execution accuracy. Furthermore, this method can achieve distributed parallel operation, improving execution efficiency.
[0093] The interpretation generation subsystem 300 is used to generate interpretations for the data cleaning operators in the pipeline based on the semantics and pipeline of the multiple data cleaning operators, using a pipeline interpretation template. The semantics of the multiple data cleaning operators can be identified by the attribute grouping subsystem 100, for example, when extracting the attribute set of the multiple data cleaning operators, the attribute grouping subsystem 100 identifies the semantics of the multiple data cleaning operators. The attribute grouping subsystem 100 provides the semantics of the multiple data cleaning operators to the interpretation generation subsystem 300, and the operator orchestration subsystem 200 provides the generated pipeline to the interpretation generation subsystem 300. Accordingly, the interpretation generation subsystem 300 can generate interpretations for each data cleaning operator in the pipeline, combining the semantics and the pipeline interpretation template.
[0094] It should be noted that the input to the pipeline generation system 10 may include multiple data cleaning operators and attribute relationships. These data cleaning operators can be fine-grained cleaning operators, and the fine-grained cleaning operators can be based on... Figure 1 The proposed solution can be implemented in this application. Figure 1 Based on the proposed solution, recommendations for parallel blocks are added to determine the range of operators that can be executed in parallel. Then, pipeline-level recommendations can be performed, such as internally sorting and dividing the determined range of parallelizable operators to generate the final data cleaning pipeline. See details... Figure 3 The system generates cleaning parameters from a large-scale dataset, then dynamically optimizes and evaluates these parameters to obtain satisfactory cleaning parameters. Based on these satisfactory parameters, fine-grained data cleaning operators can be generated. Multiple generated data cleaning operators can be used to generate a pipeline. Furthermore, attribute association analysis based on the aforementioned large-scale dataset generates attribute associations, which can be input into the pipeline generation system 10 to generate the pipeline. Further, the input to the pipeline generation system 10 can also include a knowledge base, which can store historically generated attribute associations. After generating the pipeline, the user can confirm execution to output the cleaned data.
[0095] Based on the aforementioned pipeline generation system 10, this application also provides a pipeline generation method. See [link to relevant documentation]. Figure 4The flowchart shown is a method for generating a pipeline, which includes the following steps:
[0096] S402, Pipeline generation system 10 acquires multiple data cleaning operators for generating pipelines.
[0097] Data cleaning operators are operators with data cleaning capabilities used to remove or repair outlier data in a dataset. Outlier data can be missing data, duplicate data, or erroneous data. Missing data can be data with null values or missing attribute columns in a database table. Duplicate data can be data with repeated values, and erroneous data refers to data with incorrect formatting or values.
[0098] Data cleaning operators can be of various types, such as user-defined expressions, functions, or model operators. Data cleaning operators can be fine-grained, such as functions with data cleaning capabilities, used to perform specific data cleaning tasks, including but not limited to the identification and correction of erroneous data. Furthermore, data cleaning operators can be categorized into cell-level and column-level data cleaning operators based on their cleaning granularity. Cell-level data cleaning operators can be used to repair data within cells, while column-level data cleaning operators are used to repair data in attribute columns.
[0099] A pipeline is used to orchestrate multiple data cleaning operators so that these operators are executed according to the orchestration result. The pipeline that orchestrates the data cleaning operators is also called a data cleaning pipeline, used for automated data cleaning. In specific implementation, the pipeline generation system 10 can obtain a user-defined list of operators, which includes multiple data cleaning operators. The pipeline generation system 10 can also obtain a recommended list of operators based on the data to be cleaned (such as a dataset). This list of operators includes a series of fine-grained data cleaning operators recommended based on relevant technologies.
[0100] The recommended data cleaning operators can be generated in different ways. One method can be found here. Figure 2 Specifically, it involves dynamically searching the candidate set of cleaning operations based on user-provided quality assessment metrics to determine the recommended data cleaning operators (sequences). Another approach can be found in [link to relevant documentation]. Figure 5 Specifically, it involves generating data cleaning operators based on representative samples of user-provided abnormal data. For example... Figure 5As shown, users can label outlier data and then fine-tune the model based on this labeling. The fine-tuned model generates a candidate set of potentially correct data. Features can be extracted and clustered from this candidate set, allowing data to be extracted from each category. The fine-tuned model then predicts the final cleaned data based on the extracted data. Users can decide whether to use the fine-tuned model (i.e., the model operator) as the data cleaning operator based on the final cleaned data's performance, such as accuracy or repair rate. Once the user confirms the use of the data cleaning operator, large-scale datasets can be cleaned using this operator to obtain cleaned data.
[0101] S404, the attribute association relationship of the data obtained by the pipeline generation system 10.
[0102] Attribute relationships can include dependencies between attributes, such as functional dependencies (FD) and multivalued dependencies (MD). Furthermore, based on functional dependencies, new dependency relationships have been extended, such as conditional functional dependencies (CFD).
[0103] In database theory, functional dependency refers to a specific type of relation where the value of one attribute set uniquely determines the value of another attribute set. Let R(U) be a relation schema on attribute U, X and Y be subsets of U = {A1, A2, ..., An}, and r be any relation of R. If for any two tuples u and v in r, if u[X] = v[X], then u[Y] = v[Y], then X is said to functionally determine Y, or Y is said to be functionally dependent on X, denoted as X→Y. For example, if in a database table, student ID (denoted as sno, as one attribute set) uniquely determines student name (denoted as sname, another attribute set), then student ID has a functional dependency on student name, sno→sname.
[0104] Functional dependencies can be divided into full functional dependencies and partial functional dependencies. In R(U), if X→Y, and for any proper subset X' of X, X'→>Y, then Y is said to be fully dependent on X; otherwise, if X→Y, and there exists a proper subset X' of X such that X'→Y holds, then Y is said to be partially dependent on X.
[0105] Conditional functional dependency is a rule used for data consistency checks. It extends the traditional concept of functional dependency by allowing dependencies to be defined under specific conditions. Conditional functional dependency helps identify inconsistencies and errors in data by specifying how attribute columns in a dataset behave based on the values of other attribute columns. For example, if student ID (sno) determines student name (sname), but student name determines course name (cname) only when course ID (cno) has a specific value, then the conditional functional dependency is true, denoted as (sno, cno) → sname, cname.
[0106] Multivalued dependency is another type of dependency relationship in database theory. In a relational table, if the value of one attribute set A determines the value of another attribute set B, regardless of other attributes, then B is said to have a multivalued dependency on A. This means that if two rows in a database table have the same value in A, their values in B can be interchanged without affecting other data in the table.
[0107] Let R(U) be a relation schema on a set of attributes U, where X, Y, and Z are subsets of U, and Z = UXY. If for any relation r in R, each value of r on (X, Z) corresponds to a set of values for Y, and this set of values depends only on the values of X and is independent of the values of Z, then a multivalued dependency holds, denoted as X →→ Y. It should be noted that if a set in a multivalued dependency includes only one value, the multivalued dependency becomes a functional dependency; in other words, a functional dependency can be considered a special case of a multivalued dependency.
[0108] The pipeline generation system 10 can perform attribute association analysis on the data to be cleaned, such as a large-scale dataset, to obtain attribute association relationships.
[0109] S406, the pipeline generation system 10 extracts the attribute set of multiple data cleaning operators.
[0110] The attribute set includes attributes of the input data of multiple data cleaning operators and attributes of the output data of multiple operators. Furthermore, if intermediate data is involved during the execution of the data cleaning operators, the attribute set may also include attributes of the intermediate data involved in the data cleaning operators.
[0111] The pipeline generation system 10 can extract attribute sets of multiple data cleaning operators through information extraction templates or information extraction models (such as large models). These will be explained in detail below.
[0112] In some possible implementations, the pipeline generation system 10 can obtain a system definition, construct an information extraction template based on the system definition, and then match the data cleaning operator with the information extraction template to extract the attributes of the data cleaning operator. For example, the information extraction template can define keywords, and the token following the keyword can be an attribute of the input or output data of the data cleaning operator. Taking the data cleaning operator Replace "British" with 'English' as an example, the token following the keyword replace can be an attribute of the input data, and the token following the keyword with can be an attribute of the output data.
[0113] In some other possible implementations, the pipeline generation system 10 can use a large language model (LLM) to extract attributes of the data cleaning operators. The LLM can be obtained by fine-tuning a general large model based on historical data.
[0114] S408, the pipeline generation system 10 groups the attributes in the attribute set according to the attribute association relationship, and obtains multiple attribute groups that support parallel computing and the sorting of multiple attribute groups.
[0115] Specifically, the pipelined generation system 10 can construct a dependency directed graph corresponding to the attribute set based on attribute associations and the attribute set. The dependency directed graph is a directed graph representing the dependencies between attributes in the attribute set. Vertices in the dependency directed graph represent attributes in the attribute set, and edges in the dependency directed graph represent the dependencies between attributes in the attribute set. The pipelined generation system 10 can perform connectivity analysis on the dependency directed graph to obtain multiple connected subgraphs (spanning subgraphs). A connected subgraph is a subset of the dependency directed graph. The set of multiple connected subgraphs includes all vertices in the dependency directed graph (also called the original graph), but retains some edges of the dependency directed graph; these retained edges constitute the connected subgraph. For example, the set of subgraphs G' generated from the original graph G includes the same number of vertices as the original graph G, but the set of edges E' can be any subset of the original graph G. The pipelined generation system 10 can group the attributes in the attribute set based on the multiple connected subgraphs to obtain multiple attribute groups that support parallel computation.
[0116] like Figure 6As shown, the pipeline generation system 10 can group attributes in the attribute set based on parallelism or operator magnitude after constructing a dependent directed graph and performing connectivity analysis on the dependent directed graph to obtain attribute connected subgraphs. Here, the connected subgraph is equivalent to a block operation on the attributes. The pipeline generation system 10 can first initialize the attribute set based on the connected subgraphs to obtain multiple attribute groups. Then, the pipeline generation system 10 can dynamically adjust the multiple attribute groups based on parallelism and operator magnitude to obtain multiple attribute groups that support parallel computing.
[0117] In some possible implementations, the pipeline generation system 10 can group attributes belonging to the same connected subgraph into the same attribute group, obtaining multiple attribute groups. Then, the pipeline generation system 10 can adjust at least one attribute group among the multiple attribute groups according to the parallelism and the number of operators corresponding to the multiple connected subgraphs.
[0118] To facilitate understanding, this application also provides an example. In this example, the pipeline generation system 10 performs connectivity analysis on a dependent directed graph, obtaining two connected subgraphs. One connected subgraph includes vertices A and B, and the other includes vertices C, D, E, and F. Each vertex represents an attribute. The pipeline generation system 10 can obtain the attribute groups corresponding to each connected subgraph based on the above connected subgraphs, specifically represented as {A, B} and {C, D, E, F}. The number of operators corresponding to the connected subgraphs are 47 and 103, respectively. If the corresponding operators are executed directly according to the above attribute groups, it may cause some operators to wait for other operators to complete after execution. Assuming the user-defined parallelism is 2, the pipeline generation system 10 can further divide the attribute group {C, D, E, F} into parallelizable attribute groups, such as {C, D} and {E, F}, based on the parallelism and the number of operators corresponding to the connected subgraphs.
[0119] S410, the pipeline generation system 10 determines the intra-group sorting of the operator group corresponding to at least one attribute group based on the topological sorting result of at least one attribute group among multiple attribute groups.
[0120] Topological sorting is the process of ordering all nodes in a directed acyclic graph (DAG). It is typically used to describe the execution order of a series of tasks or events with dependencies. For example, in university course scheduling, topological sorting can help determine the order in which courses are taken, as there are sequential relationships between them.
[0121] The topological sorting result satisfies the following condition: if there exists a directed edge from node A to node B in the DAG, then node A must appear before node B in the sorting result. Specifically, in this application, the pipeline generation system 10 can perform topological sorting based on the connected subgraphs corresponding to attribute groups, thereby obtaining the topological sorting result for at least one attribute group. The pipeline generation system 10 can then determine the intra-group sorting of the operator groups corresponding to at least one attribute group based on the topological sorting result. Specifically, the operator corresponding to the attribute that ranks higher in the topological sorting result can also rank higher within the operator group.
[0122] like Figure 6 As shown, considering the possibility of cyclic dependencies (also known as dependency cycles) in some connected subgraphs, the pipeline generation system 10 can perform dependency detection on the connected subgraphs corresponding to at least one of the multiple attribute groups. When a cyclic dependency exists, the pipeline generation system 10 disconnects the incoming edges of the target attribute column in the connected subgraph to obtain an acyclic attribute graph. The target attribute column can be an attribute column in the connected subgraph that meets the quality requirements; for example, the target attribute column can be the attribute column with the highest quality. Then, the pipeline generation system 10 can perform topological sorting based on the acyclic attribute graph to obtain the topological sorting result for at least one attribute group.
[0123] In the above embodiments, the quality of attribute columns can be obtained from a knowledge base or determined based on the primary key attribute set. For example, the pipeline generation system 10 can determine target attribute columns from the connected subgraph based on the quality of attribute columns recorded in the knowledge base. These target attribute columns include the attribute columns with the highest quality. Alternatively, the pipeline generation system 10 can obtain target attribute columns from the connected subgraph based on the primary key attribute set. These target attribute columns are attribute columns included in the master data within the primary key attribute set. The master data can be data shared between various systems, such as data to be shared between operational / transactional application systems and analytical systems. The primary key attribute set includes at least one primary key. The pipeline generation system 10 can use a natural key mining algorithm to mine unique and meaningful fields or combinations of fields from the dataset (the set of data to be cleaned) as candidate keys, and then determine the primary key from the candidate keys.
[0124] The pipeline generation system 10 can map attribute groups to operator groups based on the attribute set of data cleaning operators. Then, the pipeline generation system can determine the intra-group sorting of the operator groups corresponding to the attribute groups based on the topological sorting results of the attribute groups. Further details can be found in [link to relevant documentation]. Figure 6Furthermore, the pipeline generation system can also sort the data cleaning operators within an operator group by combining their contribution values. Specifically, the pipeline generation system 10 can obtain the contribution values of multiple data cleaning operators, which are used to indicate the data repair capabilities of the multiple data cleaning operators. Accordingly, the pipeline generation system 10 determines the intra-group sorting of the operator group corresponding to at least one attribute group based on the topological sorting result of at least one attribute group among multiple attribute groups and the contribution values of the multiple data cleaning operators.
[0125] S412, the pipeline generation system 10 arranges multiple data cleaning operators according to the sorting of multiple attribute groups and the intra-group sorting of at least one attribute group's corresponding operator group to obtain a pipeline.
[0126] Specifically, the pipeline generation system 10 can arrange multiple data cleaning operators according to the sorting of attribute groups and the intra-group sorting of the operators corresponding to the attribute groups. This ensures that the arranged operators satisfy the inter-group and intra-group sorting of attribute groups, thereby obtaining the final data cleaning pipeline. When executing the above data cleaning pipeline, each node in the data cleaning pipeline can be arranged to generate a parallel cleaning plan, thereby achieving more efficient parallelism and helping users quickly obtain the cleaned data.
[0127] Furthermore, the pipeline generation system 10 can also recognize the semantics of multiple data cleaning operators. For example, when the pipeline generation system 10 extracts the attributes of data cleaning operators using an information extraction template or an information extraction model, it also extracts the semantics of the data cleaning operators using the aforementioned information extraction template or model. Then, the pipeline generation system 10 can generate interpretations of the data cleaning operators in the pipeline based on the semantics of multiple data cleaning operators and the pipeline, using a pipeline interpretation template.
[0128] As described above, the pipeline generation method of this application extracts fine-grained data cleaning operators upwards to obtain multiple attribute groups supporting parallel computing and their ordering, thereby achieving coarse-grained pipeline recommendation. Furthermore, this method determines the intra-group order of operators corresponding to the attribute groups based on the topological sorting results, resulting in a more accurate operator execution order and improved repair accuracy. This method first achieves coarse-grained pipeline recommendation by grouping attributes, then intelligently orchestrates the intra-group operator execution order, significantly reducing the number of operator combinations to be searched. This solves the combinatorial explosion problem and the efficiency problem of the heuristic search stage in related technologies, improving pipeline generation efficiency. Moreover, the data cleaning pipeline generated by this method considers the dependencies and order between operators, exhibiting high accuracy. While ensuring accuracy, parallelism shortens execution time and improves execution efficiency.
[0129] To make the technical solution of this application clearer and easier to understand, the pipeline generation method of this application will be introduced below in conjunction with specific application scenarios.
[0130] In this scenario, the inputs to the pipeline generation system 10 may include fine-grained data cleaning operators, attribute relationships, and a knowledge base. The fine-grained data cleaning operators may include custom expressions, Structured Query Language (SQL) functions, or model operators.
[0131] Custom expressions can be defined using keywords, such as `Replace "British" with "English"` or `Replace "film-maker." with 'filmmaker'`. SQL is a standard programming language for managing and manipulating relational databases. SQL allows users to perform various data operations, including querying, inserting, updating, deleting, and creating and modifying database structures. A key feature of SQL is its declarative approach to expressing data operations; users specify the desired operation without needing to describe how to perform it. For example, SQL functions can include `FORMAT_DATE(date, 'yyyy / MM / DD')` and `UNIX_TIMESTAMP(unix, 'yyyy / MM / DD')`. Model operators can include operators from semantic insufficiency models, where the semantic insufficiency model can be a Bidirectional Encoder Representations from Transformers (BERT) model.
[0132] Attribute relationships can include functional dependencies (denoted as X→Y), conditional functional dependencies (denoted as (X,Z)->(Y)), or multivalued dependencies X→→Y. The knowledge base includes knowledge of dependencies and can be used to expand upon attribute relationships.
[0133] Before performing attribute group recommendations, fine-grained data cleaning operators can be parsed to obtain their semantics and attributes. Specifically, the pipeline generation system 10 can extract the semantics and attributes of the input data cleaning operators using information extraction templates or information extraction models.
[0134] Attribute group recommendations can be performed based on the attributes of the data cleaning operator and the relationships between the input attributes. The attribute group recommendation process is explained in detail below.
[0135] See Figure 7 The flowchart shown illustrates a recommendation process for attribute groups, which includes the following steps:
[0136] S702, the pipeline generation system 10 parses the attributes of the data cleaning operator according to the attribute association relationship and obtains the attribute association relationship of the data cleaning operator.
[0137] S704, the pipeline generation system 10 constructs a dependency directed graph based on the attribute association relationship of the data cleaning operator.
[0138] In this scenario, the attribute relationships of the data cleaning operator can be as follows:
[0139] Domain = [
[0140] (["zip","areacode"]->["state"]),
[0141] (["zip"]->["city"]),
[0142] (["lname","fame"]->["gender"]) ]
[0144] The dependency directed graph constructed based on the attribute relationships of the data cleaning operators is as follows: Figure 8 As shown. Figure 8 The vertices in the data are attributes, including but not limited to "zip", "lname", "fname", "area code", "state", "gender", and "city". Figure 8 Directed edges in the code represent attribute relationships. For example, a directed edge from "zip" to "city" represents ["zip"]->["city"].
[0145] S706, the pipeline generation system 10 performs connectivity analysis on the dependent directed graphs to obtain connected subgraphs.
[0146] Specifically, the pipeline generation system 10 can perform connectivity analysis on the dependent directed graph using the Tarjan algorithm or Depth-First Search (DFS) to obtain connected subgraphs. The Tarjan algorithm is an algorithm for finding strongly connected components in a graph. It takes a directed graph as input and provides a partition of its vertex set according to its strongly connected components; each node in the directed graph appears in only one strongly connected component. A specific implementation of the Tarjan algorithm can include starting a depth-first search from any node, revisiting already visited nodes, and the subtrees of the search tree constitute the strongly connected components of the directed graph, with each strongly connected component corresponding to a connected subgraph.
[0147] S708, the pipeline generation system 10 obtains multiple attribute groups based on the attribute set formed by the attributes of the connected subgraph and the data cleaning operator.
[0148] Specifically, see Figure 7 The pipeline generation system 10 can generate parallel classification sets mi based on the connected subgraph. In this example, the parallel classification sets may include m1 = ["Zip", "City", "State", "areacode"] and m2 = ["fname", "lname", "gender"]. The pipeline generation system 10 can then determine whether the attribute set Si belongs to a subset mj. If so, Si is classified into mj; otherwise, Si is subdivided until the subdivided attribute sets belong to a subset mj. Once each attribute in the attribute set has been classified, multiple attribute groups are generated based on the classified subset mj. Each attribute group includes the attributes in the attribute set that are classified into the same subset mj.
[0149] S710 and the pipeline generation system 10 adjust at least one attribute group among multiple attribute groups according to the parallelism and the number of operators corresponding to the connected subgraph to obtain the final attribute group.
[0150] The pipeline generation system 10 can adjust multiple attribute groups based on the degree of parallelism and the number of fine-grained data cleaning operators in each connected subgraph through a dynamic optimization algorithm. For example... Figure 9 As shown, the pipeline generation system 10 can split attribute group 1 into two attribute groups, such as attribute group 1-1 and attribute group 1-2. The data cleaning operators corresponding to attribute group 1-1 and attribute group 1-2 can support parallel execution. Furthermore, the pipeline generation system 10 can also generate sorts of multiple attribute groups.
[0151] The following section provides a detailed explanation of pipeline-level recommendations. See [link / reference]. Figure 10 The diagram illustrates a pipeline-level recommendation process, which includes the following steps:
[0152] S1002, the pipeline generation system 10 extracts representative samples from the data to be cleaned.
[0153] The purpose of extracting representative samples (representative sample sampling) is to provide a basis for calculating the contribution of data cleaning operators. In practice, the pipeline generation system 10 can use clustering algorithms to group the data. Figure 11 The example illustrates that points of different colors or grayscale represent different clustering results. The results of the above clustering analysis can help users understand the distribution characteristics within data blocks and ensure that samples are drawn evenly from each cluster to form a representative sample set.
[0154] S1004, the pipeline generation system 10 calculates the quality of attribute columns based on representative samples.
[0155] S1006, the pipeline generation system 10 calculates the contribution of the data cleaning operator based on representative samples.
[0156] In this application, the quality of attribute columns can be characterized by a quality score, which can be the ratio of the number of abnormal data in each column to the total number of data in the column. The contribution of the data cleaning operator can be the ratio of the number of data that the data cleaning operator can repair to the total number of data in the column.
[0157] The pipeline generation system 10 can estimate the number of outliers in each attribute column (total number of data in the column) * 100% based on existing operators and representative samples (sample data) to obtain a quality score. Similarly, the pipeline generation system 10 can estimate the number of data that the data cleaning operator can repair (total number of data in the column) * 100% based on existing operators and representative samples to obtain the contribution of the data cleaning operator.
[0158] The quality scores of attribute columns and the contribution of data cleaning operators can be used to remove cycles when cycles exist in the dependent directed graph or connected subgraph, and to effectively sort the data cleaning operators corresponding to attributes within an attribute group.
[0159] S1008, Pipeline generation system 10 determines whether there is a circular dependency in the connected subgraph. If so, execute S1010.
[0160] Specifically, the pipeline generation system 10 can determine whether a connected subgraph has cyclic dependencies based on Depth-First Search (DFS) or the Kosaraju algorithm. The Kosaraju algorithm, also known as the Kosaraju-Sharir algorithm, is used to find strongly connected components of a directed graph in linear time. When cyclic dependencies exist, the pipeline generation system 10 can execute S1010 to break the cyclic dependencies (removing duplicates from the cycle), thereby solving the topological sorting problem.
[0161] S1010, the pipeline generation system 10 cuts off circular dependencies and obtains an acyclic attribute graph.
[0162] Specifically, the pipeline generation system 10 can identify attribute columns of connected subgraphs with cyclic dependencies and determine whether the highest-quality attribute column can be obtained from the knowledge base. If so, the pipeline generation system 10 directly obtains the highest-quality attribute column from the knowledge base, which is the target attribute column. The pipeline generation system 10 can disconnect all incoming edges of the target attribute column to obtain an acyclic attribute graph. If not, the pipeline generation system 10 performs an indirect data acquisition pattern, analyzes the primary key attribute set, and determines that the attribute columns included in the master data are the highest-quality attribute columns, i.e., the target attribute columns. The pipeline generation system 10 can disconnect all incoming edges of the target attribute column to obtain an acyclic attribute graph.
[0163] S1012, the pipeline generation system 10 can perform topological sorting based on the acyclic attribute graph, and solve the intra-group sorting of the operator group corresponding to the attribute group based on the topological sorting result and the contribution of the data cleaning operator.
[0164] Topological sorting can be implemented using algorithms such as Depth-First Search (DFS). When sorting the operator groups within a group, the pipeline generation system 10 can perform a coarse sort based on the topological sorting results, prioritizing dependent data cleaning operators. Then, the pipeline generation system 10 can perform a fine sort based on the contribution of the data cleaning operators; for example, based on the topological sorting, data cleaning operators with higher contributions are prioritized. Figure 12 As shown, the pipeline generation system can sort OP1 before OP2 and OP3 before OP4 based on the topology sorting results and the contribution of the data cleaning operators.
[0165] Further, see Figure 13 The pipeline generation system 10 can clean the semantics of operators and the generated pipelines based on the extracted data, and generate interpretable pipelines through pipeline interpretation templates. In this scenario, interpretable pipelines include... Figure 14 As shown, Figure 14 Each data cleaning operator in the code corresponds to an interpretation, and the dashed lines indicate parallel relationships.
[0166] Based on the above description, the pipeline generation method of this application extracts fine-grained data cleaning operators upwards to form a data cleaning pipeline. This pipeline effectively expresses the data cleaning process, with each node containing a series of fine-grained cleaning operators with the same or similar functions. This helps users quickly locate and handle anomalies that need to be repaired by the operators. Furthermore, the pipeline generation method of this application can recommend the execution order of parallel operations based on parallel capabilities (such as parallelism), effectively improving execution efficiency. A more accurate operator execution order can be obtained based on attribute relationships, improving repair accuracy. In addition, the pipeline generation method of this application can generate explanations of data cleaning operators based on their semantics, helping users understand the role of each node in the data cleaning pipeline.
[0167] It should be noted that the pipeline generation method of this application is illustrated by data cleaning. In other possible implementations of this application, the pipeline generation method can also be applied to the fields of data development and data preparation. For example, the operators in the field of data development can be extracted upward to obtain attribute groups, thereby realizing coarse-grained pipeline generation.
[0168] Based on the aforementioned pipeline generation method, this application provides a pipeline generation system.
[0169] See Figure 3 The diagram shows a structural schematic of a production line system. The production line system 10 includes:
[0170] The attribute grouping subsystem 100 is used to obtain multiple data cleaning operators for generating the pipeline, and to obtain the attribute association relationship of the data. The pipeline is used to orchestrate multiple data cleaning operators so that the multiple data cleaning operators are executed according to the orchestration result. The data cleaning operators are used to remove or repair abnormal data in the dataset.
[0171] The attribute grouping subsystem 100 is also used to extract the attribute set of multiple data cleaning operators. The attribute set includes the attributes of the input data of multiple data cleaning operators and the attributes of the output data of multiple operators.
[0172] The attribute grouping subsystem 100 is also used to group attributes in the attribute set according to attribute associations, to obtain multiple attribute groups that support parallel computing and the sorting of multiple attribute groups;
[0173] The operator arrangement subsystem 200 is used to determine the intra-group order of the operator group corresponding to at least one attribute group based on the topological sorting result of at least one attribute group among multiple attribute groups.
[0174] The operator orchestration subsystem 200 is also used to orchestrate multiple data cleaning operators based on the sorting of multiple attribute groups and the intra-group sorting of operator groups corresponding to at least one attribute group, to obtain a pipeline.
[0175] For example, the attribute grouping subsystem 100 and the operator arrangement subsystem 200 described above can be implemented in hardware or in software.
[0176] When implemented in software, the attribute grouping subsystem 100 and the operator orchestration subsystem 200 can be applications running on computing devices. Taking the attribute grouping subsystem 100 as an example, the application can be a computing engine. The application can also be provided to users in the form of virtualization services. Virtualization services can include virtual machine (VM) services, bare metal server (BMS) services, or container services. Specifically, a VM service can be a service that uses virtualization technology to create a pool of virtual machine (VM) resources on multiple physical hosts to provide VMs for users to use on demand. A BMS service is a service that uses virtualization technology to create a pool of BMS resources on multiple physical hosts to provide BMS for users to use on demand. A container service is a service that uses virtualization technology to create a pool of container resources on multiple physical hosts to provide containers for users to use on demand. A VM is a simulated virtual computer, that is, a logical computer. A BMS is a scalable, high-performance computing service with computing performance indistinguishable from traditional physical machines and features secure physical isolation. Containers are a kernel virtualization technology that provides lightweight virtualization to isolate user space, processes, and resources. It should be understood that the VM service, BMS service, and container service mentioned above are merely specific examples. In practical applications, virtualization services can also include other lightweight or heavyweight virtualization services, which are not specifically limited here.
[0177] When implemented in hardware, the attribute grouping subsystem 100 and the operator orchestration subsystem 200 may include at least one computing device, such as a server. Alternatively, the attribute grouping subsystem 100 and the operator orchestration subsystem 200 may also be devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The aforementioned PLD may be implemented using complex programmable logical devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof.
[0178] In some possible implementations, the operator orchestration subsystem 200 is also used for:
[0179] The contribution of multiple data cleaning operators is obtained, and the contribution is used to indicate the data repair capability of the multiple data cleaning operators.
[0180] Operator orchestration subsystem 200 is specifically used for:
[0181] Based on the topological sorting result of at least one attribute group among multiple attribute groups and the contribution of multiple data cleaning operators, determine the intra-group sorting of the operator group corresponding to at least one attribute group.
[0182] In some possible implementations, the operator orchestration subsystem 200 is also used for:
[0183] Perform dependency detection on the connected subgraph corresponding to at least one of multiple attribute groups;
[0184] When a circular dependency exists, disconnect the incoming edges of the target attribute column in the connected subgraph to obtain an acyclic attribute graph. The target attribute column is the attribute column in the connected subgraph that meets the quality requirements.
[0185] Perform topological sorting based on the acyclic attribute graph to obtain the topological sorting result of at least one attribute group.
[0186] In some possible implementations, the operator orchestration subsystem is also used for:
[0187] Based on the quality of the attribute columns recorded in the knowledge base, the target attribute columns are determined from the connected subgraph, and the target attribute columns include the attribute columns with the highest quality; or,
[0188] Based on the primary key attribute set, obtain the target attribute column from the connected subgraph. The target attribute column is the attribute column included in the primary key attribute set.
[0189] In some possible implementations, the attribute grouping subsystem 100 is also used for:
[0190] Identify the semantics of multiple data cleaning operators;
[0191] The production line system 10 also includes:
[0192] An interpretation generation subsystem 300 is used to generate interpretations of the data cleaning operators in the pipeline based on the semantics of the plurality of data cleaning operators and the pipeline, using a pipeline interpretation template.
[0193] Similar to the attribute grouping subsystem 100 and the operator arrangement subsystem 200, the interpretation generation subsystem 300 can be implemented in hardware or in software.
[0194] When implemented in software, the interpretation and generation subsystem 300 can be an application running on a computing device. This application can also be provided to users as a virtualization service, such as a VM service, BMS service, or container service. When implemented in hardware, the attribute grouping subsystem 100 and the operator orchestration subsystem 200 can include at least one computing device, such as a server. Alternatively, the attribute grouping subsystem 100 and the operator orchestration subsystem 200 can also be devices implemented using ASICs or PLDs.
[0195] In some possible implementations, the attribute grouping subsystem 100 is specifically used for:
[0196] Based on the attribute associations and attribute sets, construct a dependency directed graph corresponding to the attribute set. The vertices in the dependency directed graph represent the attributes in the attribute set, and the edges in the dependency directed graph represent the dependencies between the attributes in the attribute set.
[0197] Connectivity analysis is performed on the dependent directed graph to obtain multiple connected subgraphs;
[0198] Grouping attributes in the attribute set based on multiple connected subgraphs yields multiple attribute groups that support parallel computation.
[0199] In some possible implementations, the attribute grouping subsystem 100 is specifically used for:
[0200] Attributes belonging to the same connected subgraph in multiple connected subgraphs are grouped into the same attribute group to obtain multiple attribute groups;
[0201] Adjust at least one attribute group among multiple attribute groups based on the parallelism and the number of operators corresponding to multiple connected subgraphs.
[0202] It should be noted that, Figure 3 This is merely an illustrative division of the pipeline generation system 10; the pipeline generation system 10 can also be divided into other modules or subsystems according to function. For example, the aforementioned attribute grouping subsystem 100 may include an interaction module, a parsing module, and a grouping module. The interaction module is used to obtain multiple data cleaning operators and attribute relationships; the parsing module is used to extract attributes from multiple data cleaning operators to obtain an attribute set; and the grouping module is used to group the attribute set according to attribute relationships to obtain multiple attribute groups supporting parallel computing and a sorting of the attribute groups.
[0203] This application also provides a computing device 1500. For example... Figure 15 As shown, the computing device 1500 includes a bus 1502, a processor 1504, a memory 1506, and a communication interface 1508. The processor 1504, the memory 1506, and the communication interface 1508 communicate with each other via the bus 1502. The computing device 1500 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1500.
[0204] Bus 1502 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 15 The bus 1502 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1502 may include a path for transmitting information between various components of the computing device 1500 (e.g., memory 1506, processor 1504, communication interface 1508).
[0205] Processor 1504 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0206] Memory 1506 may include volatile memory, such as random access memory (RAM). Memory 1506 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0207] The memory 1506 stores executable program code, which the processor 1504 executes to implement the aforementioned pipelined generation method. Specifically, the memory 1506 stores instructions for the pipelined generation system 10 to execute the pipelined generation method. For example, the memory 1506 may store instructions for implementing the functions of the attribute grouping subsystem 100 and the operator arrangement subsystem 200 in the pipelined generation system 10. Furthermore, the memory 1506 may also store instructions for implementing the functions of the interpretation generation subsystem 100 in the pipelined generation system 10.
[0208] The communication interface 1508 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 1500 and other devices or communication networks.
[0209] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0210] like Figure 16 As shown, the computing device cluster includes multiple computing devices 1500. The memory 1506 of the computing devices 1500 in the computing device cluster may store the same pipeline generation system 10 for executing pipeline generation methods.
[0211] In some possible implementations, one or more computing devices 1500 in the computing device cluster can also be used to execute some of the instructions of the pipeline generation system 10 for executing the pipeline generation method. In other words, a combination of one or more computing devices 1500 can jointly execute the instructions of the pipeline generation system 10 for executing the pipeline generation method.
[0212] In the case where the computing device cluster includes multiple computing devices 1500, the memory 1506 in different computing devices 1500 can store different instructions for executing some functions of the pipeline generation system 10.
[0213] Figure 17 One possible implementation is shown. For example... Figure 17 As shown, two computing devices 1500A and 1500B are connected via a communication interface 1508. The memory in computing device 1500A stores instructions for executing the functions of the attribute grouping subsystem 100. The memory in computing device 1500B stores instructions for executing the functions of the operator arrangement subsystem 200. In other words, the memories 1506 of computing devices 1500A and 1500B jointly store instructions for the pipeline generation system 10 to execute the pipeline generation method. Furthermore, computing device 1500B may also store instructions for interpreting the functions of the generation subsystem 300.
[0214] Figure 17 The connection method between the computing device clusters shown can be considered because the pipeline generation method provided in this application requires a large amount of computing power for operator sorting. Therefore, it is considered to decentralize the functions implemented by the operator orchestration subsystem 200 to different computing devices.
[0215] It should be understood that Figure 17 The functions of computing device 1500A shown can also be performed by multiple computing devices 1500. Similarly, the functions of computing device 1500B can also be performed by multiple computing devices 1500.
[0216] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 18 One possible implementation is shown. For example... Figure 18 As shown, the two computing devices 1500C and 1500D are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 1506 in computing device 1500C stores instructions for executing the functions of the attribute grouping subsystem 100. Simultaneously, the memory 1506 in computing device 1500D stores instructions for executing the functions of the operator arrangement subsystem 200. Furthermore, computing device 1500B may also store instructions for interpreting the functions of the generation subsystem 300.
[0217] It should be understood that Figure 18 The functions of the computing device 1500C shown can also be performed by multiple computing devices 1500. Similarly, the functions of the computing device 1500D can also be performed by multiple computing devices 1500.
[0218] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to execute the above-described pipeline generation method applied to the pipeline generation system 10. This application also provides another computer-readable storage medium. This computer-readable storage medium includes instructions that instruct a computing device to execute the above-described pipeline generation method applied to the pipeline generation system 10.
[0219] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the above-described pipeline generation method. This application also provides a computer program product containing instructions. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the above-described pipeline generation method.
[0220] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of pipeline generation, the method comprising: The method comprises: obtaining a plurality of data cleaning operators for generating a pipeline, and obtaining an attribute association relationship of data, the pipeline being used for scheduling the plurality of data cleaning operators so that the plurality of data cleaning operators perform according to a scheduling result, the data cleaning operator being used for removing or repairing abnormal data in a data set; extracting an attribute set of the plurality of data cleaning operators, the attribute set comprising attributes of input data of the plurality of data cleaning operators and attributes of output data of the plurality of operators; grouping attributes in the attribute set according to the attribute association relationship, obtaining a plurality of attribute groups supporting parallel computing and an order of the plurality of attribute groups; determining an intra-group order of an operator group corresponding to at least one attribute group in the plurality of attribute groups according to a topological sorting result of the at least one attribute group; scheduling the plurality of data cleaning operators according to the order of the plurality of attribute groups and the intra-group order of the operator group corresponding to the at least one attribute group, and obtaining the pipeline.
2. The method of claim 1, wherein, The method further comprises: obtaining a contribution degree of the plurality of data cleaning operators, the contribution degree being used for indicating a data repairing capability of the plurality of data cleaning operators; the determining the intra-group order of the operator group corresponding to the at least one attribute group according to the topological sorting result of the at least one attribute group, comprises: determining the intra-group order of the operator group corresponding to the at least one attribute group according to the topological sorting result of the at least one attribute group and the contribution degree of the plurality of data cleaning operators.
3. The method according to claim 1 or 2, characterized in that, The method further comprises: performing dependency detection on a connected subgraph corresponding to at least one attribute group in the plurality of attribute groups; when there is a circular dependency, disconnecting an incoming edge of a target attribute column in the connected subgraph, obtaining an acyclic attribute graph, the target attribute column being an attribute column in the connected subgraph whose quality meets a requirement; performing topological sorting according to the acyclic attribute graph, and obtaining a topological sorting result of the at least one attribute group.
4. The method of claim 3, wherein, The method further comprises: determining the target attribute column from the connected subgraph according to qualities of attribute columns recorded in a knowledge base, the target attribute column comprising an attribute column with the highest quality; or obtaining the target attribute column from the connected subgraph according to a primary key attribute set, the target attribute column being an attribute column included in primary data in the primary key attribute set.
5. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: identifying semantics of the plurality of data cleaning operators; generating an explanation of a data cleaning operator in the pipeline through a pipeline explanation template according to the semantics of the plurality of data cleaning operators and the pipeline.
6. The method according to any one of claims 1 to 5, characterized in that, The grouping attributes in the attribute set according to the attribute association relationship, and obtaining the plurality of attribute groups supporting parallel computing, comprises: constructing a dependency directed graph corresponding to the attribute set according to the attribute association relationship and the attribute set, a vertex in the dependency directed graph representing an attribute in the attribute set, and an edge in the dependency directed graph representing a dependency relationship between attributes in the attribute set; performing connectivity analysis on the dependency directed graph, and obtaining a plurality of connected subgraphs; Grouping attributes in the attribute set according to the multiple connected sub-graphs, to obtain multiple attribute groups supporting parallel computing.
7. The method of claim 6, wherein, The grouping attributes in the attribute set according to the multiple connected sub-graphs, to obtain multiple attribute groups supporting parallel computing, comprises: Dividing attributes belonging to the same connected sub-graph in the multiple connected sub-graphs into the same attribute group, to obtain multiple attribute groups; According to the parallel degree and the number of operators corresponding to the multiple connected sub-graphs, adjusting at least one attribute group in the multiple attribute groups.
8. A pipeline generation system, characterized by, The system comprises: An attribute grouping subsystem is configured to obtain multiple data cleaning operators used for generating a pipeline, and obtain an attribute association relationship of data, the pipeline is configured to arrange the multiple data cleaning operators so that the multiple data cleaning operators perform according to the arrangement result, and the data cleaning operators are configured to remove or repair abnormal data in a data set; The attribute grouping subsystem is further configured to extract an attribute set of the multiple data cleaning operators, the attribute set comprising attributes of input data of the multiple data cleaning operators and attributes of output data of the multiple operators; The attribute grouping subsystem is further configured to group attributes in the attribute set according to the attribute association relationship, to obtain multiple attribute groups supporting parallel computing and an order of the multiple attribute groups; An operator arrangement subsystem is configured to determine an intra-group order of an operator group corresponding to at least one attribute group in the multiple attribute groups according to a topological order result of the at least one attribute group; The operator arrangement subsystem is further configured to arrange the multiple data cleaning operators according to the order of the multiple attribute groups and the intra-group order of the operator group corresponding to the at least one attribute group, to obtain the pipeline.
9. The system of claim 8, wherein, The operator arrangement subsystem is further configured to: Obtain contribution degrees of the multiple data cleaning operators, the contribution degrees being used to indicate data repair capabilities of the multiple data cleaning operators; The operator arrangement subsystem is specifically configured to: Determine the intra-group order of the operator group corresponding to the at least one attribute group according to the topological order result of the at least one attribute group in the multiple attribute groups and the contribution degrees of the multiple data cleaning operators.
10. The system of claim 8 or 9, characterized in that, The operator arrangement subsystem is further configured to: Perform dependency detection on a connected sub-graph corresponding to at least one attribute group in the multiple attribute groups; When there is a cyclic dependency, disconnect an incoming edge of a target attribute column in the connected sub-graph, to obtain an acyclic attribute graph, the target attribute column being an attribute column in the connected sub-graph whose quality meets a requirement; Perform topological ordering according to the acyclic attribute graph, to obtain a topological order result of the at least one attribute group.
11. The system of claim 10, wherein, The operator arrangement subsystem is further configured to: Determine the target attribute column from the connected sub-graph according to qualities of attribute columns recorded in a knowledge base, the target attribute column comprising an attribute column with the highest quality; or Obtain the target attribute column from the connected sub-graph according to a primary key attribute set, the target attribute column being an attribute column included in primary data in the primary key attribute set.
12. The system of any one of claims 8 to 10, wherein, The attribute grouping subsystem is further configured to: Identify semantics of the multiple data cleaning operators; The system further comprises: An explanation generation subsystem configured to generate an explanation of a data cleansing operator in the pipeline based on semantics of the plurality of data cleansing operators and the pipeline by a pipeline explanation template.
13. The system of any one of claims 8 to 12, wherein, The attribute grouping subsystem is specifically configured to: construct a dependency directed graph corresponding to the attribute set according to the attribute association relationship and the attribute set, wherein a vertex in the dependency directed graph represents an attribute in the attribute set, and an edge in the dependency directed graph represents a dependency relationship between attributes in the attribute set; perform connectivity analysis on the dependency directed graph to obtain a plurality of connected subgraphs; group attributes in the attribute set according to the plurality of connected subgraphs to obtain a plurality of attribute groups that support parallel computing.
14. The system of claim 13, wherein, The attribute grouping subsystem is specifically configured to: divide attributes belonging to the same connected subgraph in the plurality of connected subgraphs into the same attribute group to obtain a plurality of attribute groups; adjust at least one attribute group in the plurality of attribute groups according to a parallel degree and a number of operators corresponding to the plurality of connected subgraphs.
15. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, the at least one computing device including at least one processor and at least one memory, the at least one memory storing computer readable instructions; the at least one processor executes the computer readable instructions, so that the computing device cluster executes the method in any one of claims 1 to 7.
16. A computer-readable storage medium, characterized in that, including computer readable instructions; the computer readable instructions are used to implement the method of any one of claims 1 to 7.
17. A computer program product, characterised in that, including computer readable instructions; the computer readable instructions are used to implement the method of any one of claims 1 to 7.