Data quality rule generation method and device, equipment and storage medium
Through the pre-constructed data quality rule mapping model and quality rule generation model, data quality rules are automatically generated, which solves the problems of inefficiency and inconsistency in the existing technology caused by relying on manual judgment, and achieves more efficient and consistent data quality rules formulation.
Patent Information
- Application Number
- CN202510281317.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art relies heavily on manual judgment when formulating data quality rules, resulting in inefficiency, inconsistent rules and lag, affecting the overall level of data quality.
Through the pre-constructed data quality rules mapping model and quality rules generation model, the quality attributes of each field in the to be processed data are automatically identified and corresponding data quality rules are generated to reduce the dependence on manual judgment.
It improves the efficiency and consistency of the formulation of data quality rules, reduces the impact of subjective bias and empirical limitations, and ensures the accuracy and unity of the rules.
Smart Images

Figure CN120216489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method, device, equipment and storage medium for generating data quality rules. Background Art
[0002] In today's data-driven era, the importance of data quality cannot be ignored. High-quality data provides a solid foundation for data processing, which can help technicians make more informed decisions, optimize business processes, improve customer experience, and drive business growth.
[0003] However, the formulation of data quality rules often relies on the knowledge and experience of data managers, and this process requires a lot of time to review and analyze and summarize internal and external evidence. The disadvantages of this method are: first, the formulation process is highly dependent on manual judgment and is easily affected by experience limitations and subjective biases, resulting in the rule may not be able to fully adapt to rapidly changing business needs. Second, the screening and analysis of information requires a lot of time, which reduces the efficiency of formulation. In addition, because data managers face complex business environments and multi-dimensional data needs, over-reliance on manual formulation may lead to inconsistency or lag in quality rules, which in turn affects the overall level of data quality.
[0004] Therefore, how to reduce reliance on manual judgment to improve the efficiency and consistency of data quality rule-making is a technical problem that technical personnel in this field urgently need to solve. Summary of the invention
[0005] Based on the above problems, the present application provides a data quality rule generation method, apparatus, device and storage medium, which can reduce the reliance on manual judgment to improve the efficiency and consistency of data quality rule formulation.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] A method for generating data quality rules, the method comprising:
[0008] Obtain metadata of the data to be processed to obtain metadata to be processed, and obtain a table name of a table where the data to be processed is located to obtain a first table name;
[0009] Input the metadata to be processed into a pre-built data quality rule mapping model to obtain a data quality rule mapping table for the metadata to be processed; the data quality rule mapping table includes each field in the metadata to be processed, data quality attributes, and a construction requirement flag for each field for each data quality attribute; the data quality attributes include accuracy, consistency, completeness, uniqueness, and precision;
[0010] Traverse each field in the metadata to be processed. For each field, based on the construction requirement flag in the data quality rule mapping table, identify the attribute to be constructed for this field, and obtain the external input item of the target attribute in the attribute to be constructed; the attribute to be constructed is the data quality attribute marked with a construction requirement for this field.
[0011] Traverse each field in the metadata to be processed. For each attribute to be constructed for each field, if the attribute to be constructed is the target attribute, input the field, the attribute to be constructed, the first table name, and the external input item of the attribute to be constructed into the pre-constructed quality rule generation model to generate the data quality rule for this field; otherwise, input the field, the attribute to be constructed, and the first table name into the pre-constructed quality rule generation model to generate the data quality rule for this field.
[0012] Integrate the data quality rules of all fields in the data to be processed to obtain the data quality rule of the data to be processed.
[0013] In a possible implementation manner, the construction process of the pre-constructed data quality rule mapping model includes:
[0014] Obtain the metadata of the first historical data to get the first historical metadata.
[0015] Perform field division on the first historical metadata to obtain multiple fields of the first historical metadata.
[0016] Determine the construction requirements between each field in the first historical metadata and the data quality attributes to obtain the data quality attribute construction mapping relationship; the construction requirements are represented by the construction requirement flag.
[0017] Input each field in the first historical metadata and the data quality attribute construction mapping relationship into a convolutional neural network for model training to obtain the data quality rule mapping model.
[0018] Among them, the convolutional neural network includes an Embedding layer, a 1DCNN layer, a pooling layer, and a fully connected layer connected in sequence.
[0019] In a possible implementation manner, the construction process of the pre-constructed quality rule generation model includes:
[0020] Obtain the metadata of the second historical data to get the second historical metadata, and obtain the table name of the table where the second historical data is located to get the second table name.
[0021] Input the second historical metadata into the data quality rule mapping model to obtain the data quality rule mapping table of the second historical data;
[0022] Traverse each field in the second historical metadata. For each field, based on the construction requirement flag in the data quality rule mapping table, identify the to-be-constructed attribute of the field, and obtain the external input item of the target attribute in the to-be-constructed attribute; the to-be-constructed attribute is the data quality attribute marked with construction requirements for this field;
[0023] For each to-be-constructed attribute of each field, determine a unique Structured Query Language (SQL) template. If the to-be-constructed attribute is a target attribute, fill the field, the second table name, and the external input item of the to-be-constructed attribute into the SQL template to obtain an SQL statement; otherwise, fill the field and the second table name into the SQL template to obtain an SQL statement;
[0024] Input the SQL statements of all fields corresponding to the second historical data into a Bidirectional Gated Recurrent Unit (BiGRU) network for model training to obtain the quality rule generation model.
[0025] In a possible implementation manner, for each to-be-constructed attribute of each field, determining a unique SQL template includes:
[0026] For each to-be-constructed attribute of each field, determine a unique SQL template corresponding to this field under this to-be-constructed attribute based on the field, the to-be-constructed attribute, and the second table name.
[0027] In a possible implementation manner, the target attributes include consistency, precision, and accuracy; the external input items of consistency include the associated tables and fields; the external input items of precision include data precision; the external input items of accuracy include verification data.
[0028] A data quality rule generation device, the device includes:
[0029] A first acquisition unit, configured to acquire the metadata of the data to be processed to obtain the metadata to be processed, and acquire the table name of the table where the data to be processed is located to obtain the first table name;
[0030] A first input unit, configured to input the metadata to be processed into a pre-constructed data quality rule mapping model to obtain the data quality rule mapping table of the metadata to be processed; the data quality rule mapping table includes each field in the metadata to be processed, data quality attributes, and the construction requirement flag of each field for each data quality attribute; the data quality attributes include accuracy, consistency, integrity, uniqueness, and precision;
[0031] A first integration unit for traversing each field in the metadata to be processed, and for each field, identifying the to-be-built attribute of the field based on the construction requirement flag in the data quality rule mapping table, and obtaining the external input item of the target attribute in the to-be-built attribute; the to-be-built attribute is a data quality attribute marked with a construction requirement for the field.
[0032] A second integration unit for traversing each field in the metadata to be processed, and for each to-be-built attribute of each field, if the to-be-built attribute is a target attribute, inputting the field, the to-be-built attribute, the first table name, and the external input item of the to-be-built attribute into the pre-constructed quality rule generation model to generate a data quality rule for the field; otherwise, inputting the field, the to-be-built attribute, and the first table name into the pre-constructed quality rule generation model to generate a data quality rule for the field.
[0033] An integration unit for integrating the data quality rules of all fields in the data to be processed to obtain the data quality rule of the data to be processed.
[0034] In a possible implementation manner, the device further includes:
[0035] A second acquisition unit for acquiring the metadata of the first historical data to obtain the first historical metadata.
[0036] A field division unit for dividing the first historical metadata into multiple fields of the first historical metadata.
[0037] A mapping relationship construction unit for determining the construction requirements between each field in the first historical metadata and the data quality attribute to obtain a data quality attribute construction mapping relationship; the construction requirement is represented by the construction requirement flag.
[0038] A second input unit for inputting each field in the first historical metadata and the data quality attribute construction mapping relationship into a convolutional neural network for model training to obtain the data quality rule mapping model.
[0039] Wherein, the convolutional neural network includes an Embedding layer, a 1D CNN layer, a pooling layer, and a fully connected layer connected in sequence.
[0040] In a possible implementation manner, the device further includes:
[0041] A third acquisition unit, configured to acquire metadata of the second historical data to obtain second historical metadata, and acquire a table name of a table where the second historical data is located to obtain a second table name;
[0042] A third input unit, configured to input the second historical metadata into the data quality rule mapping model to obtain a data quality rule mapping table of the second historical data;
[0043] A third integration unit, configured to traverse each field in the second historical metadata. For each field, based on a construction requirement flag in the data quality rule mapping table, identify a to-be-constructed attribute of the field, and acquire an external input item of a target attribute in the to-be-constructed attribute; the to-be-constructed attribute is a data quality attribute marked with a construction requirement for the field;
[0044] A fourth integration unit, configured to determine a unique SQL template for each to-be-constructed attribute of each field. If the to-be-constructed attribute is a target attribute, fill the field, the second table name, and the external input item of the to-be-constructed attribute into the SQL template to obtain an SQL statement; otherwise, fill the field and the second table name into the SQL template to obtain an SQL statement;
[0045] A fourth input unit, configured to input SQL statements of all fields corresponding to the second historical data into a BiGRU network for model training to obtain the quality rule generation model.
[0046] A data quality rule generation device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data quality rule generation method described above is implemented.
[0047] A computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is enabled to execute the data quality rule generation method described above.
[0048] Compared with the prior art, the present application has the following beneficial effects:
[0049] The present application provides a method, apparatus, device, and storage medium for generating data quality rules. Specifically, when executing the data quality rule generation method provided in the embodiments of the present application, first, the metadata of the data to be processed can be obtained to obtain the metadata to be processed, and the table name of the table where the data to be processed is located can be obtained to obtain the first table name. Subsequently, the metadata to be processed is input into a pre-constructed data quality rule mapping model to generate a data quality rule mapping table for the metadata to be processed. This mapping table includes each field in the metadata to be processed, data quality attributes (such as accuracy, consistency, integrity, uniqueness, and precision), and a construction requirement flag for each field for each data quality attribute. Then, each field in the metadata to be processed is traversed. For each field, based on the construction requirement flag in the data quality rule mapping table, the attribute to be constructed for this field is identified, and the external input item of the target attribute in the attribute to be constructed is obtained. The target attribute refers to the data quality attribute marked with a construction requirement for this field in the mapping table. Then, each attribute to be constructed for each field is continued to be traversed. If the attribute to be constructed is the target attribute, then this field, this attribute to be constructed, the first table name, and the external input item of the attribute to be constructed are input into a pre-constructed quality rule generation model for rule generation to obtain the data quality rule for this field; otherwise, this field, this attribute to be constructed, and the first table name are input into the quality rule generation model, and the data quality rule for this field is also generated. Finally, the data quality rules for all fields in the data to be processed are integrated to form the overall data quality rule for the processed data. The present application reduces the dependence on manual judgment, reduces the influence of subjective deviation and empirical limitations through the pre-constructed data quality rule mapping model and quality rule generation model, and at the same time greatly improves the speed and efficiency of rule generation through automated processing, reducing the time required for information screening and analysis. In addition, through the standardized mapping model and generation model, it is ensured that the data quality rules for each field are based on the same logic and standard, thereby improving the consistency and accuracy of the rules. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] To more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a flowchart of a method for generating a data quality rule provided in an embodiment of the present application;
[0052] Figure 2 It is a schematic diagram of the table structure of a data quality rule mapping table provided in an embodiment of the present application;
[0053] Figure 3 This is a flowchart of a method for constructing a data quality rule mapping model provided by an embodiment of the present application;
[0054] Figure 4 This is a flowchart of a method for constructing a quality rule generation model provided by an embodiment of the present application;
[0055] Figure 5 This is a schematic structural diagram of a data quality rule generation device provided by an embodiment of the present application. Detailed implementation manners
[0056] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the background art related to the embodiments of the present application will be described first below.
[0057] Data quality refers to the degree to which data meets the usage purposes of data consumers and can meet the specific requirements of business scenarios in a business environment. It is a comprehensive concept that involves multiple aspects and dimensions of data, including accuracy, integrity, consistency, timeliness, and accessibility, etc. The importance of data quality cannot be ignored in today's data-driven world. High-quality data is the key for technical personnel to make informed decisions, optimize business processes, enhance customer experience, and achieve business growth when processing data.
[0058] The formulation of current quality rules relies on the knowledge and experience of data managers. A large amount of time is required to search, consult, analyze, and summarize internal and external formulation bases during the formulation process. For example, in a financial institution, to ensure the quality of customer data, data managers need to formulate specific quality control rules according to the latest regulatory requirements, industry standards, and internal operating procedures. This not only includes referring to regulatory requirements such as the General Data Protection Regulation, but also needs to combine the company's past data management practices and customer feedback. However, this highly data manager personal knowledge and experience-dependent method is easily affected by subjective biases and time limitations, resulting in inconsistent and lagging rules, thus affecting the overall level of data quality.
[0059] To solve this problem, an embodiment of the present application provides a method, apparatus, device, and storage medium for generating data quality rules. First, the metadata of the data to be processed and the name of the table where it is located are obtained to get the metadata to be processed and the first table name. Then, the metadata to be processed is input into a pre-constructed data quality rule mapping model to generate a data quality rule mapping table. This table includes the data quality attributes of each field and their construction requirement flags, where the data quality attributes include accuracy, consistency, integrity, uniqueness, and precision. Next, for each field, the data quality rule mapping table is traversed to identify the attributes that need to be constructed, and the external input items of the target attributes are obtained. For the attributes to be constructed for each field, if the attribute is a target attribute, the field, the attribute to be constructed, the first table name, and the external input items of the attribute are input into a pre-constructed quality rule generation model to generate data quality rules; otherwise, only the field, the attribute to be constructed, and the table name are input into the model to generate rules. Finally, the generated rules for all fields are integrated to obtain the data quality rules for the data to be processed. By using the pre-constructed data quality rule mapping model and quality rule generation model in this application, the dependence on manual judgment is reduced. Compared with the traditional method that relies on experience and manual analysis, this application can automatically identify the quality attributes of each field and generate corresponding data quality rules, thus improving the efficiency. At the same time, by automatically identifying and generating rules, the time waste in the manual screening and analysis process is avoided, and the data quality rules can be formulated more quickly. In addition, this application can perform unified processing based on a preset model during rule generation, thereby ensuring the consistency of the rules and avoiding the deviation and inconsistency of manual judgment.
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0061] See Figure 1 , which is a flowchart of a method for generating data quality rules provided by an embodiment of the present application. As Figure 1 shown, the method for generating data quality rules may include steps S101-S105:
[0062] S101: Obtain the metadata of the data to be processed to get the metadata to be processed, and obtain the table name of the table where the data to be processed is located to get the first table name.
[0063] Obtaining the metadata of the data to be processed to get the metadata to be processed, and obtaining the table name of the table where the data to be processed is located to get the first table name means first extracting the information (i.e., metadata) that describes the data structure and characteristics from the data to be processed, and then obtaining the name of the database table where the data is located.
[0064] Specifically, metadata includes information such as field names, field types, data lengths, and relationships between fields, while the table name identifies the specific table where these data are located. The purpose of doing this is to provide the necessary basic information for the subsequent generation of data quality rules.
[0065] S102: Input the metadata to be processed into a pre-constructed data quality rule mapping model to obtain the data quality rule mapping table of the metadata to be processed.
[0066] After obtaining the metadata to be processed, it is necessary to input the metadata of the data to be processed into a pre-constructed data quality rule mapping model. Metadata usually contains the structural information of the data, such as field names, field types, data lengths, and relationships between fields, while the data quality rule mapping model is a rule framework for defining and managing data quality. By inputting the metadata into this model, the system will automatically generate a "data quality rule mapping table".
[0067] In the data quality rule mapping table, it contains the detailed information of each field of the data to be processed. Each field corresponds to multiple data quality attributes. Data quality attributes are the specific measurement criteria for data quality, including accuracy, consistency, integrity, uniqueness, and precision, etc. These attributes help technicians evaluate and improve the data quality from different dimensions. For example:
[0068] Accuracy: Whether the data truly reflects the situation in the real world.
[0069] Consistency: Whether the data remains consistent between different systems or at different time points.
[0070] Integrity: Whether the data is missing and whether it contains all the necessary information.
[0071] Uniqueness: Whether there are duplicate items in the data.
[0072] Precision: Whether the level of detail of the data meets the actual requirements.
[0073] In addition, in the data quality rule mapping table, each field will also have a construction requirement flag, which indicates whether it is necessary to construct or verify for a specific data quality attribute. For example, for a specific field, if it needs to ensure that there is no duplicate data, then the construction requirement flag will indicate that it is necessary to check the uniqueness of the field.
[0074] As Figure 2 shown Figure 2 in the data quality rule mapping table, there are fields such as "numeric fields, date and time fields, personal identity information fields, unique identifier fields, descriptive fields, classification and label fields, medical and health fields, security and security-related fields, association fields, status or status code fields, geographical location fields, foreign key fields, uniqueness constraint fields, non-null fields, check constraint fields, business rule-related fields, financial and accounting fields, primary key fields, natural keys, business keyword fields, user name fields, email addresses and phone numbers, serial number or identification code fields". Figure 2 The 0 and 1 in it are construction requirement flags. When the construction requirement flag is 1, it means that there is a requirement to construct the data quality attribute for this field. When the construction requirement flag is 0, it means that there is no requirement to construct the data quality attribute for this field.
[0075] Through such a mapping model and data quality rule mapping table, the system can automatically identify which data quality checks and optimizations need to be performed on each field in the data to be processed. This not only helps improve the management efficiency of data quality, but also ensures the compliance and integrity of data on different quality attributes, thereby providing higher-quality data support for subsequent data processing and analysis.
[0076] See Figure 3 , Figure 3 which is the method flow chart of a method for constructing a data quality rule mapping model provided by an embodiment of this application. Specifically, the construction of the data quality rule mapping model can be achieved through steps A1 - A4:[[]]
[0077] A1: Obtain the metadata of the first historical data to obtain the first historical metadata.
[0078] The purpose of this step is to extract metadata from the existing historical data. In this step:[[]]
[0079] The first historical data refers to the existing historical records, which are usually used for analysis and model training.
[0080] The first historical metadata is the structured description information extracted from these historical data. It can include field names, field types, data lengths, relationships between fields, etc.
[0081] A2: Perform field division on the first historical metadata to obtain multiple fields of the first historical metadata.
[0082] In this step, for the first historical metadata, field division is performed. The metadata contains multiple fields, and each field may contain different types of information:[[]]
[0083] Field division means decomposing metadata into multiple individual fields (e.g., name, age, address, etc.) according to its structure, and each field represents an independent unit in the data.
[0084] Through this division, it is possible to further understand and process each component of the data, providing a basis for subsequent data quality analysis.
[0085] A3: Determine the construction requirements between each field in the first historical metadata and the data quality attributes to obtain the data quality attribute construction mapping relationship.
[0086] The core of this step is to determine the relationship between each field and the data quality attributes and establish a mapping of this relationship:
[0087] Data quality attributes refer to the standards for measuring data quality, such as accuracy, integrity, consistency, timeliness, etc.
[0088] Construction requirements refer to the specific analysis or processing requirements for fields based on data quality requirements. By analyzing the requirements for fields and data quality attributes, it is possible to determine how to perform quality assessment and correction for each field.
[0089] The construction requirement flag is used to represent the specific types and methods of these requirements (e.g., whether to check for null values, whether to require consistency verification, etc.). These flags are used for subsequent mapping relationship construction.
[0090] A4: Input the mapping relationship between each field in the first historical metadata and the data quality attributes into a convolutional neural network for model training to obtain the data quality rule mapping model.
[0091] In this stage, based on the mapping relationship constructed from the previous fields and data quality attributes, a convolutional neural network (Convolutional Neural Network, CNN) is used for model training:
[0092] A convolutional neural network (CNN) is a deep learning model commonly used for tasks such as processing images and sequence data, capable of automatically extracting features and performing classification or regression. Here, CNN is used to learn the relationship between fields and data quality attributes from the data.
[0093] The content input into the CNN includes the field information of the first historical metadata and the corresponding quality attribute mapping relationship.
[0094] Through the training of the CNN model, a data quality rule mapping model is generated, which can predict, analyze, and optimize the data quality rules.
[0095] The structure of the convolutional neural network:
[0096] Embedding layer: This layer converts the input field information into a vector representation for subsequent processing. The embedding layer can transform discrete field information (such as strings or categorical data) into dense numerical vectors.
[0097] One-dimensional Convolutional Neural Network (1D CNN) layer: This layer extracts features from the input embedding vectors through convolution operations. The convolution operation helps capture local relationships between fields, which is very important for analyzing complex associations between fields.
[0098] Pooling layer: The pooling layer is used to compress data, reduce computational complexity, and at the same time retain important feature information. Common pooling operations include max pooling and average pooling, which help the model generalize better.
[0099] Fully connected layer: The fully connected layer integrates the features extracted from the convolutional and pooling layers for the final classification or regression task. Here, it will be used to generate a data quality rule mapping model and output data quality rules.
[0100] Through the above steps, the finally obtained data quality rule mapping model can automatically perform quality analysis on new data and optimize it according to the set data quality standards.
[0101] S103: Traverse each field in the metadata to be processed. For each field, identify the to-be-constructed attribute of the field based on the construction requirement flag in the data quality rule mapping table, and obtain the external input item of the target attribute in the to-be-constructed attribute.
[0102] To identify the quality inspection and processing requirements for each field in the data to be processed. First, it is necessary to traverse each field in the metadata to be processed: The "metadata to be processed" mentioned here refers to the structural information in the dataset, usually including metadata such as field names, field types, data sources, etc. Traversing these fields means checking each column (i.e., each field) in the dataset one by one. Then, for each field, identify the to-be-built attributes of the field based on the build requirement flags in the data quality rule mapping table: The quality requirements and inspection rules for each field are stored in a "data quality rule mapping table", which defines which fields need to perform which quality inspections. The rules for each field will indicate whether quality construction is required for the field based on the "build requirement flag". By checking this build requirement flag, it can be determined whether the field has "to-be-built attributes". "To-be-built attributes" refer to the quality characteristics that the field needs to build, such as accuracy, integrity, consistency, etc. For example, if a field is marked as needing to ensure "uniqueness", then it is a to-be-built attribute. Next, obtain the external input items of the target attribute in the to-be-built attributes: Among the to-be-built attributes, the target attribute is the most concerned quality standard. For example, the target attribute may be "accuracy" or "consistency". "External input items" refer to the external data sources or additional information required when implementing these quality attributes. For example, to verify the "uniqueness" of a field, it is necessary to consult another database or an external dataset to ensure there are no duplicate values.
[0103] Among them, the to-be-built attributes are the data quality attributes marked with build requirements for this field. That is to say, the to-be-built attributes refer to the data quality standards that need to be subject to quality inspections or improvements determined by the build requirement flags. For example, if the build requirement for a field is to ensure "integrity", then "integrity" is its to-be-built attribute.
[0104] In a possible implementation manner, the target attributes include consistency, precision, and accuracy; the external input items of the consistency include the associated tables and fields; the external input items of the precision include data precision; the external input items of the accuracy include verification data.
[0105] These target attributes are the criteria used in data quality management to evaluate and optimize data. Each target attribute has different focuses and measurement methods:
[0106] (1) Consistency: Consistency means that data should be consistent among different data sources, systems, tables, or fields, avoiding conflicts or inconsistent situations.
[0107] For example, if a certain field has the same function in two tables, their values should be the same; or multiple records of the same field in the same table should maintain consistency.
[0108] External input items: Tables and fields associated with consistency. By analyzing the relationships between these tables and fields, it is possible to determine whether there are any inconsistencies.
[0109] (2) Precision: Precision refers to the specificity and detail of data values, especially the accuracy of numerical data. For example, a price field may have different precisions, specifying the number of decimal places.
[0110] For example, an amount field in a financial system should have sufficient precision to ensure the accuracy of numerical representation.
[0111] External input item: Data precision. Data precision refers to the number of digits for data storage and representation or the precision of the decimal point.
[0112] (3) Accuracy: Accuracy refers to the degree of conformity of data values with actual or standard values. High accuracy means that the data is very close to the real world or the standard expectation.
[0113] For example, the date recorded in a date field should be consistent with the actual date of the event, or the inventory quantity of a certain product should accurately reflect the actual inventory.
[0114] External input item: Verification data. Verification data is used to compare with actual data to confirm whether the data meets the expected standards or is accurate and error-free.
[0115] External input items are the data sources that help evaluate and measure the above target attributes. By analyzing these external input items, data quality inspection and optimization can be better carried out.
[0116] Exemplarily, assume that in a data quality management system of an e-commerce platform, there are the following fields: customer ID, order ID, order date, and product price. The user hopes to perform data quality management on these fields.
[0117] And there is a data quality rule mapping table (as shown in Table 1), which lists the construction requirement flags for each field:
[0118] Table 1
[0119] Field Name Accuracy Consistency Integrity Uniqueness Precision Customer ID 1 1 0 1 1 Order ID 1 0 0 1 1 Order Date 1 1 0 1 0 Commodity Price 1 0 1 1 0
[0120] The 0 and 1 in Table 1 are construction requirement flags. A construction requirement flag of 1 indicates that the quality rule for this data quality attribute needs to be constructed, and 0 indicates that the quality rule for this data quality attribute does not need to be constructed.
[0121] Processing flow:
[0122] 1. Traverse each field and identify the construction requirement flag:
[0123] For each field, we identify the attributes to be constructed based on the construction requirement flag.
[0124] Customer ID: Accuracy (construction requirement flag is 1): It is necessary to verify the data to ensure that the value of the customer ID is consistent with the actual customer data. Consistency (construction requirement flag is 1): It is necessary to check whether there is an associated customer ID in other tables (such as the order table) for this field to ensure data consistency. Precision (construction requirement flag is 0): Precision does not need to be considered. Integrity (construction requirement flag is 1): Ensure that this field exists (is not empty) in all records. Uniqueness (construction requirement flag is 1): Ensure that each customer ID is unique in the customer information table.
[0125] Order ID: Accuracy (construction requirement flag is 1): It is necessary to ensure the accuracy of the order ID, which should match the actual order data. Consistency (construction requirement flag is 0): The association between the order ID and other tables (such as the customer table) does not need to be considered. Precision (construction requirement flag is 0): Precision does not need to be concerned. Integrity (construction requirement flag is 1): The order ID cannot be missing. Uniqueness (construction requirement flag is 1): Each order ID should be unique.
[0126] Order date: Accuracy (construction requirement flag is 1): The order date needs to be consistent with the actual order occurrence date. Consistency (construction requirement flag is 1): The order date in the order table should be consistent with other related fields (such as the delivery date). Precision (construction requirement flag is 0): Precision control is not required. Integrity (construction requirement flag is 1): Ensure that each order record has a valid order date. Uniqueness (construction requirement flag is 0): The order date itself does not require uniqueness.
[0127] Commodity price: Accuracy (construction requirement flag is 1): The commodity price needs to be consistent with the actual commodity price. Consistency (construction requirement flag is 0): The consistency between the commodity price and other tables (such as the order table) does not need to be concerned. Precision (construction requirement flag is 1): The commodity price needs to ensure sufficient precision, usually accurate to two decimal places. Integrity (construction requirement flag is 1): Ensure that the commodity price field is not empty. Uniqueness (construction requirement flag is 0): The commodity price does not need to be unique.
[0128] 2. Obtain the external input items of the attributes to be constructed:
[0129] For each field, obtain the external input items of the target attribute according to the attributes marked as 1 in the construction requirements. For example:
[0130] (1) Customer ID:
[0131] Consistency: It is necessary to check the association with the customer ID field in other tables (such as the order table);
[0132] Accuracy: It is necessary to verify the correctness of the customer ID through verification data;
[0133] (2) Order ID:
[0134] Accuracy: Confirm whether the order ID is accurate through verification data.
[0135] (3) Order date:
[0136] Accuracy: Confirm the accuracy of the order date through verification data;
[0137] Consistency: It is necessary to check whether the relationship between the order date and other fields (such as the delivery date) is consistent.
[0138] (4) Commodity price:
[0139] Accuracy: It is necessary to confirm the accuracy of the commodity price through verification data.
[0140] Precision: Check the number of decimal places of the commodity price (such as accurate to two decimal places). Integrity: Ensure that the commodity price field is not empty.
[0141] Through the above process, the data quality attributes to be constructed (accuracy, consistency, precision, etc.) can be identified, and the corresponding external input items are obtained. Based on the values of the construction requirement flags, we can perform corresponding data quality construction operations for each field to ensure that the data quality meets the expected standards.
[0142] S104: Traverse each field in the to-be-processed metadata. For each to-be-constructed attribute of each field, if the to-be-constructed attribute is the target attribute, input the field, the to-be-constructed attribute, the first table name, and the external input items of the to-be-constructed attribute into the pre-constructed quality rule generation model to generate rules to obtain the data quality rules for this field; otherwise, input the field, the to-be-constructed attribute, and the first table name into the pre-constructed quality rule generation model to generate rules to obtain the data quality rules for this field.
[0143] After obtaining the external input items of the target attribute in the construction attributes and attributes to be constructed for each field, it is possible to traverse each field in the metadata to be processed. For each attribute to be constructed for each field, different operations are performed according to whether it is the target attribute, and corresponding data quality rules are finally generated.
[0144] Specifically, first, traverse each field in the metadata to be processed:
[0145] First, check each field one by one from the metadata to be processed. For example, assume there is a customer information table containing the following fields: name, age, address, phone number, etc.
[0146] Then, determine the attributes to be constructed for the field: For each field, identify the data quality attributes that need to be constructed based on the data quality rule mapping table. For example, the name field may need to evaluate accuracy, integrity, and consistency.
[0147] Next, determine the input data for the quality rule generation model based on the attributes to be constructed: Determine whether the attribute to be constructed is the target attribute (such as, consistency, precision, and accuracy). If it is determined to be the target attribute, specific input data is used; otherwise, general input data is used.
[0148] Specific input data means that if the currently processed attribute to be constructed is the target attribute (i.e., an attribute that needs special attention in data quality management, such as "uniqueness"), then this field, this attribute, the name of the table where it is located, and any external input items related to this attribute (such as other data sources, rules, or reference values) are input into the pre-constructed data quality rule generation model. In this way, the model will generate data quality rules for this field to ensure that the quality of the field meets the standards.
[0149] General input data means that if the currently processed attribute to be constructed is not the target attribute, then only the field, the attribute, and the table name are input into the quality rule generation model, and the model will generate corresponding rules based on this information.
[0150] See Figure 4 , Figure 4 which is the flowchart of a method for constructing a quality rule generation model provided by an embodiment of the present application. Specifically, the construction process of the quality rule generation model can be implemented through steps B1 - B5:
[0151] B1: Obtain the metadata of the second historical data to obtain the second historical metadata, and obtain the table name of the table where the second historical data is located to obtain the second table name.
[0152] To construct a quality rule generation model, it is first necessary to extract the metadata of the second historical data from a database or data storage and obtain the names of the tables where these data are located, usually by querying the structural metadata of the database tables.
[0153] B2: Input the second historical metadata into the data quality rule mapping model to obtain the data quality rule mapping table of the second historical data.
[0154] Input the extracted second historical metadata into the data quality rule mapping model for matching.
[0155] The mapping model generates the data quality rules corresponding to each field according to the structure of the data table and outputs a quality rule mapping table. Each field will be associated with one or more quality attributes (such as accuracy, consistency, etc.).
[0156] B3: Traverse each field in the second historical metadata. For each field, identify the attributes to be constructed for this field based on the construction requirement flags in the data quality rule mapping table, and obtain the external input items of the target attributes in the attributes to be constructed.
[0157] Traverse each field and find the construction requirement flags of this field in the quality rule mapping table. For each field, if the flag indicates that the field needs quality construction, identify the corresponding attributes to be constructed. And obtain the external input items required for these attributes to be constructed, which are helpful for constructing the quality attributes of this field.
[0158] B4: For each attribute to be constructed of each field, determine a unique SQL template. If the attribute to be constructed is a target attribute, fill the field, the second table name, and the external input items of the attribute to be constructed into the SQL template to obtain an SQL statement; otherwise, fill the field and the second table name into the SQL template to obtain an SQL statement.
[0159] For each attribute to be constructed of each field, select an appropriate Structured Query Language (SQL) template. The SQL template is used to generate query statements to verify or fix data quality problems.
[0160] If the attribute to be constructed is a target attribute (for example, related to data accuracy or consistency), it is necessary to fill in the field name, table name, and related external input items (such as other field data, external verification conditions) to generate a complete SQL query statement.
[0161] Otherwise, if the attribute to be constructed does not require external input items, only fill in the field name and table name to generate the corresponding SQL query statement.
[0162] Exemplarily, the SQL template is as follows:
[0163] SELECT * FROM table name WHERE STR_TO_DATE(metadata field, "external input item") IS NULL.
[0164] B5: Input the SQL statements for all fields corresponding to the second historical data into the BiGRU network for model training to obtain the quality rule generation model.
[0165] Take all the generated SQL query statements (operations for constructing data quality rules) as input and input them into a Bidirectional Gated Recurrent Unit (BiGRU) network. Through training, the BiGRU network learns the structure and generation rules of these SQL statements and finally generates a quality rule generation model.
[0166] This model can be used to subsequently generate data quality rules automatically, helping to automatically check and repair data quality problems.
[0167] BiGRU is a deep learning network that can process time series data and is particularly good at learning patterns and regularities from data sequences. In this application, it is used to learn the patterns for generating data quality rules.
[0168] In a possible implementation manner, determining a unique SQL template for each to-be-constructed attribute of each field includes:
[0169] For each to-be-constructed attribute of each field, determine the unique SQL template corresponding to this field under this to-be-constructed attribute based on this field, this to-be-constructed attribute, and the second table name.
[0170] For each to-be-constructed attribute of each field, determine a unique SQL template by combining this field, the to-be-constructed attribute, and the name of the relevant table, for handling the operations of this field under a specific to-be-constructed attribute.
[0171] S105: Integrate the data quality rules for all fields in the to-be-processed data to obtain the data quality rules for the to-be-processed data.
[0172] When processing the to-be-processed data, each field has corresponding quality inspection rules, such as validity, accuracy, integrity, etc. After summarizing and organizing these rules for all fields, the data quality rule set for the entire data set can be obtained. This set can be used as a standard to comprehensively evaluate the quality of the to-be-processed data.
[0173] Based on the content of S101 - S105, first, obtain the metadata of the data to be processed and the name of the table where it is located. Then, input the metadata to be processed into the pre - constructed data quality rule mapping model to generate a data quality rule mapping table, which contains the data quality attributes of each field and their construction requirements. Next, traverse the fields in the metadata, identify the data quality attributes to be constructed for each field according to the requirement flags in the mapping table, and obtain the external input items of the target attributes among the attributes to be constructed. For each attribute to be constructed of a field, if it is a target attribute, input the field, the attribute, the name of the first table, and the external input items into the quality rule generation model to generate a data quality rule; otherwise, only input the field, the attribute, and the name of the first table for generation. Finally, integrate the data quality rules generated for all fields to obtain the overall data quality rule of the data to be processed. This process ensures the comprehensiveness and uniformity of the data quality rules, providing a solid foundation for further data processing and decision - making. Through the pre - constructed data quality rule mapping model and quality rule generation model in this application, the dependence on manual judgment is successfully reduced. Compared with the traditional method that relies on experience and manual analysis, this application can automatically identify the quality attributes of fields and generate corresponding data quality rules, effectively improving work efficiency. At the same time, automated identification and rule generation avoid the time waste in the manual screening and analysis process, accelerating the process of formulating data quality rules. In addition, in the rule generation stage, unified processing based on the preset model ensures the consistency of the rules, eliminating the possibility of human subjective bias and inconsistency.
[0174] See Figure 5 , Figure 5 is a schematic structural diagram of a data quality rule generation device provided by an embodiment of this application. As Figure 5 shown, the data quality rule generation device includes:
[0175] The first acquisition unit 501 is configured to acquire the metadata of the data to be processed to obtain the metadata to be processed, and acquire the table name of the table where the data to be processed is located to obtain the first table name;
[0176] The first input unit 502 is configured to input the metadata to be processed into the pre - constructed data quality rule mapping model to obtain the data quality rule mapping table of the metadata to be processed; the data quality rule mapping table includes each field in the metadata to be processed, data quality attributes, and construction requirement flags for each field for each data quality attribute; the data quality attributes include accuracy, consistency, integrity, uniqueness, and precision;
[0177] The first integration unit 503 is configured to traverse each field in the metadata to be processed. For each field, based on the construction requirement flag in the data quality rule mapping table, identify the attribute to be constructed for this field, and obtain the external input item of the target attribute in the attribute to be constructed; the attribute to be constructed is the data quality attribute marked with a construction requirement for this field.
[0178] The second integration unit 504 is configured to traverse each field in the metadata to be processed. For each attribute to be constructed of each field, if the attribute to be constructed is a target attribute, input this field, this attribute to be constructed, the first table name, and the external input item of the attribute to be constructed into the pre-constructed quality rule generation model for rule generation to obtain the data quality rule of this field; otherwise, input this field, this attribute to be constructed, and the first table name into the pre-constructed quality rule generation model for rule generation to obtain the data quality rule of this field.
[0179] The integration unit 505 is configured to integrate the data quality rules of all fields in the data to be processed to obtain the data quality rule of the data to be processed.
[0180] In a possible implementation, the device further includes:
[0181] The second acquisition unit is configured to acquire the metadata of the first historical data to obtain the first historical metadata.
[0182] The field division unit is configured to divide the first historical metadata into multiple fields of the first historical metadata.
[0183] The mapping relationship construction unit is configured to determine the construction requirements between each field in the first historical metadata and the data quality attribute to obtain the data quality attribute construction mapping relationship; the construction requirement is represented by the construction requirement flag.
[0184] The second input unit is configured to input each field in the first historical metadata and the data quality attribute construction mapping relationship into a convolutional neural network for model training to obtain the data quality rule mapping model.
[0185] Wherein, the convolutional neural network includes an Embedding layer, a 1D CNN layer, a pooling layer, and a fully connected layer connected in sequence.
[0186] In a possible implementation, the device further includes:
[0187] The third acquisition unit is configured to acquire the metadata of the second historical data to obtain the second historical metadata, and acquire the table name of the table where the second historical data is located to obtain the second table name.
[0188] A third input unit, configured to input the second historical metadata into the data quality rule mapping model to obtain a data quality rule mapping table of the second historical data;
[0189] A third synthesis unit, configured to traverse each field in the second historical metadata. For each field, based on the construction requirement flag in the data quality rule mapping table, identify the to-be-constructed attribute of the field, and obtain the external input item of the target attribute in the to-be-constructed attribute; the to-be-constructed attribute is a data quality attribute marked with a construction requirement for the field;
[0190] A fourth synthesis unit, configured to, for each to-be-constructed attribute of each field, determine a unique SQL template. If the to-be-constructed attribute is a target attribute, fill the field, the second table name, and the external input item of the to-be-constructed attribute into the SQL template to obtain an SQL statement; otherwise, fill the field and the second table name into the SQL template to obtain an SQL statement;
[0191] A fourth input unit, configured to input the SQL statements of all fields corresponding to the second historical data into a BiGRU network for model training to obtain the quality rule generation model.
[0192] In a possible implementation manner, the fourth synthesis unit specifically includes:
[0193] A template determination unit, configured to, for each to-be-constructed attribute of each field, determine a unique SQL template corresponding to the field under the to-be-constructed attribute based on the field, the to-be-constructed attribute, and the second table name.
[0194] In a possible implementation manner, the target attributes include consistency, precision, and accuracy; the external input item of consistency includes the associated table and field; the external input item of precision includes data precision; the external input item of accuracy includes verification data.
[0195] In addition, an embodiment of the present application further provides a data quality rule generation device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data quality rule generation method described above is implemented.
[0196] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the instructions are run on a terminal device, the terminal device is enabled to execute the data quality rule generation method described above.
[0197] An embodiment of the present application provides a data quality rule generation device. First, a first acquisition unit 301 is used to acquire mapping data from a circular buffer as data to be detected. The mapping data is combined information of system calls that has undergone rule filtering, rule supplementation, and form mapping. Then, a rule filtering unit 302 is used to perform rule filtering on the data to be detected to obtain filtered data. An input unit 303 inputs the filtered data into a pre-constructed abstract syntax tree to obtain a detection result. Then, an extraction unit 304 extracts data associated with the detection result from the mapping data as supplementary data, so that an output data supplementation unit 305 can use the supplementary data to supplement the detection result to obtain an attack behavior output result. Through the pre-constructed data quality rule mapping model and quality rule generation model in the present application, the dependence on manual judgment is significantly reduced, and the influence of subjective deviation and empirical limitations is reduced. Compared with traditional methods, the present application greatly improves the speed and efficiency of rule generation through automated processing, and reduces the time required for information screening and analysis. In addition, through the standardized mapping model and generation model, it is ensured that the data quality rules for each field are based on the same logic and standards, thereby improving the consistency and accuracy of the rules.
[0198] The above has introduced in detail a data quality rule generation method, device, equipment, and storage medium provided by the present application. The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0199] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c can be single or multiple.
[0200] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. A method for generating data quality rules, characterized in that: The method comprises: Obtain metadata of the data to be processed to obtain metadata to be processed, and obtain a table name of a table where the data to be processed is located to obtain a first table name; Input the metadata to be processed into a pre-built data quality rule mapping model to obtain a data quality rule mapping table for the metadata to be processed; the data quality rule mapping table includes each field in the metadata to be processed, data quality attributes, and a construction requirement flag for each field for each data quality attribute; the data quality attributes include accuracy, consistency, completeness, uniqueness, and precision; Traversing each field in the metadata to be processed, for each field, identifying the attribute to be constructed of the field based on the construction requirement flag in the data quality rule mapping table, and obtaining the external input item of the target attribute in the attribute to be constructed; the attribute to be constructed is the data quality attribute with construction requirement marked for the field; Traversing each field in the metadata to be processed, for each attribute to be constructed in each field, if the attribute to be constructed is a target attribute, inputting the field, the attribute to be constructed, the first table name, and the external input item of the attribute to be constructed into the pre-constructed quality rule generation model to generate rules to obtain the data quality rule of the field; otherwise, inputting the field, the attribute to be constructed, and the first table name into the pre-constructed quality rule generation model to generate rules to obtain the data quality rule of the field; The data quality rules of all fields in the data to be processed are integrated to obtain the data quality rules of the data to be processed.
2. The method according to claim 1, characterized in that The construction process of the pre-built data quality rule mapping model includes: Acquire metadata of first historical data to obtain first historical metadata; Dividing the first historical metadata into fields to obtain a plurality of fields of the first historical metadata; Determine the construction requirements between each field in the first historical metadata and the data quality attribute to obtain a data quality attribute construction mapping relationship; the construction requirements are represented by the construction requirement flag; The mapping relationship between each field in the first historical metadata and the data quality attribute is input into a convolutional neural network for model training to obtain the data quality rule mapping model; The convolutional neural network includes an embedding layer, a one-dimensional convolutional neural network 1D CNN layer, a pooling layer and a fully connected layer which are connected in sequence.
3. The method according to claim 1, characterized in that The construction process of the pre-built quality rule generation model includes: Obtain metadata of the second historical data to obtain second historical metadata, and obtain a table name of a table where the second historical data is located to obtain a second table name; Inputting the second historical metadata into the data quality rule mapping model to obtain a data quality rule mapping table for the second historical data; Traversing each field in the second historical metadata, for each field, identifying the attribute to be constructed of the field based on the construction requirement flag in the data quality rule mapping table, and obtaining an external input item of a target attribute in the attribute to be constructed; the attribute to be constructed is a data quality attribute marked with a construction requirement for the field; For each attribute to be constructed of each field, a unique structured query language SQL template is determined; if the attribute to be constructed is a target attribute, the field, the second table name, and the external input item of the attribute to be constructed are filled into the SQL template to obtain an SQL statement; otherwise, the field and the second table name are filled into the SQL template to obtain an SQL statement; The SQL statements of all fields corresponding to the second historical data are input into a bidirectional gated recurrent unit network BiGRU network for model training to obtain the quality rule generation model.
4. The method according to claim 3, characterized in that For each attribute to be constructed of each field, a unique SQL template is determined, including: For each attribute to be constructed of each field, a unique SQL template corresponding to the field under the attribute to be constructed is determined based on the field, the attribute to be constructed and the second table name.
5. The method according to any one of claims 1 to 4, characterized in that: The target attributes include consistency, precision and accuracy; the external input items of consistency include tables and fields associated therewith; the external input items of precision include data accuracy; and the external input items of accuracy include verification data.
6. A data quality rule generation device, characterized in that: The device comprises: A first acquisition unit is used to acquire metadata of the data to be processed to obtain metadata to be processed, and acquire a table name of a table where the data to be processed is located to obtain a first table name; A first input unit is used to input the metadata to be processed into a pre-constructed data quality rule mapping model to obtain a data quality rule mapping table for the metadata to be processed; the data quality rule mapping table includes each field in the metadata to be processed, data quality attributes, and a construction requirement flag for each field for each data quality attribute; the data quality attributes include accuracy, consistency, completeness, uniqueness, and precision; The first integration unit is used to traverse each field in the metadata to be processed, and for each field, identify the attribute to be constructed of the field based on the construction requirement flag in the data quality rule mapping table, and obtain an external input item of a target attribute in the attribute to be constructed; the attribute to be constructed is a data quality attribute with construction requirements marked for the field; A second synthesis unit is used to traverse each field in the metadata to be processed, and for each attribute to be constructed of each field, if the attribute to be constructed is a target attribute, the field, the attribute to be constructed, the first table name, and the external input item of the attribute to be constructed are input into the pre-constructed quality rule generation model to generate rules to obtain the data quality rule of the field; otherwise, the field, the attribute to be constructed, and the first table name are input into the pre-constructed quality rule generation model to generate rules to obtain the data quality rule of the field; The integration unit is used to integrate the data quality rules of all fields in the data to be processed to obtain the data quality rules of the data to be processed.
7. The device according to claim 6, characterized in that The device also includes: A second acquisition unit, configured to acquire metadata of the first historical data to obtain first historical metadata; A field division unit, configured to divide the first historical metadata into fields to obtain a plurality of fields of the first historical metadata; A mapping relationship construction unit, used for determining a construction requirement between each field in the first historical metadata and a data quality attribute to obtain a data quality attribute construction mapping relationship; the construction requirement is represented by the construction requirement mark; A second input unit is used to input the mapping relationship between each field in the first historical metadata and the data quality attribute into a convolutional neural network for model training to obtain the data quality rule mapping model; The convolutional neural network includes an Embedding layer, a 1D CNN layer, a pooling layer and a fully connected layer connected in sequence.
8. The device according to claim 6, characterized in that The device also includes: A third acquisition unit is used to acquire metadata of the second historical data to obtain second historical metadata, and acquire a table name of a table where the second historical data is located to obtain a second table name; A third input unit, configured to input the second historical metadata into the data quality rule mapping model to obtain a data quality rule mapping table for the second historical data; a third integration unit, configured to traverse each field in the second historical metadata, and for each field, identify the attribute to be constructed of the field based on the construction requirement flag in the data quality rule mapping table, and obtain an external input item of a target attribute in the attribute to be constructed; the attribute to be constructed is a data quality attribute with construction requirement marked for the field; a fourth integration unit, for determining a unique SQL template for each attribute to be constructed of each field, and if the attribute to be constructed is a target attribute, filling the field, the second table name, and the external input item of the attribute to be constructed into the SQL template to obtain an SQL statement; otherwise, filling the field and the second table name into the SQL template to obtain an SQL statement; The fourth input unit is used to input the SQL statements of all fields corresponding to the second historical data into the BiGRU network for model training to obtain the quality rule generation model.
9. A data quality rule generation device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating data quality rules according to any one of claims 1 to 5 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the data quality rule generating method according to any one of claims 1 to 5.
Citation Information
Cited By
Quality rule configuration method based on artificial intelligence and related device thereof
CN120849471A
E-commerce order data quality detection method and device based on NL2SQL instruction, equipment and medium
CN120950919A
AI-based data warehouse quality automatic monitoring method and system
CN122240602A
AI-based data warehouse quality automatic monitoring method and system
CN122240602B