An artificial intelligence-based quality rule configuration method and related apparatuses
By automatically configuring quality rules using artificial intelligence, the high barriers to entry and inconsistencies caused by manual configuration in data quality platforms have been resolved, achieving efficient and unified quality rule configuration.
Patent Information
- Application Number
- CN202511369953.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Current data quality platforms rely on manual configuration of quality rules, resulting in high barriers to entry, low efficiency, and inconsistent quality.
By using an AI-based quality rule configuration method, the system automatically acquires field data features of structured data using a master assignment model and a natural language model, and quickly matches and generates quality rules based on training results of historical structured data.
It lowers the professional skill requirements for users, improves configuration efficiency, avoids omissions of quality rules, ensures configuration consistency, and enhances data quality assurance capabilities.
Smart Images

Figure CN120849471B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a quality rule configuration method and related apparatus based on artificial intelligence. Background Technology
[0002] The current mainstream data quality platforms operate by requiring users to manually define quality rules for each data table and each data field. Common quality rules include null value validation rules, enumeration validation rules, and uniqueness constraint rules. However, this method of relying on manual configuration of quality rules has significant drawbacks:
[0003] First, the configuration threshold is high: users need to have both database expertise and a deep understanding of the business in order to accurately select the appropriate quality rules.
[0004] Second, low configuration efficiency: When faced with a large number of newly added data tables, manually configuring rules one by one is not only labor-intensive, but also prone to missing quality rules.
[0005] Third, inconsistent quality: Different operators have different understandings of business and quality rules, resulting in inconsistent quality rule configuration and ultimately uneven data quality assurance capabilities. Summary of the Invention
[0006] The purpose of this application is to provide a quality rule configuration method and related apparatus based on artificial intelligence, so as to solve the problems of high threshold, low efficiency and inconsistent quality caused by the current data quality platform relying on manual configuration of quality rules.
[0007] To achieve the above objectives, this application provides the following technical solution:
[0008] Firstly, this application proposes a technical solution for a quality rule configuration method based on artificial intelligence, which includes:
[0009] Based on the current structured data, obtain the first field data; the current structured data is obtained in advance; the first field data is any field data in the current structured data that needs to be configured with quality rules;
[0010] Based on the first field data, at least one first data feature is obtained; the first data feature is any one of structural features, semantic features, and distribution features.
[0011] Based on each first data feature, the matching degree between the first field data and each quality rule is obtained through the main allocation model; the main allocation model is pre-trained based on historical structured data; each quality rule is pre-set.
[0012] Based on the matching degree of each quality rule, an interpretable output is obtained through a natural language model;
[0013] Configure quality rules for the first field data based on the interpretability output and each matching degree.
[0014] Secondly, this application proposes a technical solution for an artificial intelligence-based quality rule configuration device, which includes:
[0015] The reading module is used to obtain the first field data based on the current structured data; the current structured data is obtained in advance; the first field data is any field data in the current structured data that needs to be configured with quality rules;
[0016] The processing module is configured to obtain at least one first data feature based on the first field data; the first data feature is any one of structural features, semantic features, and distribution features;
[0017] Furthermore, based on each first data feature, the matching degree between the first field data and each quality rule is obtained through the main allocation model; the main allocation model is pre-trained based on historical structured data; each quality rule is pre-set.
[0018] Furthermore, based on the matching degree of each quality rule, interpretable output is obtained through a natural language model;
[0019] In addition, quality rules are configured for the first field data based on the interpretability output and each matching degree.
[0020] Thirdly, this application proposes a technical solution for a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the quality rule configuration method based on artificial intelligence as described in the first aspect.
[0021] Compared with the prior art, the beneficial effects of this application are:
[0022] This application achieves automatic quality rule matching by acquiring the first field data and its various data features (structural features, semantic features, and distribution features, etc.) and utilizing the training results of the master allocation model and historical structured data. This eliminates the need for users to have in-depth knowledge of databases and the details of quality rules, reducing the professional skill requirements and enabling ordinary users to participate in quality rule configuration. Based on current structured data, this application automatically acquires the first field data and data features, quickly calculates the matching degree using a pre-trained master allocation model, and then combines this with a natural language model to generate interpretable output, automatically completing the quality rule configuration. This significantly reduces the workload of manual configuration, greatly improving configuration efficiency. When processing a large number of newly accessed data tables, it can quickly complete the quality rule configuration, avoiding omissions and improving user understanding of the quality rules. This application trains the master allocation model using historical structured data, performing quality rule matching based on a unified model and data. It is unaffected by subjective factors of operators, ensuring consistency in quality rule configuration and avoiding technical problems caused by inconsistent configuration interpretations due to differences in understanding among different operators, thus improving data quality assurance capabilities. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a quality rule configuration method based on artificial intelligence proposed in an embodiment of this application.
[0024] Figure 2 This is a schematic diagram of the structure of a quality rule configuration device based on artificial intelligence proposed in an embodiment of this application. Detailed Implementation
[0025] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects (e.g., first field data and second field data are represented as different field data, and so on), and are not necessarily used to describe a specific order or sequence. It should be understood that such names can be used interchangeably where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules in the embodiments of this application is merely a logical division; in actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be omitted or not performed. Additionally, the shown or discussed mutual coupling or direct coupling or communication connection may be through some interface, indirect coupling between modules, or electrical or other similar forms of communication connection, none of which are limited in the embodiments of this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed among multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.
[0026] The solutions provided in this application involve technologies such as Artificial Intelligence (AI) and Machine Learning (ML), and are specifically illustrated through the following embodiments:
[0027] AI, or Artificial Intelligence, refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, Artificial Intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine capable of reacting in a manner similar to human intelligence. Artificial Intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0028] AI technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0029] Machine learning is an important branch of AI that focuses on developing algorithms that enable computers to automatically learn and improve from data without being explicitly programmed to perform specific tasks. Simply put, the core of machine learning is to enable machines to "learn" patterns from experience (data) and then use these patterns to make predictions or decisions about new, unknown data.
[0030] In order for those skilled in the art to understand the embodiments of this application, it is also necessary to clearly understand the definitions of the following technical terms:
[0031] Quality rules are the specifications and guidelines that ensure data, processes, products, or services conform to preset standards. Their core function is to guarantee the reliability, consistency, and effectiveness of "quality" at every stage through systematic inspection and verification. Their specific functions can be elaborated from multiple dimensions, including data, business, and management.
[0032] First, ensure data quality to lay the foundation for decision-making and analysis.
[0033] In data-driven application scenarios, quality rules are a core tool for data governance, and their main functions include:
[0034] Identify data anomalies: Use quality rules (e.g., format validation, range validation, and logical validation) to identify invalid values (e.g., a 12-digit mobile phone number), outliers (e.g., a person who is 200 years old), and mismatched data (e.g., the "gender" field contains both "male" and "1"), thus preventing dirty data from flowing into subsequent analysis or the system.
[0035] Ensure data consistency: Verify the consistency of the same data in different scenarios (e.g., user ID format is consistent in the order table and user table), ensure logical consistency of related data (e.g., "order amount = unit price × quantity"), and avoid data conflicts.
[0036] Maintaining data integrity: By using rules such as non-empty validation and required field validation, ensure that key fields (such as "amount" and "time" in transaction records) are not missing, thus guaranteeing data availability.
[0037] Enhancing data credibility: Verified data sources are more reliable, and analytical conclusions derived from them (such as business reports and machine learning models) have decision-making value, avoiding "garbage data leading to garbage decisions".
[0038] Second, standardize business processes to reduce operational errors and risks.
[0039] In business execution scenarios, quality rules can standardize operations and reduce human error. Their main functions include:
[0040] Constrain operational behavior: Standardize employee operations through quality rules (e.g., "order amount must not be negative", "approval process requires signature at each level") to avoid business deviations caused by negligence or non-standard operations (e.g., incorrect filling of customer information and unauthorized approval).
[0041] Mitigating business risks: Set up verification rules for high-risk processes (e.g., "a single transfer exceeding 50,000 yuan requires secondary verification" and "contract terms must include force majeure clauses") to identify potential risks in advance (e.g., fraud and compliance issues).
[0042] Ensuring process compliance: In regulated industries such as finance and healthcare, quality rules can ensure that business operations comply with laws and regulations (e.g., "user information collection must be clearly authorized" or "drug release must verify batch number and expiration date"), thus avoiding compliance risks.
[0043] Third, improve product or service quality to meet user and standard requirements.
[0044] In product manufacturing or service delivery applications, quality rules are the core means of quality control, and their main functions include:
[0045] Ensuring product compliance: In manufacturing, quality rules (such as "part size error ≤ 0.1mm" or "product compressive strength ≥ 100MPa") are used to inspect whether products meet design standards and prevent unqualified products from entering the market.
[0046] Unified service standards: In the service industry, service behavior is standardized through quality rules (such as "responding to customer complaints within 24 hours" or "the service process must include 3 customer confirmations") to ensure consistent service quality in different scenarios and by different personnel.
[0047] Continuous optimization and improvement: The results of the execution of the quality rules (e.g., "percentage of non-conforming data" and "product defect rate") can be used as quality indicators to help companies locate problem areas (e.g., frequent failures in the verification of a certain type of data may be due to a malfunction in the data collection tool), and drive continuous quality improvement.
[0048] In other words, the essence of quality rules is to "prevent deviations with clear standards and ensure reliability with systematic checks." Their role spans the entire chain of data, business, products, and management, ultimately achieving the core goals of "reducing errors, mitigating risks, improving efficiency, and ensuring value." Whether in data governance, business operations, or product services, quality rules are the fundamental guarantee for ensuring "doing things according to standards and producing reliable results."
[0049] Structured data is a highly organized data format that follows a fixed model, making it easily identifiable, stored, and processed by computers. It is typically presented in tabular form, with clear relationships between data points, making it suitable for efficient querying, analysis, and statistics. Structured data is fundamental to data analysis and is widely used in enterprise business systems (e.g., ERP or CRM), government data statistics, e-commerce platform operations, and other scenarios. Due to its structured nature, it directly supports data visualization (e.g., reports and charts) and business decision-making. Structured data comes in various types, including data in relational database tables, comma-separated values (CSV), tab-separated values (TSV), and data in Excel spreadsheets.
[0050] Field data refers to the specific value or content (hereinafter referred to as data value) corresponding to each "field" in a structured data record. It is the basic building block of structured data, used to describe a specific attribute of a record. For example, structured data can be imagined as a table. The column headers of the table (e.g., "Name," "Age," and "Gender") are "fields," representing the attribute categories of the data. Each row in the table corresponds to one record, and the specific content of the corresponding column in each row (e.g., "Zhang San," "28," and "Male") is the field data for that field. Field data is the core of the value of structured data. All queries, statistics, and analyses based on structured data (e.g., "Average age of the Computer Technology Department" or "Filtering employees who joined after 2021") are essentially processing and operations on field data.
[0051] Natural Language Model (NLP): A NLP is a model in the field of artificial intelligence that focuses on understanding, generating, and processing human natural language. It learns from large amounts of text data to capture the patterns, grammatical structures, semantic relationships, and even contextual logic of language, thereby achieving intelligent processing of natural language.
[0052] Clustering: A cluster refers to a grouping of samples or data points in a dataset that share similar characteristics, achieved through algorithms. Samples or data points within the same cluster have a high degree of similarity, while samples or data points in different clusters differ significantly.
[0053] To address the technical problems of high barriers to entry, low efficiency, and inconsistent quality caused by the reliance on manual configuration of quality rules in current data quality platforms, as mentioned in the background, this application proposes an artificial intelligence-based method for configuring quality rules. Figure 1 As shown, the AI-based quality rule configuration method includes steps 100 to 500.
[0054] Step 100: Based on the current structured data, obtain the data of the first field.
[0055] In this embodiment, the current structured data is obtained in advance. In this embodiment, there are no restrictions on the source of the current structured data; for example, the current structured data can be entirely new structured data, or it can be structured data modified from historical structured data.
[0056] In a specific embodiment, entirely new structured data can refer to structured data created and recorded for the first time due to the addition of a new business module, organizational unit, or new scenario, without any historical data as a basis. For example, a company newly establishes an "Artificial Intelligence R&D Department" and creates a separate employee information table for this department. The employee information table contains fields such as "department name, employee name, position, and date of employment," and all records are entered for the first time; therefore, this employee information table is entirely new structured data. Or, an e-commerce platform adds a "cross-border direct mail" business and designs a separate order table for this business. Fields include "cross-border order number, overseas warehouse address, customs clearance status, and international logistics tracking number," etc. This type of order data has never existed before and is entirely new structured data. Or, a school adds an "Innovation and Entrepreneurship College" and creates a dedicated grade table for the students of this college, containing fields such as "college code, student ID, course name, and final grade." This dedicated grade table is entirely new structured data generated by the new entity.
[0057] In a specific embodiment, the current structured data obtained by modifying historical structured data can refer to new data formed on the basis of existing structured data through operations such as adding fields and updating field values, relying on the historical data framework. For example, a company originally had an "Employee Basic Information Table" (including fields: employee number, name, department, and salary, etc.). Due to policy requirements, a "Social Security Contribution Base" field was added, and data for this field was added for all historical employees. The updated table is the current structured data obtained by modifying historical structured data. Or, a hospital originally had a "Patient Medical Record Table" (including fields: medical record number, date of visit, and attending physician, etc.). To optimize services, a "Patient Satisfaction Rating" field was added. Subsequent medical records need to be filled in with this rating, and the rating data of the past 3 months of historical records is retrospectively supplemented. The resulting record table containing the new field is the current structured data obtained by modifying historical structured data. Alternatively, a supermarket may have an existing "Product Inventory Table" (containing fields such as product number, product name, and inventory quantity). Due to the introduction of a "near-expiry product management" mechanism, a "shelf life expiration date" field is added. After this information is added to all products, the updated inventory table is the current structured data obtained by modifying the historical structured data.
[0058] It should be noted that the AI-based quality rule configuration method proposed in this application has numerous application scenarios, making it impossible to exhaustively list them all. The embodiments in this application only list the application scenarios based on the described requirements (e.g., employee information tables, order tables, or dedicated performance tables mentioned above), and do not imply that the AI-based quality rule configuration method proposed in this application is only applicable to the exemplified application scenarios. It should be understood that the AI-based quality rule configuration method proposed in this application is applicable to all application scenarios similar to the listed scenarios, which will not be elaborated upon further.
[0059] In this embodiment, the first field data is any field data in the current structured data that requires quality rule configuration. It should be noted that if the current structured data is entirely new, all field data needs to be configured with quality rules; if the current structured data is structured data updated from historical structured data, only the newly added fields need to have quality rules configured.
[0060] Step 200: Based on the first field data, obtain at least one first data feature.
[0061] In this embodiment, the first data feature can be any data feature possessed by the first field data. For example, the data feature can be any one of the structural features, semantic features, and distribution features of the first field data.
[0062] In this embodiment, the first data feature in the first field data can be obtained in any suitable manner. For example, the method of obtaining the first data feature is at least as shown in Embodiments 1 to 3.
[0063] Example 1 of obtaining the first data feature
[0064] In this embodiment, the first data feature is a structural feature. Step 200 involves obtaining at least one first data feature based on the first field data, including steps 201 to 204.
[0065] Step 201: Based on the data in the first field, obtain the data definition language.
[0066] It is important to understand that Data Definition Language (DDL) is an important component of database languages. It is mainly used to define and manage data structures in a database, including the creation, modification, and deletion of objects such as databases, tables, views, and indexes.
[0067] Step 202: Based on the data definition language, obtain the field relationship graph through a graph neural network.
[0068] It is important to understand that a Graph Neural Network (GNN) is a deep learning model specifically designed for processing graph-structured data. It learns the feature representations of nodes or the entire graph by simulating the relationships between nodes (i.e., the first field in this embodiment) and other nodes (i.e., the fields other than the first field in this embodiment), thereby solving graph-related tasks. In other words, in this embodiment, the field relationship graph is the relationship graph formed by the first field and the remaining fields in the current structured data.
[0069] Step 203: Obtain the feature vector based on the field relationship graph.
[0070] In this embodiment, the feature vector is used to reflect at least the strength of the relationship between the first field and the remaining fields in the current structured data. For example, in a specific embodiment of this application, an e-commerce order table includes fields such as order number, user number, product number, payment amount, and order time, with the first field being the order number. Through the field relationship diagram, the feature vectors corresponding to the order number with the user number, product number, payment amount, and order time are 0.95, 0.80, 0.70, and 0.30, respectively. Based on the above feature vectors, it can be seen that the feature vector of "order number" and "user number" is 0.95, meaning that there is a strong correlation between "order number" and "user number". In other words, each order must be bound to a unique user, making it a core correlation field. The feature vector of "order number" and "product number" is 0.80, meaning that "order number" and "product number" have a relatively strong correlation. In other words, an order needs to record the purchased goods, which are closely related but may contain multiple product numbers. The feature vector of "order number" and "payment amount" is 0.70, meaning that there is a moderate correlation between "order number" and "payment amount". In other words, the order amount is a key attribute of the order, but it does not directly determine the uniqueness of the order. The feature vector of "order number" and "order time" is 0.30, meaning there is a weak correlation between "order number" and "order time". In other words, time is supplementary information and does not affect the logical binding of the order with other core fields. In another specific embodiment of this application, an employee information table includes fields such as employee number, department number, name, position, and direct supervisor number, with the first field being the department number. Through the field relationship diagram, the feature vectors of the department number with the employee number, name, direct supervisor number, and position are 0.90, 0.60, 0.85, and 0.20, respectively. Based on the above feature vectors, we can see that "Department ID" and "Employee ID" are strongly correlated (feature vector is 0.90), meaning that each employee must belong to a unique department, making it a mandatory correlation; "Department ID" and "Direct Supervisor ID" are relatively strongly correlated (feature vector is 0.85), meaning that supervisors and employees usually belong to the same department, with a close logical connection; "Department ID" and "Position" are moderately correlated (feature vector is 0.60), meaning that positions are related to departments, but there is a possibility of the same position across departments; "Department ID" and "Name" are weakly correlated (feature vector is 0.20), meaning that names and departments are not directly logically bound, but are only indirectly correlated through employee IDs.
[0071] It should be noted that the relationship strengths described above, such as strong association, relatively strong association, moderate association, and weak association, are merely illustrative examples in this embodiment. They do not mean that this embodiment can only divide relationship strength into four levels; in other embodiments, relationship strength can be divided into three or six levels, etc. In this embodiment, the threshold range for each relationship strength level can be set according to requirements. Taking the relationship strength levels of strong association, relatively strong association, moderate association, and weak association mentioned above as examples, assuming the value range of the feature vector is greater than or equal to 0 and less than or equal to 1, if the feature vector between any two fields is greater than or equal to 0.9 and less than or equal to 1.0, then these two fields can be considered strongly associated; if the feature vector between any two fields is greater than or equal to 0.8 and less than 0.9, then these two fields can be considered relatively strongly associated; if the feature vector between any two fields is greater than or equal to 0.7 and less than 0.8, then these two fields can be considered moderately associated; and if the feature vector between any two fields is greater than or equal to 0 and less than 0.7, then these two fields can be considered weakly associated.
[0072] Step 204: Based on the feature vector, obtain the structural features of the first field data.
[0073] In this embodiment, the feature vector can be directly used as the structural feature (i.e., the first data feature) of the first field data.
[0074] It's important to note that because feature vectors reflect the strength of relationships between fields, they can represent the structural logic (i.e., structural features) of a data table, thereby improving the accuracy of data quality rule recommendations. Field relationship strength is a key manifestation of a data table's structural features. In a data table, fields do not exist in isolation; the relationships between primary keys and foreign keys, and the dependencies between core and auxiliary fields, directly determine the business importance of a field and the requirements of quality rules. For example, a primary key field with high relationship strength with multiple foreign key fields indicates that it is a core identifier of the data in the table, typically requiring strict rules such as NOT NULL and uniqueness; while ordinary fields with low relationship strength (e.g., memo fields) may have more lenient rule requirements. Feature vectors representing relationship strength can help distinguish the business role of a field. For example, in an e-commerce user table, the user ID has strong relationships with fields in the order and payment tables, and its relationship strength feature vector value is significantly higher than that of ordinary fields (e.g., user avatar). This difference allows the primary allocation model described below to identify the user ID as a core business field, requiring priority for strict rule recommendations, while the user avatar, as an auxiliary field, can have more flexible rules, thus solving the technical problem of inaccurate quality rule recommendations caused by neglecting structural relationships.
[0075] Example 2 of obtaining the first data feature
[0076] In this embodiment, the first data feature is a semantic feature. Step 200 involves obtaining at least one first data feature based on the first field data, including steps 205 to 207.
[0077] Step 205: Obtain the field name from the first field data based on the language model.
[0078] In this embodiment, there are no restrictions on the language model. For example, the language model can be a large language model such as GPT-3 or LLaMA. To reduce resource consumption, in this embodiment, the language model can also be a lightweight language model (LLM). A lightweight language model, compared to a large language model, optimizes parameter size, computational efficiency, and deployment costs, making it easier to run in resource-constrained environments while maintaining core language understanding and generation capabilities. Its core design principle is to find a balance between "model performance" and "resource consumption" to meet the needs of efficiency, flexibility, and deployment scenarios in practical applications.
[0079] Step 206: Obtain the similarity between the field name and each semantic tag in the domain thesaurus.
[0080] In this embodiment, the domain-specific lexicon is pre-set. A domain-specific lexicon is a specialized vocabulary collection built for a specific discipline, industry, scenario, or theme. It contains core terms, professional expressions, industry jargon, abbreviations, and related semantic information within that domain, serving as a fundamental tool for supporting information processing, knowledge management, and language understanding within that domain. For example, in an e-commerce platform, processing data in the "user_info" table involves numerous fields. The "phone" field can be directly matched to the business meaning of "phone number" using the domain lexicon. For a vague field like "reg_time," combining a language model and referring to the annotation "user registration time," the final output is the semantic tag "registration time," with a similarity of 93%. Alternatively, when a financial institution processes transaction data, the "transaction_amount" field, based on the domain lexicon, is related to the transaction amount. Through in-depth analysis using a language model, the semantic tag "transaction amount" is output from the field name and possible annotations, with a similarity of 95%.
[0081] Step 207: Based on each similarity, obtain the semantic features of the first field data.
[0082] In this embodiment, the semantic tag with the highest similarity to the field name in the domain thesaurus can be used as the semantic feature (i.e., the first data feature) of the first field data. For example, in the e-commerce domain, the field name obtained through step 205 is "logistics tracking number". The semantic tags in the domain thesaurus include e-commerce logistics terminology tags such as express tracking number, waybill number, logistics code, and delivery tracking number. After step 206, the similarity between the logistics tracking number and the express tracking number, waybill number, logistics code, and delivery tracking number are calculated to be 0.88, 0.83, 0.65, and 0.60, respectively. Since the express tracking number has the highest similarity to the logistics tracking number, "express tracking number" can be used as the semantic feature of the first field data.
[0083] Example 3 of obtaining the first data feature
[0084] In this embodiment, the first data feature is a distribution feature. Step 200 involves obtaining at least one first data feature based on the first field data, including steps 208 to 210.
[0085] Step 208: Based on the data in the first field, obtain multiple rows of data.
[0086] As mentioned above, field data refers to the specific value or content corresponding to each "field" in a structured data record. In other words, in this embodiment, the multiple rows of data are the specific values or content of the multiple rows corresponding to the first field.
[0087] In this embodiment, the specific values or content of all rows in the first field data can be extracted as multiple rows of data. It should be noted that in application scenarios with extremely large amounts of data (e.g., order data from large e-commerce platforms, user behavior data from telecommunications operators, and user log data from internet companies), extracting the specific values or content of all rows in the first field data as multiple rows of data (e.g., millions of rows of data) will increase the computational burden and consume more computing resources and time.
[0088] In order to avoid excessive data volume of multiple rows of data, that is, to avoid increasing the computational burden, in one embodiment of this application, step 208, based on the first field data, obtains multiple rows of data, including steps 208a and 208b.
[0089] Step 208a: Based on the data in the first field, obtain the total number of rows.
[0090] In this embodiment, the total number of rows refers to the total number of rows of data in the first field.
[0091] Step 208b: If the total number of rows is less than or equal to the first preset value, then all rows of data in the first field data are taken as the multi-row data; otherwise, multiple rows of data are randomly selected from the first field data, and the number of rows of the multi-row data is equal to the first preset value.
[0092] In this embodiment, the first preset value can be a fixed value that can be set according to requirements. For example, the first preset value can be 5000 rows, 10000 rows, or 20000 rows, etc. In this embodiment, based on the total number of rows of data in the first field, the range of multiple rows of data used to extract distribution features is reasonably determined, which ensures data representativeness while taking into account processing efficiency, and provides a suitable sample basis for subsequent calculation of key statistics (e.g., null value rate and number of unique values).
[0093] It's important to note that in applications with extremely large datasets, extracting a fixed number (i.e., a first preset value) of rows from the first field may not fully reflect the distribution characteristics of the first field, thus affecting the accuracy of distribution feature extraction. For example, if the first field is a product category, containing data values for electronics, clothing, and food, with electronics accounting for 99%, clothing and food each accounting for 0.5%, and the total number of rows in the first field is 100,000, and the first preset value is 1,000 rows, during random sampling, due to the large proportion of electronics, it may be impossible to extract data values corresponding to clothing or food. Alternatively, if the first field is an order amount, with a total of 500,000 rows, most data is concentrated between 100 and 1,000 yuan, but there are 100 extreme upper limit data values (i.e., data values exceeding 1,000 yuan) and 100 extreme lower limit data values (i.e., data values below 100 yuan). If the first preset value is 5000 rows, the randomly selected multiple rows of data may not contain any extreme data values, resulting in significant deviations between the distribution characteristics of the "minimum value", "maximum value" and "data range" in the sample and the actual data, and failing to reflect the true distribution of extreme values in the monetary data.
[0094] To avoid the situation where extracting a fixed number of rows of data from the first field data fails to fully reflect the distribution characteristics of the first field data, in one embodiment of this application, the first preset value can be dynamically changed. That is, the first preset value can be dynamically updated. Based on this, after randomly extracting multiple rows of data from the first field data in step 208b, the method further includes steps 208c and 208d.
[0095] Step 208c: Obtain feedback information based on randomly selected data.
[0096] In this embodiment, the feedback information includes at least whether the randomly selected data lacks a required data value. The feedback information can be information actively entered by the user after reviewing the multiple rows of data extracted in step 208b (e.g., "minimum value not selected" and "not selected for payment or returned goods" hereinafter). To reduce workload, feedback information can be automatically generated by a computer. For example, in one embodiment of this application, a data value set can be pre-set, containing at least one required data value, and then feedback information can be automatically generated based on the data value set and multiple rows of data. For example, if the minimum price of a product in a table is 10 yuan and the maximum price is 100 yuan, the corresponding data value set could be "10 yuan, 100 yuan". After step 208b is completed, if there is no 10 yuan or 100 yuan in the multiple rows of data, then step 208d needs to be executed to update the first preset value. Alternatively, if a table contains information on the sales progress of goods, such as pending payment, paid, shipped, and returned, and the corresponding data value set is "pending payment, paid, shipped, and returned", after step 208b is completed, if none of the "pending payment, paid, shipped, and returned" values are found in the multiple rows of data, then step 208d needs to be executed to update the first preset value.
[0097] In this embodiment, there are no restrictions on the type of feedback information. For example, if the minimum price of a product in a table is 10 yuan and the maximum price is 100 yuan, and the extracted rows of data must include both the maximum and minimum prices, and the value of 10 yuan is not found in the extracted rows of data in step 208b, then the feedback information could be a proactive notification that "the minimum price was not selected." Alternatively, if the sales progress of a product in a table includes items such as pending payment, paid, shipped, and returned, and the extracted rows of data must include "pending payment, paid, shipped, and returned," and the items "pending payment" and "returned" are not found in the extracted rows of data in step 208b, then the feedback information could be a proactive notification that "pending payment and returned items were not selected."
[0098] To facilitate the computer's understanding of whether "the multi-row data is missing a data value that must be extracted," in one embodiment of this application, any data value in the data value set (hereinafter referred to as the first data value) can be compared with each data value in the multi-row data (hereinafter referred to as the second data value). If the first data value is present in each of the second data values, the counter value (initially 0) can be incremented by 1. After traversing all data values in the data value set, if the difference between the number of data values in the data value set and the counter value is not 0, it indicates that a data value that must be extracted is missing; otherwise, a data value that must be extracted is not missing. In another embodiment of this application, any data value in the data value set (hereinafter referred to as the first data value) can be compared with each data value in the multi-row data (hereinafter referred to as the second data value). If the first data value is present in each of the second data values, the first data value is deleted from the data value set. After traversing all data values in the data value set, if the data value set is not empty, it indicates that a data value that must be extracted is missing; otherwise, a data value that must be extracted is not missing.
[0099] Step 208d: If the feedback information meets the first preset condition, then the randomly selected data is used as the multi-row data; otherwise, the first preset value is updated, and data is randomly selected and feedback information is obtained again based on the updated first preset value, until the re-obtained feedback information meets the first preset condition; the first preset condition is that the multi-row data does not lack the data values that must be selected.
[0100] In this embodiment, the updated first preset value is greater than the original first preset value, and the first preset value is less than or equal to the total number of rows in the first field data. This embodiment optimizes the extracted multi-row data by dynamically updating the first preset value, ensuring that it can fully reflect the distribution characteristics of the first field data. This solves the technical problem of insufficient representativeness of the extracted multi-row data caused by a small initial first preset value, providing a reliable data foundation for accurately extracting distribution characteristics.
[0101] Step 209: Based on the multiple rows of data, obtain key statistics.
[0102] In this embodiment, the key statistics include at least one or a combination of several of the following: null value rate, number of unique values, maximum value, minimum value, and proportion of high-frequency values.
[0103] In this embodiment, the null value rate refers to the proportion of rows containing null values to the total number of rows in the data set; the number of unique values refers to the total number of values that appear only once in the data set; the maximum value refers to the maximum value in the data set; the minimum value refers to the minimum value in the data set; and the high-frequency value ratio refers to the proportion of rows containing that value to the total number of rows in the data set. In the field of data statistics technology, the null value rate, the number of unique values, the maximum value, the minimum value, and the high-frequency value ratio are all technical terms and will not be explained in detail here.
[0104] Step 210: Based on the key statistics, obtain the distribution characteristics of the first field data.
[0105] In this embodiment, the distribution characteristics of the first field data can be any value among the key statistics (e.g., the null value rate, the number of unique values, the maximum value, the minimum value, and the proportion of high-frequency values mentioned above).
[0106] This concludes the description of Embodiment 3 for obtaining the first data feature. Step 300 will be described below.
[0107] Step 300: Based on each first data feature, obtain the matching degree between the first field data and each quality rule through the master allocation model.
[0108] In this embodiment, the primary assignment model can be any model capable of matching the first field data with various quality rules based on the first data features. For example, the primary assignment model can be a neural network model, a support vector machine model, or a logistic regression model.
[0109] It is important to note that the primary allocation model can match a corresponding quality rule for the first field data based on a certain first data feature. For example, if the feature vector (i.e., the first data feature) between the first field data and the field data corresponding to the primary key is 0.95 (meaning the first field data and the field data corresponding to the primary key are strongly correlated), then a uniqueness constraint rule can be matched for the first field data; if the first data feature is a null value rate of 0.3%, then a null value check rule can be matched for the first field data.
[0110] In this embodiment, the primary allocation model is pre-trained based on historical structured data, and each quality rule is pre-set. To more accurately match the corresponding quality rule to field data (e.g., the first field data), in one embodiment of this application, the primary allocation model includes multiple sub-allocation models, each corresponding one-to-one with a quality rule. It is worth noting that the advantage of this one-to-one correspondence between sub-allocation models and quality rules is that, since each sub-allocation model corresponds to only one quality rule, the irrelevant influence of other quality rules can be reduced when calculating the matching degree between field data and that quality rule. This makes the matching degree result more accurately reflect the true fit between the field data and the quality rule, providing a more reliable basis for subsequent configuration. Furthermore, when the definition or applicable scenario of a quality rule changes, only the corresponding sub-allocation model needs to be updated specifically, without adjusting the entire primary allocation model, reducing the complexity of model maintenance. This allows the primary allocation model to quickly respond to dynamic changes in quality rules, ensuring the timeliness of the configuration logic.
[0111] In this embodiment, the master allocation model is trained based on historical structured data, including steps 310 to 360.
[0112] Step 310: Based on the historical structured data, obtain multiple second field data.
[0113] It should be noted that, based on the aforementioned historical structured data, multiple second field data can be obtained by referring to step 200, which will not be elaborated here.
[0114] Step 320: Traverse all quality rules and obtain the first quality rule.
[0115] In this embodiment, the first quality rule is any quality rule among the various quality rules that has not been trained to obtain the corresponding sub-assignment model. That is, in this embodiment, the training method of the sub-assignment model corresponding to each quality rule is the same as the training method of the sub-assignment model of the first quality rule.
[0116] In this embodiment, for each first quality rule, steps 330 to 350 below are executed until a corresponding sub-assignment model is trained based on each quality rule.
[0117] Step 330: Based on the first quality rule, obtain the current sub-assignment model.
[0118] In this embodiment, the current sub-assignment model corresponds to the first quality rule, and the current sub-assignment model is updated from the historical sub-assignment model. The historical sub-assignment model includes the initial sub-assignment model. In this embodiment, there are no restrictions on the type of sub-assignment model; for example, the sub-assignment model can be any one of a neural network model, a support vector machine model, and a logistic regression model.
[0119] It is important to note that neural network models, support vector machine models, and logistic regression models rely on learning the feature distribution of the data itself, lacking a mechanism to actively focus on key differential features, and have a weak ability to capture subtle feature differences. In other words, neural network models, support vector machine models, and logistic regression models are typically trained based on the overall sample distribution, with insufficient targeted optimization for locally misclassified samples, making them prone to repeating the same errors. To avoid the inability to match the correct quality rules to data in a certain field due to unclear or complex data features, in one embodiment of this application, the sub-assignment model includes at least one first assignment network.
[0120] In this embodiment, the allocation network refers to any allocation model capable of classifying training samples based on judgment rules (which can be preset). For example, the allocation network can be an allocation model based on the judgment rule of "whether a field in the field data is a primary key". If a field in the field data is a primary key, then in this allocation network, the value corresponding to the field data can be "1" (of course, in other embodiments, it can also be other values); if a field in the field data is not a primary key, then in this allocation network, the value corresponding to the field data can be "0" (of course, in other embodiments, it can also be other values). The allocation network can also be an allocation model based on the judgment rule of "whether the null value rate is less than a threshold (e.g., 3% or 5%)". If the null value rate of a field data is less than the threshold, then in this allocation network, the value corresponding to the field data can be "1"; if the null value rate of a field data is greater than or equal to the threshold, then in this allocation network, the value corresponding to the field data can be "0".
[0121] In one embodiment of this application, the current sub-assignment model is updated from the historical sub-assignment model in the following manner:
[0122] Step 331: Based on the historical sub-allocation model, obtain multiple third field data and multiple fourth field data from each second field data.
[0123] In this embodiment, the third field data refers to any field data in each of the second field data that should have been configured with the first quality rule, but was not configured with the first quality rule by the historical sub-assignment model. The fourth field data refers to any field data in each of the second field data that should not have been configured with the first quality rule, but was configured with the first quality rule by the historical sub-assignment model. That is to say, in this embodiment, both the third field data and the fourth field data are field data that were misjudged by the historical sub-assignment model.
[0124] Step 332: Obtain the second allocation network based on the data in each of the third and fourth fields.
[0125] In this embodiment, the second allocation network is any allocation network capable of correctly configuring the first quality rule for each second field data, and the allocation accuracy of the second allocation network is greater than 50%. For example, in one specific embodiment, the first quality rule is an "enumeration verification rule," and the judgment rule corresponding to the allocation network is "whether the unique value is less than 3%." The allocation accuracy is found to be 67%, so the allocation network corresponding to the "whether the unique value is less than 3%" judgment rule can be used as the second allocation network. In another specific embodiment, the first quality rule is an "enumeration verification rule," and the judgment rule corresponding to the allocation network is "whether the null value rate is less than 3%." The allocation accuracy is found to be 27%, so the allocation network corresponding to the "whether the null value rate is less than 3%" judgment rule cannot be used as the second allocation network.
[0126] In this embodiment, one allocation network can be randomly selected from the allocation networks that meet the above conditions (i.e., the allocation accuracy is greater than 50%) as the second allocation network, and the second allocation network is different from the previous first allocation networks.
[0127] In order to obtain an allocation network that can accurately configure the first quality rule for each third field data and each fourth field data to the greatest extent possible during each iteration of the sub-allocation model, so that the sub-allocation model can converge quickly, in one embodiment of this application, step 332, obtaining the second allocation network based on each third field data and each fourth field data, includes steps 332a to 332d.
[0128] Step 332a: Based on the data of each third field and each fourth field, obtain the first feature set and the second feature set.
[0129] In this embodiment, each data feature in the first feature set originates from the third field data. Each data feature in the second feature set originates from the fourth field data.
[0130] Step 332b: Based on the clustering algorithm, obtain multiple first clusters from the first feature set and multiple second clusters from the second feature set.
[0131] It's important to understand that clustering algorithms are an unsupervised learning method. Their core objective is to divide samples (i.e., data features) in a dataset (i.e., the feature set in this example) into several different "clusters" based on a certain similarity metric. Samples within the same cluster have high similarity, while samples in different clusters differ significantly. It doesn't require prior knowledge of the data's category labels; instead, it automatically discovers potential structures and patterns based on the data's own distribution characteristics. In other words, classifying samples in a dataset using clustering algorithms is a mature technique, which will not be elaborated upon here.
[0132] In this embodiment, the data features in each cluster are of the same type. For example, in one embodiment of this application, the data features in each cluster (e.g., the first cluster or the second cluster) can be of the type of "unique value rate", "null value rate", or "association strength", etc.
[0133] Step 332c: Based on each first cluster and each second cluster, obtain the third cluster and the fourth cluster.
[0134] In this embodiment, the third cluster is any cluster among the first clusters, and the fourth cluster is a cluster of the same type as the third cluster among the second clusters. Furthermore, the data characteristics of the third cluster and the fourth cluster have the greatest difference.
[0135] In one embodiment of this application, the calculation formula for the difference in data features between a first cluster and a second cluster of the same type (hereinafter referred to as the first calculation formula) can be as follows:
[0136]
[0137] in, This indicates the differences in data characteristics between the first cluster and the second cluster; Indicates the upper limit of data characteristics; This represents the minimum value of the data feature in the first cluster; This represents the minimum value of the data feature in the second cluster; This indicates that the absolute value is being calculated.
[0138] In a specific embodiment of this application, it is assumed that the first quality rule is a "uniqueness constraint rule," and the data feature types are only "unique value rate" and "null value rate." When the data feature type is "unique value rate," the minimum value of the first cluster is 97%, the minimum value of the second cluster is 20%, and the data feature difference is 77% (derived from the first calculation formula). When the data feature type is "null value rate," the minimum value of the first cluster is 3%, the minimum value of the second cluster is 5%, and the data feature difference is 2% (derived from the first calculation formula). In this embodiment, since the data feature difference corresponding to the "unique value rate" is greater than the data feature difference corresponding to the "null value rate," the first cluster based on the "unique value rate" can be used as the third cluster, and the second cluster based on the "unique value rate" can be used as the fourth cluster.
[0139] In another embodiment of this application, the formula for calculating the differences in data features between the first cluster and the second cluster of the same type can also be as follows:
[0140]
[0141] in, This indicates the differences in data characteristics between the first cluster and the second cluster; Indicates the upper limit of data characteristics; This represents the average value of the data features in the first cluster; This represents the average value of the data features in the second cluster; This indicates that the absolute value is being calculated.
[0142] It is important to note that in most application scenarios, the data features within clusters are numerical. To avoid situations where the data features of some first and second clusters of the same type are non-numerical, thus making it impossible to obtain the differences in data features between these clusters using the aforementioned calculation formula, in the embodiments of this application, these data features can be mapped to numerical values. For example, if the data feature corresponding to a certain data field is "primary key," then that data feature (i.e., the "primary key" data feature) can be mapped to the value "1"; if the data feature corresponding to a certain data field is "non-primary key," then that data feature (i.e., the "non-primary key" data feature) can be mapped to the value "0." Of course, in the embodiments of this application, these mapped values can all be preset.
[0143] Step 332d: Obtain the second allocation network based on the third cluster and the fourth cluster.
[0144] In this embodiment, since the data features in the third and fourth clusters differ the most, and the data features in the third and fourth clusters are of the same type, the allocation network with the same or similar data feature type as the third and fourth clusters can more accurately configure the first quality rule for each third field data and each fourth field data. For example, in a specific embodiment of this application, assuming the first quality rule is a "uniqueness constraint rule," and the data feature type of the third and fourth clusters is a "unique value rate," then the allocation network corresponding to the judgment rule related to the "unique value rate" (e.g., whether the unique value rate is greater than 95%) can be selected as the second allocation network. In this embodiment, the second allocation network can maximize the use of the key information of the misjudged samples (i.e., the third and fourth clusters), which not only improves the allocation accuracy but also reduces the interference of redundant features on the learning of the sub-allocation model, ultimately achieving rapid convergence of the sub-allocation model.
[0145] Step 333: Based on the second allocation network, update each first weight in the historical sub-allocation model.
[0146] In this embodiment, the first weight is the weight of the first allocation network in the historical sub-allocation model. That is, each first allocation network has a corresponding weight. In this embodiment, the weight of the allocation network is closely related to the quality rules. For example, if the judgment rule corresponding to the allocation network is "whether the field in the field data is a primary key" and the quality rule is "uniqueness constraint rule", then because the judgment rule "whether the field in the field data is a primary key" has a greater impact on the judgment of the "uniqueness constraint rule", the allocation network corresponding to the judgment rule "whether the field in the field data is a primary key" has a larger weight when configuring the "uniqueness constraint rule". If the judgment rule corresponding to the allocation network is "whether the field in the field data is a primary key" and the quality rule is "enumeration verification rule", then because the judgment rule "whether the field in the field data is a primary key" has a very small impact on the judgment of the "enumeration verification rule", the allocation network corresponding to the judgment rule "whether the field in the field data is a primary key" has a very small weight when configuring the "enumeration verification rule".
[0147] In this embodiment, there are no restrictions on the iterative update method for each first weight. For example, in one embodiment of this application, the allocation accuracy corresponding to each allocation network (e.g., the first allocation network or the second allocation network) can be used as the weight of that allocation network. As mentioned above, the third field data and the fourth field data are both field data that were misjudged by the historical sub-allocation model (i.e., field data that were incorrectly configured with the first quality rule). The second allocation network is the allocation network that can most accurately determine whether the first quality rule should be configured for the third field data and the fourth field data. That is, the second allocation network can correct the misjudgment results of the historical sub-allocation model. In order to highlight the value of the newly added second allocation network in correcting historical misjudged samples during the iteration process, in one embodiment of this application, step 333 updates each first weight in the historical sub-allocation model based on the second allocation network, including steps 333a to 333c.
[0148] Step 333a: Based on the second allocation network, obtain the number of iteration updates.
[0149] It is important to understand that obtaining the number of iterations during the iterative update of the model (i.e., the sub-assignment model) is a mature technique and will not be elaborated upon here.
[0150] Step 333b: Obtain the second weight based on the number of iterations.
[0151] In this embodiment, the second weight is the weight corresponding to the second allocation network, and the second weight is positively correlated with the number of iterations.
[0152] In this embodiment, any reasonable method can be used to obtain the second weight based on the number of iterations. For example, in one embodiment of this application, step 333b, the calculation formula for obtaining the second weight based on the number of iterations is as follows:
[0153]
[0154] in, Indicates the second weight; This represents a preset constant, which can be any positive number, such as 0.5 or 1.0. Indicates the number of iterations; This represents the normalization function, used to normalize the values within the parentheses to the range [0, 1]. In this embodiment, the larger the number of iterations, the larger the corresponding second weight, meaning the greater the role of the second allocation network in the current sub-allocation model formed in step 334.
[0155] In another embodiment of this application, step 333b, based on the number of iterations, obtains the following formula for calculating the second weight:
[0156]
[0157] in, Indicates the second weight; This represents a preset constant, which can be any positive number, such as 0.5 or 1.0. Indicates the number of iterations; An exponential function with the natural constant e as its base. In this embodiment, the larger the number of iterations, the larger the corresponding second weight, which means the second allocation network plays a greater role in the current sub-allocation model formed in step 334.
[0158] Step 333c: Update each of the first weights in the historical sub-assignment model based on the second weight.
[0159] In this embodiment, the second weight is positively correlated with the number of iterations. That is, as the number of iterations increases, the weight of the second allocation network (the allocation network used to correct misjudgments in the historical sub-allocation model) introduced in step 334 will increase. In other words, the updated historical sub-allocation model (i.e., the current sub-allocation model) has a higher value in correcting the overall performance of the historical sub-allocation model because the second allocation network can more accurately handle complex or insufficiently covered misjudgment scenarios. Compared to directly using the allocation accuracy as the second weight, this embodiment fully considers the temporal value of the second allocation network in the iteration process, and leverages the role of the second allocation network in repairing the defects of the historical sub-allocation model by increasing the weight of the second allocation network.
[0160] It is important to note that the second weight corresponding to the second allocation network is closely related to its own allocation accuracy. That is, if the allocation accuracy of the second allocation network is higher, the second weight should also be higher; if the allocation accuracy of the second allocation network is lower, the second weight should also be lower. Based on this, in one embodiment of this application, before step 333c, which updates each of the first weights in the historical sub-allocation model based on the second weight, the method further includes steps 333d to 333f.
[0161] Step 333d: Configure the first quality rule for each second field data based on the second allocation network, and obtain the allocation accuracy.
[0162] In this embodiment, the formula for calculating the assignment accuracy is as follows:
[0163]
[0164] in, Indicates the accuracy of the allocation; This indicates the number of times the first quality rule can be correctly configured for each second field data based on the second allocation network; This indicates the total number of data in each of the second fields.
[0165] Step 333e: Based on the allocation accuracy, obtain the correction coefficient.
[0166] In this embodiment, the correction coefficient is positively correlated with the allocation accuracy. As will be explained below, the correction coefficient is used to correct the second weight obtained in step 333b, and can be obtained based on the allocation accuracy in any reasonable manner. For example, the allocation accuracy can be directly used as the correction coefficient; or, in step 333e, the formula for calculating the correction coefficient based on the allocation accuracy can be as follows:
[0167]
[0168] in, Indicates the correction factor; Indicates the accuracy of the allocation; This represents an exponential function with the natural constant e as its base.
[0169] Step 333f: Adjust the second weight based on the correction coefficient.
[0170] In this embodiment, the second weight can be adjusted based on the correction coefficient in any reasonable manner. For example, in step 333f, the calculation formula for adjusting the second weight based on the correction coefficient can be as follows:
[0171]
[0172] in, This indicates the second weight after correction; This indicates the second weight before the correction; Indicates the correction factor; This represents a normalization function used to normalize the values within the parentheses to the range [0, 1]. Alternatively, in step 333f, based on the correction coefficient, the calculation formula for the second weight can be modified as follows:
[0173]
[0174] in, This indicates the second weight after correction; This indicates the second weight before the correction; Indicates the correction factor; This represents a normalization function used to normalize the values within the parentheses to the range [0, 1].
[0175] Step 334: Add the second allocation network to the historical sub-allocation model to obtain the current sub-allocation model.
[0176] In other words, in this embodiment, the current sub-assignment model always has one more assignment network (i.e., a second assignment network) than the historical sub-assignment model. In other words, since the current sub-assignment model has added a second assignment network and the sum of the weights is 1, it is necessary to update the weights of the remaining assignment networks of the current sub-assignment model (i.e., the first weights in the historical sub-assignment model in step 333c).
[0177] This concludes the introduction to the current method for obtaining the sub-assignment model.
[0178] Step 340: Assign the first quality rule to each second field data based on the current sub-assignment model, and obtain the current loss value.
[0179] It is important to understand that the higher the error rate of the current sub-assignment model in assigning the first quality rule to each second field of data, the greater the current loss value; conversely, the lower the error rate, the smaller the current loss value. In this embodiment, the error rate of the current sub-assignment model in assigning the first quality rule to each second field of data can be directly used as the current loss value.
[0180] Step 350: If the current loss value does not meet the second preset condition, update the current sub-assignment model and re-acquire the current loss value until the re-acquired current loss value meets the second preset condition. Then, take the sub-assignment model whose current loss value meets the second preset condition as the trained sub-assignment model.
[0181] In this embodiment, the second preset condition can be that the current loss value is less than a certain preset value (e.g., 0.02 or 0.03); or, the second preset condition can be that the number of iterations of the current loss value exceeds a certain preset value (e.g., 10 times or 20 times).
[0182] Step 360: Based on each trained sub-assignment model, obtain the main assignment model.
[0183] In this embodiment, the main allocation model is composed of each trained sub-allocation model.
[0184] It is important to understand that this embodiment extracts misjudged samples (i.e., third and fourth field data) from the historical sub-assignment model and constructs a second assignment network that can correctly handle these samples. It dynamically integrates the original first assignment networks with the newly added second assignment network and optimizes the performance by adjusting the weights. Compared with neural networks, support vector machines, and logistic regression models, it can more accurately repair the defects of the historical sub-assignment model, thereby improving the accuracy and adaptability of quality rule configuration.
[0185] Step 400: Based on the matching degree of each quality rule, obtain interpretable output through a natural language model.
[0186] In this embodiment, the natural language model can be a lightweight language model or a large language model (e.g., Doubao or GPT). It is important to understand that the interpretability output should at least include the corresponding quality rule and the reason for selecting that quality rule. For example, in one embodiment of this application, the interpretability output might be as follows: "Since this field is a primary key field and the null value rate is 0.1%, a uniqueness constraint rule is configured for this field"; in another embodiment of this application, the interpretability output might be as follows: "Since this field has few unique values and high-frequency values are concentrated, an enumeration validation rule is configured for this field."
[0187] In the embodiments of this application, the quality parameters of the quality rules can be fixed. Quality parameters refer to specific quantitative indicators or conditions corresponding to a particular quality rule, used to define the execution standard or judgment basis of that quality rule, and are the core elements for the quality rule to be implemented. For example, in a specific embodiment, the quality rule is a null value validation rule (used to check whether the proportion of null values in field data meets the standard), and its corresponding quality parameter can be a null value rate of less than or equal to 5%, indicating that when the null value rate of the corresponding field exceeds 5%, it is judged as data quality abnormality.
[0188] Step 500: Configure quality rules for the first field data based on the interpretability output and each matching degree.
[0189] In this embodiment, quality rules can be configured for the first field data based on the interpretability output and each matching degree in any reasonable manner. For example, quality rules can be configured manually based on the interpretability output and each matching degree. Alternatively, quality rules can be configured automatically based on each matching degree. For example, one or more quality rules with the highest corresponding matching degrees (arranged in descending order of matching degree) can be selected as the quality rules for the first field data; or, all quality rules with matching degrees greater than a third preset value can be configured as the quality rules for the first field data.
[0190] In this embodiment, the third preset value can be set according to requirements. For example, the third preset value can be a fixed value (e.g., 0.9 or 0.8) or a dynamically changing value (e.g., the third preset value can be 0.8 times or 0.9 times the maximum value among the various matching degrees).
[0191] It is important to note that if the quality parameters are fixed, after the quality rules are configured, the data characteristics may change over time or in different business scenarios, or the initially configured quality parameters may not fully adapt to the actual data distribution. This could lead to numerous incorrect judgments in the quality rules, i.e., frequent false alarms. To enable dynamic adjustment of the quality parameters and thus avoid frequent false alarms, in one embodiment of this application, after configuring quality rules for the first field data based on the interpretability output and each matching degree in step 500, the method further includes steps 600 to 900.
[0192] Step 600: Based on the first field data, obtain the second quality rule.
[0193] In this embodiment, the second quality rule is any quality rule configured in the first field data. That is, in this embodiment, the dynamic adjustment method of the quality parameters of any quality rule configured in the first field data is the same as the dynamic adjustment method of the quality parameters of the second quality rule.
[0194] Step 700: Obtain the first quality parameter based on the second quality rule.
[0195] In this embodiment, the first quality parameter is the quality parameter corresponding to the second quality rule. For example, in the application scenario of the order transaction table, if the second quality rule corresponding to the field "transaction amount" is a null value validation rule, then the corresponding first quality parameter can be a null value rate of less than or equal to 5%, that is, when the null value rate of the corresponding field exceeds 5%, it is determined to be a data quality anomaly. If the second quality rule corresponding to the field "transaction amount" is a range validation rule, then the corresponding first quality parameter can be a value between [18, 60], and when the value exceeds this range, it is determined to be a data quality anomaly. In the application scenario of the student information table, if the second quality rule corresponding to the field "student number" is a uniqueness constraint rule, then the corresponding first quality parameter can be a duplicate value count of 0, that is, when a duplicate value appears in the field, it is determined to be a data quality anomaly. In the application scenario of the student information table, if the second quality rule corresponding to the field "gender" is an enumeration validation rule, then the corresponding first quality parameter can be that the value of the field data must belong to a preset enumeration list (i.e., "male" and "female"). If a value outside the enumeration list appears, it is determined to be a data quality anomaly.
[0196] Step 800: Obtain the false positive rate based on the first quality parameter.
[0197] In this embodiment, the false positive rate can be obtained periodically. For example, for tables with fast data update speeds, the false positive rate can be obtained every 10 or 20 minutes; for tables with slow data refresh speeds, the false positive rate can be obtained every 3 or 4 hours.
[0198] Step 900: If the misjudgment rate is greater than the second preset value, then update the first quality parameter to obtain the second quality parameter, and update the quality parameter corresponding to the second quality rule from the first quality parameter to the second quality parameter.
[0199] This embodiment can dynamically optimize and iteratively upgrade the configured quality rules. By monitoring the false judgment rate and adjusting the quality parameters accordingly, it can continuously improve the adaptability and accuracy of the quality rules.
[0200] In this embodiment, a second preset value can be set according to requirements, for example, the second preset value can be 2% or 3%. In this embodiment, the first quality parameter can be updated in any reasonable way to obtain the second quality parameter. For example, in a specific embodiment, the transaction amount of a certain product is between 30 yuan and 60 yuan. The second quality rule corresponding to the "transaction amount" field in the formed "order transaction table" is a range verification rule, and the first quality parameter corresponding to the range verification rule is that the transaction amount is greater than or equal to 30 and less than or equal to 60. Subsequently, the product participates in a promotional activity, and the price drops to 15 yuan. Then, the "transaction amount" in the "order transaction table" will have a large number of data values that exceed the range [30, 60] (i.e., 15). By verifying the data values in the "transaction amount" through the range verification rule, it is found that the false judgment rate reaches 0.5% (the second preset value is 0.1%). In this embodiment, in order to avoid frequent errors, the first quality parameter (i.e., the transaction amount is greater than or equal to 30 and less than or equal to 60) can be modified to the second quality parameter (e.g., the transaction amount is greater than or equal to 15 and less than or equal to 60).
[0201] It is important to note that the core function of quality rules is to ensure that data conforms to preset standards, and their parameter settings must match the reasonable boundaries of the business scenario. Directly modifying quality parameters may cause the quality rules to deviate from their original design purpose. For example, the transaction amount value in the quality parameters is greater than or equal to 30 and less than or equal to 60 to prevent erroneous transaction records below cost price. If the quality parameters are directly expanded to a transaction amount value greater than or equal to 15 and less than or equal to 60, then after the promotional activity, truly abnormal transactions below cost price (e.g., a 20 yuan erroneous entry in a non-promotional scenario) may be misjudged as normal, weakening the quality rules' role in ensuring the rationality of the business. Based on this, in one embodiment of this application, step 900, if the misjudgment rate is greater than a second preset value, updates the first quality parameter to obtain a second quality parameter, including steps 910 to 930.
[0202] Step 910: Based on the first field data, obtain multiple incorrect judgment samples and multiple correct judgment samples.
[0203] In this embodiment, assuming the transaction amount of a certain product is between 30 and 60 yuan, the second quality rule corresponding to the "Transaction Amount" field in the generated "Order Transaction Table" is a range verification rule, and the first quality parameter corresponding to the range verification rule is that the transaction amount value is between [30, 60]. Subsequently, the product participates in a promotional activity, and the price drops to 15 yuan. Then, the "Transaction Amount" in the "Order Transaction Table" will show a large number of data values exceeding the range [30, 60] (i.e., 15). In this embodiment, samples with transaction amounts between [30, 60] are correctly judged samples; samples with transaction amounts not falling within the range [30, 60] are incorrectly judged samples.
[0204] Step 920: Based on each incorrectly judged sample and each correctly judged sample, obtain the distinguishing data features.
[0205] In this embodiment, the distinguishing data feature is a data feature present in the incorrect judgment sample but not present in the correct judgment sample. For example, in the embodiment of step 910, the "Remarks" field corresponding to the incorrect judgment sample has the data feature "Promotion"; while the "Remarks" field in the correct judgment sample is empty or has the data feature "Normal Sales". That is to say, in this embodiment, "Promotion" in the "Remarks" field can be used as the distinguishing data feature.
[0206] In this embodiment, there are many algorithms that can acquire distinguishable data features, such as frequency threshold algorithm, feature selection algorithm based on cluster consistency and Fisher exact test algorithm.
[0207] Step 930: Based on the distinguishing data features, update the first quality parameter and obtain the second quality parameter.
[0208] In this embodiment, if the first quality parameter is a sales amount greater than or equal to 30 yuan and less than or equal to 60 yuan, then the second quality parameter can be: if the "Remarks" field contains "Promotion," then the sales amount is greater than or equal to 15 yuan and less than or equal to 60 yuan; if the "Remarks" field does not contain "Promotion," then the sales amount is greater than or equal to 30 yuan and less than or equal to 60 yuan. This embodiment updates the first quality parameter based on distinguishing data features to obtain the second quality parameter, relaxing the constraints only when specific conditions are met (i.e., the "Remarks" field contains "Promotion"), thus limiting the impact of the adjustment to a specific range. This avoids both frequent false alarms and the inability to identify genuine abnormal data (e.g., a 15 yuan product in a non-promotional scenario) caused by directly expanding the constraints (i.e., directly adjusting the transaction amount range from [30,60] to [15,60]).
[0209] The AI-based quality rule configuration method proposed in this application automatically matches quality rules by acquiring the first field data and its various data features (structural features, semantic features, and distribution features, etc.) and using the training results of the master allocation model and historical structured data. This eliminates the need for users to have in-depth knowledge of databases and the details of quality rules, reducing the professional skills required and enabling ordinary users to participate in quality rule configuration. This application automatically acquires the first field data and data features based on current structured data, quickly calculates the matching degree using a pre-trained master allocation model, and then generates interpretable output using a natural language model, automatically completing the quality rule configuration. This significantly reduces the workload of manual configuration, greatly improving configuration efficiency. When processing a large number of newly accessed data tables, it can quickly complete the quality rule configuration, avoiding omissions and improving user understanding of the quality rules. This application trains the master allocation model using historical structured data, performing quality rule matching based on a unified model and data. It is unaffected by subjective factors of operators, ensuring consistency in quality rule configuration and avoiding technical problems caused by inconsistent configuration interpretations due to differences in understanding among different operators, thus improving data quality assurance capabilities.
[0210] Having introduced the AI-based quality rule configuration method proposed in the embodiments of this application, the following describes an embodiment of the AI-based quality rule configuration device proposed in this application, such as... Figure 2 As shown, the AI-based quality rule configuration device 10 includes:
[0211] The reading module 11 is used to obtain first field data based on the current structured data; the current structured data is obtained in advance; the first field data is any field data in the current structured data that needs to be configured with quality rules;
[0212] Processing module 12 is used to obtain at least one first data feature based on the first field data; the first data feature is any one of structural features, semantic features, and distribution features;
[0213] Furthermore, based on each first data feature, the matching degree between the first field data and each quality rule is obtained through the main allocation model; the main allocation model is pre-trained based on historical structured data; each quality rule is pre-set.
[0214] Furthermore, based on the matching degree of each quality rule, interpretable output is obtained through a natural language model;
[0215] In addition, quality rules are configured for the first field data based on the interpretability output and each matching degree.
[0216] As a specific embodiment of this application, the processing module 12 is further configured to obtain a data definition language based on the first field data;
[0217] Furthermore, based on the data definition language, a field relationship graph is obtained through a graph neural network; the field relationship graph is a relationship graph formed by the first field and the remaining fields in the current structured data;
[0218] Furthermore, based on the field relationship graph, a feature vector is obtained; the feature vector is used to reflect at least the relationship strength between the first field and the remaining fields in the current structured data.
[0219] Furthermore, based on the feature vector, the structural features of the first field data are obtained.
[0220] As a specific embodiment of this application, the processing module 12 is further configured to obtain the field name from the first field data based on the language model;
[0221] In addition, the similarity between the field name and each semantic tag in the domain thesaurus is obtained; the domain thesaurus is preset.
[0222] Furthermore, based on each similarity, the semantic features of the first field data are obtained.
[0223] As a specific embodiment of this application, the processing module 12 is further configured to obtain multiple rows of data based on the first field data;
[0224] Furthermore, based on the multiple rows of data, key statistics are obtained; the key statistics include at least one or more of the following: null value rate, number of unique values, maximum value, minimum value, and proportion of high-frequency values.
[0225] Furthermore, based on the key statistics, the distribution characteristics of the first field data are obtained.
[0226] As a specific embodiment of this application, the processing module 12 is further configured to obtain the total number of rows based on the first field data; the total number of rows is the total number of rows of data in the first field data;
[0227] Furthermore, if the total number of rows is less than or equal to a first preset value, then all rows of data in the first field data are taken as the multi-row data; otherwise, multi-row data is randomly extracted from the first field data, and the number of rows of the multi-row data is equal to the first preset value.
[0228] As a specific embodiment of this application, the processing module 12 is further configured to obtain feedback information based on randomly selected data; the feedback information includes at least whether the multiple rows of data are missing data values that must be selected;
[0229] Furthermore, if the feedback information meets the first preset condition, the randomly selected data will be used as the multi-row data; otherwise, the first preset value will be updated, and data will be randomly selected and feedback information will be obtained again based on the updated first preset value until the re-obtained feedback information meets the first preset condition; the first preset condition is that the multi-row data does not lack the data values that must be selected.
[0230] As a specific embodiment of this application, the main allocation model includes multiple sub-allocation models, each sub-allocation model corresponding to a quality rule; the processing module 12 is further used to obtain multiple second field data based on the historical structured data;
[0231] And, iterate through each quality rule to obtain the first quality rule; the first quality rule is any quality rule among the various quality rules that has not been trained to obtain the corresponding sub-assignment model;
[0232] Furthermore, for each first quality rule, the following steps are performed until a corresponding sub-assignment model is trained based on each quality rule:
[0233] Furthermore, based on the first quality rule, a current sub-assignment model is obtained; the current sub-assignment model corresponds to the first quality rule, and the current sub-assignment model is updated from the historical sub-assignment model; the historical sub-assignment model includes the initial sub-assignment model;
[0234] And, based on the current sub-assignment model, assign the first quality rule to each second field data and obtain the current loss value;
[0235] Furthermore, if the current loss value does not meet the second preset condition, the current sub-assignment model is updated, and the current loss value is re-acquired until the re-acquired current loss value meets the second preset condition. The sub-assignment model whose current loss value meets the second preset condition is then used as the trained sub-assignment model.
[0236] Furthermore, based on each trained sub-assignment model, the main assignment model is obtained.
[0237] As a specific embodiment of this application, the sub-allocation model includes at least one first allocation network; each first allocation network has a corresponding weight; the processing module 12 is further configured to obtain multiple third field data and multiple fourth field data based on the historical sub-allocation model from each second field data; the third field data is any field data in each second field data that should have been configured with the first quality rule, but was not configured with the first quality rule by the historical sub-allocation model; the fourth field data is any field data in each second field data that should not have been configured with the first quality rule, but was configured with the first quality rule by the historical sub-allocation model;
[0238] Furthermore, based on the data in each of the third and fourth fields, a second allocation network is obtained; the allocation accuracy of the second allocation network is greater than 50%.
[0239] Furthermore, based on the second allocation network, each first weight in the historical sub-allocation model is updated; the first weight is the weight in the historical sub-allocation model corresponding to the first allocation network.
[0240] In addition, the second allocation network is added to the historical sub-allocation model to obtain the current sub-allocation model.
[0241] As a specific embodiment of this application, the processing module 12 is further configured to obtain a first feature set and a second feature set based on each third field data and each fourth field data; each data feature in the first feature set comes from the third field data; each data feature in the second feature set comes from the fourth field data;
[0242] Furthermore, based on the clustering algorithm, multiple first clusters are obtained from the first feature set, and multiple second clusters are obtained from the second feature set; the data features in each cluster are of the same type.
[0243] Furthermore, based on each first cluster and each second cluster, a third cluster and a fourth cluster are obtained; the third cluster is any cluster among each first cluster, and the fourth cluster is a cluster of the same type as the third cluster among each second cluster, and the data features of the third cluster and the fourth cluster have the greatest difference.
[0244] Furthermore, a second allocation network is obtained based on the third cluster and the fourth cluster.
[0245] As a specific embodiment of this application, the processing module 12 is further configured to obtain the number of iteration updates based on the second allocation network;
[0246] Furthermore, based on the number of iterations, a second weight is obtained; the second weight is the weight corresponding to the second allocation network, and the second weight is positively correlated with the number of iterations.
[0247] And, based on the second weight, update each of the first weights in the historical sub-assignment model.
[0248] As a specific embodiment of this application, the processing module 12 is further configured to configure the first quality rule for each second field data based on the second allocation network, and obtain the allocation accuracy.
[0249] Furthermore, a correction coefficient is obtained based on the allocation accuracy; the correction coefficient is positively correlated with the allocation accuracy.
[0250] And, based on the correction coefficient, the second weight is corrected.
[0251] As a specific embodiment of this application, the processing module 12 is further configured to obtain a second quality rule based on the first field data; the second quality rule is any quality rule configured in the first field data;
[0252] And, based on the second quality rule, a first quality parameter is obtained; the first quality parameter is the quality parameter corresponding to the second quality rule;
[0253] And, based on the first quality parameter, obtain the false positive rate;
[0254] Furthermore, if the false positive rate is greater than the second preset value, the first quality parameter is updated, the second quality parameter is obtained, and the quality parameter corresponding to the second quality rule is updated from the first quality parameter to the second quality parameter.
[0255] As a specific embodiment of this application, the processing module 12 is further configured to obtain multiple incorrect judgment samples and multiple correct judgment samples based on the first field data;
[0256] Furthermore, based on each incorrect judgment sample and each correct judgment sample, distinguishing data features are obtained; the distinguishing data features are data features present in the incorrect judgment samples but not present in the correct judgment samples.
[0257] Furthermore, based on the distinguishing data features, the first quality parameter is updated, and the second quality parameter is obtained.
[0258] The AI-based quality rule configuration device proposed in this application acquires the first field data and its various data features (structural features, semantic features, and distribution features, etc.), and uses the training results of the master allocation model and historical structured data to achieve automatic matching of quality rules. This eliminates the need for users to have in-depth knowledge of databases and the details of quality rules, reducing the professional skills required and enabling ordinary users to participate in quality rule configuration. This application automatically acquires the first field data and data features based on current structured data, quickly calculates the matching degree using a pre-trained master allocation model, and then combines it with a natural language model to generate interpretable output, automatically completing the quality rule configuration. This significantly reduces the workload of manual configuration, greatly improving configuration efficiency. When processing a large number of newly accessed data tables, it can quickly complete the quality rule configuration, avoiding omissions and improving user understanding of quality rules. This application trains the master allocation model using historical structured data, performing quality rule matching based on a unified model and data. It is unaffected by subjective factors of operators, ensuring consistency in quality rule configuration and avoiding technical problems caused by inconsistent configuration interpretations due to differences in understanding among different operators, thus improving data quality assurance capabilities.
[0259] Having introduced the AI-based quality rule configuration apparatus proposed in the embodiments of this application, the following describes an embodiment of a computer-readable storage medium proposed in this application. This computer-readable storage medium stores a computer program, which, when executed by a processor, implements the AI-based quality rule configuration method as described in any of the above embodiments.
[0260] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0261] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.
Claims
1. A quality rule configuration method based on artificial intelligence, characterized in that, include: Based on the current structured data, retrieve the data for the first field; The current structured data is obtained in advance; The first field data is any field data in the current structured data that needs to be configured with quality rules; Based on the first field data, at least one first data feature is obtained; the first data feature is any one of structural features, semantic features, and distribution features. Based on each first data feature, the matching degree between the first field data and each quality rule is obtained through the main allocation model; the main allocation model is pre-trained based on historical structured data; each quality rule is pre-set. Based on the matching degree of each quality rule, an interpretable output is obtained through a natural language model; Configure quality rules for the first field data based on the interpretability output and each matching degree; The main allocation model includes multiple sub-allocation models, and each sub-allocation model corresponds one-to-one with a quality rule. The master assignment model is trained based on historical structured data, including: Based on the aforementioned historical structured data, multiple second field data are obtained; Iterate through each quality rule to obtain the first quality rule; the first quality rule is any quality rule among the various quality rules that has not been trained to obtain the corresponding sub-assignment model; For each first quality rule, perform the following steps until a corresponding sub-assignment model has been trained based on each quality rule: Based on the first quality rule, a current sub-assignment model is obtained; the current sub-assignment model corresponds to the first quality rule, and the current sub-assignment model is updated from the historical sub-assignment model; the historical sub-assignment model includes the initial sub-assignment model; Based on the current sub-assignment model, the first quality rule is assigned to each second field data, and the current loss value is obtained; If the current loss value does not meet the second preset condition, then the current sub-assignment model is updated, and the current loss value is re-acquired until the re-acquired current loss value meets the second preset condition. The sub-assignment model whose current loss value meets the second preset condition is then taken as the trained sub-assignment model. Based on each trained sub-assignment model, the main assignment model is obtained; The sub-assignment model includes at least one first assignment network; each first assignment network has corresponding weights; the current sub-assignment model is updated from the historical sub-assignment model in the following manner: Based on the historical sub-assignment model, multiple third field data and multiple fourth field data are obtained from each second field data; the third field data are any field data in each second field data that should have been configured with the first quality rule, but were not configured with the first quality rule by the historical sub-assignment model; the fourth field data are any field data in each second field data that should not have been configured with the first quality rule, but were configured with the first quality rule by the historical sub-assignment model. Based on the data from each of the third and fourth fields, a second allocation network is obtained; the allocation accuracy of the second allocation network is greater than 50%. Based on the second allocation network, update each first weight in the historical sub-allocation model; the first weight is the weight in the historical sub-allocation model corresponding to the first allocation network; Add the second allocation network to the historical sub-allocation model to obtain the current sub-allocation model.
2. The quality rule configuration method based on artificial intelligence according to claim 1, characterized in that, The step of obtaining at least one first data feature based on the first field data includes: Based on the data in the first field, obtain the data definition language; Based on the data definition language, a field relationship graph is obtained through a graph neural network; the field relationship graph is the relationship graph formed by the first field and the remaining fields in the current structured data; Based on the field relationship graph, a feature vector is obtained; the feature vector is used to reflect at least the relationship strength between the first field and the remaining fields in the current structured data; Based on the feature vector, obtain the structural features of the first field data; And / or, The step of obtaining at least one first data feature based on the first field data includes: Based on the language model, the field name is obtained from the first field data; Obtain the similarity between the field name and each semantic tag in the domain thesaurus; the domain thesaurus is pre-set; Based on each similarity score, the semantic features of the first field data are obtained; And / or, The step of obtaining at least one first data feature based on the first field data includes: Based on the data in the first field, obtain multiple rows of data; Based on the multiple rows of data, key statistics are obtained; the key statistics include at least one or a combination of multiple of the following: null value rate, number of unique values, maximum value, minimum value, and proportion of high-frequency values. Based on the key statistics, the distribution characteristics of the first field data are obtained.
3. The quality rule configuration method based on artificial intelligence according to claim 2, characterized in that, The step of obtaining multiple rows of data based on the first field data includes: Based on the data in the first field, obtain the total number of rows; the total number of rows is the total number of rows of data in the first field. If the total number of rows is less than or equal to a first preset value, then all rows in the first field data are taken as the multi-row data; otherwise, multiple rows are randomly selected from the first field data, and the number of rows in the multi-row data is equal to the first preset value.
4. The quality rule configuration method based on artificial intelligence according to claim 3, characterized in that, The method of randomly extracting multiple rows of data from the first field further includes: Based on randomly selected data, feedback information is obtained; the feedback information includes at least whether the multiple rows of data are missing data values that must be selected. If the feedback information meets the first preset condition, the randomly selected data will be used as the multi-row data; otherwise, the first preset value will be updated, and data will be randomly selected and feedback information will be obtained again based on the updated first preset value until the re-obtained feedback information meets the first preset condition; the first preset condition is that the multi-row data does not lack the data values that must be selected.
5. The quality rule configuration method based on artificial intelligence according to claim 1, characterized in that, The process of obtaining the second allocation network based on the data in each third field and each fourth field includes: Based on the data in each third field and each fourth field, a first feature set and a second feature set are obtained; each data feature in the first feature set comes from the data in the third field; each data feature in the second feature set comes from the data in the fourth field. Based on the clustering algorithm, multiple first clusters are obtained from the first feature set, and multiple second clusters are obtained from the second feature set; the data features in each cluster are of the same type; Based on each first cluster and each second cluster, a third cluster and a fourth cluster are obtained; the third cluster is any cluster among each first cluster, and the fourth cluster is a cluster of the same type as the third cluster among each second cluster, and the data characteristics of the third cluster and the fourth cluster have the greatest difference. Based on the third cluster and the fourth cluster, a second allocation network is obtained; The step of updating each first weight in the historical sub-allocation model based on the second allocation network includes: Based on the second allocation network, the number of iteration updates is obtained; Based on the number of iterations, a second weight is obtained; the second weight is the weight corresponding to the second allocation network, and the second weight is positively correlated with the number of iterations. Based on the second weight, update each of the first weights in the historical sub-assignment model; Before updating each of the first weights in the historical sub-assignment model based on the second weight, the method further includes: The first quality rule is configured for each second field data based on the second allocation network, and the allocation accuracy is obtained; Based on the allocation accuracy, a correction coefficient is obtained; the correction coefficient is positively correlated with the allocation accuracy. The second weight is adjusted based on the aforementioned correction coefficient.
6. The quality rule configuration method based on artificial intelligence according to any one of claims 1 to 5, characterized in that, After configuring quality rules for the first field data based on the interpretability output and each matching degree, the method further includes: Based on the first field data, a second quality rule is obtained; the second quality rule is any quality rule configured in the first field data. Based on the second quality rule, a first quality parameter is obtained; the first quality parameter is the quality parameter corresponding to the second quality rule. Based on the first quality parameter, the false positive rate is obtained; If the false positive rate is greater than the second preset value, then update the first quality parameter, obtain the second quality parameter, and update the quality parameter corresponding to the second quality rule from the first quality parameter to the second quality parameter; If the false positive rate is greater than a second preset value, then update the first quality parameter and obtain the second quality parameter, including: Based on the data in the first field, obtain multiple incorrect judgment samples and multiple correct judgment samples; Based on each incorrect judgment sample and each correct judgment sample, distinguishing data features are obtained; the distinguishing data features are data features that are present in the incorrect judgment samples but not in the correct judgment samples. Based on the distinguishing data features, the first quality parameter is updated, and the second quality parameter is obtained.
7. A quality rule configuration device based on artificial intelligence, characterized in that, include: The read module is used to retrieve the first field data based on the current structured data; The current structured data is obtained in advance; The first field data is any field data in the current structured data that needs to be configured with quality rules; The processing module is configured to obtain at least one first data feature based on the first field data; the first data feature is any one of structural features, semantic features, and distribution features; Furthermore, based on each first data feature, the matching degree between the first field data and each quality rule is obtained through the main allocation model; the main allocation model is pre-trained based on historical structured data; each quality rule is pre-set. Furthermore, based on the matching degree of each quality rule, interpretable output is obtained through a natural language model; In addition, quality rules are configured for the first field data based on the interpretability output and each matching degree; The main allocation model includes multiple sub-allocation models, each of which corresponds one-to-one with a quality rule; the processing module is also used to obtain multiple second field data based on the historical structured data. And, iterate through each quality rule to obtain the first quality rule; the first quality rule is any quality rule among the various quality rules that has not been trained to obtain the corresponding sub-assignment model; Furthermore, for each first quality rule, the following steps are performed until a corresponding sub-assignment model is trained based on each quality rule: Furthermore, based on the first quality rule, a current sub-assignment model is obtained; the current sub-assignment model corresponds to the first quality rule, and the current sub-assignment model is updated from the historical sub-assignment model; the historical sub-assignment model includes the initial sub-assignment model; And, based on the current sub-assignment model, assign the first quality rule to each second field data and obtain the current loss value; Furthermore, if the current loss value does not meet the second preset condition, the current sub-assignment model is updated, and the current loss value is re-acquired until the re-acquired current loss value meets the second preset condition. The sub-assignment model whose current loss value meets the second preset condition is then used as the trained sub-assignment model. And, based on each trained sub-assignment model, obtain the main assignment model; The sub-assignment model includes at least one first assignment network; each first assignment network has corresponding weights; the processing module is further configured to obtain multiple third field data and multiple fourth field data based on the historical sub-assignment model from each second field data; the third field data are any field data in each second field data that should have been configured with the first quality rule, but were not configured with the first quality rule by the historical sub-assignment model; the fourth field data are any field data in each second field data that should not have been configured with the first quality rule, but were configured with the first quality rule by the historical sub-assignment model; Furthermore, based on the data in each of the third and fourth fields, a second allocation network is obtained; the allocation accuracy of the second allocation network is greater than 50%. Furthermore, based on the second allocation network, each first weight in the historical sub-allocation model is updated; the first weight is the weight in the historical sub-allocation model corresponding to the first allocation network. In addition, the second allocation network is added to the historical sub-allocation model to obtain the current sub-allocation model.
8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the artificial intelligence-based quality rule configuration method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Generating rules for data processing values of data fields from semantic tags of data fields
CN115380281A
Data quality machine learning model
US20230177379A1